Three-dimensional molecule generation in latent voxelization space
By encoding and decoding the voxelized representation of the input molecule, and using a molecular design computational model for denoising and sampling iteration, the problem of lack of systematicity and three-dimensional structure capture in the generation of three-dimensional molecules in the prior art is solved, and three-dimensional molecules exhibiting the desired properties are generated efficiently.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to efficiently generate three-dimensional molecules that exhibit desired properties, especially in molecular space exploration where the lack of systematic and accurate capture of three-dimensional structures prevents conventional methods from discovering molecules with optimal properties.
A molecular design computational model is used to encode and decode the voxelized representation of the input molecule. The model is trained to approximate the data distribution of the desired properties. The voxelized representation is used for denoising and sampling iteration to generate the output molecule, thus avoiding indiscriminate search of the molecular space.
This method enables the generation of three-dimensional molecules exhibiting desired properties in a potential voxelized space, improving the systematic nature of molecule generation and the accuracy of three-dimensional structures, thereby enhancing the efficiency and quality of generated molecules.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 502,529 entitled “THREE-DIMENSIONAL MOLECULE GENERATION BY DENOISING VOXEL GRIDS”, filed May 16, 2023; U.S. Provisional Application No. 63 / 586,263 entitled “THREE-DIMENSIONAL MOLECULE GENERATION BY DENOISING VOXEL GRIDS”, filed September 28, 2023; and U.S. Provisional Application No. 63 / 623,062 entitled “THREE-DIMENSIONAL MOLECULE GENERATION BY DENOISING VOXEL GRIDS”, filed January 19, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The topics discussed in this article generally relate to generative artificial intelligence, and more specifically to machine learning-enabled techniques for generating representations of three-dimensional molecules in discrete and potentially voxelized spaces. Background Technology
[0004] A molecule is a group of two or more atoms linked together by chemical bonds. Molecules form the smallest identifiable unit, and a pure substance can be broken down into such units while still retaining its composition and chemical properties. An example of a molecule is a small molecule, which is a low-weight compound with a molecular weight between approximately 100 and 1000 Daltons. Due to its many compelling advantages, small molecule therapeutics that modulate biochemical processes to diagnose, treat, and prevent a range of diseases have become a cornerstone of modern pharmacology. For example, small molecule drugs can penetrate cell membranes to reach intracellular targets. Furthermore, small molecule drugs are suitable for a wide variety of therapeutic applications. For example, small molecule drugs can be formulated as pills and capsules, intravenous or subcutaneous injections, inhaled medications, or suppositories. The development of small molecule drugs can be further extended to tailoring various pharmacokinetic properties, including release, absorption, distribution, metabolism, potency, efficacy, phenotypic effects, and excretion.
[0005] In contrast, macromolecules (also known as biopharmaceuticals, biologicals, or biologics) can range in molecular weight from approximately 3,000 Daltons to 150,000 Daltons. Macromolecular drugs are typically derivatives of natural human proteins that regulate many important cellular functions, such as enzymatic reactions, molecular transport, regulation and execution of numerous biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, and intercellular communication. A single macromolecule typically has more than 1,300 amino acid residues linked by peptide bonds to form one or more polypeptides. Due to their size and complexity, macromolecular drugs are generated through engineered cellular recombination rather than chemical synthesis, as is the case with most small molecule drugs. Furthermore, because oral administration is ineffective, macromolecular therapeutics are typically delivered by injection or infusion. The development of macromolecular drugs may require designing one or more sequences of amino acid residues capable of binding to targets (e.g., proteins, nucleic acids, etc.) that possess sufficient specificity and are free from undesirable properties such as immunogenicity, self-association, and instability. Summary of the Invention
[0006] Systems, methods, and articles of manufacture for generating three-dimensional molecules in a latent voxelized space are provided, including computer program products. In one aspect, a system for enabling machine learning-based three-dimensional molecule generation is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that provides operation when executed by the at least one processor. The operation may include: encoding a voxelized representation of an input molecule to generate an embedding of the input molecule having fewer features than the voxelized representation of the input molecule; applying a molecular design computational model to update the embedding of the input molecule, wherein the molecular design computational model has been trained to approximate a data distribution of molecules exhibiting one or more desired properties by: taking in a corrupted embedding of a voxelized representation of a sample molecule exhibiting the one or more desired properties as input, and recovering the embedding of the voxelized representation of the sample molecule from the corrupted embedding, the molecular design computational model updating the embedding of the input molecule to increase the likelihood that the resulting updated embedding is within the data distribution; and generating a voxelized representation of an output molecule by decoding at least the resulting updated embedding.
[0007] On the other hand, a method for enabling machine learning to generate three-dimensional molecules is provided. The method may include: encoding a voxelized representation of an input molecule to generate an embedding of the input molecule having fewer features compared to the voxelized representation of the input molecule; applying a molecular design computational model to update the embedding of the input molecule, wherein the molecular design computational model has been trained to approximate a data distribution of molecules exhibiting one or more desired properties by: taking in a corrupted embedding of a voxelized representation of a sample molecule exhibiting the one or more desired properties as input, and recovering the embedding of the voxelized representation of the sample molecule from the corrupted embedding, the molecular design computational model updating the embedding of the input molecule to increase the likelihood that the resulting updated embedding is within the data distribution; and generating a voxelized representation of an output molecule by decoding at least the resulting updated embedding.
[0008] On the other hand, a computer program product for enabling machine learning to generate three-dimensional molecules is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations. These operations may include: encoding a voxelized representation of an input molecule to generate an embedding of the input molecule having fewer features than the voxelized representation of the input molecule; applying a molecular design computational model to update the embedding of the input molecule, wherein the molecular design computational model has been trained to approximate a data distribution of molecules exhibiting one or more desired properties by: taking in a corrupted embedding of a voxelized representation of a sample molecule exhibiting the one or more desired properties as input, and recovering the embedding of the voxelized representation of the sample molecule from the corrupted embedding, the molecular design computational model updating the embedding of the input molecule to increase the likelihood that the resulting updated embedding is within the data distribution; and generating a voxelized representation of an output molecule by at least decoding the resulting updated embedding.
[0009] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.
[0010] In some variations, the data distribution can be a noisy data distribution filled with noisy embeddings of voxelized representations of the molecule exhibiting one or more desired properties. The voxelized representation of the output molecule can be further generated by denoising the noisy voxelized representation of the output molecule, which is generated by decoding the resulting updated embedding.
[0011] In some variations, a vector quantization variational autoencoder (VQ-VAE) can be applied to encode the voxelized representation of the input molecule and decode the resulting updated embedding.
[0012] In some variations, the molecule's embedding can be a discrete latent embedding vector generated by quantizing the corresponding continuous latent embedding. This quantization may include matching the corresponding continuous latent embedding with a vector in the codebook of embeddings using a nearest neighbor search.
[0013] In some variations, the voxelized representation of the input molecule can be encoded by at least the following: compressing multiple atomic density values including the voxelized representation of the input molecule, such that the embedding of the input molecule includes fewer features compared to the voxelized representation of the input molecule.
[0014] In some variations, the voxelized representation of a molecule may include multiple voxels organized into a three-dimensional voxel grid. Each atom in the molecule may be represented as a continuous density across one or more voxels in the three-dimensional voxel grid.
[0015] In some variations, the continuous density of each atom in the molecule can be centered on the center of each atom. A first voxel located further away from any atom in the molecule can be associated with a lower atomic density value compared to a second voxel located near the center of an atom in the molecule.
[0016] In some variations, each voxel in the three-dimensional voxel grid can be associated with a value indicating the atomic density at the corresponding location.
[0017] In some variations, the voxelized representation of a molecule may include one or more channels. Each channel may correspond to a type of atom present in the molecule.
[0018] In some variations, the voxelized representation of a molecule can collectively represent the type and location of one or more atoms present in the molecule.
[0019] In some variations, the embedding of the input molecule can be updated based at least on a function parameterized by multiple parameters of the molecular design computational model. This function can output a value indicating the likelihood of the resulting updated embedding within the data distribution.
[0020] In some variations, this function can be a scoring function. The value output by this function can be a score that indicates the local variation in the density of the noisy data distribution at the location of each updated noisy embedding generated by updating the noisy embedding of the input molecule.
[0021] In some variations, the molecular design computational model may update the embedding of the input molecule by at least the following: applying the molecular design computational model to update the embedding of the input molecule to generate a first updated embedding; applying the molecular design computational model to update the embedding of the input molecule to generate a second updated embedding; applying a function parameterized by a plurality of parameters of the molecular design computational model to determine: (i) a first value indicating a first local change in the density of the data distribution at a first position occupied by the first updated embedding, and (ii) a second value indicating a second local change in the density of the data distribution at a second position occupied by the second updated embedding; and applying the molecular design computational model to further update the first updated embedding rather than the second updated embedding based at least on the first value and the second value.
[0022] In some variations, the molecular design computational model can be applied to further update the first updated embedding until one or more criteria are met. The one or more criteria may include at least one of the following: (i) the embedding of the input molecule has been updated by a threshold amount, (ii) the first value of the first updated embedding reaches one or more thresholds, and (iii) an output molecule of the threshold amount has been generated.
[0023] In some variations, the molecular design computational model can be applied to further modify the first updated embedding rather than the second updated embedding based at least on the following: the first value and the second value indicate that the first updated embedding is more likely to be within the data distribution than the second updated embedding.
[0024] In some variations, the molecular design computational model can be applied to further modify the first updated embedding rather than the second updated embedding based at least on the following: the first value and the second value indicate that the first updated embedding was sampled from a higher density region of the data distribution compared to the second updated embedding.
[0025] In some variations, the voxelized representation of the output molecule can be converted into a one-dimensional representation of the output molecule and / or a two-dimensional representation of the output molecule.
[0026] In some variations, the voxelized representation of the output molecule can be transformed by at least the following: determining the location of one or more atoms in the output molecule by detecting at least one or more peaks among a plurality of atomic density values including the voxelized representation of the output molecule, and determining one or more interconnecting bonds based at least on the location of the one or more atoms.
[0027] Systems, methods, and articles of manufacture for generating three-dimensional molecules in a latent voxelized space are provided, including computer program products. In one aspect, a system for enabling machine learning-enabled three-dimensional molecule generation is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that provides operation when executed by the at least one processor. The operation may include: generating a training dataset comprising a plurality of training samples, each training sample in the training dataset comprising a corrupted embedding generated by at least adding noise to a noisy voxelized representation of a sample molecule exhibiting one or more desired properties; training a molecular design computational model at least based on the training dataset to approximate a data distribution of molecules exhibiting the one or more desired properties, the training comprising: applying the molecular design computational model to recover an uncorrupted embedding of the noisy voxelized representation of the sample molecule from the corrupted embedding of the noisy voxelized representation of the sample molecule; and optionally, applying the molecular design computational model to generate an output molecule by at least: denoising the embedding of the voxelized representation of the input molecule and decoding the updated embedding generated therefrom to generate a voxelized representation of the output molecule.
[0028] On the other hand, a method for enabling machine learning to generate three-dimensional molecules is provided. The method may include: generating a training dataset comprising a plurality of training samples, each training sample in the training dataset comprising a corrupted embedding generated by adding at least noise to a noisy voxelized representation of a sample molecule exhibiting one or more desired properties; training a molecular design computational model, at least based on the training dataset, to approximate a data distribution of molecules exhibiting the one or more desired properties, the training comprising: applying the molecular design computational model to recover an uncorrupted embedding of the noisy voxelized representation of the sample molecule from the corrupted embedding of the noisy voxelized representation of the sample molecule; and optionally, applying the molecular design computational model to generate an output molecule by at least: denoising the embedding of the voxelized representation of the input molecule and decoding the updated embedding generated therefrom to generate a voxelized representation of the output molecule.
[0029] On the other hand, a computer program product for enabling machine learning to generate three-dimensional molecules is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation. The operation may include: generating a training dataset comprising a plurality of training samples, each training sample in the training dataset comprising a corrupted embedding generated by at least adding noise to a noisy voxelized representation of a sample molecule exhibiting one or more desired properties; training a molecular design computational model, at least based on the training dataset, to approximate a data distribution of molecules exhibiting the one or more desired properties, the training comprising: applying the molecular design computational model to recover an uncorrupted embedding of the noisy voxelized representation of the sample molecule from the corrupted embedding of the noisy voxelized representation; and optionally, applying the molecular design computational model to generate an output molecule by at least: denoising the embedding of the voxelized representation of the input molecule and decoding the updated embedding generated therefrom to generate a voxelized representation of the output molecule.
[0030] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.
[0031] In some variations, the noisy voxelization representation of the sample molecule may include multiple voxels organized into a three-dimensional voxel grid. Each atom in the sample molecule may be represented as a continuous density across one or more voxels in the three-dimensional voxel grid.
[0032] In some variations, the continuous density of each atom in the sample molecule can be centered on the center of each atom. A first voxel located further away from any atom in the sample molecule is associated with a lower atomic density value compared to a second voxel located near the center of an atom in the sample molecule.
[0033] In some variations, each voxel in the three-dimensional voxel grid can be associated with a value indicating the atomic density at the corresponding location.
[0034] In some variations, the noisy voxelization representation of the sample molecule may include one or more channels. Each channel may correspond to the type of atom present in the sample molecule.
[0035] In some variants, the noisy voxelization representation of the sample molecule can collectively represent the type and location of one or more atoms present in the sample molecule.
[0036] In some variations, training the molecular design computation model may include adjusting multiple parameters of the molecular design computation model to reduce the difference between the recovered embedding generated by the molecular design computation model and the undamaged embedding of the noisy voxelized representation of the sample molecule.
[0037] In some variations, the multiple parameters of the molecular design computational model can be parameterized to a function. These multiple parameters can be adjusted such that the function outputs values indicating local variations in the density of the data distribution of molecules exhibiting one or more desired properties.
[0038] In some variations, the molecular design computational model can denoise the embedding of the voxelized representation of the input molecule by updating at least one value of a plurality of atomic density values present in the embedding that are present in the voxelized representation of the input molecule.
[0039] In some variations, updating the atomic density value of one or more voxels in the at least one channel of the voxelization representation of the input molecule may correspond to updating at least one of the types and / or positions of one or more atoms present in the input molecule.
[0040] In some variants, the molecular design computational model can denoise the embedding of the voxelized representation of the input molecule through multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until one or more criteria are met.
[0041] In some variations, the one or more criteria may include at least one of the following: (i) an iteration of the threshold amount of gradient-based Markov chain Monte Carlo (MCMC) sampling, (ii) an updated embedding obtained from sampling a region with threshold density, and (iii) an output molecule of the threshold amount has been generated.
[0042] In some variations, the molecular design computational model can generate the voxelized representation of the output molecule by at least the following: applying a first update to the embedding of the voxelized representation of the input molecule to generate a first updated embedding; applying a second update to the embedding of the voxelized representation of the input molecule to generate a second updated embedding; and further updating the first updated embedding instead of the second updated embedding when it is determined that the first updated embedding was sampled from a higher density region of the data distribution compared to the second updated embedding.
[0043] In some variations, the voxelized representation of the output molecule can be converted into a one-dimensional and / or two-dimensional representation of the output molecule.
[0044] In some variations, training the molecular design computation model may include: applying the molecular design computation model with a first adjustment to generate a first recovered embedding of the noisy voxelized representation of the sample molecule; determining a first mean squared error (MSE) to quantify a first difference between the first recovered embedding of the noisy voxelized representation of the sample molecule and the undamaged embedding; applying the molecular design computation model with a second adjustment to generate a second recovered embedding of the noisy voxelized representation of the sample molecule; determining a second mean squared error (MSE) to quantify a second difference between the second recovered embedding of the noisy voxelized representation of the sample molecule and the undamaged embedding; and further adjusting the molecular design computation model with the first adjustment instead of the molecular design computation model with the second adjustment when it is determined that the first mean squared error (MSE) is less than the second mean squared error (MSE).
[0045] In some variations, the molecular design computational model can be further tuned until one or more criteria are met. These criteria include at least one of the following: (i) an iteration of the threshold amount used to tune the molecular design computational model has been performed, and (ii) a recovered embedding exhibiting a threshold mean squared error (MSE) value has been generated.
[0046] In some variations, an autoencoder that includes an encoder and a decoder can be trained. The training of the autoencoder may include: training the encoder to encode the noisy voxelized representation of the sample molecule such that the decoder is able to recover the voxelized representation of the sample molecule from the resulting embedding of the noisy voxelized representation of the sample molecule.
[0047] In some variations, the autoencoder can be a vector quantization variational autoencoder (VQ-VAE), in which the encoder generates a continuous latent embedding of the sample molecule and then quantizes the continuous latent embedding of the sample molecule into a discrete latent embedding by matching the vectors in the codebook of the embeddings via nearest neighbor lookup, for decoding by the decoder.
[0048] Implementations of the present subject matter may include, but are not limited to, methods consistent with the descriptions provided herein, and articles of art comprising a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to perform operations implementing one or more of the described features. Similarly, computer systems comprising one or more processors and one or more memories coupled to the one or more processors are also described. Memory that may include a non-transitory computer-readable or machine-readable storage medium may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implementations consistent with one or more implementations of the present subject matter may be implemented by one or more data processors existing in a single computing system or multiple computing systems. Such multiple computing systems may be interconnected and may exchange data and / or commands or other instructions via one or more connections, including, for example, direct connections between one or more of the multiple computing systems via a network (e.g., the Internet, wireless wide area network, local area network, wide area network, wired network, etc.).
[0049] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Further features and advantages of the subject matter described herein will become apparent from the description, drawings, and claims. While some features of the subject matter currently disclosed are described for purposes of illustrative purposes relating to the computational design of molecules, including drug molecules, it should be readily understood that these features are not intended to constitute limitation. The claims following this disclosure are intended to define the scope of the protected subject matter. Attached Figure Description
[0050] The accompanying drawings, incorporated in and forming part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the specification, help explain some principles associated with the disclosed embodiments. In the drawings,
[0051] According to some exemplary embodiments, FIG1A depicts a system diagram illustrating an example of a molecular design system;
[0052] Figure 1B depicts a system diagram illustrating another example of a molecular design system according to some exemplary embodiments;
[0053] Figure 2 depicts a flowchart illustrating an example of a process for machine learning-enabled 3D molecular generation in voxelized space according to some exemplary embodiments;
[0054] Figure 3A depicts a flowchart illustrating an example of a process for training a computational model for molecular design to generate three-dimensional molecules in voxelized space, according to some exemplary embodiments.
[0055] Figure 3B depicts a flowchart illustrating an example of a process for applying a molecular design computational model to generate three-dimensional molecules in voxelized space, according to some exemplary embodiments.
[0056] Figure 3C depicts a flowchart illustrating an example of a process for applying a molecular design computational model to generate three-dimensional molecules in voxelized space, according to some exemplary embodiments.
[0057] Figure 4 illustrates an example of a voxelized representation of a molecule according to some exemplary embodiments;
[0058] Figure 5A depicts a schematic diagram illustrating an example of a process for training a denoising engine to denoise a noisy voxelized representation of a molecule, according to some exemplary embodiments.
[0059] Figure 5B depicts a schematic diagram illustrating an example of a walk-and-jump sampling scheme according to some exemplary embodiments;
[0060] Figure 5C depicts a schematic diagram illustrating an example of a computational model for molecular design, according to some exemplary embodiments, generating a three-dimensional molecule by denoising a voxelized molecular representation.
[0061] Figure 5D depicts a schematic diagram illustrating an example of a process for generating other molecular representations from a voxelized representation of a molecule, according to some exemplary embodiments.
[0062] Figure 6 illustrates an example of a computational model for molecular design, according to some exemplary embodiments, generating a voxelized representation of a molecule by operating in a noisy latent voxelization space.
[0063] Figure 7 depicts a graph illustrating the effect of noise levels on the generative performance of molecular design computational models according to some exemplary embodiments;
[0064] Figure 8 illustrates a schematic diagram showing the effect of the number of sampling iterations in Markov chain Monte Carlo (MCMC) sampling on the generative performance of a molecular design computational model, according to some exemplary embodiments.
[0065] Figure 9A depicts an example of a voxelized representation of a molecule generated by a molecular design computational model trained on the QM9 molecular dataset, according to some exemplary embodiments.
[0066] Figure 9B depicts an example of a voxelized representation of a molecule generated by a computational model of molecular design trained on the Geometry Integration of Molecular (GEOM) drug dataset, according to some exemplary embodiments.
[0067] Figure 10A depicts a graph of the cumulative distribution function (CDF) of strain energy for molecules in the QM9 molecular dataset according to some exemplary embodiments, molecules generated by a conventional generative model, and molecules generated by a molecular design computational model trained on the QM9 molecular dataset.
[0068] Figure 10B depicts a graph comparing the empirical distribution of the number of atoms per molecule in the QM9 molecular dataset according to some exemplary embodiments with the empirical distribution of the number of atoms in molecules generated by a molecular design computational model trained on the QM9 molecular dataset.
[0069] Figure 11A depicts a graph of the cumulative distribution function (CDF) of strain energy for molecules in the Geometry Integration of Molecular (GEOM) drug dataset according to some exemplary embodiments, molecules generated by a conventional generative model, and molecules generated by a molecular design computational model trained on the GEOM drug dataset.
[0070] Figure 11B depicts a graph comparing the empirical distribution of the number of atoms per molecule in a Geometry of Molecular Integration (GEOM) drug dataset according to some exemplary embodiments with the empirical distribution of the number of atoms in molecules generated by a molecular design computational model trained on the GEOM drug dataset; and
[0071] Figure 12A depicts a schematic diagram showing a comparison of seed generation of a geometrically integrated molecular (GEOM) drug in discrete voxelization space and potential voxelization space according to some exemplary embodiments;
[0072] Figure 12B depicts a schematic diagram showing a comparison of seed generation of PubChem drugs in discrete voxelization space and potential voxelization space according to some exemplary embodiments.
[0073] Figure 12C depicts a molecular diagram of an additional example of a molecule generated by seed generation of an actual drug in a potential voxelized space.
[0074] Figure 12D depicts a molecular diagram of additional examples of molecules generated by de novo generation of an actual drug in a potential voxelization space.
[0075] Figure 13 depicts a graph comparing seed generation of geometrically integrated (GEOM) drugs in discrete voxelization space and potential voxelization space according to some exemplary embodiments;
[0076] Figure 14 depicts a block diagram illustrating an example of a computing system according to some exemplary embodiments.
[0077] In practical applications, similar reference numerals indicate similar structures, features, or elements. Detailed Implementation
[0078] Generating new molecules with desired properties is a crucial task in chemistry with applications in many scientific fields. In the context of drug discovery, conventional computational techniques for generating molecules with drug-like properties require searching the molecular space (or chemical space) occupied by every possible chemical compound (e.g., every possible combination of atoms of two or more chemical elements). For example, some search-based methods may include scoring and ranking different molecules in the molecular space based on one or more drug-like properties, such as affinity, specificity, bioactivity, and exploitability. However, the aforementioned molecular space (estimated to contain 10...) 60 The number of possible chemical compounds in molecular space is enormous and scales exponentially with molecular size (e.g., the number of constituent atoms). Even a tiny fraction of molecular space can contain billions to trillions of molecules. Using the computational resources of existing technologies, conventional search-based methods can only explore a small fraction of molecular space, such as small regions of molecular space selected based on prior domain knowledge. This limitation in search scope means that conventional search-based methods may overlook molecules with optimal properties. Furthermore, conventional search-based methods do not explore molecular space in a principled manner, which prevents generative processes from being conditioned on specific properties.
[0079] Furthermore, whether a molecule exhibits certain desired properties can depend on its conformation (or three-dimensional structure). For example, the binding affinity between a drug molecule and a target molecule (e.g., a protein, nucleic acid, etc.) may depend on the ability of the drug molecule to adopt a conformation complementary to that of the target molecule (or three-dimensional structure). Moreover, molecules are flexible, meaning that a single molecule can present one of many possible conformations (or three-dimensional structures). In some cases, a group of the same molecules can exist as an integration of many different conformations in equilibrium, but not every possible conformation is associated with the desired property. For example, in the context of binding affinity, the biologically active conformation of a molecule can be one or more of the conformations it exhibits in solution or a new conformation induced by interaction with the target molecule. However, one-dimensional representations of molecules (e.g., simplified linear input canonical (SMILES) strings) or two-dimensional representations (e.g., molecular diagrams) cannot adequately capture the conformation (or three-dimensional structure) of a molecule. Therefore, when a molecular design computational model operates on a one-dimensional or two-dimensional representation of the input molecule, the resulting output molecule may not exhibit the conformation (or three-dimensional structure) associated with that one or more desired properties.
[0080] Various exemplary embodiments of this disclosure can improve the computational resources of the prior art by providing a molecular design computational model that generates output molecules by exploring molecular space (or chemical space) in a principled manner, rather than by indiscriminately searching a finite portion of molecular space. For example, in some cases, the molecular design computational model can be trained to approximate a data distribution of molecules exhibiting one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, exploitability, etc.). Training the molecular design computational model may include determining the parameters of a function (e.g., a scoring function) such that the output of the function is a value indicating a change in density across the data distribution. In some cases, the molecular design computational model may sample the data distribution to generate output molecules that also exhibit the one or more desired properties. For example, in some cases, the molecular design computational model may sample the data distribution by denoising the input molecules (such as a voxelized representation of the input molecules) through multiple sampling iterations. During each sampling iteration, the molecular design computational model may update the input molecules to remove a portion of the noise present in the input molecules. This generates updated molecules (e.g., voxelized representations of the updated molecules) that constitute samples selected from the data distribution. As described in more detail below, sampling can be guided by a function such that each successive sample (or updated molecule) is selected from regions of increasing density in the data distribution, regions that are more likely to be occupied by molecules exhibiting one or more of the properties.
[0081] In some exemplary embodiments, a molecular design computational model that operates on a three-dimensional representation of the input molecule can increase (or maximize) the likelihood that the output molecule will exhibit one or more desired properties. For example, in some cases, the molecular design computational model can generate the output molecule by denoising the three-dimensional representation of the input molecule at least, for example, through multiple sampling iterations. In other cases, the molecular design computational model can generate the output molecule by denoising a voxelized representation of the input molecule instead of its conventional three-dimensional representation. A conventional three-dimensional representation of the input molecule (such as a point cloud representation) can specify the conformation (or three-dimensional structure) of the input molecule by specifying at least the coordinates of the constituent atoms (e.g., in Euclidean space). However, a conventional three-dimensional representation of the input molecule can impose many constraints on the generative process. For example, in order for the molecular design computational model to operate on a conventional three-dimensional representation of the input molecule, the number of atoms in the output molecule to be generated must be known a priori. In order for molecular design computational models to approximate the distribution of atom types in the output molecule (which forms a discrete distribution, while the locations of atoms in the output molecule (e.g., atomic coordinates in Euclidean space) form a continuous distribution), denoising the conventional 3D representation of the input molecule may also require some workarounds. Furthermore, the conventional 3D representation of the input molecule may not adequately capture long-range dependencies existing across multiple atoms, especially as the number of constituent atoms increases.
[0082] In some exemplary embodiments, a voxelized representation of a molecule (e.g., an input molecule) can avoid the aforementioned limitations by representing the input molecule as a continuous distribution of atomic density across a voxel grid centered on the atomic coordinates of each individual atom present in the molecule. For example, in a graph network representation of a molecule, the dependency between two adjacent atoms can be represented by interconnecting edges. However, these edges may not adequately capture longer-range dependencies, such as those between non-adjacent atoms. In contrast, a voxelized representation of a molecule captures long-range dependencies between distant atoms better, even when the input molecule contains a large number of atoms. Furthermore, a molecular design computational model can operate on the voxelized representation of the input molecule to generate an output molecule without any prior knowledge of the number of atoms present in the output molecule. This is because the molecular design computational model can freely add or remove different types of atoms by updating the distribution of atomic density across the voxel grid. The voxelized representation of the input molecule also collectively represents the type and location of atoms in the input molecule, thus avoiding workarounds for reconciling these two different types of data distributions (e.g., a discrete distribution for atom type and a continuous distribution for atom location).
[0083] In some exemplary embodiments, a voxelized molecular representation of a molecule (such as an input molecule) can represent each atom in the molecule (e.g., the input molecule) as a continuous (e.g., Gaussian-like) density across one or more voxels in a voxel grid. In this context, the voxel grid is a three-dimensional voxel grid organized into continuous layers of rows and columns. Various instances of the voxel grid described herein may contain multiple voxels, each of which is a volume element (e.g., a three-dimensional cube) at the intersection of rows and columns. Each volume element may have a predetermined size, which may or may not be the same for all voxels in the voxel grid. In the case where the input molecule is a drug molecule, the voxelized representation of the input molecule may include a voxel grid containing n×n×n voxels (e.g., 32×32×32 voxels, 64×64×64 voxels, etc.). In some cases, each voxel in the voxel grid may be associated with a value indicating the atomic density at the corresponding location. For example, a first voxel associated with a higher atomic density is more likely to be part of an atom than a second voxel associated with a lower atomic density. It should be understood that the volume of a single atom can span one or more voxels. In some cases, atomic density can also exist centered on the atom in the input molecule, meaning that the atomic density of a single atom spanning multiple voxels can include the voxel centered on the atom. A voxel with an atomic density of 0 can be located away from any atom in the input molecule, while a voxel with an atomic density of 1 can be located at the center of the atom in the input molecule. Furthermore, in some cases, the voxelized representation of a molecule (e.g., the input molecule) can include multiple channels, each of which corresponds to a type of atom that can exist in the input molecule. “Atom type” can refer to the individual chemical element to which the atom belongs. The voxelized representation can include multiple channels, one for each type of atom present in the molecule or at least one for each type of heavy atom. For example, in some cases, the voxelized representation of a molecule (e.g., the input molecule) can include a first channel corresponding to a first type of atom that can exist in the input molecule (e.g., carbon (C) atom) and a second channel corresponding to a second type of atom that can exist in the input molecule (e.g., nitrogen (N) atom). Each voxel in the first channel can be associated with a value indicating the density of atoms of a first atom type at the corresponding position, while each voxel in the second channel can be associated with a value indicating the density of atoms of a second atom type at the corresponding position. Therefore, as described in more detail below, a molecular design computational model can denoise the voxelized representation of an input molecule by at least updating (e.g., over multiple sampling iterations) the atomic densities of one or more voxels in at least one channel of the voxelized representation of the input molecule. That is, in some cases, the term "denoising" refers to updating the voxelized representation of the input molecule, which may include updating the atomic density of at least one voxel in the voxelized representation of the input molecule.In some cases, updating the atomic density of a voxel in a channel of the voxel representation of the input molecule can change the possibility that the voxel is part of the type of atoms associated with that channel.
[0084] In some exemplary embodiments, the molecular design computational model may denoise the input molecule (e.g., a voxelized representation of the input molecule) through multiple sampling iterations, wherein each sampling iteration generates an updated voxelized representation different from the voxelized representation of the input molecule. In some cases, each updated voxelized representation may include a sample of a data distribution (e.g., a voxelized representation of a molecule) of molecules exhibiting one or more desired properties. In this context, the term "data distribution" may refer to a population of different molecular compositions and conformations (or three-dimensional structures). Those molecules exhibiting the one or more desired properties may cluster in higher-density regions of the data distribution, meaning that the molecular design computational model should sample each updated voxelized representation from these higher-density regions. However, such a data distribution may be too high-dimensional to be directly approximated. For example, calculating the probability density function (PDF) characterizing the probability of different molecules in the data distribution requires a normalization constant. In the case of molecular design, this normalization constant may correspond to the total number of molecules in the data distribution, and estimating the total number of molecules may be impractical. Therefore, in some cases, molecular design computational models can be trained to approximate the data distribution by determining a function (such as a scoring function) that estimates the gradient (or density variation) across the data distribution. As described in more detail below, molecular design computational models can use functions to guide sampling from the data distribution to an updated voxelized representation, such that each successive sample is selected from regions of gradually increasing density in the data distribution.
[0085] As described above, in some cases, denoising an input molecule through multiple sampling iterations (which includes continuously updating the voxelized representation of the input molecule) is equivalent to selecting successive samples from the data distribution of the molecule (e.g., the data distribution of the voxelized representation of the molecule), where each sample corresponds to an updated voxelized representation that is different from the voxelized representation of the input molecule. For example, in some cases, the voxelized representation of the input molecule can be denoised by at least updating the atomic density of one or more voxels in at least one channel of the voxel grid forming the voxelized representation of the input molecule. In some cases, molecules in higher-density regions of the data distribution may exhibit one or more desired properties, including, for example, drug-like properties such as affinity, specificity, bioactivity, exploitability, etc. Operating on a three-dimensional representation of the input molecule (such as the voxelized representation of the input molecule) increases the possibility that the conformation (or three-dimensional structure) of the resulting output molecule is selected from a higher-density region of the data distribution and thus exhibits the one or more desired properties.
[0086] In some cases, a molecular design computation model can be trained to approximate a data distribution by training it with a training dataset of known molecules exhibiting one or more desired properties (e.g., the PubChem dataset, the QM9 molecular dataset, the Geometry Integration of Molecular (GEOM) drug dataset, etc.). For example, in some cases, the molecular design computation model can be trained to approximate a data distribution by determining, at least for example, a function (e.g., a scoring function) that approximates different densities across the data distribution through Bayesian inference. In some cases, the function can be parameterized by the molecular design computation model, meaning that the parameters of the function (e.g., the scoring function) are parameters of the molecular design computation model that are adjusted when the molecular design computation model is trained to approximate the data distribution. In some cases, high-density regions of the data distribution may be filled with molecules similar to known molecules exhibiting one or more desired properties, while low-density regions of the data distribution may be filled with molecules dissimilar to known molecules exhibiting one or more desired properties. The scoring function of the data distribution can indicate transitions between regions of different densities in the data distribution, including, for example, transitions between regions of higher density and regions of lower density in the data distribution. Therefore, once trained, the molecular design computational model can sample the data distribution based on a scoring function, so that each consecutive sample (or molecule) is selected from regions of increasing density in the data distribution.
[0087] In some exemplary embodiments, a molecular design computational model can be trained to denoise corrupted 3D representations of known molecules from a training dataset and recover the initial 3D representation of the known molecules. For example, in some cases, a corrupted 3D representation of a known molecule can be generated by corrupting the 3D representation of the known molecule with noise (e.g., Gaussian noise, such as isotropic Gaussian noise). Training the molecular design computational model may include adjusting one or more parameters of the molecular design computational model (e.g., weights, biases, etc.) to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the recovered 3D representation of the known molecule and the initial 3D representation of the known molecule.
[0088] In some exemplary embodiments, to avoid overfitting the molecular design computation model to known molecules in the training dataset, the molecular design computation model may be trained to recover a noisy form of the 3D representation of the known molecules in the training dataset instead of the initial 3D representation. That is, the 3D representation of each known molecule in the training dataset may be doped with additional noise, but this noise should not be confused with the noise that the molecular design computation model is trained to remove from the corrupted 3D representation of each known molecule in the training dataset. In other words, in some cases, the molecular design computation model may be trained based on a training dataset that includes noisy 3D representations of known molecules and their corrupted forms. As described in more detail below, in some cases, a noisy 3D representation of a known molecule may be generated by doping the 3D representation (e.g., voxelized representation) of the known molecule with a first amount of noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to smooth the density of the data distribution of the known molecules while still retaining at least a portion of the conformation (e.g., 3D structure) of the known molecules, thereby obtaining a noisy representation of the known molecules. Then, a second amount of noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) can be used to further corrupt the noisy 3D representation of the known molecule to generate a corrupted 3D representation. In some cases, the molecular design computation model can be trained to denoise the corrupted 3D representation of the known molecule (e.g., by removing the second amount of noise) and recover the noisy 3D representation of the known molecule (which still includes the first amount of noise). Furthermore, in some cases, the training of the molecular design computation model may include gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.), in which the parameters of the molecular design computation model are adjusted through successive sampling iterations to increase the similarity between the 3D representation of each known molecule recovered by the molecular design computation model from the corresponding corrupted 3D representation of the known molecule in the training dataset and the noisy 3D representation of the sample molecules in the training dataset (e.g., to reduce the mean squared error (MSE)). The scoring function derived in this way can capture data distributions with smoother density transitions, which mitigates mode collapse, in which molecular design computational models are less robust and can only generate a limited selection of output molecules (e.g., those in the immediate vicinity of known molecules in the data distribution).
[0089] As described in more detail below, during inference, a trained molecular design computational model can be applied to generate one or more output molecules by denoising the three-dimensional representation of the input molecule. In some cases, the input molecule can be a random molecule (e.g., a molecule with a random selection of atom types and / or locations) or a known molecule with one or more undesirable properties. This means that the three-dimensional representation of the input molecule may include at least some noise that needs to be removed so that the three-dimensional representation of the output molecule generated from it is consistent with those exhibiting one or more desired properties. The molecular design computational model can achieve this by traversing a smoothed density of the noisy data distribution of the noisy three-dimensional representation of the molecule towards regions with gradually increasing density, for example, through one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo, etc.). Each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling may include updating the three-dimensional representation of the input molecule, which is equivalent to selecting noisy three-dimensional representations of one or more molecules from different locations in the noisy data distribution. Although selected from noisy data distributions, the molecules corresponding to these noisy 3D representations may be less distorted than the initial 3D representation of the input molecule. Such molecules corresponding to these noisy 3D representations may be more consistent with molecules exhibiting the one or more desired properties compared to the input molecule. In some cases, the noisy 3D representations of molecules selected from noisy data distributions may undergo further denoising to recover the corresponding molecule by mapping the noisy 3D representation of each molecule from the noisy data distribution to the corresponding clean 3D representation of the molecule in the real data distribution exhibiting the one or more desired properties. It should be understood that sampling from noisy data distributions can provide several advantages over sampling from the real data distribution exhibiting the one or more desired properties. For example, sampling from a noisy data distribution exhibiting the one or more desired properties (which exhibits a smoother density transition) may be less susceptible to mode collapse than sampling from the real data distribution. In some cases, this may be because there are fewer steep or less abrupt gradients in the noisy data distribution than in the real data distribution, where steep gradients constrain sampling to the immediate vicinity of known molecules characterizing the real data distribution. In other words, while molecular design computational models can generate outputs with finite variation when sampled from real-world data distributions (e.g., the aforementioned phenomenon is referred to as "mode collapse"), sampling from noisy distributions can increase the variability of the model output. Furthermore, sampling from noisy data distributions filled with noisy voxelized representations of molecules offers additional advantages over sampling from noisy distributions filled with noisy conventional 3D representations of molecules (such as point cloud representations of molecules).For example, in some cases, computational molecular design models can be trained to operate on voxelized molecular representations and generate large drug-like molecules with greater ease of use, effectiveness, expressiveness, and scalability. Unlike when operating on conventional three-dimensional molecular representations (e.g., point cloud representations), operating on voxelized molecular representations allows the disclosed computational molecular design models to function without specifying the number of atoms present in the output molecule and without requiring workarounds to reconcile discrete distributions for atom types with continuous distributions for the atom locations associated with each molecule.
[0090] Despite the aforementioned advantages, operating on voxelized molecular representations can impose a significant computational burden that scales exponentially with molecular size (e.g., the number of constituent atoms). For example, while a small molecule containing 10 heavy atoms already requires a [32×32×32] voxel grid with 32,000 features (or atomic density values) per molecule, larger, more realistic drug-like molecules may require voxel grids with an exponentially larger number of points (e.g., a [64×64×64] voxel grid with 260,000 features (or atomic density values) per molecule). In some cases, applying molecular design computational models to operate on voxelized representations of larger, more realistic drug-like molecules can be a challenging task, as is the case when training molecular design computational models on large training datasets (e.g., training datasets containing millions of known molecular voxelized representations) to learn more diverse molecular (or chemical) spaces. Furthermore, in practice, most of the candidate molecules generated by molecular design computational models may not be successfully synthesized in the laboratory, even if the candidate molecules are realistic and valid. Therefore, it may be desirable to apply molecular design computational models to generate tens of thousands or even millions of candidate molecules. The computational burden associated with generating molecules (especially large molecules or large quantities of molecules) can be reduced by using molecular design computational models that operate on lower-dimensional embeddings of voxelized molecular representations. For example, in some cases, molecular design computational models can be trained to generate output molecules by denoising the embeddings of the three-dimensional representations of the input molecules. As described in more detail below, in some cases, molecular design computational models can be trained based on a training dataset comprising corrupted embeddings of sample molecules exhibiting one or more desired properties, each of which is generated by encoding a noisy three-dimensional representation (e.g., a voxelized representation) of a known molecule exhibiting the one or more desired properties, and then corrupting the resulting embedding by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.).
[0091] In some exemplary embodiments, encoding a three-dimensional representation of a molecule (such as a voxelized representation of the molecule) can project the three-dimensional representation of the molecule (e.g., a voxelized molecular representation) from a high-dimensional discrete space filled with the three-dimensional representation of the molecule (e.g., a discrete voxelized space filled with the voxelized molecular representation) into a lower-dimensional representation in a lower-dimensional latent space filled with corresponding molecular embeddings. In other words, encoding a three-dimensional representation of a molecule (such as a voxelized representation of the molecule) can take the three-dimensional representation of the molecule (e.g., a voxelized molecular representation) in a filled high-dimensional discrete space as input and produce a lower-dimensional representation of the molecule (i.e., a molecular embedding corresponding to the input three-dimensional representation) in a lower-dimensional latent space as output. Encoding a three-dimensional representation of a molecule (such as a voxelized representation of the molecule) can be performed using a machine learning model trained to identify the latent space representation of the input molecule from which the three-dimensional representation of the input molecule can be recovered. In some cases, each embedding in the latent space can be a latent space representation of the voxelized representation of the corresponding molecule. Therefore, the embedding of a voxelized representation of a molecule can have a different dimension or number of features than the voxelized representation itself. For example, in some cases, encoding the voxelized representation of a molecule can reduce the number of features present in the voxelized representation. Thus, the computational burden of denoising the voxelized representation of a molecule to generate one or more output molecules can be reduced by using a molecular design computational model that operates on the voxelized representation of a molecule by embedding it rather than directly on the voxelized representation, at least because the embedding contains fewer features. Furthermore, in some cases, a molecular design computational model can be trained to approximate a noisy data distribution of molecules exhibiting one or more desired properties, so that the one or more output molecules can be generated by sampling from it. This noisy data distribution (which can be filled with noisy embeddings of the voxelized representation of the molecule) can exhibit a smoother density transition than the corresponding true data distribution. Therefore, noisy data distributions can support more efficient sampling (e.g., via gradient-based Markov chain Monte Carlo (MCMC) sampling, such as Langevin Markov chain Monte Carlo, etc.), at least because noisy data distributions can exhibit fewer steep gradient changes, which would prevent molecular design computational models from fully exploring the data distribution when sampling from it.
[0092] Figures 1A and 1B depict system diagrams illustrating different instances of a molecular design system 100 according to some exemplary embodiments. Referring to Figures 1A and 1B, in some cases, the molecular design system 100 may include a molecular design engine 110, a training engine 120, and a client device 130. In the examples of the molecular design system 100 shown in Figures 1A and 1B, the molecular design engine 110, the training engine 120, and the client device 130 may be communicatively coupled via a network 140. The client device 130 may be a processor-based device, including, for example, a workstation, desktop computer, laptop computer, smartphone, tablet computer, wearable device, etc. The network 140 may be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. In the examples shown in Figures 1A and 1B, the molecular design computational model 115 may include a denoising model 117 trained to generate an output molecule 162 by denoising at least the input molecule 152. The denoising model is a machine learning model trained to take a damaged 3D representation of a molecule or its lower-dimensional embedding as input (where the damaged 3D representation of a molecule is a 3D representation of a molecule to which noise has been added) and to use training data including 3D representations of multiple known molecules and corresponding damaged 3D representations to produce a corresponding denoised 3D representation of the molecule or its lower-dimensional embedding as output. The known molecules in the training data may include multiple molecules exhibiting one or more desired properties. The denoising model may be an artificial neural network (ANN). The denoising model may be a deep learning model. The denoising model may be an encoder-decoder 3D convolutional neural network (CNN). For example, in some cases, the molecular design computational model 115 may apply a denoising model 117 to denoise the three-dimensional representation of the input molecule 152, or alternatively, to denoise the embedding 154 of the three-dimensional representation of the input molecule 152 in order to generate the output molecule 162. In some cases, the denoising model 117 may denoise the three-dimensional representation (or its embedding 154) of the input molecule 152 through multiple successive sampling iterations, wherein at each sampling iteration, a portion of the noise present in the three-dimensional representation (or its embedding 154) of the input molecule 152 is removed.
[0093] As described in more detail below, denoising the three-dimensional representation of input molecule 152 can alter the composition and / or conformation (or three-dimensional structure) of input molecule 152 such that the composition and conformation (or three-dimensional structure) of the resulting output molecule 162 is consistent with those of molecules exhibiting one or more desired properties. For example, in the case that output molecule 162 is a drug molecule, the one or more desired properties may include drug-like properties such as affinity, specificity, bioactivity, and exploitability. In some cases, whether output molecule 162 exhibits certain desired properties may depend on the conformation (or three-dimensional structure) exhibited by output molecule 162. Therefore, in some cases, applying denoising model 117 to the three-dimensional representation (or its embedding 154) of input molecule 152, rather than the one-dimensional or two-dimensional representation of input molecule 152, increases the likelihood that the resulting output molecule 162 exhibits a conformation (or three-dimensional structure) consistent with the one or more desired properties.
[0094] In some exemplary embodiments, a molecular design computational model 115 (including a denoising model 117) may be trained to learn or approximate a data distribution of molecules exhibiting one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, and exploitability). For example, in some cases, the molecular design computational model 115 may be trained based on a training dataset of known molecules exhibiting one or more desired properties (e.g., the PubChem dataset, the QM9 molecular dataset, the Geometry Integration of Molecular (GEOM) drug dataset, etc.) to approximate a data distribution of molecules exhibiting one or more desired properties. As described in more detail below, the molecular design computational model 115 may be trained to approximate a noisy data distribution filled with noisy three-dimensional representations of known molecules exhibiting one or more desired properties. Furthermore, the molecular design computational model 115 can be trained to operate in a discrete space filled with three-dimensional representations (e.g., voxelized representations) of molecules exhibiting the one or more desired properties, or alternatively in a latent space filled with embeddings of three-dimensional representations of molecules as lower-dimensional representations of molecules (e.g., voxelized representation embeddings).
[0095] To further illustrate, Figure 1A depicts an example of the molecular design engine 110 of Figure 1A, where the molecular design computational model 115 is trained to operate in a discrete space (three-dimensional voxelized representation), while the molecular design computational model 115 included in the example of the molecular design engine 110 shown in Figure 1B can be trained to operate in a latent space. The latent space may include continuous values (embeddings) for multiple features. In either instance of the molecular design engine 110, the molecular design computational model 115 can be trained to approximate a noisy data distribution, for example, by being trained based on a noisy three-dimensional representation of a molecule exhibiting one or more desired properties. For example, the example of the molecular design computational model 115 shown in Figure 1A can be trained to approximate a noisy discrete distribution filled with noisy three-dimensional representations of molecules, while the example of the molecular design computational model 115 shown in Figure 1B can be trained to approximate a noisy latent distribution filled with noisy embeddings of three-dimensional representations of molecules.
[0096] Referring first to Figure 1A, in some cases, the molecular design engine 110 may include a molecular design computational model 115 and a recovery model 118. Alternatively, Figure 1B depicts another instance of the molecular design engine 110, which may include an encoder 111 and a decoder 119 in addition to the molecular design computational model 115 and the recovery model 118. In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 115 to generate an output molecule 162 based at least on an input molecule 152. For example, in the example of the molecular design engine 110 shown in Figure 1A, the molecular design computational model 115 may operate on a three-dimensional representation of the input molecule 152, which in some cases may be a voxelized representation of the input molecule 152. In doing so, the molecular design computational model 115 may operate in a discrete voxelized space filled with noisy voxelized representations of different molecules (e.g., molecules exhibiting the one or more desired properties). Alternatively, in the example of the molecular design engine 110 shown in Figure 1B, the molecular design computational model 115 can operate on the embedding 154 of the input molecule 152. As described in more detail below, the embedding 154 of the input molecule 152 can be generated by an encoder 111 that encodes a three-dimensional representation (e.g., a voxelized representation) of the input molecule 152. In this variant of the molecular design engine 110, the molecular design computational model 115 can operate in a latent space filled with noisy embeddings of three-dimensional representations (e.g., voxelized representations) of different molecules.
[0097] In some exemplary embodiments, the molecular design computational model 115 may include a denoising model 117 trained to denoise the three-dimensional representation of the input molecule 152 based on a function 175, such that the resulting three-dimensional representation of the output molecule 162 is sampled from higher-density regions of the data distribution of molecules exhibiting the one or more desired properties. Figure 1A depicts an example of the denoising model 117 trained to denoise the three-dimensional representation of the input molecule 152, while Figure 1B depicts another example of the denoising model 117 trained to denoise the embedding 154 of the three-dimensional representation of the input molecule 152. In some cases, the denoising model 117 may denoise the three-dimensional representation of the input molecule 152 (e.g., the voxelized representation of the input molecule 152) or alternatively its embedding 154 over multiple time steps. In some cases, denoising performed at each time step may be equivalent to selecting one or more samples (e.g., intermediate molecules) from different locations in the data distribution. In some cases, function 175 can be a scoring function that outputs a value (e.g., a score) indicating the local density change at a specific location within the data distribution (e.g., a location occupied by a specific molecule). Therefore, the input molecule 152 can be denoised based on the output of the scoring function, such that each successive sample (or molecule) is selected from regions of gradually increasing density within the data distribution.
[0098] In some exemplary embodiments, denoising the input molecule 152 may include updating a three-dimensional representation (e.g., a voxelized representation) (or its embedding 154) of the input molecule 152 that represents its composition and conformation (or three-dimensional structure) to increase the likelihood that the resulting output molecule 162 will exhibit the desired properties in a data distribution. In some cases, the molecular design computational model 115 may apply a denoising model 117 to modify the three-dimensional representation (or its embedding 154) of the input molecule 152 through one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Markov chain Monte Carlo (MCMC) with Langevin dynamics, etc.). For example, in some cases, each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling may include selecting samples (or molecules) from the data distribution that include one or more modifications to the three-dimensional representation (or its embedding 154) of the input molecule 152. As described above, in some cases, sampling from the data distribution can be guided by function 175 (e.g., a scoring function). For example, when function 175 is a scoring function, for each sample (or molecule) selected from the data distribution, function 175 can output a value (e.g., a score) corresponding to the density change observed at the location in the data distribution occupied by the sample (or molecule). Therefore, in some cases, sampling from the data distribution can be guided by function 175 such that each successive sample is selected from a region of gradually increasing density in the data distribution.
[0099] To further illustrate, in some cases, a denoising model 117 can be applied to update the three-dimensional representation (e.g., voxelized representation) (or its embedding 154) of the input molecule 152 by selecting, for example, a first sample and a second sample from the data distribution. It should be understood that each of the first and second samples may correspond to a modified three-dimensional representation of the input molecule 152 in the example shown in FIG1A, or, in the case of FIG1B, to a modified embedding of the three-dimensional representation of the input molecule 152. In the case where function 175 is a scoring function, function 175 may assign a first value (e.g., a first score) to the first sample, indicating a more positive local change (e.g., an increase, or a minor decrease) in the density of the data distribution at a first location of the first sample, and assign a second value to the second sample, indicating a less positive local change (e.g., a minor increase, or a decrease) in the density of the data distribution at a second location of the second sample. In some cases, the molecular design computational model 115 may apply a denoising model 117 to select a third sample from the data distribution by further modifying the first sample to sample a third sample from a region of higher density than the first and second samples (e.g., another modified 3D representation or another modified embedding).
[0100] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to denoise the voxelized representation (or its embedding 154) of the input molecule 152, rather than its conventional three-dimensional representation (such as a point cloud representation of the input molecule 152). For example, in some cases, the voxelized representation of the input molecule 152 may represent the types and locations of atoms present in the input molecule 152 as continuous (e.g., Gaussian-like) densities across a three-dimensional voxel grid. To indicate the location of atoms present in the input molecule 152, each voxel in the voxel grid may be associated with a value indicating the atomic density at the corresponding location. In some cases, the atomic density associated with a particular voxel in the voxel grid may correspond to the probability that the voxel is part of an atom at that location. For example, a first voxel with a higher atomic density may be more likely to be part of an atom forming the input molecule 152 than a second voxel with a lower atomic density. Therefore, the voxelized representation of input molecule 152 can be represented as the location of atoms of input molecule 152 in the voxel grid that distinguish between voxels forming part of atoms in input molecule 152 and voxels not forming part of atoms in input molecule 152, based on the atomic density associated with each voxel in the voxel grid. In some cases, atoms forming input molecule 152 can be positioned at the locations of those voxels associated with atomic densities reaching one or more thresholds.
[0101] In some exemplary embodiments, the voxelized representation of input molecule 152 may include one or more channels, each of which corresponds to a type of atom that may be present in input molecule 152. For example, in some cases, the voxelized representation of input molecule 152 may include a separate channel for each type of heavy atom that may be present in input molecule 152. Including different channels for different atom types in the voxelized representation of input molecule 152 avoids the discrete distribution typically associated with atom types found in conventional three-dimensional representations (e.g., point cloud representations, etc.). Instead, the voxelized representation of input molecule 152 may represent the type and location of atoms in input molecule 152 as one or more continuous (e.g., Gaussian-like) densities across the aforementioned three-dimensional voxel grid. For example, the voxelized representation of input molecule 152 may include a first channel representing a first type of atom (e.g., carbon (C) atom) that may be present in input molecule 152. The presence and corresponding positions of atoms of the first type in the input molecule 152 can be represented by a first continuous (e.g., Gaussian-like) density of a first channel in the voxelized representation of the input molecule 152. It should be noted that the density is continuous in the sense that the values associated with each voxel can take continuous values (e.g., values within a continuous, bounded, or unbounded distribution). In some cases, the voxelized representation of the input molecule 152 may further include a second channel representing atoms of the second type (e.g., nitrogen (N) atoms) that may be present in the input molecule 152. The presence and corresponding positions of atoms of the second type can be represented by a second continuous (e.g., Gaussian-like) density of a second channel in the voxelized representation of the input molecule.
[0102] Unlike a conventional three-dimensional representation (e.g., point cloud representation, etc.) of the input molecule 152 that represents the types and locations of atoms in the input molecule 152 as two independent types of distributions (e.g., a discrete distribution for atom types and a continuous distribution for atom locations), a voxelized representation of the input molecule 152 can represent the types and locations of atoms in the input molecule 152 together as one or more continuous (e.g., Gaussian-like) distributions in the manner described above. Therefore, the molecular design computational model 115 can apply a denoising model 117 to operate on the voxelized representation of the input molecule 152 without requiring the workarounds necessary to reconcile the two different types of distributions in the case of a conventional three-dimensional representation of the input molecule 152. The voxelized representation of the input molecule 152 can also represent the conformation (or three-dimensional structure) of the input molecule 152 better than its conventional three-dimensional representation. For example, the voxelized representation of the input molecule 152 can capture long-range dependencies between distant atoms, even when the input molecule 152 contains a large number of atoms. Furthermore, the molecular design computational model 115 can apply a denoising model 117 to denoise the voxelized representation of the input molecule 152 and generate the output molecule 162 without any prior knowledge of the molecular weight present in the output molecule 162.
[0103] In some exemplary embodiments, training engine 120 may be trained to generate one or more training samples to be included in a training dataset. Figure 1A depicts an example of training engine 120, wherein damage engine 121 generates each training sample by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to a noisy three-dimensional representation 182 of a sample molecule (e.g., a known molecule exhibiting one or more desired properties), thereby generating a damaged three-dimensional representation 184 of the sample molecule. It should be understood that damage engine 121 may add noise to the existing noisy three-dimensional representation 182 of the sample molecule so that molecular design computation model 115 is trained to approximate a noisy data distribution of molecules exhibiting the one or more desired properties, which has a smoother density transition compared to the true data distribution of molecules exhibiting the one or more properties. Thus, in some cases, molecular design computation model 115 (including denoising model 117) may be trained to denoise the damaged three-dimensional representation of the sample molecule and recover the corresponding noisy three-dimensional representation of the sample molecule for each training sample, rather than a clean (or initial) three-dimensional representation of the sample molecule. This is equivalent to sampling from the noisy data distribution of molecules that exhibit one or more desired properties, rather than from the true distribution of the molecules.
[0104] Alternatively, Figure 1B depicts another instance of the training engine 120, where encoder 111 first encodes a noisy 3D representation 182 of a sample molecule to generate its embedding 186, and then corruption engine 121 adds noise to generate a corrupted embedding 188 of the noisy 3D representation 182 of the sample molecule. In this variant of the molecular design engine 110, the molecular design computational model 115 can be trained to approximate a noisy but discrete distribution of the noisy latent distribution of the 3D representation of molecules exhibiting one or more desired properties, rather than the noisy but discrete distribution approximated by the molecular design computational model 115 in the instance shown in Figure 1A. To achieve this result, training engine 120 can generate each training sample in the training dataset that includes the corrupted embedding 188. For example, Figure 1B shows that the training engine may further include an encoder 111, which may first encode a noisy three-dimensional representation 182 of a sample molecule (e.g., a known molecule exhibiting one or more desired properties) to generate an embedding 186, and then a corruption engine 121 adds noise to this embedding to generate a corrupted embedding 188. In some cases, a molecular design computational model 115 (including a denoising model 117) may be trained to denoise the corrupted embedding 188 and recover the embedding 186 from it. The denoising model 117, the encoder 111, and the corresponding decoder may be trained together.
[0105] In some exemplary embodiments, training the molecular design computational model 115 may include adjusting one or more parameters (e.g., weights, biases, etc.) of the denoising model 117. In the example shown in FIG1A, the parameters of the denoising model 117 may be adjusted such that the denoising model 117 is able to recover the corresponding noisy three-dimensional representation 182 of the sample molecule from the damaged three-dimensional representation 184 of the sample molecule. For example, as described in more detail below, the one or more parameters of the denoising model 117 may be adjusted through multiple iterations to increase (or maximize) the similarity between the noisy three-dimensional representation 182 of the sample molecule recovered by the denoising model 117 from the damaged three-dimensional representation 184 of the sample molecule and the initial noisy three-dimensional representation 182 of the sample molecule (e.g., to reduce (or minimize) the mean squared error (MSE)). Alternatively, the example in Figure 1B illustrates that the parameters of the denoising model 117 can be adjusted such that the denoising model 117 can recover the embedding 186 generated by encoding the noisy 3D representation 182 of the sample molecule from the corrupted embedding 188. For example, in the example shown in Figure 1B, the parameters of the denoising model 117 can be adjusted to increase (or maximize) the similarity between the embedding 186 recovered by the denoising model 117 from the corrupted embedding 188 and the initial uncorrupted embedding 186 of the 3D representation 182 of the sample molecule (e.g., to reduce (or minimize) any loss function that quantifies the difference (e.g., mean squared error (MSE)) between the embedding 186 recovered by the denoising model 117 from the corrupted embedding 188 and the initial uncorrupted embedding 186 of the 3D representation 182 of the sample molecule).
[0106] In some exemplary embodiments, training the molecular design computational model 115 may include determining a function 175, which can be parameterized by parameters of the denoising model 117 (e.g., weights, biases, etc.). For example, in some cases, training the molecular design computational model 115 (which includes adjusting the one or more parameters of the denoising model 117) may be determined by adjusting at least the corresponding parameters of function 175 (e.g., a scoring function). In some cases, function 175 may approximate different densities across a data distribution of molecules exhibiting the one or more desired properties, wherein molecules exhibiting the one or more properties are more likely to occupy regions of higher density in the data distribution. In the example shown in Figure 1A, the data distribution may be a noisy data distribution filled with noisy three-dimensional representations of molecules. Alternatively, in Figure 1B In the example of the molecular design engine 110 shown, the data distribution can be a noisy latent distribution filled with noisy embeddings of a three-dimensional representation of molecules.
[0107] As described above, in some exemplary embodiments, overfitting of the molecular design computation model 115 (including the denoising model 117) to known molecules in the training dataset can be avoided by training the denoising model 117 to approximate a noisy data distribution filled with noisy three-dimensional representations of molecules (such as noisy voxelized representations). The training dataset. Figure 1A illustrates such an example where the molecular design computation model 115 is trained to recover a noisy three-dimensional representation 182 of a sample molecule, rather than a clean (or initial) three-dimensional representation of the sample molecule. As described in more detail below, once trained, the molecular design computation model 115 can generate a three-dimensional representation of the output molecule 158 by sampling at least one updated three-dimensional representation 160 by traversing the smoothed density of the noisy data distribution (i.e., iteratively sampling different regions of the noisy data distribution). The noisy data distribution can be filled with noisy three-dimensional representations of molecules exhibiting the one or more desired properties. Therefore, the updated 3D representation 160 may include at least some noise, which can be removed by applying the recovery model 118 to denoise the updated 3D representation 160. The recovery model 118 may be a machine learning model that has been trained to take the noisy 3D representation 160 of the molecule as input and produce a corresponding denoised 3D representation as output. This generates a 3D representation of an output molecule 152 that occupies the true data distribution of clean 3D representations of molecules exhibiting the one or more desired properties.
[0108] Alternatively, Figure 1B depicts another instance of the molecular design engine 110, wherein a molecular design computational model 115 is trained to approximate a noisy latent distribution filled with embeddings of noisy 3D representations of molecules exhibiting one or more desired properties. In this variation of the molecular design computational model 115, a denoising engine 117 can be trained to recover embeddings 186 of the noisy 3D representation 182 of the sample molecule from a corrupted embedding 188. Once trained, the molecular design computational model 115 can apply the denoising model 117 to denoise the embeddings 154 generated by the encoder 111 encoding the 3D representation of the input molecule 152. Denoising may include the denoising model 117 updating the embeddings 154 through multiple successive sampling iterations to generate at least one updated embedding 156 during each sampling iteration. Doing so is equivalent to the molecular design computational model 115 selecting samples from the noisy latent distribution, and the molecular design computational model 115 continuing to select samples from it until one or more criteria are met. In some cases, such as before the noisy 3D representation 158 obtained by recovery model 118 is denoised to generate a 3D representation of output molecule 162, the updated embedding 156 may be decoded by decoder 119. As described in more detail below, in this instance of molecular design engine 110, molecular design computation 115 may operate in a noisy latent space filled with the noisy embedding of the 3D representation of the molecule rather than the noisy 3D representation found in the discrete voxelized space. Furthermore, it should be understood that although the denoising model 117 and recovery model 118 may share the same architecture (e.g., artificial neural network (ANN) etc.) in some cases, the two models are trained to remove different noises. For example, denoising model 117 may be trained to denoise the 3D representation of input molecule 152 or its embedding 154 such that the resulting 3D representation of output molecule 162 is consistent with the composition and / or conformation of a molecule exhibiting the one or more desired properties (e.g., drug-like properties). In contrast, recovery model 118 can be trained to remove noise that is added to smooth the density of known molecules that can be used to train molecular design computational model 115.
[0109] Referring again to Figure 1B, the embedding 154 of the three-dimensional representation of the input molecule 152 can be generated by an encoder 111 that encodes the three-dimensional representation of the input molecule 152. In some cases, the embedding 154 can be a lower-dimensional representation of the three-dimensional representation of the input molecule 152 generated by the encoder 111, which reduces the dimension of the three-dimensional representation of the input molecule 152. For example, in some cases, the encoder 111 and the decoder 119 can form an autoencoder, including, for example, a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQ-VAE). The encoder 111 can generate the embedding 154 by at least reducing the dimension of the three-dimensional representation of the input molecule 152 (e.g., a voxelized representation). In this context, reducing the dimensionality of the three-dimensional representation of the input molecule 152 may include downsampling, compressing, or reducing the dimensionality of the three-dimensional representation of the input molecule 152, for example, by compressing at least some of the features (e.g., atomic density values) present in the three-dimensional representation of the input molecule 152, such that the resulting embedding 154 includes fewer features compared to the initial three-dimensional representation of the input molecule 152, but these features still capture the same (or similar) information conveyed in the initial three-dimensional representation of the input molecule 152. For a voxelized representation of the input molecule 152, each feature present therein may correspond to an atomic density value associated with each voxel included in the voxel grid representing the input molecule 152. For example, in the case where the voxelized representation of the input molecule 152 includes a [32×32×32] voxel grid, the voxelized representation of the input molecule 152 may include 32,000 features (or atomic density values). When generating embedding 154, at least some of those 32,000 features can be compressed by encoder 111. In doing so, embedding 154 can include fewer features (such as 4×4×4=64 features) to allow denoising model 117 to operate on when generating output molecule 162.
[0110] As described above, the embedding 154 of the three-dimensional representation of the input molecule 152 shown in FIG. 1B may include fewer features compared to the initial three-dimensional representation of the input molecule 152. Furthermore, the embedding 154 may be generated by the encoder 111 downsampling, compressing, or reducing the dimensionality of the three-dimensional representation of the input molecule 152. Doing so is equivalent to the encoder 111 mapping the three-dimensional representation of the input molecule 152 from a high-dimensional discrete voxel space to a lower-dimensional latent space. In some cases, the molecular design computational model 115 (e.g., a denoising engine 117) may denoise the embedding 154, for example, through one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling. During each iteration of gradient-based Markov Monte Carlo (MCMC) sampling, the molecular design computational model 115 (e.g., the denoising model 117) may sample at least one updated embedding 156 from the lower-dimensional latent space. The fact that embedding 154 can include several orders of magnitude fewer features compared to the initial three-dimensional representation of input molecule 152 means that molecular design computational model 115 can apply denoising model 117 to operate on embedding 154 and generate output molecule 162 faster and with higher computational efficiency, while achieving comparable or better performance in both qualitative and quantitative aspects. This can be particularly advantageous in applications that require the generation of a large number of candidate molecules in a short period of time, such as computational drug design. In some cases, reducing the dimensionality of the initial three-dimensional representation of input molecule 152 can enable denoising model 117 to operate on and generate larger-scale molecules (e.g., molecules containing more than 200 atoms) and a greater number of molecules.
[0111] In some cases, the compactness of the embedding 154 relative to the three-dimensional representation of the input molecule 152 also means that fewer computational resources are required when operating on the embedding 154. For example, in order for the denoising model 117 to operate directly on the three-dimensional representation of the input molecule 152 (e.g., a [32×32×32] voxel grid with 32,000 features), the denoising model 117 can be implemented with a large number of trainable parameters (e.g., 100 million parameters). In contrast, if the denoising model 117 is instead applied to operate on the embedding 154, it can be implemented with far fewer trainable parameters. Implementing the denoising model 117 with fewer parameters improves its performance, as a larger number of parameters can reduce its generalizability. For example, the denoising model 117 may be particularly prone to overfitting when it includes a large number of features but relatively few known molecules available for training. When the denoising model 117 overfits to known molecules in the training dataset (meaning it is trained too well on the training dataset—i.e., trained to learn a data distribution too tightly centered on molecules in the training dataset)), it may fail to generalize. In this context, "generalization" refers to the ability to accurately denoise the input molecule 152 even when it is not one of the known molecules in the training dataset. Therefore, overfitting the denoising model 117 prevents it from accurately denoising any molecule that is not one of the known molecules in the training data.
[0112] Referring again to Figure 1B, in the example of the molecular design engine 110 shown therein, the updated embedding 156 generated by sampling from the latent voxelized space can be decoded by the decoder 119 before the noisy 3D representation 158 obtained by the recovery engine 118 is denoised. While the decoding performed by the decoder 119 maps the updated embedding 156 from the latent space to the discrete space, the subsequent denoising performed by the recovery model 118 can constitute a jump from the noisy data distribution back to the true data distribution of the molecule. For example, during each round of gradient-based Markov chain Monte Carlo (e.g., Langevin Markov chain Monte Carlo, etc.), the molecular design computational model 115 can apply a denoising model 117 to sample from the noisy latent distribution of the molecule exhibiting one or more desired properties. In this context, sampling from the noisy latent distribution may include updating the embedding 154 of the 3D representation (e.g., voxelized representation) of the input molecule 152. This is equivalent to selecting at least one updated embedding 156 from the data distribution, which is then decoded by decoder 119 and denoised by recovery model 118 to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162. As described in more detail below, molecular design computational model 115 can further update the noisy embedding 156, for example, through multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, until one or more criteria are met.
[0113] Once one or more criteria are met, the molecular design engine 110 can recover the output molecule 162 based at least on its three-dimensional representation. For example, in some cases, the molecular design engine 110 can recover the locations (e.g., coordinates) of atoms present in the output molecule 162 and one or more bonds between them based at least on its three-dimensional representation. Recovering the locations of atoms present in the output molecule 162 can be done by identifying the maximum value (e.g., peak) of the predicted density included in the three-dimensional representation of the output molecule 162. In doing so, the molecular design engine 110 can determine another representation of the output molecule 162, including, for example, a one-dimensional representation of the output molecule 162 (e.g., a simplified linear input specification (SMILES) string), a two-dimensional representation of the output molecule 162 (e.g., a molecular graph), etc. It should be understood that the output molecule 162 generated in this way is more likely to exhibit the one or more desired properties of the molecule in the data distribution. Specifically, the denoising model 117 can generate an output molecule 152 that exhibits a composition and / or conformation (or three-dimensional structure) consistent with the one or more desired properties. By operating on an embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152, the denoising model 117 can generate the output molecule 162 faster and with less computational burden.
[0114] As described above, in some exemplary embodiments, the molecular design computational model 115 can be trained to recover a noisy three-dimensional representation 182 of a sample molecule (e.g., a noisy voxelized representation of the sample molecule), which is generated by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to the three-dimensional representation of the sample molecule. In the example shown in FIG1A, the molecular design computational model 115 can be trained to denoise the damaged three-dimensional representation 184 of the sample molecule by at least modifying the damaged three-dimensional representation 184 of the sample molecule in order to recover the noisy three-dimensional representation 182 of the sample molecule. Alternatively, FIG1B shows an example of a molecular design engine 110 in which the molecular design computational model 115 is trained to recover the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule. As shown in Figure 1B, in some cases, the molecular design computational model 115 can denoise the damaged embedding 188 (which can be generated by the damaged engine 121 adding noise (e.g., Gaussian noise, etc.) to the embedding 184 of the three-dimensional representation 182 of the sample molecule) by at least modifying the damaged embedding 186. As mentioned above, the embedding 184 can be generated by the encoder 111 by downsampling or reducing the dimension of the noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation). However, as mentioned above, downsampling of the noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation) can be optional, which may be the case when the encoder 111 implements the identity function. Therefore, in some cases, it is possible that the embedding 184 includes the same amount of features as the initial noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation). In those cases, embedding 154 can capture the same information present in the initial three-dimensional representation (e.g., voxelized representation) of the input molecule 152 without compressing the features present therein. In other words, encoding the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 can be an optional operation, even in the case of an encoder 111 included in the instance of the molecular design system 100 shown in FIG. 1B.
[0115] Figure 2 depicts a flowchart illustrating an example of a process 200 for machine learning-enabled 3D molecule generation in voxelized space according to some exemplary embodiments. Referring to Figures 1A through 1B and Figure 2, process 200 may be performed by a molecular design engine 110 to train and apply a molecular design computational model 115 to generate an output molecule 162 by denoising at least a 3D representation (such as a voxelized representation) of an input molecule 152. For example, in some cases, the molecular design computational model 115 may be trained on a noisy 3D representation 182 of a sample molecule (which is generated by adding noise to an initial 3D representation of the sample molecule), such that the molecular design computational model 115 is trained to approximate a noisy data distribution with a smoother density transition. Figure 1A illustrates a variation in which the molecular design computational model 115 is trained on a corrupted 3D representation 184 of a sample molecule, which can be generated by adding additional noise to the noisy 3D representation 182 of the sample molecule without any downsampling or compression. Alternatively, Figure 1B illustrates another variation in which the molecular design computational model 115 is trained on a corrupted embedding 186 of a noisy three-dimensional representation 182 of the sample molecule. This corrupted embedding 188 can be generated by adding noise to an embedding 154 of the noisy three-dimensional representation 182 of the sample molecule, which is generated by an encoder 111 downsampling (or compressing) features present in the noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation). In other words, it should be understood that the molecular design computational model 115 can be trained to operate in a noisy discrete voxelized space filled with the noisy three-dimensional representation of the molecule, or alternatively in a noisy latent voxelized space filled with the embeddings of the noisy three-dimensional representation of the molecule.
[0116] At 202, training engine 120 may generate a training dataset comprising multiple damaged sample molecules. In some exemplary embodiments, generating the training dataset may include: training engine 120 generating a training dataset comprising multiple damaged sample molecules. This training dataset can then be used to train molecular design computational model 115 to approximate a data distribution of molecules exhibiting one or more desired properties (e.g., drug-like properties). In some cases, each damaged sample molecule may be a noisy three-dimensional representation of a known molecule, which is further corrupted by additional noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.). For example, Figure 1A illustrates such an instance where the corruption engine 121 generates a corrupted three-dimensional representation 184 of a sample molecule by adding additional noise to the noisy three-dimensional representation 182 of the sample molecule. As described in more detail below, molecular design computational model 115 (e.g., denoising model 117) may be trained to recover the noisy three-dimensional representation 182 of the sample molecule from the corrupted three-dimensional representation 184 of the sample molecule to approximate a noisy data distribution with a smoother density transition. Alternatively, Figure 1B illustrates another example where training engine 120 generates a corrupted embedding for each corrupted sample molecule in the training dataset that includes a noisy 3D representation 182 of the sample molecule. For example, in some cases, corruption engine 121 can generate a corrupted embedding 188 by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to the embedding 186 of the noisy 3D representation 182 of the sample molecule (e.g., a noisy voxelized representation). In some cases, training engine 120 can further enhance the training dataset that can be enhanced by applying one or more transformations to the noisy 3D representation 182 of the sample molecule (e.g., a voxelized representation), including, for example, translation (e.g., shifting the center of the sample molecule in each of the three dimensions by sampling a uniform offset), rotation (e.g., uniformly sampling the three Euler angles), reflection, etc.
[0117] Training a molecular design computational model 115 based on a noisy 3D representation 182 of sample molecules can reduce the incidence of overfitting and mode collapse, which typically occur when the molecular design computational model 115 is trained on a disproportionately small number of known molecules (e.g., PubChem dataset, QM9 molecular dataset, Geometric Integration of Molecular (GEOM) drug dataset, etc.) to fit high-dimensional data distributions (e.g., 10). 60 This occurs when the molecular space of a possible chemical compound is approximated.
[0118] In the examples of the molecular design computational model 115 shown in Figures 1A and 1B, the molecular design computational model 115 may include a denoising model 117. In this case, training the molecular design computational model 115 may include training the denoising model 117 to approximate a noisy data distribution, or in some cases, to approximate a noisy latent distribution, either of which exhibits a smoother density transition and is more efficient to sample from than the true data distribution. In some cases, the denoising model 117 may be an artificial neural network (ANN), in which case training the denoising model 117 may include adjusting one or more parameters of the artificial neural network (ANN) (e.g., weights, biases, etc.). Doing so also determines the parameters of function 175 such that function 175 outputs a value indicating the probability that a molecule exhibiting the one or more desired properties is located at a specific position within the data distribution. For example, in some cases, function 175 can be a scoring function whose output is a value indicating the transition between different density regions of the data distribution (e.g., a score), including, for example, a transition between higher density regions occupied by molecules more likely to exhibit the one or more desired properties and lower density regions of the data distribution less likely to be occupied by molecules exhibiting the one or more desired properties. In some cases, denoising model 117 can be trained to recover noisy three-dimensional representations 182 or embeddings 186 of the sample molecules in order to avoid overfitting denoising model 117, for example, to the relatively few known molecules in the training dataset that can be used to characterize the data distribution.
[0119] To further illustrate, let p(x) represent the true data distribution of a voxelized representation of a molecule exhibiting one or more desired properties, and let p(y) represent the corresponding noisy data distribution, which can exhibit a smoother energy landscape and from which sampling is more efficient than from an unknown data distribution p(x). In some cases, the true data distribution p(x) may be unknown, meaning that the denoising model 117 can be trained on a training dataset of known molecules from the true data distribution to approximate the true data distribution p(x). To avoid overfitting the denoising model 117 to the training dataset, instead, the denoising model 117 can be trained to approximate the noisy data distribution p(y). In some cases, the noisy data distribution p(y) can be approximated using a voxelized representation of a molecule with a known covariance σ. 2 I dThis is obtained by convolving the real data distribution p(x) with a Gaussian kernel (e.g., an isotropic Gaussian kernel). This is equivalent to generating a noisy voxelized molecular representation y (e.g., y = x + ϵ, where x ~ p(x), ϵ ~ N(0, σ)) by adding noise ϵ to the voxelized molecular representation x from the real data distribution p(x). 2 I d Given the above formulation, the noisy voxelized molecular representation Y can be sampled from the noisy data distribution p(y) expressed as follows:
[0120]
[0121] Transforming the real data distribution p(x) in this way smooths the density of the real data distribution p(x) while still preserving some of the structural information present in the clean (or initial) voxelized representation x, without any added noise ϵ. If the noise ϵ added to the clean voxelized molecular representation x is Gaussian (e.g., isotropic Gaussian), it can be estimated by applying the least squares estimator as shown in the following equation (1). To directly recover the clean voxelized molecular representation x from the corresponding noisy voxelized molecular representation y. It should be understood that the least squares estimator... It can act as a denoiser and restore the clean voxelized molecular representation x by removing the noise ϵ present in the noisy voxelized molecular representation y.
[0122] (1)
[0123] in This corresponds to the scoring function g(y) of the noisy data distribution p(y). Equation (1) shows that if the noisy data distribution p(y) is known up to the normalization constant (and thus the corresponding scoring function g(y)), the clean voxelized molecular representation x can be estimated from its noisy counterpart y. Correspondingly, the scoring function g(y) of the noisy data distribution p(y) can also be based on the least squares estimator of the true data distribution p(x). To derive. As described in more detail below, when determining the scoring function g(y), or in some cases when determining the corresponding scoring function, the molecular design computational model 115 can apply the denoising model 117 to sample from the noisy data distribution p(y) based on the scoring function g(y) (or the corresponding scoring function).
[0124] As described in more detail below, once the denoising model 117 is trained, the molecular design computational model 115 can apply the denoising model 117 to perform a “walk-and-jump” generative process to generate output molecules that exhibit one or more desired properties of molecules in the true data distribution p(x). For example, in some cases, the denoising model 117 may sample from the noisy data distribution p(y) through multiple successive sampling iterations, each of which includes: the denoising model 117 selecting at least one sample from the noisy data distribution p(y) by denoising at least the noisy voxelized molecular representation y. In some cases, the sampling of the noisy data distribution p(y) may be guided by a scoring function g(y) such that the sample selected during one sample iteration originates from a different position in the noisy data distribution p(y) than the sample selected during another sample iteration. This traversal of the noisy data distribution p(y) is referred to as the “walk” portion of the generative process.
[0125] In some cases, the scoring function g(y) can constrain the sampling of the molecule y to a certain region p(y|c) within the noisy data distribution based on a condition c (e.g., the gradient of the classifier), rather than sampling freely from the entire noisy data distribution p(y). Therefore, in some cases, each consecutive sample can be selected from a region of gradually increasing density within the noisy data distribution p(y), as indicated by the score output by the scoring function g(y). Similarly, as mentioned above, traversing the noisy data distribution p(y) guided by the scoring function g(y) can be considered a “walk” through the noisy data distribution p(y). Recovering the clean voxelized molecular representation x from the corresponding noisy voxelized molecular representation y can constitute a “jump” from the noisy data distribution p(y) back to the true data distribution p(x). In some cases, the "leap" from a noisy data distribution p(y) back to the true data distribution p(x) can be achieved by applying a denoiser (such as denoising engine 117) to remove the noise ϵ present in the noisy voxelized molecular representation y and recovering the corresponding clean voxelized molecular representation x. For example, in some cases, the least squares estimator can be transformed by denoising engine 117. It is applied to noisy voxelized molecular representation y to recover clean voxelized molecular representation x.
[0126] In some exemplary embodiments, training engine 120 can generate each corrupted sample molecule in the training dataset by adding noise to the embedding of the noisy three-dimensional representation of the sample molecule. For example, Figure 1B shows that in some cases, corrupting engine 121 can generate corrupted embedding 188 by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation) instead of directly adding noise to the noisy three-dimensional representation 182 of the sample molecule. Figure 1B further shows that embedding 186 can be generated by encoder 111 downsampling (or compressing) the noisy three-dimensional representation 182 of the sample molecule. Downsampling (or compressing) the noisy three-dimensional representation 182 of the sample molecule can reduce the dimensionality of the noisy three-dimensional representation 182 of the sample molecule. For example, in the case where the noisy 3D representation 182 of a sample molecule is a [32×32×32] voxel grid with 32,000 features (or atomic density values), downsampling (or compression) can produce a [4×4×4] voxel grid with 64 features (or atomic density values). Therefore, downsampling (or compression) of the noisy 3D representation 182 of the sample molecule can increase the overall speed and efficiency of the generative process. In some cases, a large number of candidate molecules (such as tens of thousands or even millions of candidate molecules) can be generated in a short period of time to support low-yield applications, such as computational drug design, where most of the candidate molecules cannot be successfully synthesized in the laboratory. Downsampling (or compression) of the 3D representation 182 of the sample molecule also enables the generation of larger molecules (e.g., molecules containing more than 200 atoms) that would be too cumbersome to manipulate if left in their initial 3D representation without any downsampling (or compression).
[0127] At 204, the molecular design engine 110 can train the molecular design computation model 115 by at least the following: applying the molecular design computation model 115 to recover the 3D representation of each damaged sample molecule in the training dataset from the damaged 3D representation of the sample molecule. In some exemplary embodiments, the training step 204 of the molecular design computation model 115 may include: training a denoising model 117, at least based on the training dataset, to approximate the noisy data distribution (or noisy latent distribution) of molecules exhibiting the one or more desired properties. For example, in the example shown in FIG1A, the denoising model 117 may be trained to recover the noisy 3D representation 182 (e.g., voxelized representation) of the sample molecule from the damaged 3D representation 184. In the example shown in FIG1B, the denoising model 117 may be trained to recover the embedding 186 of the noisy 3D representation 182 of the sample molecule from the damaged embedding 188. It should be understood that, in any instance, the denoising engine 117 may be trained on the noisy three-dimensional representation 182 of the sample molecules rather than the clean three-dimensional representation of the sample molecules, so that the denoising engine 117 approximates the noisy data distribution that exhibits a smoother density transition than the true data distribution.
[0128] In the example shown in Figure 1A, training the molecular design computational model 115 may include adjusting one or more parameters (e.g., weights, biases, etc.) of the denoising model 117 to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the noisy three-dimensional representation 182 of the sample molecule recovered by the denoising model 117 and the initial noisy three-dimensional representation 182 of the sample molecule. Furthermore, training the denoising model 117 may include determining a function 175, which is parameterized by the parameters (e.g., weights, biases, etc.) of the denoising model 117. For example, in some cases, function 175 may be a scoring function that outputs values (e.g., scores) indicating local variations in the density (or gradient) of the data distribution. Thus, in some cases, function 175 may output: a first value (e.g., a first score) indicating a first local variation in the density of the data distribution at a first location occupied by a first molecule, and a second value (e.g., a second score) indicating a second local variation in the density of the data distribution at a second location occupied by a second molecule. In some cases, a first value (e.g., a first score) may indicate a more positive local change in the density of the data distribution at a first position of a first molecule (e.g., an increase, or a slight decrease), while a second value (e.g., a second score) may indicate a less positive local change in the density of the data distribution at a second position of a second molecule (e.g., a slight increase, or a decrease). When function 175 is a scoring function, sampling of the data distribution may be guided by the value (e.g., the score) output by function 175. As described in more detail below, sampling of the data distribution may be guided by function 175 such that samples (or molecules) are selected from regions where the density of the data distribution gradually increases, regions more likely to be occupied by molecules exhibiting the desired one or more properties.
[0129] In the example shown in Figure 1B, training the denoising model 117 may include adjusting one or more parameters of the denoising model 117 (e.g., weights, biases, etc.) to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the initial undamaged embedding 186 and the embedding 186 of the noisy 3D representation 182 of the sample molecule recovered by the denoising model 117 from the damaged embedding 188. Doing so may also adjust the parameters of the function 175. Similarly, in the case where the function 175 is a scoring function, the parameters of the function 175 may be adjusted such that the function 175 outputs a higher value (e.g., a higher score) for the first molecule occupying a first position in the data distribution that exhibits a more positive local change in the density of the data distribution (e.g., a positive gradient indicating a transition from a lower density region to a higher density region in the data distribution) than for the second molecule occupying a second position in the data distribution that exhibits a less positive local change in the density of the data distribution.
[0130] In some exemplary embodiments, the molecular design engine 110 may train a molecular design computational model 115 (including a denoised model 117) to approximate the function 175 by at least performing gradient-based Markov chain Monte Carlo (MCMC) sampling (such as Markov chain Monte Carlo (MCMC) sampling using Langevin dynamics). In some cases, the function 175 may output a value (e.g., a score) indicating a transition between different density regions of the data distribution. For example, as a scoring function, the value (e.g., a score) output by the function 175 for each molecule may indicate a local change in density (or gradient) at the corresponding location in the data distribution. In the example shown in Figure 1A, determining the gradient-based Markov chain Monte Carlo (MCMC) sampling of function 175 may include: undergoing multiple iterations to adjust the parameters of the denoising model 117 (e.g., weights, biases, etc.) and those parameters of function 175 to increase (or maximize) the similarity between the noisy 3D representation 182 of the sample molecule recovered by the denoising model 117 and the initial 3D representation 182 of the sample molecule (e.g., by reducing (or minimizing) the mean squared error (MSE)). For the example shown in Figure 1B, one or more parameters of the denoising model 117 and those parameters of the function 175 may be adjusted through multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) to increase (or maximize) the similarity between the embedding 186 of the noisy 3D representation 182 of the sample molecule recovered by the denoising engine 117 from the damaged embedding 188 and the initial undamaged embedding 186 (e.g., by reducing (or minimizing) the mean squared error (MSE)).
[0131] As described above, the denoising model 117 can be trained to recover the noisy three-dimensional representation 182 of the sample molecules (e.g., a noisy voxelized representation) to avoid overfitting the denoising model 117 to known molecules that can be used to train it. When relatively few known molecules are available to characterize the high-dimensional data distribution, training the denoising model 117 directly based on known molecules can produce an overly jagged energy landscape, with sharp gradient changes between regions filled by known molecules. Sampling from the data distribution guided by steep gradients can hinder sufficient exploration of the data distribution, at least because the steepness of the gradient may limit sampling to regions within the immediate vicinity of known molecules. In contrast, training the denoising model 117 based on the noisy three-dimensional representation 182 of the sample molecules (e.g., a noisy voxelized representation) produces a smoother density transition, where the gradient of the function 175 is more asymptotic, allowing for more efficient exploration of the data distribution when sampling from it.
[0132] At 206, the molecular design engine 110 may apply a trained molecular design computational model 115 to generate an output molecule by denoising at least the voxelized representation of the input molecule. In some exemplary embodiments, the molecular design computational model 115 may use a denoising model 117 to generate an output molecule 162 by updating the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 at least when guided by function 175, or in some cases updating its embedding 154. For example, the molecular design computational model 115 of FIG. 1A may directly update the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 without any downsampling or compression. Alternatively, the molecular design computational model 115 of FIG. 1B may update the embedding 154 of the three-dimensional representation of the input molecule 152, which may be generated by encoder 111 by downsampling (or compressing) the three-dimensional representation (e.g., voxelized representation) of the input molecule 152. This reduces the dimensionality (or number of features) of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152, making the resulting embedding 154 more compact than the initial (or uncompressed) three-dimensional representation of the input molecule 152. For example, while the initial three-dimensional representation (e.g., voxelized representation) of the input molecule 152 may include a [32×32×32] voxel grid containing 32,000 features (or atomic density values), the embedding 154 may include a [4×4×4] voxel grid containing 64 features (or atomic density values). It should be understood that the encoder 111 may be trained to downsample (or compress) the voxelized representation of the input molecule 152 such that the resulting embedding 154 conveys the same (or similar) information as the voxelized representation of the input molecule 152 in its initial (or uncompressed) form. In some cases, encoder 111 may be part of an autoencoder (e.g., a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQ-VAE)) that also includes decoder 119. In some cases, encoder 111 may be trained to generate embedding 154 such that decoder 119 is able to recover the initial voxelized representation of input molecule 152 by decoding at least the embedding 154.
[0133] In some cases, the molecular design computational model 115 can denoise the input molecule 152 by at least updating the three-dimensional representation of the input molecule 152, or alternatively by updating the embedding 154 of the three-dimensional representation of the input molecule 152. As described above, in some cases, the three-dimensional representation of the input molecule 152 can be a voxelized representation of the input molecule 152, wherein the types and locations of atoms present in the input molecule 152 are represented as a continuous (e.g., Gaussian-like) atomic density centered on the atoms. For example, in some cases, the voxelized representation of the input molecule 152 may include an [n×n×n] voxel grid containing n 3 The input molecule 152 consists of a number of voxels, each associated with a value indicating the atomic density at the corresponding location. In some cases, the atomic density associated with a single voxel may have a value between a range of values (such as [0,1]), where the atomic density at the lower end of the range indicates that the voxel is farther from any atom in the input molecule 152, and the atomic density at the upper end of the range indicates that the voxel is closer to the center of the atoms in the input molecule 152. Furthermore, in some cases, the voxelized representation of the input molecule 152 may include multiple channels, where each channel corresponds to a different type of atom that may be present in the input molecule 152. Thus, in some cases, the voxelized representation of the input molecule 152 may represent the types of atoms present in the input molecule 152 and the location of the atoms as a continuous (e.g., Gaussian-like) atomic density across one or more channels.
[0134] In some exemplary embodiments, denoising the input molecule 152 may include updating the three-dimensional representation of the input molecule 152, or alternatively updating the embedding 154 of the input molecule 152. When the molecular design computational model 115 operates on the three-dimensional representation of the input molecule 152, or when the embedding 154 is generated without any downsampling (or compression) of the three-dimensional representation of the input molecule 152, denoising may include updating the atomic density of one or more voxels in at least one channel of the noisy voxelized representation of the input molecule 152. Doing so may be equivalent to adding, removing, and / or repositioning one or more atoms of different atom types in the input molecule 152. For example, increasing (or decreasing) the atomic density of one or more voxels in one channel of the noisy voxelized representation of the input molecule 152 may be equivalent to adding (or removing) atoms of the corresponding type to the input molecule 152. Alternatively and / or additionally, reducing the atomic density of the first voxel while increasing the atomic density of the second voxel is equivalent to repositioning the atom from the first position of the first voxel to the second position of the second voxel.
[0135] Alternatively, when the molecular design computational model 115 operates on an embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152, denoising may include updating the values of voxels present in the embedding 154. As described above, the embedding 154 may be generated by the encoder 111 by compressing at least some of the features (e.g., atomic density values) present in the voxelized representation of the input molecule 152. The embedding 154 may include fewer features compared to the initial voxelized representation of the input molecule 152, but still conveys the same (or similar) information as the initial three-dimensional representation of the input molecule 152. Therefore, denoising the embedding 154 may include updating one or more values present in the embedding 154, at least some of which may represent multiple features (or atomic density values) from the initial voxelized representation of the input molecule 152.
[0136] Updating the three-dimensional representation of the input molecule 152 or its embedding 154 in the manner described above may include selecting samples (or updated molecules) from a noisy data distribution (or noisy latent distribution) of molecules exhibiting the one or more desired properties. In the case of gradient-based Markov chain Monte Carlo (MCMC) sampling, this update may be guided by the output of function 175 (e.g., a score output by function 175) such that the samples (or updated molecules) selected during each successive sampling iteration originate from regions of noisy data distribution with progressively increasing density that are more likely to be filled by molecules exhibiting the one or more desired properties.
[0137] To further illustrate, in some cases, the three-dimensional representation of the input molecule 152 or its embedding 154 may undergo a first update and a second update. Doing so is equivalent to selecting a first sample (or a first updated molecule) and a second sample (or a second updated molecule) from a noisy data distribution. In some cases, when selecting the first sample (or the first updated molecule) and the second sample (or the second updated molecule) from a noisy data distribution (or a noisy latent distribution), the molecular design computational model 115 may apply a function 175 to determine a value (e.g., a score, etc.) indicating the probability of each sample (or updated molecule) within the noisy data distribution (or noisy latent distribution). In the case where function 175 is a scoring function, for example, a higher value (e.g., a lower score) may indicate that the sample (or updated molecule) was selected from a noisy data distribution region exhibiting a larger positive local density change (e.g., an increase, or a smaller decrease), or similarly, that the sample (or updated molecule) has a higher probability within the noisy data distribution. Therefore, in some cases, when selecting a first sample (or a first updated molecule) and a second sample (or a second updated molecule), the molecular design computational model 115 may apply a denoising model 117 to continue updating the three-dimensional representation or its embedding 154 of the input molecule 152 in order to select additional samples (or further updated molecules) from regions of gradually increasing density in the noisy data distribution, until, for example, a sample (or updated molecule) exhibiting a threshold probability within the noisy data distribution (or noisy latent distribution) is selected. For example, in some cases, if the three-dimensional representation or embedding 154 of the first updated input molecule 152 (or the first updated molecule) is selected from a region of higher density in the data distribution, the denoising model 117 may be applied to further modify the three-dimensional representation or embedding 154 of the first updated input molecule 152 (or the first updated molecule) instead of the second updated (or the second updated molecule). Doing so is similar to “traversing” the noisy data distribution (or the noisy latent distribution) to sample from regions of gradually increasing density in the noisy data distribution. In the case where the denoising model 117 is modifying the embedding 154 of the three-dimensional representation of the input molecule 152, the denoising model 117 can operate in a noisy latent space, in which the distance between two or more embeddings reflects the similarity (or dissimilarity) of the types and locations of atoms in different molecules. Drastic shifts in the density of the true data distribution of molecules exhibiting one or more desired properties can be smoothed by adding noise. Since the denoising model 117 is trained to approximate the data distribution of molecules exhibiting certain desired properties (e.g., drug-like properties), the update to the embedding 154 when denoising the input molecule 152 can be consistent with the types and locations of atoms found in molecules exhibiting one or more desired properties.Therefore, the same desired properties can also exist in the output molecule 162 generated by the molecular design computation model 115, which applies the denoising model 117 to denoise the input molecule 152.
[0138] Figure 3A depicts a flowchart illustrating an example of a process 300 for training a molecular design computational model 115 according to some exemplary embodiments. Referring to Figures 1 through 2 and Figure 3A, process 300 implements operation 204 of process 200 shown in Figure 2. In some cases, process 300 may be performed by molecular design engine 110 to train molecular design computational model 115 (including, for example, a denoised model 117) to approximate a noisy data distribution of a noisy three-dimensional representation (e.g., a noisy voxelized representation) of a molecule exhibiting one or more desired properties. As described in more detail below, in some cases, molecular design computational model 115 (including denoised model 117) may be trained to approximate a noisy data distribution rather than a true data distribution to avoid overfitting molecular design computational model 115 to known molecules that can be used to train molecular design computational model 115. Furthermore, in some cases, the molecular design computational model 115 (including the denoised model 117) can be trained by gradient-based Markov chain Monte Carlo (MCMC) sampling (including, for example, Markov chain Monte Carlo (MCMC) sampling with Langevin dynamics).
[0139] At 302, the molecular design engine 110 may apply a molecular design computational model with a first adjustment to denoise the damaged sample molecule and generate a first updated molecule. In some exemplary embodiments, step 302 may include: the molecular design engine 110 training a molecular design computational model 115 (including, for example, a denoising model 117) to approximate a data distribution of three-dimensional representations (e.g., voxelized representations) of molecules exhibiting one or more desired properties, such that candidate molecules exhibiting the same desired properties can be generated by sampling from them. In some cases, the molecular design computational model 115 may be trained based on a training dataset of damaged sample molecules to approximate the aforementioned data distribution, each of which is generated based on a noisy three-dimensional representation (e.g., a voxelized representation) of sample molecules (e.g., known molecules) from the data distribution. An example of such an example is shown in Figure 1A, where the molecular design computational model 115 (e.g., a denoising model 117) is trained to recover a noisy three-dimensional representation 182 of the sample molecule from a damaged three-dimensional representation 184 of the sample molecule generated by the damage engine 121. In some cases, the molecular design computational model 115 is not trained to directly recover the noisy three-dimensional representation (e.g., voxelized representation) of the sample molecule, but may be trained based on the corrupted embeddings of those three-dimensional representations (e.g., voxelized representations). This is illustrated in Figure 1B, where the molecular design computational model 115 (e.g., denoised model 117) is trained to recover the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule from the corrupted embedding 188 generated by the corruption engine 121.
[0140] In some exemplary embodiments, training the molecular design computation model 115 may include applying a denoising model 117 to denoise the damaged 3D representation (e.g., voxelized representation) of each sample molecule, or alternatively denoising its damaged embeddings. Figure 1A illustrates an example where training the molecular design computation model 115 includes adjusting parameters (e.g., weights, biases, etc.) of the denoising model 117 to, for example, gradually reduce the difference (e.g., mean squared error (MSE)) between the noisy 3D representation (e.g., voxelized representation) of the sample molecule and those recovered by the denoising model 117 from the damaged 3D representation of the sample molecule. Alternatively, in the example shown in Figure 1B, the molecular design computation model 115 may be trained by adjusting the parameters (e.g., weights, biases, etc.) of the denoising model 117 to gradually reduce the difference (e.g., mean squared error) between the embeddings of the noisy 3D representation of the sample molecule and the embeddings recovered by the denoising model 117 from the corresponding damaged embeddings. In some cases, the parameters of the denoising model 117 (e.g., weights, biases, etc.) may undergo different adjustments, and then further adjustments may be made to those that produce lower errors (e.g., mean squared error (MSE)). For example, in some cases, the parameters of the denoising model 117 (e.g., weights, biases, etc.) may be first adjusted, and then the denoising model 117 with the first adjustment may be applied to denoise the damaged 3D representation of the sample molecule or the damaged embedding of the 3D representation of the sample molecule, and at least a first updated molecule may be generated. In some cases, the first updated molecule may be an updated 3D representation of the first molecule (e.g., a voxelized representation), or alternatively an updated embedding of the 3D representation of the first molecule (e.g., a voxelized representation).
[0141] In some cases, a damaged 3D representation of a sample molecule or a damaged embedding of that 3D representation can be denoised by updating one or more atomic density values representing the types and locations of atoms present in the sample molecule. When a damaged embedding is generated by downsampling (or compressing) the 3D representation of the sample molecule, at least some of the values being updated can compress multiple features (or atomic density values) from the initial 3D representation (e.g., a voxelized representation) of the sample molecule. As described in more detail below, a denoising model 117 with a second adjustment (instead of a first adjustment) can be applied to denoise a damaged 3D representation (e.g., a damaged voxelized representation) or a damaged embedding of that 3D representation of the sample molecule, and generate at least a second updated molecule. The denoising model 117 with the first or second adjustment can be further adjusted. Doing so trains the denoising model 117 to approximate a noisy data distribution, or in some cases, a noisy latent distribution, exhibiting a smoother density transition to support more efficient sampling, since there are no steep gradient changes that restrict sampling to the immediate vicinity of the sample molecules forming the basis of the training dataset.
[0142] In some exemplary embodiments, training the denoising model 117 may further include determining a function 175. As described above, in some cases, the function 175 may be a scoring function parameterized by the parameters of the denoising model 117 (e.g., weights, biases, etc.). Therefore, in some cases, training the molecular design computation model 115 (which includes adjusting the parameters of the denoising model 117) may also include adjusting the parameters of the function 175. For example, in some cases, the function 175 may be determined by performing gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.) to approximate the gradient of the noisy data distribution (or noisy latent distribution). Doing so may include adjusting the parameters of the function 175 through one or more iterations such that the function 175 outputs a value (e.g., a score) indicating local density changes in the noisy data distribution (or noisy latent distribution). When function 175 is a scoring function, the parameters of function 175 can be adjusted such that function 175 assigns higher values (e.g., higher scores) to samples from locations exhibiting more positive local changes in density (e.g., increases, or small decreases) than to samples from locations exhibiting less positive local changes in density (e.g., decreases, or small increases). Therefore, once the denoising model 117 is trained, function 175 can output values (e.g., scores) that distinguish between samples from higher-density regions of a noisy data distribution (or a noisy latent distribution) (e.g., 3D representations, embeddings of 3D representations, etc.) and those sampled from lower-density regions of a noisy data distribution.
[0143] At 304, the molecular design engine 110 may apply a molecular design computational model with a second adjustment to denoise the damaged sample molecule and generate a second updated molecule. In some exemplary embodiments of step 304, when the denoising model 117 with a first adjustment is applied to generate at least a first updated molecule, the denoising model 117 with a second adjustment may be applied to generate at least a second updated molecule, such as, for example, an updated three-dimensional representation (e.g., an updated voxelized representation of the second molecule or an updated embedding of the three-dimensional representation of the second molecule). It should be understood that the first and second adjustments may include different changes to the parameters of the denoising model 117 (e.g., weights, biases, etc.). Therefore, applying the denoising model 117 with the second adjustment to denoise the damaged 3D representation 184 of the sample molecule or the damaged embedding 186 of the noisy 3D representation 182 of the sample molecule can produce an updated molecule that is different from the denoising model 117 with the first adjustment to denoise the damaged 3D representation 182 of the sample molecule or the damaged embedding 186 of the noisy 3D representation 182 of the same sample molecule. As described in more detail below, training the denoising model 117 may include further adjusting the denoising model 117 with the first adjustment or the second adjustment based on the difference (e.g., mean squared error (MSE)) present in the noisy 3D representation 182 of the sample molecule (FIG. 1A) or the embedding 184 of the noisy 3D representation 182 of the sample molecule recovered by the denoising model 117.
[0144] At 306, the molecular design engine 110 may determine that the first updated molecule is more similar to the sample molecule than the second updated molecule. In some exemplary embodiments, step 306 may include: if the first updated molecule generated by the first adjusted denoising model 117 is more similar to the noisy 3D representation (or its embedding) of the sample molecule (e.g., exhibiting a lower mean square error (MSE)) than the second updated molecule generated by the second adjusted denoising model 117, the molecular design engine 110 selects the first adjusted denoising model 117 instead of the second adjusted denoising model 117 for further adjustments during subsequent iterations. For example, in FIG1A, the first updated molecule may be an updated 3D representation of a first molecule that has a smaller difference (e.g., a lower mean square error (MSE)) relative to the noisy 3D representation 182 of the sample molecule compared to the second updated molecule. In Figure 1B, the first updated molecule can be the updated embedding of the three-dimensional representation of the first molecule that has a smaller difference (e.g., a lower mean square error (MSE)) in the embedding 186 of the noisy three-dimensional representation 182 relative to the sample molecule compared to the second updated molecule.
[0145] The fact that the first updated molecule is more similar to the noisy three-dimensional representation 182 (or its embedding 186) of the sample molecule than the second updated molecule suggests that the denoising model 117 with the first adjustment is better than the denoising model 117 with the second adjustment in recovering the noisy three-dimensional representation 182 (or its embedding 186) of the sample molecule. Therefore, the denoising model 117 with the first adjustment may better approximate the noisy data distribution (or noisy latent distribution) of molecules exhibiting one or more desired properties than the denoising model 117 with the second adjustment. Therefore, in some cases, the molecular design engine 110 may choose the denoising model 117 with the first adjustment instead of the denoising model 117 with the second adjustment to undergo one or more additional adjustment iterations.
[0146] At 308, the molecular design engine 110 may further adjust the molecular design computation model with the first adjustment instead of the molecular design computation model with the second adjustment until one or more criteria are met. In some exemplary embodiments, if the first updated molecule generated by the denoised model 117 with the first adjustment is more similar to the noisy three-dimensional representation 182 (or its embedding 186) of the sample molecule (e.g., with a lower mean squared error (MSE)) than the second updated molecule generated by the denoised model 117 with the second adjustment, the molecular design engine 110 may further adjust the denoised model 117 with the first adjustment instead of the denoised model 117 with the second adjustment. For example, during subsequent iterations of the adjustment, the molecular design engine 110 may further adjust the parameters (e.g., weights, biases, etc.) of the denoised model 117 with the first adjustment, and then apply the further adjusted denoised model 117 to generate one or more additional updated molecules. In some cases, the denoising model 117 may be further tuned to further increase the similarity (or reduce the mean squared error (MSE)) between the updated molecules generated by the denoising model 117 and the noisy 3D representations of sample molecules (or their embeddings) in the training dataset. In some cases, the molecular design engine 110 may continue to tune the denoising model 117 until one or more criteria are met. For example, in some cases, the molecular design engine 110 may continue to tune the parameters of the denoising model 117 (e.g., weights, biases, etc.) until the molecular design engine 110 has performed threshold-based iterative tuning. Alternatively and / or additionally, the molecular design engine 110 may continue to tune the parameters of the denoising model 117 (e.g., weights, biases, etc.) until the similarity (e.g., mean squared error (MSE)) between the updated molecules generated by the denoising model 117 and the noisy 3D representations (or their embeddings) of sample molecules in the training dataset reaches one or more thresholds. In some cases, the molecular design engine 110 may continue to adjust the parameters of the denoising model 117 (e.g., weights, biases, etc.) until the updated molecules generated by the denoising model 117 exhibit a threshold probability in the data distribution of molecules exhibiting the one or more desired properties in the training dataset.
[0147] As described in more detail below, once one or more of the criteria are met, the trained denoising model 117 can be applied to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162 by denoising at least the three-dimensional representation (e.g., a voxelized representation) of the input molecule 152. As shown in Figures 1A and 1B, the trained denoising model 117 can generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162 by at least the following: sampling from a noisy data distribution filled with noisy three-dimensional representations (e.g., voxelized representations) of molecules exhibiting one or more desired properties (e.g., drug-like properties) based on function 175, or alternatively sampling from a noisy latent distribution filled with embeddings of noisy three-dimensional representations (e.g., voxelized representations) of molecules. Sampling may include one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo, etc.), which may be guided by function 175 such that each sampling iteration includes: selecting one or more samples (or molecules) from regions of gradually increasing density of a noisy data distribution (or a noisy latent distribution).
[0148] Figure 3B depicts a flowchart illustrating an example of a process 325 for applying a molecular design computational model to generate a three-dimensional molecule in voxelized space according to some exemplary embodiments. Referring to Figures 1A, 1B, 2, and 3B, process 325 can implement operation 206 of process 200 shown in Figure 2. In some cases, process 325 can be performed by a molecular design engine 110. For example, in some cases, molecular design engine 110 can apply a molecular design computational model 115 (e.g., a denoising model 117) to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule by denoising at least the three-dimensional representation (e.g., a voxelized representation) of the input molecule. In some cases, the input molecule can be a random molecule (e.g., a molecule with random selection of atom types and / or locations) or a known molecule with one or more desired properties. Therefore, the three-dimensional representation (e.g., voxelized representation) of the input molecule may include noise that needs to be removed by the molecular design computational model 115 so that the resulting three-dimensional representation (e.g., voxelized representation) of the output molecule is consistent with a molecule exhibiting one or more desired properties (e.g., drug-like properties). The molecular design computational model 115 can denoise the three-dimensional representation (e.g., voxelized representation) of the input molecule by at least the following: sampling from a noisy data distribution (or a noisy latent distribution), which is more efficient because the smoother density transitions present therein allow for a thorough exploration of the data distribution. As described in more detail below, once the molecular design computational model 115 generates the three-dimensional representation (e.g., voxelized representation) of the output molecule, the molecular design engine 110 can further generate one or more other representations of the output molecule, including, for example, a one-dimensional representation of the output molecule, a two-dimensional representation of the output molecule, etc. The output molecule is generated by manipulating the three-dimensional representation (e.g., voxelization representation) of the input molecule to capture the conformation (or three-dimensional structure) of the input molecule, which means that the conformation (or three-dimensional) structure of the output molecule is more likely to be consistent with one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, exploitability, etc.).
[0149] At 332, the molecular design engine 110 may update the three-dimensional representation of the input molecule to generate an updated three-dimensional representation. In some exemplary embodiments, updating the three-dimensional representation may include: the molecular design engine 110 applying a molecular design computational model 115 to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162 by denoising at least the three-dimensional representation (e.g., a voxelized representation) of the input molecule 152. An example of such a process is illustrated in Figure 1A, where a denoising engine 117 denoises the three-dimensional representation of the input molecule 152 to generate an updated three-dimensional representation 160. In some cases, the input molecule 152 may be a noisy molecule (e.g., a molecule with a random selection of atom types and / or locations) or a known molecule with one or more undesirable properties. This means that the three-dimensional representation of the input molecule 152 may include at least some noise that makes it inconsistent with the three-dimensional representation of a molecule exhibiting one or more desired properties (e.g., drug-like properties). Therefore, in some cases, the denoising engine 117 can be trained to update the three-dimensional representation of the input molecule 152 such that the resulting updated three-dimensional representation 160 is consistent with the three-dimensional representation of the molecule exhibiting the one or more desired properties.
[0150] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to update the three-dimensional representation of the input molecule 152 based on a function 175. In some cases, the function 175 may be a scoring function that outputs a value (e.g., a score) indicating the likelihood of a sample (or molecule) within the noisy data distribution for each sample (or molecule) selected from the noisy data distribution. For example, in some cases, the value output by function 175 for a particular sample (or molecule) may indicate a local variation in the density of the sample (or molecule) from the location where it was selected. The denoising model 117 may update the three-dimensional representation of the input molecule 152 over multiple successive sampling iterations, at least based on the value output by function 175. During each sampling iteration, the denoising model 117 may be applied to further update the three-dimensional representation of the input molecule 152, such that the resulting updated three-dimensional representation 160 is selected from regions of higher density within the noisy data distribution than the regions selected during one or more previous sampling iterations.
[0151] In some exemplary embodiments, the molecular design computational model 115 may perform gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling) of a noisy data distribution, wherein the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 is updated over multiple successive sampling iterations. In some cases, each iteration may include: the molecular design computational model 115 further updating the three-dimensional representation of the input molecule 152 to sample from regions with progressively increasing density of the noisy data distribution. Furthermore, in some cases, updates to the three-dimensional representation of the input molecule 152 may accumulate over multiple successive iterations. For example, in some cases, the three-dimensional representation of the input molecule 152 may undergo a first update and a second update. The molecular design computational model 115 may apply a function 175 to determine: a first value (e.g., a first score, etc.) for the three-dimensional representation of the input molecule 152 with the first update, and a second value (e.g., a second score, etc.) for the three-dimensional representation of the input molecule 152 with the second update. During subsequent iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, if the first and second values indicate that the three-dimensional representation of the input molecule 152 with the first update was sampled from a higher-density region of the noisy data distribution compared to the three-dimensional representation with the second update, and exhibits a higher probability of being within the noisy data distribution, then the denoising model 117 may be applied to further update the three-dimensional representation of the input molecule 152 with the first update.
[0152] In some cases, one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling may be performed, wherein the molecular design computational model 115 applies a denoising model 117 to further modify the three-dimensional representation of the input molecule 152 until one or more criteria are met. For example, in some cases, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until a threshold amount of sampling iterations is performed. Alternatively and / or additionally, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until the function 175 outputs a value (e.g., a score, etc.) for the updated three-dimensional representation 160 that reaches one or more thresholds. The value (e.g., a score, etc.) associated with the updated three-dimensional representation 160 reaching one or more thresholds may indicate that the updated three-dimensional representation 160 is selected from a region with a sufficiently high density of noisy data distribution, and that the probability of the updated three-dimensional representation 160 being within the noisy data distribution reaches one or more thresholds. In some cases, the one or more criteria may also include output molecules that have generated a threshold amount exhibiting the one or more desired properties (e.g., at least one output molecule exhibiting one or more drug-like properties (such as affinity, specificity, bioactivity, extensibility, etc.) at the threshold level).
[0153] At 336, the molecular design engine 110 can denoise the updated 3D representation to generate a 3D representation of the output molecule. In some exemplary embodiments of step 336, the molecular design computation model 115 can denoise the 3D representation of the input molecule 152 by sampling from a noisy data distribution occupied by noisy 3D representations of molecules exhibiting one or more desired properties. As described above, the molecular design computation model 115 (including the denoising model 117) can be trained to approximate the noisy data distribution (rather than the true data distribution) by at least the following: being trained to denoise the damaged 3D representation 184 of the sample molecule to recover the noisy 3D representation 182 of the sample molecule rather than the clean 3D representation of the sample molecule. Furthermore, such a noisy data distribution can exhibit a smoother density transition, and therefore sampling from it is more efficient. Sampling the updated 3D representation 160 from the noisy data distribution means that the updated 3D representation 160 can undergo additional denoising. For example, Figure 1A illustrates that the molecular design engine 110 can apply a recovery model 118 to denoise the updated 3D representation 160 and generate a 3D representation (e.g., a voxelized representation) of the output molecule 162. In some cases, the recovery model 118 can be trained to denoise the updated 3D representation 160 to map the updated 3D representation 160 from a noisy data distribution back to the true data distribution of molecules exhibiting the one or more desired properties (e.g., drug-like properties). It should be understood that this denoising differs from the denoising performed by the denoising model 117, which includes updating the 3D representation of the input molecule 152 to sample from higher-density regions of the noisy data distribution that are more likely to be occupied by molecules exhibiting the one or more desired properties.
[0154] At 338, the molecular design engine 110 may generate one or more other representations of the output molecule based at least on the three-dimensional representation of the output molecule. In some exemplary embodiments, the three-dimensional representation of the output molecule 162 (e.g., a voxelized representation) may be further transformed into one or more other representations of the output molecule 162, which are generated by a recovery model 118 that denoises an updated three-dimensional representation 160 sampled from a noisy data distribution by a molecular computation model 115. For example, in some cases, the molecular design engine 110 may recover the locations (e.g., coordinates) of atoms present in the output molecule 162 and one or more bonds between them based at least on the three-dimensional representation of the output molecule 162 (e.g., a voxelized representation). In doing so, the molecular design engine 110 may determine another representation of the output molecule 162, including, for example, a one-dimensional representation of the output molecule 162 (e.g., a simplified molecular linear input specification (SMILES) string), a two-dimensional representation of the output molecule 162 (e.g., a molecular graph), etc. In some cases, the molecular design engine 110 can recover the localization of atoms present in the output molecule 162 by applying peak detection techniques. These peak detection techniques determine the atom localization (e.g., coordinates) based on one or more peaks in the atomic density included in the three-dimensional representation (e.g., voxelized representation) of the output molecule 162, and then determine one or more interconnecting bonds based on the atom localization. Alternatively, the molecular design engine 110 can apply a machine learning model trained to transform the voxelized representation of the output molecule 162 into one or more other representations.
[0155] Figure 3C depicts a flowchart illustrating an example of a process 350 for applying a molecular design computational model to generate a three-dimensional molecule in voxelized space, according to some exemplary embodiments. Referring to Figures 1 through 2 and Figure 3C, process 350 may implement operation 206 of process 200 shown in Figure 2. In some cases, process 350 may be performed by a molecular design engine 110. For example, in some cases, molecular design engine 110 may apply a molecular design computational model 115 (e.g., a denoising model 117) to generate a three-dimensional representation of an output molecule, such as a voxelized representation of the output molecule, by denoising at least the three-dimensional representation (e.g., a voxelized representation) of the input molecule. In some cases, the three-dimensional representation of the input molecule may be denoised by updating the embedding of the three-dimensional representation (e.g., the voxelized representation) of the input molecule instead of directly updating the three-dimensional representation of the input molecule, at least because the embedding is more compact and computationally more efficient to operate on. In some cases, the embedding of the 3D representation of the input molecule can be generated by downsampling (or compressing) the 3D representation of the input molecule, although it is also possible to generate the embedding without any downsampling (or compression) of the 3D representation of the input molecule. In the former case, the embedding of the 3D representation of the input molecule can occupy a latent voxelization space, while in the latter case, the embedding of the 3D representation of the input molecule can remain in the same discrete voxelization space as the initial 3D representation of the input molecule. It should be understood that the latent voxelization space can have a lower dimension than the discrete voxelization space, allowing operations on the embedding of the 3D representation of the input molecule to improve the speed and computational efficiency of the generative process, while achieving comparable or better generative performance.
[0156] It should be understood that the molecular design computational model 115 can denoise the embedding of the three-dimensional representation of the input molecule by sampling from a noisy latent distribution. That is, as described above, the molecular design computational model 115 can be trained to approximate a noisy latent distribution rather than the true data distribution to at least avoid the steep density transitions present in the true data distribution. In other words, the updated embedding generated by the molecular design computational model 115, which updates the embedding of the three-dimensional representation of the input molecule, can still occupy the noisy latent distribution. This noisy latent distribution can be sampled from more effectively because the smoother density transitions of the noisy latent distribution support a full exploration of the data distribution. As described in more detail below, the updated embedding can undergo decoding and further denoising in order to “jump” back to the true data distribution. Furthermore, in some cases, the molecular design engine 110 can generate one or more other representations of the output molecule based on the three-dimensional representation of the output molecule resulting from the decoding and denoising of the updated embedding, including, for example, a one-dimensional representation of the output molecule, a two-dimensional representation of the output molecule, etc. The output molecule is generated by manipulating the three-dimensional representation (e.g., voxelization representation) of the input molecule to capture the conformation (or three-dimensional structure) of the input molecule, which means that the conformation (or three-dimensional) structure of the output molecule is more likely to be consistent with one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, exploitability, etc.).
[0157] At 352, the molecular design engine 110 can encode a three-dimensional representation of the input molecule to generate an embedding of the input molecule. In some exemplary embodiments, the encoder 111 can encode a three-dimensional representation (e.g., a voxelized representation) of the input molecule 152 to generate an embedding 154 of the input molecule 152. Such an example is shown in Figure 1B. In the case of “seed generation,” the input molecule 152 can be a known molecule (e.g., a molecule from a validation set derived from the PubChem dataset, the QM9 molecular dataset, the Geometry Integration of Molecular (GEOM) drug dataset, etc.). In some cases, the known molecule may exhibit one or more undesirable properties. When a known molecule is used as the input molecule 152, the generative process can be initialized with a voxel grid having an atomic density distribution corresponding to the types and locations of atoms expected to be found in the known molecule. Alternatively, the molecular design computational model 115 can be generated de novo, in which case the input molecule 152 can be a noisy molecule whose atomic types and locations correspond to pure noise (e.g., uniform noise, etc.). When a noisy molecule is used as input molecule 152, the generative process can be initialized across the entire voxel grid without any expectation of atom type and / or location. In either case, the atom type and / or location in input molecule 152 may be inconsistent with the atom type and / or location of molecules exhibiting the one or more desired properties (e.g., drug-like properties). Therefore, molecular design computational model 115 can be applied to update the 3D representation of input molecule 152 by at least updating embedding 154 and generating an updated embedding 156 such that the corresponding 3D representation of output molecule 162 is more consistent with those of molecules exhibiting the one or more desired properties.
[0158] In some exemplary embodiments, encoder 111 may encode a three-dimensional representation (e.g., a voxelized representation) of input molecule 152 by downsampling or compressing at least a three-dimensional representation of input molecule 152. This may include compressing at least some of the features present in the three-dimensional representation of input molecule 152, which reduces the dimension (or amount of features) present in the three-dimensional representation of input molecule 152. For example, in the case where the three-dimensional representation of input molecule 152 comprises a [32×32×32] voxel grid containing 32,000 features (or atomic density values), encoder 111 may compress at least some of these 32,000 features (or atomic density values) to generate a [4×4×4] voxel grid containing 64 features as an embedding 154 of input molecule 152.
[0159] In some exemplary embodiments, encoder 111 may generate an embedding 154 of the input molecule 152 with or without downsampling or compression of its three-dimensional representation (e.g., a voxelized representation). In some cases, encoder 111 may implement an identity function, meaning that embedding 154 may include the same amount of features (e.g., atomic density values) present in the three-dimensional representation of the input molecule 152. Alternatively, where embedding 154 is generated by downsampling the voxelized representation of the input molecule 152, this projects the voxelized representation of the input molecule 152 from a higher-dimensional discrete voxelized space to a lower-dimensional latent space. Sampling from the lower-dimensional latent space imposes less computational burden than sampling directly from the higher-dimensional discrete voxelized space. For example, in resource-intensive tasks such as sampling from discrete voxelization space, such as when the input molecule 152 is large (e.g., containing between 80 and 200 atoms) or when a large number of candidate molecules are being generated from it, the molecular design engine 110 can sample from the latent voxelization space by applying a molecular design computational model 115 to operate on an embedding 154 of the three-dimensional representation of the input molecule 152. It should be understood that even when the input molecule 152 is relatively large (e.g., containing more than 200 atoms) or when a large number of candidate molecules are being generated, sampling from the latent voxelization space can impose a modest computational overhead.
[0160] In some exemplary embodiments, encoder 111 may be part of an autoencoder (e.g., a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQ-VAE)) together with decoder 119. In some cases, encoder 111 may be trained to encode a voxelized representation of input molecule 152, such that decoder 119 is able to recover a three-dimensional representation (e.g., a voxelized representation) of input molecule 152 from the resulting embedding 154. To further illustrate, let x represent the voxelized representation of the molecule, f e Indicates encoder 111, f d This indicates decoder 119, and z e This indicates embedding 154. For example, encoder f can be used. e (x) Encode the voxelized molecular representation x (such as the voxelized molecular representation of input molecule 152) to generate a continuous latent embedding z according to the following equation (2). e (x).
[0161] (2)
[0162] According to equation (3), the continuous latent embedding ze Each of (x) can be quantized into a discrete latent embedding z by matching one of the k vectors in the learned embedding shared codebook e via nearest neighbor search.
[0163] (3)
[0164] This allows for the quantization of the latent embedding z q (x) via decoder f d The initial voxelized molecular representation x is reconstructed according to the following equation (4).
[0165] (4)
[0166] The latent embedding space can be represented as , where K is the number of discrete latent vectors in the learned codebook, and d is the dimension of each latent embedding vector in the codebook. It should be understood that K and d are hyperparameters that can be chosen experimentally.
[0167] In some cases, there may not be a gradient constrained for the nearest neighbor lookup in the codebook for each latent embedding because the operation is non-differentiable. Instead, the nearest neighbor lookup in the codebook replaces each quantized latent embedding z with one of the learned codebook embeddings of the same dimension. q (x). The stop gradient (sg) operation transfers the gradient from the input to the decoder f. d (θ d The quantized latent embedding z q (x) is copied to the encoder f before quantization. e (θ e The continuous latent embedding z of the output e (x). The stopping gradient (sg) operation can act as an identity function in the positive direction by copying the variable without making any changes. However, when updating the encoder f... e (θ e During the backward pass of the gradient of ), the stop gradient (sg) operation prevents the gradient from flowing through the gradient update for the specific term to which the operation is applied, at least because the gradient cannot be computed for that term.
[0168] In some exemplary embodiments, the encoder f that forms the autoencoder (e.g., variational autoencoder (VAE) etc.) e and decoder f d Training may include: adjusting the encoder f e and decoder f dThis is to reduce (or minimize) the three independent losses or loss terms. The first loss term may include reconstruction loss (e.g., mean squared error (MSE) reconstruction loss), which is related to the loss caused by the encoder f. e Ingestion to generate embedded z e The voxelized molecular representation x and the decoder f d Based on embedding z e The generated reconstruction The difference between them corresponds. The second loss term can be obtained by changing the embedding vector e. i Towards the encoder f e Output continuous latent embedding z e (x) The shift forces the learning of the embedding codebook e used to quantize the latent space. The third loss term quantizes the commitment loss, which ensures that the encoder f e To embed z e (x) promises that its output will not grow arbitrarily. This third loss term can be associated with the commitment cost weight β, which can also be a hyperparameter set experimentally. The following equation (5) is used to train the encoder f. e and decoder f d An example of the overall loss function L.
[0169] (5)
[0170] At 354, the molecular design engine 110 can generate an updated embedding by at least updating the embedding of the three-dimensional representation of the input molecule. In some exemplary embodiments, the molecular design engine 110 may apply a molecular design computation model 115 (e.g., a denoising model 117) to denoise the three-dimensional representation of the input molecule 152 by at least updating the embedding 154 of the three-dimensional representation of the input molecule 152 and generating an updated embedding 156. For example, in some cases, the three-dimensional representation of the input molecule 152 may include noise that causes inconsistencies between the types and / or locations of atoms present in the input molecule 152 and those types and / or locations of atoms in molecules exhibiting the one or more desired properties (e.g., drug-like properties). In other words, the molecular design computation model 115 may update the embedding 154 of the three-dimensional representation of the input molecule 152 to increase the likelihood that the resulting output molecule 162 exhibits the one or more desired properties. As described above, the noise being removed from the embedding 154 by the denoising model 117 should not be confused with noise that projects the three-dimensional representation of the input molecule 152 from the real data distribution, which exhibits a jagged density transition, to a noisy data distribution exhibiting a smoother density transition for more efficient sampling (e.g., gradient-based Markov chain Monte Carlo (MCMC) sampling, etc.). As described in more detail below, by updating the embedding 154 of the three-dimensional representation of the input molecule 152, the molecular design computational model 115 (e.g., the denoising model 117) can traverse the smoother density of the noisy data distribution to sample the updated embedding 156 from regions of gradually increasing density of the noisy data distribution, and then choose to “jump” back to the real data distribution when a sample exhibits a threshold probability within the noisy data distribution.
[0171] In some exemplary embodiments, the denoising model 117 may apply updates corresponding to changes in the type and / or location of atoms present in the input molecule 152 to an embedding 154 of a three-dimensional representation of the input molecule 152. Where the encoder 111 implements the identity function and the embedding 154 is generated without any downsampling (or compression) of the underlying three-dimensional representation (e.g., voxelized representation) of the input molecule 152, the denoising model 117 may update the embedding 154 by updating the atomic density of one or more voxels in at least one channel of the embedding 154. Alternatively, where the generation of the embedding 154 includes downsampling (or compression) of the underlying three-dimensional representation (e.g., voxelized representation) of the input molecule 152, the denoising model 117 may update the embedding 154 by updating at least one or more values present in the embedding 154, at least some of which are compressed multiple atomic density values included in the three-dimensional representation (e.g., voxelized representation) of the input molecule 152.
[0172] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to update the embedding 154 of the input molecule 152 based on a function 175. In some cases, the function 175 may output a value (e.g., a score, etc.) indicating the likelihood of a sample (or molecule) in the noisy data distribution for each sample (or molecule) selected from the noisy data distribution. For example, in some cases, the value output by the function 175 for a particular sample (or molecule) may indicate a local variation in the density of the sample (or molecule) from its selected location. The denoising model 117 may update the embedding 154 over multiple successive sampling iterations, at least based on the value output by the function 175. During each sampling iteration, the denoising model 117 may be applied to further update the embedding 154, such that the resulting updated embedding 156 is selected from regions of the noisy data distribution with higher density than the regions in previous sampling iterations.
[0173] In some exemplary embodiments, the molecular design computational model 115 may perform gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling) of a noisy data distribution, wherein the embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 is updated over multiple successive sampling iterations, wherein each iteration samples from regions of gradually increasing density in the noisy data distribution to increase the likelihood that the resulting updated embedding 156 is present in the noisy data distribution. Furthermore, in some cases, updates to the embedding 154 of the input molecule 152 may accumulate over multiple successive iterations. To further illustrate, consider an instance where the embedding 154 of the three-dimensional representation of the input molecule 152 undergoes a first update and a second update. The molecular design computational model 115 may apply a function 175 to determine: a first value (e.g., a first score, etc.) of the embedding 154 with the first update, and a second value (e.g., a second score, etc.) of the embedding 154 with the second update. During subsequent iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, if the first and second values indicate that the embedding 154 with the first update was sampled from a higher-density region of the noisy data distribution compared to the embedding 154 with the second update, and exhibits a higher probability of being within the noisy data distribution, then the denoising model 117 can be applied to further update the embedding 154 with the first update. In some cases, one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling may be performed, wherein the molecular design computational model 115 applies the denoising model 117 to further modify the embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 until one or more criteria are met. For example, in some cases, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until a threshold amount of sampling iterations is performed. Alternatively and / or additionally, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until the function 175 outputs a value (e.g., a score, etc.) for the updated embedding 156 that reaches one or more thresholds. The value (e.g., a score, etc.) associated with the updated embedding 156 reaching these one or more thresholds may indicate that the updated embedding 156 is selected from a region of noisy data distribution with a sufficiently high density, and that the probability of the updated embedding 156 being within the noisy data distribution reaches one or more thresholds.In some cases, the one or more criteria may also include output molecules that have generated a threshold amount exhibiting the one or more desired properties (e.g., at least one output molecule exhibiting one or more drug-like properties (such as affinity, specificity, bioactivity, extensibility, etc.) at the threshold level).
[0174] At 356, the molecular design computational model 115 can decode the updated embedding to generate a noisy three-dimensional representation of the output molecule. In some exemplary embodiments, the molecular design engine 110 can apply the decoder 119 to decode the updated embedding 156 and generate a noisy three-dimensional representation 158 of the output molecule 162 when the molecular design computational model 115 (e.g., denoising model 117) has been applied to update the embedding 154 of the three-dimensional representation of the input molecule 152 and an updated embedding 156 has been generated. Decoding the updated embedding 156 can map the updated embedding 156 from a latent voxelized space filled with embeddings of three-dimensional representations of various molecules to a latent discrete space. However, as described in more detail below, the latent discrete space can be a noisy latent space, which means that the noisy three-dimensional representation 158 generated by decoding the updated embedding 156 by the decoder 119 may need further denoising in order to project the noisy three-dimensional representation 158 back to the true data distribution of the molecule exhibiting the one or more desired properties.
[0175] In some exemplary embodiments, the decoder 119 of the molecular design engine 110 can generate a noisy three-dimensional representation 158 by decoding at least an updated embedding 156 generated by the molecular design computation model 115 (e.g., a denoising engine 117). As described above, in some cases, the decoder 119 may form part of an autoencoder (e.g., a variational autoencoder, such as a vector quantization variational autoencoder (VQ-VAE)) together with the encoder 111. In some cases, the encoder 111 and the decoder 119 may be trained in tandem, wherein the encoder 111 may be trained to generate embeddings of the three-dimensional representation of the molecule (e.g., a voxelized representation), such as embedding 154 of the three-dimensional representation of the input molecule 152, which enables the decoder 119 to recover the initial three-dimensional representation (e.g., the voxelized representation) from it. Therefore, when generating the updated embedding 156, the decoder 119 may be applied to recover the noisy three-dimensional representation 158 of the output molecule 162 from it.
[0176] In some cases, decoding the updated embedding 156 may include upsampling (or decompressing) the updated embedding 156, which may project the updated embedding 156 from the latent voxelization space back to the discrete voxelization space. The noisy 3D representation 158 of the output molecule 162 (e.g., a noisy voxelization representation) may exhibit the same dimension (or amount of features) as the 3D representation (e.g., a voxelization representation) of the input molecule 152 taken up by the molecular design engine 110 at operation 352. For example, in some cases, the 3D representation of the input molecule 162 may include a [32×32×32] voxel grid, meaning that the 3D representation of the input molecule 152 may include 32,000 features (or atomic density values). Meanwhile, each of the embeddings 154 operated on by the molecular design computation model 115 and the resulting updated embedding 156 may include a [4×4×4] voxel grid with 64 features. In some cases, decoder 119 can decode the updated embedding 156 by upsampling (or decompressing) the included [4×4×4] voxel grid to generate a [32×32×32] voxel grid for a noisy three-dimensional representation 158 (e.g., a noisy voxelized representation) of the output molecule 162. It should be understood that this upsampling (or decompression) recovers 32,000 features (or atomic density values) in the noisy three-dimensional representation 158 (e.g., a voxelized representation) of the output molecule 162. As described above, these 32,000 features (or atomic density values) indicate the location of various atoms present in the output molecule 162. Furthermore, these 32,000 features (or atomic density values) may span one or more channels, each of which corresponds to a type of atom that may be present in the output molecule 162.
[0177] At 358, the molecular design engine 110 can denoise the noisy 3D representation of the output molecule to generate a 3D representation of the output molecule. In some exemplary embodiments, the denoising engine 117 of the molecular design computation model 115 can generate a 3D representation (e.g., a voxelized representation) of the output molecule 162 by at least the following: denoising the noisy 3D representation 158 generated by decoding the updated embedding 156 by the decoder 119. As described above, in some cases, the molecular design computation model 115 (e.g., the denoising model 117) can undergo one or more iterations of gradient-based Markov chain Monte Carlo (e.g., Langevin Markov chain Monte Carlo, etc.) to generate the updated embedding 156. In doing so, the molecular design computational model 115 may traverse the noisy latent distribution based at least on the output of function 175 (e.g., the score output by function 175) to sample the updated embedding 156 from higher-density regions of the noisy latent distribution filled with embeddings of 3D representations of molecules more likely to exhibit the one or more desired properties (e.g., drug-like properties). However, decoding the updated embedding 156 only maps the updated embedding 156 from the latent voxelization space to the discrete voxelization space, but the noisy 3D representation 158 still occupies the noisy data distribution rather than the true data distribution of molecules exhibiting the one or more desired properties. Therefore, in some cases, the recovery model 118 may be applied to map the noisy 3D representation 158 from the noisy data distribution to the true data distribution. In some cases, this may constitute a “jump” back to the true data distribution, meaning that the 3D representation of the output molecule 162 generated from it occupies the true data distribution.
[0178] In some cases, the recovery model 118 may share the same architecture (e.g., an artificial neural network (ANN) or the like) as the denoising model 117, which is trained to traverse noisy latent distributions to denoise the embedding 154 of the three-dimensional representation of the input molecule 152 and generate an updated embedding 156. However, as mentioned above, the recovery model 118 may be trained to remove different types of noise. Therefore, in some cases, the recovery engine 118 may be trained based on a training dataset to denoise the noisy three-dimensional representation 182 of the sample molecule and recover the initial three-dimensional representation 182 from it. In contrast, the denoising engine 117 may be trained to recover the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule from the corrupted embedding 188. In this context, training the denoising engine 117 may include adjusting one or more parameters of the denoising engine 117 (e.g., an artificial neural network (ANN) or the like) to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the initial three-dimensional representation of the sample molecule and the three-dimensional representation of the sample molecule recovered by the denoising engine 117 from the noisy three-dimensional representation of the sample molecule.
[0179] To further illustrate, consider the quantized latent embedding z described in operation 252. q (x). As described above, the quantization of the latent embedding z q (x) can be encoded by an encoder f that encodes the voxelized molecular representation x. e (x) Generation. In some cases, noise ϵ (e.g., Gaussian noise, such as isotropic Gaussian noise) can be added to the quantized latent embedding z. q (x). For example, in some cases, noise ϵ with a unit covariance matrix scaled to a fixed large noise level σ can be added according to the following equation (6).
[0180] (6)
[0181] A denoising engine 117, which can be represented as a latent model ζ(ϕ), can be trained on the latent embedding z. q (x) performs denoising and restoration, while simultaneously improving the initial latent embedding z before (or without) adding noise ϵ. q (x) and the denoised latent embedding generated by denoising engine 117 The reconstruction loss between these parameters (e.g., mean squared error (MSE) reconstruction loss) is reduced (or minimized). This is achieved by using the latent model ζ(ϕ) to generate the denoised latent embedding. The denoising is shown in Equation (7) below. Meanwhile, Equation (8) shows the loss function L used to train the latent model ζ(ϕ), which includes: [the loss function L is then applied to the initial latent embedding z before (or without) adding the noise ϵ]. q (x) with the denoised latent embedding The difference between them (e.g., mean squared error (MSE)) is reduced (or minimized).
[0182] (7)
[0183] (8)
[0184] At 360°, the molecular design engine 110 may generate one or more other representations of the output molecule based at least on a three-dimensional representation of the output molecule. In some exemplary embodiments, the molecular design engine 110 may generate one or more other representations of the output molecule 162 based at least on a voxelized representation of the output molecule 162, including, for example, a one-dimensional representation of the output molecule 162 (e.g., a simplified molecular linear input specification (SMILES) string), a two-dimensional representation of the output molecule 162 (e.g., a molecular graph), etc. For example, in some cases, the molecular design computational model 110 may recover the locations (e.g., coordinates) of atoms present in the output molecule 162 and the bonds between them from the voxelized representation of the output molecule 162. In some cases, the molecular design engine 110 may apply a peak detection technique that determines the locations (e.g., coordinates) of atoms present in the output molecule 162 based on one or more peaks in the atomic density included in the voxelized representation of the output molecule 162, and then determines one or more interconnecting bonds based on the atomic locations. Alternatively, the molecular design engine 110 may apply a machine learning model trained to transform the voxelized representation of the output molecule 162 into one or more other representations.
[0185] As described above, in some exemplary embodiments, the molecular design computational model 115 (including the denoising model 117) may operate on a three-dimensional representation of the molecule rather than a one-dimensional or two-dimensional representation, at least because realistic and efficient molecules exhibiting certain desired properties are more likely to be generated based on a molecular representation that captures both the molecule's composition (e.g., constituent atoms) and its conformation (or three-dimensional structure). In some cases, the molecular design computational model 115 (including the denoising model 117) may operate on a voxelized representation of the molecule. Unlike conventional three-dimensional representations of molecules (e.g., point cloud representations, etc.), voxelized representations of molecules may represent both atom types and locations as one or more continuous (e.g., Gaussian-like) distributions across a voxel grid centered on the atomic coordinates of individual atoms. Therefore, unlike conventional three-dimensional representations of molecules (e.g., point cloud representations, etc.), the molecular design computational model 115 can apply a denoising model 117 to operate on the voxelized representation of the input molecule without requiring any workarounds to reconcile different types of data distributions (e.g., discrete distributions for atom types and continuous distributions for atom locations) and without requiring any prior knowledge of the number of atoms present in the output molecule produced from it.
[0186] To further illustrate, Figure 4 depicts examples of voxelized representations and corresponding two-dimensional representations of different molecules according to some exemplary embodiments. For example, Figure 4 shows a voxelized representation 400 and a two-dimensional representation 450 of the molecule. In some exemplary embodiments, the voxelized representation 400 of the molecule can be generated by partitioning (or discretizing) the three-dimensional space surrounding the constituent atoms into a voxel grid 410, wherein each type of atom (or element) present in the molecule is represented by a different grid channel. This partitioning (or discretization) can generate n voxelized molecules. , , , where l represents the length of each grid edge, and c represents the number of channels in the dataset (e.g., the amount of atoms (or elements) of different types).
[0187] In some cases, the voxel grid 410 can be a three-dimensional voxel grid organized into continuous layers of rows and columns. Each voxel in the voxel grid 410 can be a volume element formed at the intersection of rows and columns, such as a three-dimensional cube. Furthermore, each voxel in the voxel grid 410 can be associated with a value indicating the atomic density at the corresponding location (e.g., having a value [0,1]). For a single molecule, the corresponding voxelized representation can be a box around the center of the molecule, which is then divided into voxels. To generate a voxelized representation 400 of the molecule, each constituent atom can be converted to a three-dimensional continuous (e.g., Gaussian-like) density according to the following equation (9). For example, an example of the voxel grid 410 shown in FIG4 may include a first atomic density 415a representing a first atom of a first type and a second atomic density 415b representing a second atom of a second type.
[0188] (9)
[0189] Where V α It is defined as having a radius r at a distance d from the atom center. α The fraction of the volume occupied by atom α. Different types of atoms (or elements) can have different radii or the same radius (e.g., r). α =.5Å). According to the following equation (10), the occupancy of each voxel in the voxel grid can be calculated by integrating the occupancy generated by each atom in the molecule. cc .
[0190] (10)
[0191] Where N α α represents the number of atoms in a molecule. n For the nth th One atom, C i,j,k Let x be the coordinates (i,j,k) in the voxel grid, and x n The coordinates represent the center of atom n.
[0192] As mentioned above, in some cases, the voxelization of a molecule represents an atomic density of 400 that can exist centered on atoms within the molecule. Therefore, the occupancy rate O ccThe value can be maximized at the atom center (e.g., a value of 1) and decreases to a minimum (e.g., a value of 0) as the distance from the atom center increases. Each channel in the voxel grid can be independent. That is, channels do not interact or share volume contributions. In some cases, the size of the voxel grid 410 included in the voxelized representation 400 of the molecule can correspond to the size of the represented molecule (e.g., the amount of constituent atoms). For example, in some cases, if the molecule has fewer atoms (e.g., the QM9 molecular dataset), the voxel grid 410 can be a [32×32×32] voxel grid, or if the molecule has more atoms (e.g., the Geometry Integration of Molecular (GEOM) drug dataset), the voxel grid can be a [64×64×64] voxel grid. Furthermore, in some cases, the number of channels in the voxelized representation 400 of the molecule can correspond to the number of atom types (or elements) present in the molecule. For example, the voxelized representation of molecules in the QM9 molecular dataset may include five channels for the five types of atoms that make up those molecules (e.g., carbon (C), hydrogen (H), oxygen (O), nitrogen (N), and fluorine (F)). Similarly, the voxelized representation of molecules in the Geometry Integral (GEOM) drug dataset may include eight channels for the eight types of atoms present in those molecules (e.g., carbon (C), hydrogen (H), oxygen (O), nitrogen (N), fluorine (F), sulfur (S), chlorine (Cl), and bromine (Br)). Therefore, the voxelized representation of each molecule in the QM9 molecular dataset may include R... 5×32×32×32 Voxel meshes, while the voxelized representation of each molecule in the Geometry Integration of Molecular Elements (GEOM) drug dataset can include R... 8×64×64×64 Voxel grid.
[0193] As described above, in some exemplary embodiments, the molecular design computational model 115 (including the denoising model 117) may be trained to approximate and subsequently sample from a noisy data distribution of a noisy voxelized representation of a molecule, or in some cases, to approximate and subsequently sample from a noisy embedding of a voxelized representation of a molecule, rather than approximate and subsequently sample from a true data distribution of a voxelized representation of a molecule that has not yet been subjected to any noise. Training the denoising model 117 to approximate a noisy data distribution of a molecule (such as a noisy data distribution of a noisy voxelized representation of a molecule exhibiting certain desired properties (e.g., drug-like properties) may include: determining a function 175 such that the function 175 outputs a value indicating the density at a corresponding position in the noisy data distribution for each voxelized representation (or its noisy embedding) of a molecule sampled from the noisy data distribution. In the case that the function 175 is a scoring function, the function 175 may output a score corresponding to a local change in the density (or gradient) of the noisy data distribution. Therefore, when function 175 is a scoring function, the score output by function 175 for the noisy voxelized representation of the molecule (or its noisy embedding) can indicate the local variation in density at the corresponding position in the noisy data distribution.
[0194] In some cases, the denoising engine 117 can be trained to denoise noisy voxelized representations of molecules, or in other cases, to denoise noisy embeddings of voxelized representations of molecules generated by a molecular design computational model 115 (e.g., denoising model 117). To further illustrate, Figure 5A depicts a schematic diagram illustrating instances of training the denoising engine 117 to denoise noisy voxelized representations of molecules according to some exemplary embodiments. As shown in Figure 5A, the training dataset used to train the denoising engine 117 can be generated as a plurality of training samples, each corresponding to a sample molecule. For example, Figure 5A shows sample molecule 500, which can be a known molecule from the PubChem dataset, the QM9 molecular dataset, the Geometry Integration of Molecular (GEOM) drug dataset, etc. Sample molecule 500 can be represented in one-dimensional (e.g., a simplified molecular linear input canonical (SMILES) string) or two-dimensional (e.g., a molecular graph) form, neither of which is sufficient to capture the conformation (or three-dimensional structure) of sample molecule 500. Therefore, in some cases, to generate training samples for inclusion in the training dataset, the one-dimensional or two-dimensional representation of sample molecule 500 can be converted into a three-dimensional representation of sample molecule 500. For example, in some cases, the one-dimensional or two-dimensional representation of sample molecule 500 can be converted into the voxelized representation x shown in Figure 5A. i Voxelized representation of sample molecule 500 xi The types and locations of atoms present in sample molecule 500 can be collectively represented as one or more continuous (e.g., Gaussian-like) densities across a voxel grid centered on individual atoms present in sample molecule 500.
[0195] Referring again to Figure 5A, in some cases, the voxelized representation of sample molecule 500 is x i It can be doped with noise ϵ (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) that can have a noise level σ in order to generate a noisy voxelized representation y i The addition of noise ϵ can voxelize the representation of x. i The data is projected from the real data distribution p(x) filled with clean (or initial) voxelized representations of molecules to the noisy data distribution p(y) filled with noisy voxelized representations of molecules. As described above, if the molecular design computational model 115 directly projects the clean (or initial) voxelized representations of molecules from the real data distribution p(x) (such as the voxelized representation x of sample molecule 500)... i If the jagged energy landscape of the real data distribution p(x) is used, the molecular design computational model 115 may prevent it from fully exploring the real data distribution p(x) when sampling from it. In contrast, the noisy data distribution p(y) can exhibit a smoother energy landscape with more asymptotic gradient changes, which means that the molecular design computational model 115 can sample from the noisy data distribution p(y) to produce greater diversity of the resulting output molecules. Therefore, in some cases, the denoising engine 117 can be trained to handle noisy voxelized representations of molecules (such as the noisy voxelized representation x of sample molecule 500). i The denoising engine 117 is applied to denoise the noisy voxelized representation of molecules generated by sampling from the noisy data distribution p(y) by the molecular design computational model 115. As described in more detail below, in some cases, the voxelized representation x of sample molecules 500... i Downsampling (or compression) can be performed before adding noise ε, which means that the denoising engine 117 can be trained to denoise the noisy embedding of the voxelized representation of the molecule, rather than the noisy voxelized representation of the molecule as shown in Figure 5A.
[0196] Referring again to Figure 5A, the denoising engine 117 can be trained to handle noisy voxelized representations y. i Denoising is performed. In some cases, the denoising engine 117 can be trained to recover at least the corresponding clean voxelized representation x from it. i To represent noisy voxelization y iDenoising is performed. For example, in some cases, the denoising engine 117 can be an encoder-decoder 3D convolutional neural network (CNN) trained to denoise the noisy voxelized representation y. i The noisy voxels in the model are mapped to their corresponding clean voxels. In doing so, the denoising engine 117 can generate clean voxel representations x. i Perform an approximate denoised voxel representation. For example, in some cases, training the denoising engine 117 may include adjusting the parameters of the denoising engine 117 to improve the denoised voxel representation. and the corresponding clean voxelization representation x i The difference between them (e.g., mean squared error (MSE)) can be reduced (or minimized). In some cases, the noise level σ (whose determination can be added to the voxelized molecular representation x) can be reduced (or minimized). i The noise level ε is set as a hyperparameter of the denoising engine 117. Furthermore, in some cases, the noise level σ can be kept fixed (or constant) during the training of the denoising engine 117, which reduces the complexity of the training process compared to the diffusion model. It should be understood that single-step denoising (as opposed to diffusion over multiple time steps) is sufficient to reconstruct the initial voxelized representation x. i This is due to the voxelization representation of x i Due to its properties, unlike natural images, this voxelization representation contains more structural information about the sample molecules than just texture information.
[0197] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to generate a voxelized representation of the output molecule by denoising the noisy voxelized representation of the input molecule through at least one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.). In some cases, the denoising model 117 may sample from a noisy data distribution p(y), which includes traversing the noisy data distribution p(y) towards regions with progressively increasing density, these regions being filled with molecules exhibiting one or more desired properties (e.g., drug-like properties). To further illustrate, Figure 5B shows that traversing across the noisy data distribution p(y) includes selecting a sample (or molecule) y at sampling iteration k-1. k-1 Choose y at sampling iteration k k And choosing y at sampling iteration k+1 k+1 In some cases, the traversal of the noisy data distribution p(y) can be guided by function 175, such that the sample y kThe ratio of the sample y to the distribution p(y) of the noisy data k-1 Higher density areas were sampled, while sample y k+1 The ratio of the sample y to the distribution p(y) of the noisy data k Even higher density regions are sampled. In some cases, each iteration of gradient-based Markov chain Monte Carlo (MCMC) may include further modification of the samples (or molecules) selected during previous iterations. Thus, as shown below, the sample y selected during previous sampling iteration k can be modified... k To generate samples y selected from the noisy data distribution p(y) during sampling iteration k+1. k+1 The following equation (10) expresses the traversal of the noisy data distribution p(y).
[0198]
[0199] (10)
[0200] Among them B t R represents d The standard Brownian motion in the model is represented by γ and u, where γ and u are hyperparameters (friction and inverse mass, respectively). Discretization techniques (such as Algorithm 1 in Table 1 below) can be applied to generate samples y. k It includes the discretization step δ.
[0201] Referring again to Figure 5B, in some cases, when the denoising engine 117 applies the corresponding noisy voxelized representation y selected from the noisy data distribution p(y)... k During denoising, a voxelized representation of the molecules can be generated. As mentioned above, for noisy voxelized representation y k Denoising can represent noisy voxels. k Projecting back to the true data distribution p(x), for example, by applying the least squares estimator. This constitutes the "jump" shown in Figure 5B. Furthermore, in the example shown in Figure 5B, a "jump" back to the true data distribution p(x) can be made at each sampling iteration, while the denoising model 117 is applied to traverse the noisy data distribution p(y) and select samples from it. For example, when a sample y is selected from the noisy data distribution p(y) during sampling iteration k+1... k+1 When denoised and projected back to the true data distribution p(x), molecules can be generated. And when the sample y is selected from the noisy data distribution p(y) during the subsequent sampling iteration k+1, kWhen denoised and projected back to the true data distribution p(x), molecules can be generated. .
[0202] Table 1
[0203]
[0204] In some exemplary embodiments, the denoising model 117 may continue to traverse the noisy data distribution p(y) and select samples from it until one or more criteria are met. For example, the denoising model 117 may continue to traverse the energy landscape of the noisy data distribution until sampling iteration k+1 (if a sampling iteration of a threshold amount is performed at that point). Alternatively and / or additionally, the denoising model 117 may continue to traverse the energy landscape of the noisy data distribution p(y) until sample y k+1 Selected (if sample y) k+1 This demonstrates the threshold probability in the noisy data distribution p(y). To further illustrate, Figure 5C shows that a denoising model 117 is applied to select multiple consecutive samples from the noisy data distribution p(y), including, for example, samples 510a to 510f. In the example shown in Figure 5C, sampling (e.g., gradient-based Markov chain Monte Carlo (MCMC) sampling) may begin with the molecular design computational model 115 applying the denoising model 117 to select a first sample y0 from the noisy data distribution p(y). In some cases, the selection of the first sample y0 may include: the denoising model 117 updating the noisy voxelized representation (or its noisy embedding) of the corresponding molecule.
[0205] As shown in Figure 5C, the first sample y0 can be denoised to generate the corresponding voxelized representation. This denoising operation can constitute a "leap" from the noisy data distribution p(y) back to the true data distribution p(x). Each subsequent sampling iteration may include: the denoising model 117 being applied to further update the noisy voxelized representation of the molecules selected during the previous sampling iteration. In the example shown in Figure 5C, the molecular design computational model 115 may continue to apply the denoising model 117 until k consecutive samples have been selected from the noisy data distribution p(y). The kth sample y k For example, the denoising engine 117 can be used to denoise the data to generate the corresponding voxelized representation. This will allow the k-th sample y to be processed. kThe process involves projecting back from the noisy data distribution p(y) to the true data distribution p(x). It should be understood that the value of k determines the number of sampling iterations and the number of samples selected from the noisy data distribution p(y). Increasing the value of k increases the updates made to the initial input molecule (e.g., the "seed" molecule). Higher values of k increase the difference between the initial input molecule (e.g., the "seed" molecule) and the final output molecule, as well as the novelty of the final output molecule.
[0206] In some exemplary embodiments, the k-th sample y is selected from the noisy data distribution p(y). k And for the k-th sample y k Denoising is performed to generate the corresponding voxelized representation. At that time, the molecular design engine 110 can be based on voxel representation. To generate one or more other representations. For example, in some cases, the molecular design engine 110 may be based at least on voxelized representations. To generate a one-dimensional representation (e.g., a simplified linear input specification (SMILES) string) and / or a two-dimensional representation (e.g., a molecular graph) of the corresponding molecule.
[0207] Figure 5D depicts a voxelized representation according to some exemplary embodiments. A schematic diagram illustrating examples of processes for generating other molecular representations. In the example shown in Figure 5D, the molecular design engine 110 can generate representations by at least identifying voxelized representations. The peak values (e.g., atomic density values reaching one or more thresholds) in the molecular design engine determine the atoms present in the corresponding molecule. Furthermore, the molecular design engine 110 can determine one or more bonds between interconnecting atoms present in the molecule. A one-dimensional or two-dimensional representation of the molecule can be generated based at least on the atoms and interconnecting bonds. Alternatively, in some cases, the molecular design engine 110 can apply a machine learning model trained to process the voxelized representation. Convert into one or more other representations of the corresponding molecule.
[0208] In some exemplary embodiments, the molecular design computational model 115 may operate in a noisy latent voxelization space, rather than in a noisy discrete voxelization space, as shown, for example, in Figures 5A to 5D. For example, in some cases, a denoising model 117 may be applied to the voxelization representation x. i Denoising is performed on the noisy embeddings, rather than on the molecular design computational model 115. A denoising model 117 is applied to the noisy voxelized representation y. iDenoising is performed. To further illustrate, Figure 6 depicts a schematic example of a process by which a molecular design computational model 115, according to some exemplary embodiments, generates a voxelized representation of a molecule by operating in a noisy latent voxelized space. Referring to Figure 6, an input molecule 600 (which may be represented in one-dimensional or two-dimensional form) can be transformed into a three-dimensional representation of an input molecule 152. In some cases, the three-dimensional representation of the input molecule 152 can be a voxelized representation of the input molecule 152 that represents the types and locations of atoms in the input molecule 152 together as one or more continuous distributions of atomic densities across a voxel grid. In some cases, instead of applying a denoising model 117 to directly operate on the noisy voxelized representation of the input molecule 152, the encoder 111 may first generate an embedding 154 of the voxelized representation of the input molecule 152 and then add noise ε to the embedding 154 of the voxelized representation of the input molecule 152. The resulting noisy embedding 156 may undergo one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.). For example, each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling may include: applying a denoising model 117 to the molecular design computational model 115 to denoise the noisy embedding 156 by at least updating the noisy embedding 156. As described above, updating the noisy embedding 156 in this way is equivalent to selecting one or more samples from a noisy data distribution filled with noisy embeddings representing voxelized representations of molecules exhibiting one or more desired properties. Sampling may be guided by a function 175 (e.g., a scoring function, etc.) such that successive samples are selected from regions of increasing density in the noisy data distribution, regions that are more likely to be filled with noisy embeddings representing voxelized representations of molecules exhibiting the one or more desired properties.
[0209] Referring again to Figure 6, the molecular design computational model 115 can generate an updated embedding 156 by updating the voxelized representation of the input molecule 152's embedding 154 through at least one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling. As shown in Figure 6, the embedding 154 can be denoised, for example, by a denoising engine 117 to generate the updated embedding 156. Denoising the embedding 154 may include sampling the updated embedding 156 from a noisy latent distribution of molecules exhibiting the one or more desired properties. Furthermore, as shown in Figure 6, the decoder 119 can decode the updated embedding 156 to generate a voxelized representation of the corresponding output molecule 162. Decoding the updated embedding 156 can project the updated embedding 156 from the latent voxelized space back into the discrete voxelized space. The resulting voxelized representation of the output molecule 162 can be further transformed into the reconstructed molecule 650. It should be understood that the reconstructed molecule 650 may correspond to a one-dimensional representation of the output molecule (e.g., a simplified linear input specification (SMILES) string) or a two-dimensional representation (e.g., a molecular graph).
[0210] In some exemplary embodiments, the generative performance of the molecular design computational model 115 can be evaluated based on a variety of metrics, some examples of which are described in Table 2 below.
[0211] Table 2
[0212]
[0213] In some exemplary embodiments, the generative performance of the molecular design computational model 115 may depend on one or more factors, including, for example, the noise level σ, the difference ∆k between the number of sampling iterations, and the radius of the atomic density in the voxelized molecular representation. Figure 7 depicts a graph illustrating the effect of the noise level σ on the stability and uniqueness (Figure 7(a)), total atomic variation and total bond variation (Figure 7(b)), and valence W1 and bond angle W1 (Figure 7(c)) of the molecules generated by the molecular design computational model 115 when noise ϵ (e.g., Gaussian noise, such as isotropic Gaussian noise) of different levels σ is added to the voxelized representation of the molecules operated by the molecular design computational model 115. As described above, unlike diffusion models, the noise level σ may be fixed during training and sampling according to the various exemplary embodiments described herein. Furthermore, it should be understood that the noise level σ is a hyperparameter that imposes a trade-off between the quality of sampling (e.g., gradient-based Markov chain Monte Carlo (MCMC) sampling) and denoising (e.g., denoising of empirical Bayesian frameworks). In some cases, the molecular design engine 110 can determine the noise level σ at which the denoising engine 117 can still learn to perform denoising, which corresponds to the maximum amount of noise ϵ added to the voxelized representation of the molecule. For example, in some cases, the molecular design computational model 115 and the denoising engine 117 can be trained on the QM9 molecular dataset with different noise levels σ = {.6,.7,…,1.2}, while other hyperparameters are kept constant. The graphs in Figures 7(a), (b), and (c) show that while some metrics improve at higher noise levels σ, molecular stability and valence W1 deteriorate with increasing noise level σ. For the QM9 molecular dataset, the best overall performance across all metrics is achieved at a noise level σ of 0.9.
[0214] In some exemplary embodiments, the number of sampling iterations ∆k performed as part of a gradient-based Markov chain Monte Carlo (MCMC) can affect the novelty of molecules generated by the molecular design computational model 115. This phenomenon is illustrated in Figure 8, which depicts the molecules output by the molecular design computational model 115 (trained on the Geometry of Molecular Integration (GEOM) drug dataset) after different number of sampling iterations k, updating noisy molecules (for de novo generation) and known molecules (for seed generation). For example, Figure 8 shows the voxelized representation of the first molecule 810 generated by the molecular design computational model 115 after k=10 sampling iterations to denoise the noisy molecule (for de novo generation), the voxelized representation of the first molecule 820 generated by the molecular design computational model 115 after k=50 sampling iterations to denoise the noisy molecule (for de novo generation), and the voxelized representation of the third molecule 830 generated by the molecular design computational model 115 after k=100 sampling iterations to denoise the noisy molecule (for de novo generation).
[0215] Besides the novelty of molecules generated by the molecular design computational model 115, adjusting the number of sampling iterations k can also affect other aspects of the generative performance of the molecular design computational model 115. Table 3 below compares the generative performance of the molecular design computational model 115 with that of the conventional generative model EDM at different sampling iteration numbers ∆k, the latter performing 1,000 diffusion steps. The results in Table 3 show that the molecular design computational model 115 performs better in some metrics as the number of sampling iterations ∆k increases. As expected, the average time (in seconds) consumed to generate each molecule increases linearly with the number of sampling iterations ∆k. However, even at 500 sampling iterations, the molecular design computational model 115 is still faster than EDM. Notably, at only 50 sampling iterations, the molecular design computational model 115 already outperforms EDM in most metrics, while being on average an order of magnitude faster.
[0216] Table 3
[0217]
[0218] In some exemplary embodiments, the generative performance of the molecular design computational model 115 can also be affected by the size of the atomic radii in the voxelized representation operated by the molecular design computational model 115. It should be understood that the size of the atomic radii can be varied, while the resolution of the voxel grid remains fixed (e.g., at .25 Å). The generative performance of the molecular design computational model 115 can peak at certain atomic radii even with different hyperparameters. For example, when the molecular design computational model 115 is applied to operate on voxelized representations with atomic radii of .25, .5, .75, and 1.0, a fixed radius of .5 consistently outperforms the other values, even when the hyperparameters of the molecular design computational model 115 are varied.
[0219] In some exemplary embodiments, the generative performance of the molecular design computation model 115 can be compared with existing generative models that operate on conventional three-dimensional molecular representations, such as GSchNet (a point cloud autoregressive model) and EDM (a point cloud diffusion-based model). Each model is applied to generate 10,000 samples, which are then evaluated based on atomic stability, molecular stability, validity, uniqueness, total atomic variation (TV), total bond variation (TV), valence W1, bond length W1, and bond angle W1. Table 4 below shows the results of samples generated by the molecular design computation model 115 (MDCM) trained on the QM9 molecular dataset, with mean and standard deviation across three runs. Figure 9A depicts voxelized representations of molecules generated by the molecular design computation model 115 trained on the QM9 molecular dataset and some examples of corresponding molecular graphs. Figure 10A, plotted in graph 1000, shows the cumulative distribution function (CDF) of strain energy for molecules generated by the molecular design computational model 115 trained on the QM9 molecular dataset, compared to molecules in the QM9 molecular dataset and those generated by the conventional generative model EDM. Figure 10B, plotted in graph 1050, compares the empirical distribution of the number of atoms per molecule in the QM9 molecular dataset with the empirical distribution of the number of atoms in molecules generated by the molecular design computational model 115 trained on the QM9 molecular dataset.
[0220] Table 4
[0221]
[0222] In some cases, the molecular design computational model 115 was also trained on the Geometry of Molecular Integration (GEOM) drug dataset and then applied to generate 10,000 samples. Table 5 below shows a comparison of those samples with 10,000 samples generated by a conventional generative model EDM, with mean and standard deviation across three independent runs. Figure 9B Voxelized representations of molecules generated by the Molecular Design Computational Model 115 (MDCM) trained on the Geometry Integration of Molecular (GEOM) Drugs dataset and some examples of corresponding molecular graphs are depicted. Figure 11A, plot 1100, shows the cumulative distribution function (CDF) of strain energy for molecules generated by the Molecular Design Computational Model 115 (MDCM) trained on the GEOM Drugs dataset, compared to molecules in the GEOM Drugs dataset and those generated by a conventional generative model EDM. Figure 11B, plot 1150, compares the empirical distribution of the number of atoms per molecule in the GEOM Drugs dataset with the empirical distribution of the number of atoms in molecules generated by the Molecular Design Computational Model 115 (MDCM) trained on the GEOM Drugs dataset.
[0223] Table 5
[0224]
[0225] When the Molecular Design Computational Model 115 (MDCM) is trained on the QM9 dataset, it exhibits generative performance comparable to the conventional generative model EDM. However, when trained on the Geometry Integration of Molecular (GEOM) drug dataset (a more challenging and realistic drug-like dataset than QM9), MDCM outperforms EDM by a considerable margin in eight out of nine metrics. For example, molecules generated by MDCM trained on the GEOM drug dataset show significantly lower median strain energies than those generated by EDM. The results in Tables 3 and 4 also show that augmenting the training dataset with rotation and translation improves the generative performance of MDCM (e.g., MDCM). 无旋转(Compared to MDCM). Overall, the molecular design computational model 115 is a more expressive model that scales better with the data. In particular, the molecular design computational model 115 is better able to capture many patterns present in large-scale data distributions, such as the Geometry of Molecular Integration (GEOM) drug dataset.
[0226] Figure 12A depicts a schematic comparison of seed generation of Geometry Integration of Molecular Drugs (GEOM) in discrete voxelization space and potential voxelization space according to some exemplary embodiments. Inset Figure 1210 shows molecular graphs of molecules generated at steps (or sampling iterations) 10, 20, 50, 100, and 200, generated by the molecular design computational model 115 operating in the potential voxelization space and updating the embeddings of voxelized representations of seed molecules from the GEOM drug dataset. Inset Figure 1220 shows the corresponding voxelized representations of these molecules. Inset Figure 1215 shows molecular graphs of molecules generated at steps (or sampling iterations) 5, 10, 50, 100, and 200, generated by the molecular design computational model 115 operating in discrete voxelization space and updating the voxelized representations of seed molecules from the GEOM drug dataset. Inset Figure 1225 shows the corresponding voxelized representations of these molecules. As shown in Figures 12A and 12B, the molecular design computation model 115 is able to generate stable, effective, and unique molecules, regardless of whether it operates in the latent voxelization space or the discrete voxelization space. These molecules are also highly similar to the seed molecules from the Geometry Integration of Molecular (GEOM) drug dataset.
[0227] Table 6 below further illustrates the seed generation results on the Geometry of Molecular Integration (GEOM) drug dataset (average 5 repetitions).
[0228] Table 6
[0229]
[0230] Figure 12B depicts a schematic comparison of seed generation of PubChem drugs in discrete voxelization space and potential voxelization space according to some exemplary embodiments. Inset 1450 shows molecular graphs of molecules generated at steps (or sampling iterations) 10, 20, 50, 100, and 200 by the molecular design computational model 115 operating in the potential voxelization space and updating the embeddings of voxelized representations of seed molecules from the PubChem dataset. Inset 1260 shows the corresponding voxelized representations of these molecules. Inset 1255 shows molecular graphs of molecules generated at steps (or sampling iterations) 5, 10, 50, 100, and 200 by the molecular design computational model 115 operating in discrete voxelization space and updating the voxelized representations of seed molecules from the PubChem dataset. Inset 1265 shows the corresponding voxelized representations of these molecules. As shown in Figure 12B, the molecular design computation model 115 is still able to generate stable, efficient and unique molecules, regardless of whether it operates in the latent voxelization space or the discrete voxelization space. These molecules are also highly similar to the seed molecules from the PubChem dataset.
[0231] Table 7 below further illustrates the seed generation results on the PubChem dataset (averaging 5 repetitions).
[0232] Table 7
[0233]
[0234] Figure 12C depicts molecular graphs of additional instances of molecules generated at steps (or sampling iterations) 10, 20, 50, 100, and 200, generated by the molecular design computational model 115 operating in the latent voxelization space and updating the embeddings of the voxelized representations of two actual drug seed molecules. Figure 12D shows molecular graphs of some exemplary molecules generated at random selections of steps (or sampling iterations), generated by the molecular design computational model operating in the latent voxelization space and updating the embeddings of random molecules (e.g., molecules with random selections of atom types and / or locations).
[0235] Table 8 below depicts the seed generation results on five real drugs (averaging 5 replicates).
[0236] Table 8
[0237]
[0238] Figure 13 depicts a graph 1300 comparing the number of stable, effective, and unique molecules generated over time by a molecular design computational model 115 operating in a latent voxelization space, a molecular design computational model 115 operating in a discrete voxelization space, and a generative model of the prior art. As shown in Figure 13, regardless of whether operating in the latent or discrete voxelization space, the molecular design computational model 115 is able to generate a much larger number of stable, effective, and unique molecules than the generative model of the prior art. Furthermore, the molecular design computational model 115 generates a larger number of stable, effective, and unique molecules when operating in the latent voxelization space than when operating in the discrete voxelization space.
[0239] Table 9 below shows the computational model 115 (MDCM) for de novo molecular design in the potential voxelization space. 潜在 ), Molecular design computational model 115 (MDCM) generated de novo in discrete voxel space 离散 A comparison of the generative performance of existing generative models EDM on geometric integration of molecules (GEOM) drugs (averaging the generation of 10,000 molecules over 3 repetitions).
[0240] Table 9
[0241]
[0242] Table 10 below shows the computational model 115 (MDCM) for de novo molecular design in the potential voxelization space. 潜在 ), Molecular design computational model 115 (MDCM) generated de novo in discrete voxel space 离散 A comparison of the generative performance of existing generative models GSchNet and EDM on QM9 drugs (averaging the generation of 10,000 molecules over 3 repetitions).
[0243] Table 10
[0244]
[0245] Figure 14 depicts a block diagram illustrating an example of a computing system 1400 according to some exemplary embodiments. Referring to Figures 1 through 14, the computing system 1400 can be used to implement a molecular design engine 110, a training engine 120, a client device 130, and / or any of its components.
[0246] As shown in Figure 14, the computing system 1400 may include a processor 1410, a memory 1420, a storage device 1430, and an input / output device 1440. The processor 1410, memory 1420, storage device 1430, and input / output device 1440 may be interconnected via a system bus 1450. The processor 1410 is capable of processing instructions for execution within the computing system 1400. Such executed instructions may implement one or more components, such as a molecular design engine 110, an analysis engine 120, a client device 130, etc. In some exemplary embodiments, the processor 1410 may be a single-threaded processor. Alternatively, the processor 1410 may be a multi-threaded processor. The processor 1410 is capable of processing instructions stored on the memory 1420 and / or storage device 1430 to display graphical information for a user interface provided via the input / output device 1440.
[0247] Memory 1420 is a computer-readable medium, such as a volatile or non-volatile computer-readable medium, that stores information within computing system 1400. For example, memory 1420 may store a data structure representing a configuration object database. Storage device 1430 provides persistent storage for computing system 1400. Storage device 1430 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, or other suitable persistent storage device. Input / output device 1440 provides input / output operations for computing system 1400. In some exemplary embodiments, input / output device 1440 includes a keyboard and / or a pointing device. In various embodiments, input / output device 1440 includes a display unit for displaying a graphical user interface.
[0248] According to some exemplary embodiments, input / output device 1440 may provide input / output operations for network devices. For example, input / output device 1440 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0249] In some exemplary embodiments, the computing system 1400 can be used to execute various interactive computer software applications that can be used to organize, analyze, and / or store data in various formats. Alternatively, the computing system 1400 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generating, managing, and editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications may include various additional functions or may be standalone computing products and / or functions. Once activated within the application, the functions can be used to generate a user interface provided via the input / output device 1440. The user interface can be generated by the computing system 1400 and presented to the user (e.g., on a computer screen monitor, etc.).
[0250] One or more aspects or features of the subject matter described herein can be implemented as digital electronic circuits, integrated circuits, specially designed ASICs, field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These aspects or features may be implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor (which may be dedicated or general-purpose, coupled to receive and send data and instructions), a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Typically, clients and servers are remotely configured to each other and generally interact via a communication network. Relationships between clients and servers arise from computer programs running on their respective computers and the client-server relationships between them.
[0251] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, including machine instructions for a programmable processor, and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. Machine-readable media may (such as, for example, non-transitory solid-state memory or magnetic hard disk drive or any equivalent storage medium) store such machine instructions non-transitory. Machine-readable media may (such as, for example, a processor cache or other random access memory associated with one or more physical processor cores) optionally or additionally store such machine instructions transiently.
[0252] To provide interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or a light-emitting diode (LED) monitor for displaying information to the user) and a keyboard and pointing device (such as, for example, a mouse or trackball, through which the user can provide input to the computer). Other kinds of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; input from the user can be received in any form, including sound, speech, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical indicators, digital image capture devices, and associated interpretation software, etc.
[0253] In the foregoing description and claims, phrases such as “at least one” or “one or more” may appear, followed by a list of combinations of elements or features. The term “and / or” may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to mean any element or feature listed alone, or any other recounted element or feature in combination with any other recounted element or feature. For example, the phrases “at least one of A and B”; “one or more of A and B”; and “A and / or B” are each intended to mean “A alone, B alone, or A and B together”. A similar interpretation applies to lists comprising three or more items. For example, the phrases “at least one of A, B, and C”; “one or more of A, B, and C”; and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together”. The use of the term "based on" in the above and claims is intended to mean "at least partially based on," so that undescribed features or elements are also permissible.
[0254] Depending on the desired construction, the subject matter described herein can be embodied in systems, apparatuses, methods, and / or articles of manufacture. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations may be provided in addition to those features and / or variations set forth herein. For example, the above embodiments may be applicable to various combinations and sub-combinations of the disclosed features and / or combinations and sub-combinations of several further features disclosed above. Furthermore, the logical flows depicted in the drawings and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method, comprising: The voxelized representation of the input molecule is encoded to generate an embedding of the input molecule that has fewer features compared to the voxelized representation of the input molecule; A molecular design computational model is applied to update the embedding of the input molecule. The molecular design computational model has been trained to approximate the data distribution of molecules exhibiting one or more desired properties by: The corrupted embedding of the voxelized representation of a sample molecule exhibiting one or more of the desired properties is taken as input, and Recover the voxelized representation of the embedding from the damaged embedding. The molecular design computational model updates the embedding of the input molecule to increase the likelihood that the resulting updated embedding is within the data distribution; as well as A voxelized representation of the output molecule is generated by decoding at least the obtained updated embedding.
2. The method of claim 1, wherein the data distribution is a noisy data distribution filled with noisy embeddings of voxelized representations of the molecules exhibiting the one or more desired properties, and wherein the voxelized representations are further generated by denoising the noisy voxelized representations of the output molecules, the noisy voxelized representations of the output molecules being generated by decoding the resulting updated embeddings.
3. The method according to any one of claims 1 to 2, further comprising: A vector quantization variational autoencoder (VQ-VAE) is applied to encode the voxelized representation of the input molecule and to decode the resulting updated embedding.
4. The method of claim 3, wherein the embedding of the molecule is a discrete latent embedding vector generated by quantizing the corresponding continuous latent embedding, and wherein the quantization includes: The corresponding consecutive latent embeddings are matched with vectors in the codebook of the embeddings by nearest neighbor search.
5. The method according to any one of claims 1 to 4, wherein the voxelized representation of the input molecule is encoded by at least the following: compressing a plurality of atomic density values comprising the voxelized representation of the input molecule such that the embedding of the input molecule comprises fewer features compared to the voxelized representation of the input molecule.
6. The method according to any one of claims 1 to 5, wherein the voxelization of the molecule represents a plurality of voxels organized into a three-dimensional voxel grid, and wherein each atom in the molecule is represented as a continuous density across one or more voxels in the three-dimensional voxel grid.
7. The method of claim 6, wherein the continuous density of each atom in the molecule is centered on the center of each atom, and wherein a first voxel located far from the center of any atom in the molecule is associated with a lower atomic density value compared to a second voxel located adjacent to the center of an atom in the molecule.
8. The method according to any one of claims 6 to 7, wherein each voxel in the three-dimensional voxel grid is associated with a value indicating the atomic density at the corresponding location.
9. The method according to any one of claims 1 to 8, wherein the voxelization representation of the molecule comprises one or more channels, and wherein each channel corresponds to a type of atom present in the molecule.
10. The method according to any one of claims 1 to 9, wherein the voxelization representation of the molecule collectively represents the type and location of one or more atoms present in the molecule.
11. The method according to any one of claims 1 to 10, wherein the embedding of the input molecule is updated based at least on a function parameterized by a plurality of parameters of the molecular design computational model, and wherein the function outputs a value indicating the probability that the resulting updated embedding is within the data distribution.
12. The method of claim 11, wherein the function is a scoring function, and wherein the value output by the function is a score, the score indicating a local variation in the density of the noisy data distribution at the location of each updated noisy embedding generated by updating the noisy embedding of the input molecule.
13. The method according to any one of claims 1 to 12, wherein the molecular design computational model updates the embedding of the input molecule by at least the following: The molecular design computational model is applied to update the embedding of the input molecule, thereby generating a first updated embedding. The molecular design computational model is applied to update the embedding of the input molecule, thereby generating a second updated embedding. A function parameterized by multiple parameters of the molecular design computational model is applied to determine: (i) a first value indicating a first local change in the density of the data distribution at a first location occupied by the first updated embedding, and (ii) a second value indicating a second local change in the density of the data distribution at a second location occupied by the second updated embedding. The molecular design computational model is applied to further update the first updated embedding rather than the second updated embedding, based at least on the first and second values.
14. The method of claim 13, wherein the molecular design computational model is applied to further update the first updated embedding until one or more criteria are met, and wherein the one or more criteria include at least one of the following: (i) the embedding of the input molecule has been updated iteratively by a threshold amount, (ii) the first value of the first updated embedding reaches one or more thresholds, and (iii) an output molecule of the threshold amount has been generated.
15. The method of any one of claims 13 to 14, wherein the molecular design computational model is applied to further modify the first updated embedding rather than the second updated embedding based at least on the following: the first value and the second value indicate that the first updated embedding is more likely to be within the data distribution than the second updated embedding.
16. The method of any one of claims 13 to 15, wherein the molecular design computational model is applied to further modify the first updated embedding rather than the second updated embedding based at least on the following: the first value and the second value indicate that the first updated embedding was sampled from a higher density region of the data distribution compared to the second updated embedding.
17. The method according to any one of claims 1 to 16, further comprising: The voxelized representation of the output molecule is converted into a one-dimensional representation of the output molecule and / or a two-dimensional representation of the output molecule.
18. The method of claim 17, wherein the voxelized representation of the output molecule is converted by at least the following: The localization of one or more atoms in the output molecule is determined by detecting at least one or more peaks among a plurality of atomic density values including the voxelized representation of the output molecule, and One or more interconnecting bonds are determined based at least on the positioning of the one or more atoms.
19. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 1 to 18.
20. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 1 to 18.
21. A computer-implemented method, comprising: Generate a training dataset comprising multiple training samples, each training sample in the training dataset comprising a corrupted embedding generated by adding at least noise to a noisy voxelized representation of a sample molecule exhibiting one or more desired properties; The molecular design computation model is trained, at least based on the training dataset, to approximate the data distribution of molecules exhibiting the one or more desired properties. The training includes: applying the molecular design computation model to recover the undamaged embedding of the noisy voxelized representation of the sample molecule from the damaged embedding of the noisy voxelized representation of the sample molecule. as well as Optionally, the molecular design computational model is applied to generate an output molecule by at least the following: denoising the embedding of the voxelized representation of the input molecule and decoding the updated embedding generated therefrom to generate the voxelized representation of the output molecule.
22. The method of claim 21, wherein the noisy voxelization representation of the sample molecule comprises a plurality of voxels organized into a three-dimensional voxel grid, and wherein each atom in the sample molecule is represented as a continuous density across one or more voxels in the three-dimensional voxel grid.
23. The method of claim 22, wherein the continuous density of each atom in the sample molecule is centered on the center of each atom, and wherein a first voxel located far from the center of any atom in the sample molecule is associated with a lower atomic density value compared to a second voxel located adjacent to the center of the atom in the sample molecule.
24. The method according to any one of claims 22 to 23, wherein each voxel in the three-dimensional voxel grid is associated with a value indicating the atomic density at the corresponding location.
25. The method according to any one of claims 21 to 24, wherein the noisy voxelization representation of the sample molecule comprises one or more channels, and wherein each channel corresponds to a type of atom present in the sample molecule.
26. The method according to any one of claims 21 to 25, wherein the noisy voxelization representation of the sample molecule collectively represents the type and location of one or more atoms present in the sample molecule.
27. The method according to any one of claims 21 to 26, wherein the training of the molecular design computational model comprises: Several parameters of the molecular design computation model are adjusted to reduce the difference between the recovered embedding generated by the molecular design computation model and the undamaged embedding represented by the noisy voxelization of the sample molecule.
28. The method of claim 27, wherein the plurality of parameters of the molecular design computational model are parameterized to a function, and wherein the plurality of parameters are adjusted such that the function outputs values indicating local variations in the density of the data distribution of molecules exhibiting the one or more desired properties.
29. The method of claim 1, wherein the molecular design computational model denoises the embedding of the voxelized representation of the input molecule by at least the following: updating at least one value present in the embedding representing a plurality of atomic density values present in the voxelized representation of the input molecule.
30. The method of claim 29, wherein updating the atomic density value of one or more voxels in the at least one channel of the voxelization representation of the input molecule corresponds to updating at least one of the types and / or locations of one or more atoms present in the input molecule.
31. The method according to any one of claims 21 to 30, wherein the molecular design computational model denoises the embedding of the voxelized representation of the input molecule through multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until one or more criteria are met.
32. The method of claim 31, wherein the one or more criteria comprise at least one of the following: (i) an iteration of a threshold amount of gradient-based Markov chain Monte Carlo (MCMC) sampling has been performed, (ii) an updated embedding obtained from sampling a region having a threshold density, and (iii) an output molecule of the threshold amount has been generated.
33. The method according to any one of claims 21 to 32, wherein the molecular design computational model generates the voxelized representation of the output molecule by at least the following: The first update is applied to the embedding of the voxelized representation of the input molecule to generate the first updated embedding. The second update is applied to the embedding of the voxelized representation of the input molecule to generate the second updated embedding, and When it is determined that the first updated embedding was sampled from a higher density region of the data distribution compared to the second updated embedding, the first updated embedding is further updated instead of the second updated embedding.
34. The method according to any one of claims 21 to 33, further comprising: The voxelized representation of the output molecule is converted into a one-dimensional and / or two-dimensional representation of the output molecule.
35. The method according to any one of claims 21 to 34, wherein the training of the molecular design computational model comprises: The molecular design computational model with a first adjustment is applied to generate a first recovered embedding of the noisy voxelized representation of the sample molecule. A first mean squared error (MSE) is determined to quantify the first difference between the first recovered embedding and the undamaged embedding in the noisy voxelized representation of the sample molecule. The molecular design computational model with a second adjustment is applied to generate a second recovered embedding of the noisy voxelized representation of the sample molecule. A second mean squared error (MSE) is determined to quantify the second difference between the noisy voxelized representation of the sample molecule and the undamaged embedding. When it is determined that the first mean square error (MSE) is less than the second mean square error (MSE), the molecular design calculation model with the first adjustment is further adjusted instead of the molecular design calculation model with the second adjustment.
36. The method of claim 35, wherein the molecular design computational model is further adjusted until one or more criteria are met, and wherein the one or more criteria include at least one of the following: (i) an iteration of a threshold amount of adjustment to the molecular design computational model has been performed, and (ii) a recovered embedding exhibiting a threshold mean square error (MSE) value has been generated.
37. The method according to any one of claims 21 to 36, further comprising: The training includes an autoencoder comprising an encoder and a decoder, the training of the autoencoder comprising: training the encoder to encode the noisy voxelized representation of the sample molecule such that the decoder is able to recover the voxelized representation of the sample molecule from the resulting embedding of the noisy voxelized representation of the sample molecule.
38. The method of claim 37, wherein the autoencoder is a vector quantization variational autoencoder (VQ-VAE), wherein the encoder generates a continuous latent embedding of the sample molecule and then quantizes the continuous latent embedding of the sample molecule into a discrete latent embedding by means of nearest neighbor lookup and vector matching in the codebook of the embeddings, for decoding by the decoder.
39. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 21 to 38.
40. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 21 to 38.