Three-dimensional molecule generation

The use of continuous molecular occupancy field representations and neural fields addresses the limitations of conventional methods by efficiently generating three-dimensional molecules with desired properties, enhancing expressivity and scalability for larger molecules.

WO2025244930A9PCT designated stage Publication Date: 2026-02-05GENENTECH INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029612
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2025-05-15
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional methods for generating three-dimensional molecules are limited by large molecular spaces that are computationally prohibitive and lack principled exploration, failing to capture desired properties due to inadequate three-dimensional representations, and are inefficient for larger molecules.

Method used

Utilizing continuous molecular occupancy field representations and neural fields to encode and decode molecular structures, enabling structure-conditioned generation of three-dimensional molecules through low-dimensional latent codes, allowing efficient and expressive modeling of molecular structures.

Benefits of technology

Enables efficient generation of molecules with desired properties by capturing complex three-dimensional structures, improving generalization and reducing computational resources required, particularly for larger molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029612_05022026_PF_FP_ABST
    Figure US2025029612_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A method for generating an output molecule may include determining a three-dimensional target structure representation of a target structure and a noisy latent code. A molecule design computation model may be applied to generate a modified latent code encoding an output molecule molecular occupancy field (or atomic density field) of the output molecule. The molecule design computation model may generate the modified latent code by modifying the noisy latent code such that a three-dimensional structure of the output molecule is consistent with a three-dimensional structure of the target structure. The modified latent code may be decoded to determine the output molecule molecular occupancy field (or atomic density field) of the output molecule. The decoding may include determining, based on the modified latent code, an atomic density exhibited by the output molecule at one or more points in three-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1THREE-DIMENSIONAL MOLECULE GENERATIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No.63 / 650,621, entitled “SCORE-BASED THREE-DIMENSIONAL MOLECULE GENERATION WITH NEURAL FIELDS” and fded on May 22, 2024, and U.S. Provisional Application No. 63 / 776,879, entitled “STRUCTURE-CONDITIONED THREE-DIMENSIONAL MOLECULE GENERATION” and fded on March 24, 2025, the disclosures of which are incorporated herein by reference in their entireties.TECHNICAL FIELD

[0002] The subject matter described herein relates generally to generative artificial intelligence and more specifically to machine learning enabled techniques for structure- conditioned generation of three-dimensional molecules using continuous molecular occupancy field representations of the three-dimensional molecular structures.INTRODUCTION

[0003] A molecule is a group of two more atoms held together by chemical bonds. Molecules form the smallest identifiable unit into which a pure substance can be divided while still retaining the composition and chemical properties of that substance. One example of a molecule is a small molecule, which is a low-weight compound having a molecular weight between approximately 100 Daltons and 1000 Daltons. Small molecule therapeutics, which modulate biochemical processes to diagnose, treat, and prevent a gamut of illnesses, have been a cornerstone in modern pharmacology due to a number of compelling advantages. For example, small molecule drugs are capable of penetrating cell membranes to reach intracellular targets.Moreover, small molecule drugs are adaptable to a wide variety of therapeutic applications. ForAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 instance, a small molecule drug may be formulated as pills and capsules, intravenous or subcutaneous injectables, inhalational medicines, or suppositories. The development of the small molecule drug may further extend to tailoring various pharmacokinetic properties including liberation, absorption, distribution, metabolism, potency, efficacy, phenotypic effects, and excretion.

[0004] By contrast, large molecules (also known as biopharmaceuticals, biologicals, or biologies) can range between approximately 3000 Daltons and 150,000 Daltons in molecular weight. Some modalities between small and large molecules, such as macro cyclic peptides, also exist. Large molecule drugs are often derivatives of natural human proteins, which modulate many essential cellular functions such as enzymatic reactions, transport of molecules, regulation and execution of a number of biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. It is common for a single large molecule to have more than 1,300 amino acid residues, which are linked by peptide bonds to form one or more polypeptide. Some large molecules may be formed from canonical amino acid residues, which are the 20 standard amino acids naturally incorporated into proteins during translation, but others may include non-canonical amino acid residues (or those that are not part of the standard genetic code) as well. Due to their size and complexity, large molecule drugs are recombinantly produced by engineered cells instead of being chemically synthesized like the majority of small molecule drugs. Moreover, large molecule therapeutics are usually delivered through injection or infusion due to the ineffectiveness of oral administration. The development of a large molecule drug may entail designing one or more sequences of amino acid residues capable of binding to a target (e.g., a protein, a nucleic acid, and / or the like) with sufficientAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 specificity and absent undesired traits such as immunogenicity, self-association, instability, and / or the like.SUMMARY

[0005] Systems, methods, and articles of manufacture, including computer program products, are provided for generating three-dimensional molecules using continuous molecular occupancy field representations of the three-dimensional molecule. In some cases, the three- dimensional structure of a molecule (e.g., a small molecule, a peptide, a protein molecule, and / or the like) may be represented as a molecular occupancy field (or atomic density field) encoding atomic occupancy. In some cases, the molecular occupancy field (or atomic density field) may be a continuous function mapping three-dimensional coordinates to atomic densities. A neural field may be a neural network that is used to approximate (or parameterize) the molecular occupancy fields representative of the three-dimensional structure (or conformation) of various molecules. For example, in some cases, an encoder (e.g., a three-dimensional convolutional neural network (CNN)) may be trained to encode the voxelized molecular occupancy field (or atomic density field) of a molecule into a low dimensional latent code (or modulation code). The latent code (or modulation code) may be decoded by the neural field, which may recover the molecular occupancy field (or atomic density field) of the molecule by at least mapping points in three-dimensional space to the atomic densities of the molecule at the corresponding locations. In some cases, three-dimensional molecules may be generated by operating on the low dimensional latent code, which is tantamount to sampling from a low-dimensional latent space populated by low dimensional latent codes. Low dimensional latent codes sampled from the latent space may be decoded by the neural field to recover the corresponding molecular occupancy fields, thus generating one or more output molecules. In some cases, the generatingAttomey Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 of the three-dimensional molecules may be structure-conditioned. For instance, the generating of the three-dimensional molecules may be conditioned on the structure of at least a portion of another molecule, such as that of an epitope, a binding pocket, and / or the like. Alternatively and / or additionally, the generating of the three-dimensional molecules may be conditioned on a substructure, such as a scaffold (of a chemical compound), a backbone (or a protein molecule), and / or the like.

[0006] In one aspect, there is provided a system for using continuous molecular occupancy field representations to generate three-dimensional molecules. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that cause operations when executed by the at least one data processor. The operations may include: receiving an input molecule, the input molecule being associated with a molecular occupancy field representative of a three-dimensional structure of the input molecule, where the molecular occupancy field of the input molecule maps a point in three-dimensional space to an atomic density of the input molecule at a corresponding location in three-dimensional space; generating a latent code of the input molecule by at least applying a computation model to encode, into the latent code of the input molecule, the molecular occupancy field of the input molecule; and generating an output molecule by at least modifying the latent code of the input molecule.

[0007] In another aspect, there is provided a computer-implemented method for using continuous molecular occupancy field representations to generate three-dimensional molecules. The method may include: receiving an input molecule, the input molecule being associated with a molecular occupancy field representative of a three-dimensional structure of the input molecule, where the molecular occupancy field of the input molecule maps a point in three-Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 dimensional space to an atomic density of the input molecule at a corresponding location in three-dimensional space; generating a latent code of the input molecule by at least applying a computation model to encode, into the latent code of the input molecule, the molecular occupancy field of the input molecule; and generating an output molecule by at least modifying the latent code of the input molecule.

[0008] In another aspect, there is provided a computer program product for using continuous molecular occupancy field representations to generate three-dimensional molecules. The computer program product may include a non-transitory computer readable medium storing instructions. The instructions may cause operations when executed by the at least one data processor. The operations may include: receiving an input molecule, the input molecule being associated with a molecular occupancy field representative of a three-dimensional structure of the input molecule, where the molecular occupancy field of the input molecule maps a point in three-dimensional space to an atomic density of the input molecule at a corresponding location in three-dimensional space; generating a latent code of the input molecule by at least applying a computation model to encode, into the latent code of the input molecule, the molecular occupancy field of the input molecule; and generating an output molecule by at least modifying the latent code of the input molecule.

[0009] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0010] In some variations, he atomic density of the input molecule at the corresponding location in three-dimensional space comprises a distance between that location and a center of an atom forming the input molecule.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0011] In some variations, the molecular occupancy field comprises a continuous function mapping, based at least on the latent code of the input molecule, the point in three- dimensional space to the atomic density of the input molecule at the corresponding location in three-dimensional space.

[0012] In some variations, the computation model includes an encoder that generates the latent code of the input molecule by at least encoding the molecular occupancy field of the input molecule.

[0013] In some variations, the computation model further comprises a decoder that decodes the latent code of the input molecule to recover the molecular occupancy field of the input molecule. The decoder recovers the molecular occupancy field of the input molecule by at least determining, based at least on the latent code of the input molecule and a coordinate of the point in three-dimensional space, the atomic density of the input molecule at the corresponding location in three-dimensional space.

[0014] In some variations, a noisy latent code is generated by at least adding noise to the latent code of the input molecule. A molecule design computation model is applied to modify the noisy latent code of the input molecule.

[0015] In some variations, a molecule design computation model is applied to modify the noisy latent code over one or more successive iterations.

[0016] In some variations, the molecule design computation model is trained to approximate a noisy data distribution of molecules exhibiting one or more desired properties, wherein the molecule design computation model is trained based at least on a plurality of noisy latent codes of sample molecules exhibiting one or more desired properties. The modifying ofAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 the noisy latent code comprises sampling one or more noisy latent codes from the noisy data distribution.

[0017] In some variations, the molecule design computation model is applied to modify the noisy latent code and sample one or more noisy latent codes from the noisy data distribution until one or more criteria are met.

[0018] In some variations, a molecular occupancy field of the output molecule is determined by at least applying the computation model to decode a modified latent code resulting from the modifying of the latent code of the input molecule.

[0019] In some variations, a three-dimensional structure of the output molecule is determined based at least on the molecular occupancy field of the output molecule.

[0020] In some variations, the three-dimensional structure of the output molecule is determined by at least rendering, based at least on the molecular occupancy field of the output molecule, a voxel grid comprising a plurality of voxels, wherein each voxel of the plurality of voxels is associated with an atomic density value at a corresponding location in three- dimensional space; identifying one or more peak atomic density values in a plurality of atomic density values across the voxel grid; and determining, based at least on one or more peak atomic density values, a quantity and a location of one or more atoms present in the output molecule.

[0021] In some variations, the rendering the voxel grid includes applying the computation model to determine, for each voxel in the voxel grid, the atomic density value at the corresponding location in three-dimensional space.

[0022] In some variations, the computation model comprises a neural field or a conditional neural field.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0023] In some variations, the computation model comprises an autoencoder. The autoencoder includes an encoder that encodes the molecular occupancy field of the input molecule. The autoencoder further includes a decoder that recovers a molecular occupancy field of the output molecule by at least decoding a modified latent code resulting from the modifying of the latent code of the input molecule.

[0024] In some variations, the encoder comprises a three-dimensional convolutional neural network and the decoder comprises a multiplicative filter network.

[0025] In some variations, the computation model encodes the molecular occupancy field of the input molecule such that the latent code of the input molecule is unique to the input molecule and captures one or more structural features unique to the input molecule.

[0026] In some variations, the computation model is trained to model one or more common molecular features including one or more of bonds, angles, valencies, and symmetries.

[0027] In some variations, the molecular occupancy field of the input molecule comprises a voxel grid having a plurality of voxels. Each voxel of the plurality of voxels is associated with an atomic density value of the input molecule at a corresponding location in three-dimensional space.

[0028] In some variations, the molecular occupancy field of the input molecule includes a plurality of channels. Each channel of the plurality of channels corresponds to a different type of atom present in the input molecule.

[0029] In another aspect, there is provided a system for structure-conditioned generation of three- dimensional molecules. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that cause operations when executed by the at least one data processor. The operations may include: determining aAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 three-dimensional target structure representation of a target structure; generating a noisy latent code; applying a molecule design computation model to generate a modified latent code, wherein the modified latent code encodes an output molecule molecular occupancy field of the output molecule, wherein the molecule design computation model generates the modified latent code by at least modifying, based at least on the three-dimensional target structure representation, the noisy latent code, and wherein the molecule design computation model modifies the noisy latent code such that a three-dimensional structure of the output molecule encoded by the modified noisy latent code is consistent with a three-dimensional structure of the target structure; and decoding the modified latent code to determine the output molecule molecular occupancy field of the output molecule, wherein the decoding includes determining, based at least on the modified latent code, an atomic density exhibited by the output molecule at one or more points in three-dimensional space.

[0030] In another aspect, there is provided a computer-implemented method for structure-conditioned generation of three- dimensional molecules. The method may include: determining a three-dimensional target structure representation of a target structure; generating a noisy latent code; applying a molecule design computation model to generate a modified latent code, wherein the modified latent code encodes an output molecule molecular occupancy field of the output molecule, wherein the molecule design computation model generates the modified latent code by at least modifying, based at least on the three-dimensional target structure representation, the noisy latent code, and wherein the molecule design computation model modifies the noisy latent code such that a three-dimensional structure of the output molecule encoded by the modified noisy latent code is consistent with a three-dimensional structure of the target structure; and decoding the modified latent code to determine the output moleculeAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 molecular occupancy field of the output molecule, wherein the decoding includes determining, based at least on the modified latent code, an atomic density exhibited by the output molecule at one or more points in three-dimensional space.

[0031] In another aspect, there is provided a computer program product for structure- conditioned generation of three- dimensional molecules. The computer program product may include a non-transitory computer readable medium storing instructions. The instructions may cause operations when executed by the at least one data processor. The operations may include: determining a three-dimensional target structure representation of a target structure; generating a noisy latent code; applying a molecule design computation model to generate a modified latent code, wherein the modified latent code encodes an output molecule molecular occupancy field of the output molecule, wherein the molecule design computation model generates the modified latent code by at least modifying, based at least on the three-dimensional target structure representation, the noisy latent code, and wherein the molecule design computation model modifies the noisy latent code such that a three-dimensional structure of the output molecule encoded by the modified noisy latent code is consistent with a three-dimensional structure of the target structure; and decoding the modified latent code to determine the output molecule molecular occupancy field of the output molecule, wherein the decoding includes determining, based at least on the modified latent code, an atomic density exhibited by the output molecule at one or more points in three-dimensional space.

[0032] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0033] In some variations, the target structure includes at least a portion of another molecule. The molecule design computation model modifies the noisy latent code such that aAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 three-dimensional structure of the output molecule is complementary to a three-dimensional structure of the target structure.

[0034] In some variations, the target structure comprises an epitope or a binding pocket.

[0035] In some variations, the target structure includes a substructure. The molecule design computation model modifies the noisy latent code while conditioned on the target structure such that a three-dimensional structure of the output molecule encoded by the modified noisy code conforms to a three-dimensional structure of the target structure.

[0036] In some variations, the substructure comprises a scaffold of a chemical compound or a backbone of a protein molecule.

[0037] In some variations, a decoder is applied to decode the modified latent code. The decoder comprises a neural field trained to approximate the output molecule molecular occupancy field by at least mapping, based at least on the modified latent code, the one or more points in three-dimensional space to an atomic density of the input molecule at a corresponding location in three-dimensional space.

[0038] In some variations, the neural field comprises a neural network approximating a continuous function that maps the one or more points to the atomic density of the output molecule at one or more corresponding locations in three-dimensional space.

[0039] In some variations, the neural field maps the one or more points to the atomic density of the output molecule at any arbitrary resolution.

[0040] In some variations, a location near a center of an atom in the output molecule has a higher atomic density than a location far away from all atoms in the output molecule.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0041] In some variations, the three-dimensional target structure representation is encoded to generate a target structure embedding. The noisy latent code is encoded to generate a latent code embedding. A joint embedding is generated by at least combining the target structure embedding and the latent code embedding. The molecule design computation model is applied to generate the modified latent code by at least operating on the joint embedding.

[0042] In some variations, the target structure embedding and the latent code embedding are encoded to occupy a common embedding space and have a same spatial dimensions.

[0043] In some variations, the noisy latent code is generated to include random noise.

[0044] In some variations, the molecule design computation model samples noisy latent codes from a noisy data distribution by modifying the noisy latent code.

[0045] In some variations, the molecule design computation model samples noisy latent codes from the noisy data distribution while guided by a function that outputs a value indicative of a local change in a density of the noisy data distribution at a location of each noisy latent code sampled by the molecule design computation model.

[0046] In some variations, the molecule design computation model is applied to sample, based on the value output by the function, noisy latent codes from incrementally higher density regions of the noisy data distribution.

[0047] In some variations, higher density regions of the noisy data distribution are populated with noisy latent codes of molecules more likely to exhibit one or more desired properties while lower density regions of the noisy data distribution are populated by noisy latent codes of molecules less likely to exhibit the one or more desired properties.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0048] In some variations, the molecule design computation model is applied to modify the noisy latent code to generate a first noisy latent code and a second noisy latent code. The function is applied to determine a first value indicative of a first local change in the density of the noisy data distribution at a first location occupied by the first noisy latent code. The function is applied to determine a second value indicative of a second local change in the density of the noisy data distribution at a second location occupied by the second noisy latent code. The molecule design computation model is applied to further modify, when the first value and the second value indicates the first noisy latent code is sampled from a higher density region of the noisy data distribution than the second noisy latent code, the first noisy latent code and not the second noisy latent code.

[0049] In some variations, the modified latent code is denoised prior to decoding the modified latent code.

[0050] In some variations, a three-dimensional structure of the output molecule is determined based at least on the output molecule molecular occupancy field of the output molecule. The determining the three-dimensional structure of the output molecule includes determining a type and a location of each atom in the output molecule.

[0051] In some variations, the determining the three-dimensional structure of the output molecule includes determining, based at least on the output molecule molecular occupancy field, a voxel grid in which each constituent voxel is associated with a value corresponding to an atomic density at a corresponding location; identifying one or more peak atomic density values across the voxel grid; and determining, based at least on the one or more peak atomic density values, one or more coordinates of each atom in the output molecule.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0052] In some variations, the determining the three-dimensional structure of the output molecule further includes refining the one or more coordinates of each atom in the output molecule around a neighborhood of each atom.

[0053] In some variations, the determining the three-dimensional structure of the output molecule further includes inferring one or more of a bond between atoms in the output molecule and an identity of an amino acid residue comprising the output molecule.

[0054] In some variations, the output molecular occupancy field includes a plurality of channels, wherein each channel of the plurality of channel corresponds to a different atom type, and wherein the one or more peak atomic density values are identified for each channel.

[0055] In some variations, a training dataset is generated to include, for each training sample in the training dataset, a sample target structure and a corrupted latent code encoding a corrupted molecular occupancy field of a sample molecule whose three-dimensional structure is consistent with a three-dimensional structure of the sample target molecule. The molecule design computation model is trained to recover an uncorrupted latent code encoding a molecular occupancy field of the sample molecule included with each training sample.

[0056] In some variations, the corrupted latent code is generated by the addition of a first quantity of noise to corrupt a molecular occupancy field of the sample molecule and a second quantity of noise projecting the corrupted latent code to a noisy data distribution. The molecule design computation model is trained to approximate the noisy data distribution by at least removing the first quantity of noise but not the second quantity of noise.

[0057] In some variations, the first quantity of noise and the second quantity of noise comprise Gaussian noise.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0058] In some variations, an encoder is applied to generate the corrupted latent code by at least encoding the corrupted molecular occupancy field of the sample molecule.

[0059] In some variations, the encoder is trained to generate the corrupted latent code by at least adjusting one or more parameters of the encoder such that the encoder generates, for each training molecular occupancy field, a latent code matching a ground truth latent code of the training molecular occupancy field.

[0060] In some variations, the encoder is trained to generate the corrupted latent code along with a decoder trained to decode the modified latent code. The training of the encoder and the decoder includes adjusting one or more parameters of the encoder and the decoder such that the encoder generates, for each training molecular occupancy field, a latent code from which the decoder is able to recover a molecular occupancy field matching the training molecular occupancy field.

[0061] In some variations, the three-dimensional target structure representation of the target molecule comprises a voxel grid representation of the target structure in which one or more atoms forming the target structure are represented by discrete atomic density values across a voxel grid.

[0062] In another aspect, there is provided a system for using continuous molecular occupancy field representations to generate three-dimensional molecules. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that cause operations when executed by the at least one data processor. The operations may include: receiving a voxel grid representative of a three-dimensional structure of a molecule, wherein the voxel grid includes a plurality of voxels, and wherein each voxel of the plurality of voxels is associated with an atomic density value of the molecule at a correspondingAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 location in three-dimensional space; determining a molecular occupancy field representative of the three-dimensional structure of the molecule, wherein the molecular occupancy field of the molecule comprises a continuous function mapping each point in three-dimensional space to a corresponding atomic density value of the molecule at a location of the point in three- dimensional space; and generating a latent code of the molecule by at least encoding the molecular occupancy field of the molecule.

[0063] In another aspect, there is provided a computer-implemented method for using continuous molecular occupancy field representations to generate three-dimensional molecules. The method may include: receiving a voxel grid representative of a three-dimensional structure of a molecule, wherein the voxel grid includes a plurality of voxels, and wherein each voxel of the plurality of voxels is associated with an atomic density value of the molecule at a corresponding location in three-dimensional space; determining a molecular occupancy field representative of the three-dimensional structure of the molecule, wherein the molecular occupancy field of the molecule comprises a continuous function mapping each point in three- dimensional space to a corresponding atomic density value of the molecule at a location of the point in three-dimensional space; and generating a latent code of the molecule by at least encoding the molecular occupancy field of the molecule.

[0064] In another aspect, there is provided a computer program product for using continuous molecular occupancy field representations to generate three-dimensional molecules. The computer program product may include a non-transitory computer readable medium storing instructions. The instructions may cause operations when executed by the at least one data processor. The operations may include: receiving a voxel grid representative of a three- dimensional structure of a molecule, wherein the voxel grid includes a plurality of voxels, andAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 wherein each voxel of the plurality of voxels is associated with an atomic density value of the molecule at a corresponding location in three-dimensional space; determining a molecular occupancy field representative of the three-dimensional structure of the molecule, wherein the molecular occupancy field of the molecule comprises a continuous function mapping each point in three-dimensional space to a corresponding atomic density value of the molecule at a location of the point in three-dimensional space; and generating a latent code of the molecule by at least encoding the molecular occupancy field of the molecule.

[0065] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0066] In some variations, a computation model is used to encode the molecular occupancy field of the molecule to generate the latent code of the molecule.

[0067] In some variations, the computation model is used to decode the latent code of the molecule to recover the molecular occupancy field of the molecule. The computation model decodes the latent code of the molecule by at least mapping, based at least on the latent code of the molecule, one or more point in three-dimensional space to atomic densities of the molecule at corresponding locations in three-dimensional space.

[0068] In some variations, the three-dimensional structure of the molecule is determined based at least on the molecular occupancy field of the molecule.

[0069] In some variations, the three-dimensional structure of the molecule is determined by at least rendering, based at least on the molecular occupancy field of the molecule, the voxel grid, wherein the rendering of the voxel grid includes applying the computation model to determine, for each voxel in the voxel grid, the atomic density value at the corresponding location in three-dimensional space; identifying one or more peak atomic densityAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 values in a plurality of atomic density values across the voxel grid; and determining, based at least on one or more peak atomic density values, a quantity and a location of one or more atoms present in the molecule.

[0070] In some variations, the computation model is trained to generate the latent code of the molecule to capture one or more structural features unique to the molecule.

[0071] In some variations, the computation model is trained to is trained to model one or more common molecular features including one or more of bonds, angles, valencies, and symmetries.

[0072] In some variations, the computation model comprises a neural field or a conditional neural field.

[0073] In some variations, the molecular occupancy field of the molecule includes a plurality of channels, and wherein each channel of the plurality of channels corresponds to a different type of atom present in the input molecule.

[0074] In some variations, one or more additional molecules are generated by at least modifying the latent code of the molecule. The generating of the one or more additional molecules includes determining, based at least on a modified latent code resulting from the modifying of the latent code of the molecule, a molecular occupancy field of each additonal molecule, and determining, based at least on the molecular occupancy field of each additional molecule, a three-dimensional structure of each additional molecule.

[0075] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non- transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.

[0076] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to the computational design of drug molecules, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS

[0077] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, togetherAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,

[0078] FIG. 1 depicts a system diagram illustrating an example of a molecule design system, in accordance with some example embodiments;

[0079] FIG. 2 depicts a flowchart illustrating an example of a process for three- dimensional molecule generation using a neural molecular field representation of molecules, in accordance with some example embodiments;

[0080] FIG. 3A depicts a schematic diagram illustrating an example of a neural-field based autoencoder, in accordance with some example embodiments;

[0081] FIG. 3B depicts a schematic diagram illustrating an example of a process for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments;

[0082] FIG. 3C depicts a schematic diagram illustrating another example of a process for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments;

[0083] FIG. 3D depicts a schematic diagram illustrating another example of a process for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments;

[0084] FIG. 4A depicts a flowchart illustrating an example of a process for structure- conditioned generation of three-dimensional molecules, in accordance with some example embodiments, in accordance with some example embodiments;Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0085] FIG. 4B depicts a flowchart illustrating an example of a process for training a molecule design computation model to perform structure-conditioned generation of three- dimensional molecules, in accordance with some example embodiments;

[0086] FIG. 5A depicts a schematic diagram illustrating an example of a neural field decoding a latent code, in accordance with some example embodiments;

[0087] FIG. 5B depicts a block diagram illustrating an example of a conditional multiplicative filter network (MFN) for implementing a neural field, in accordance with some example embodiments;

[0088] FIG. 5C depicts a block diagram illustrating an example of a multiplicative block in a conditional multiplicative filter network (MFN), in accordance with some example embodiments;

[0089] FIG. 6 depicts a schematic diagram illustrating an example of a process for sampling latent (or modulation) codes populated from a noisy data distribution populated by noisy latent codes, in accordance with some example embodiments;

[0090]

[0091] FIG. 7A depicts a schematic diagram illustrating an example of an output molecule whose generation is conditioned on a target structure that includes at least a portion of another molecule, in accordance with some example embodiments;

[0092] FIG. 7B depicts a schematic diagram illustrating another example of an output molecule whose generation is conditioned on a target structure that includes a substructure, in accordance with some example embodiments;

[0093] FIG. 8 depicts the results of sampling in a protein pocket relaxed around a macro-cyclic peptide seed, in accordance with some example embodiments; andAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0094] FIG. 9 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.

[0095] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION

[0096] Generating new molecules with desired properties is a task of critical importance across many industries. In the context of drug discovery, conventional computational techniques for generating molecules with drug-like properties require conducting a search of the molecular space (or chemical space) occupied by every possible chemical compound (e.g., every possible combination of atoms of two or more chemical elements). For example, some search-based approaches may include scoring and ranking different molecules in the molecular space based on one or more drug-like properties, such as affinity, specificity, biological activity, and developability. However, the aforementioned molecular space, which is estimated to contain 1060possible chemical compounds, is prohibitively large and scales exponentially with molecule size (e.g., the number of constituent atoms). Even a very small portion of the molecular space can contain on the order of billions and trillions of molecules. With state-of-the-art computational resources, conventional search-based approaches are capable of exploring only a small fraction of the molecular space, such as small regions of the molecular space selected based on prior domain knowledge. This limitation in search scope means that conventional search-based approaches are likely to overlook molecules with more optimal properties. Moreover, conventional search-based approaches do not explore the molecular space in a principled manner, which prevents the generative process from being conditioned upon specific properties.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0097] In addition, whether a molecule exhibits certain desired properties may be contingent on the three-dimensional structure (or conformation) of the molecule. For example, the binding affinity between a drug molecule and a target molecule (e.g., a protein, a nucleic acid, and / or the like) may depend on the ability of the drug molecule to adopt a three- dimensional structure that is complementary to that of the target molecule. Furthermore, molecules are flexible, meaning that a single molecule may assume one of numerous possible three-dimensional structures. For instance, while conformers of a molecule may exhibit the same chemical composition, their three-dimensional structures may differ via rotations about one or more intramolecular bonds. As such, in some cases, a population of the same molecule can exist as an ensemble of many different three-dimensional structures, such as conformers, in equilibrium with one another. Nevertheless, not every possible three-dimensional structure of the molecule is associated with desired properties. In the context of binding affinity, for instance, the biologically active conformation of a molecule may be one or more of the three- dimensional structures exhibited by the molecule in solution or a new three-dimensional structure that is induced by interactions with the target molecule. A one-dimensional representation (e.g., a simplified molecular-input line-entry system (SMILES) string) or a two- dimensional representation (e.g., a molecular graph) of a molecule do not adequately capture the three-dimensional structure of the molecule. Thus, in cases where the molecule design computation model operates on a one-dimensional representation or a two-dimensional representation of the input molecule, the resulting output molecule may not necessarily exhibit a three-dimensional structure that is associated with the one or more desired properties.

[0098] The likelihood that the output molecule exhibits one or more desired properties may be increased (or maximized) by the molecule design computation modelAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 operating on a three-dimensional representation of the input molecule. For example, in some cases, the molecule design computation model may modify the three-dimensional representation of the input molecule to generate a three-dimensional representation of the output molecule. The molecule design computation model may be trained to apply modifications that are consistent with the three-dimensional structure of molecules exhibiting the one or more desired properties. However, conventional three-dimensional representations of molecules are compatible with machine learning architectures with limited expressivity, such as graph neural networks (GNNs). A machine learning mode with limited expressivity cannot adequately capture the complexity and variability in the sample molecules used for training the machine learning model. Such a machine learning model will fail to generalize to new, unseen data after training. That is, the trained machine learning model is unlikely to achieve satisfactory performance when operating on data not encountered during training. In addition , the resources required to operate on conventional three-dimensional representations of molecules tend to scale exponentially with the size of molecules, meaning that conventional three-dimensional representation of molecules are too computationally cumbersome for larger sized molecules.

[0099] One example of a conventional three-dimensional representation of a molecule is a point cloud representation in which the atoms in the molecule are represented as individual points in three-dimensional space. For example, the point cloud representation of a molecule may include, for each atom in the molecule, the coordinates (e.g., x, y,z) coordinates) of the corresponding location in three-dimensional space. Point cloud representations of molecules are typically processed by graph neural networks (GNN), whose equivariant architectures tend to be less expressive than other machine learning architectures due to the message passing formalism. The computational resources required to operate on point cloudAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 representation of molecules scale quadratically with the number of atoms, which renders point cloud representations of molecules impractical for generating larger sized molecules, such as peptides, proteins, and / or the like. Voxel grids are another example of a conventional three- dimensional representation of a molecule. The voxel grid representation of a molecule may include, for each voxel in a voxel grid, an atomic density at the corresponding location. Despite voxel grid representations of molecules being compatible with more expressive machine learning architectures (e.g., convolutional neural networks, transformers, and / or the like), the computational resources required to operate on voxel grid representations of molecules still scale cubically with the size of molecules (e.g., the volume of the voxel grid occupied by the molecules). As such, similar to point cloud representations of molecules, voxel grid representations of molecules are also impractical for generating larger sized molecules. These deficiencies in expressivity and scalability limit the applicability of the molecule design computation model when the molecule design computation model is applied to operate on conventional three-dimensional representations of molecules. For instance, the molecule design computation model may be unable to operate on conventional three-dimensional representations of larger sized molecules as the computational resources required to do so are likely to exceed what is available.

[0100] Various example embodiments of the present disclosure overcome the limitations associated with conventional three-dimensional representation of molecules, such as point cloud representations, voxel grid representations, and / or the like. For example, in some example embodiments, the three-dimensional structure of an input molecule may be represented as a neural molecular field. In some cases, the neural molecular field representation of a molecule may be a molecular occupancy field (or atomic density field) encoded as a lowAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 dimensional latent code (or modulation code) by a computation model. Tn some cases, the computation model may be a neural field or, in some cases, a conditional neural field. As used herein, the term “field” may refer to a function that maps a location in three-dimensional space to a physical quantity. A molecular occupancy field may be a function mapping a location in three- dimensional space (e.g., as identified the coordinates (x, y, z)) to the atomic density at the location. This molecular occupancy field may be a called a “neural molecular occupancy field” or a “neural field” if it is parameterized by a neural network.

[0101] Unlike a voxel grid representation with a discrete grid of atomic densities, a molecular occupancy field (or atomic density field) may be a continuous function that maps three-dimensional coordinates to the atomic densities at the corresponding locations. These atomic densities may be centered around the individual atoms, meaning that a location nearer the center of an atom may exhibit a higher atomic density than a location farther away from the center of the atom. For example, a location at the center of an atom may have an atomic density of 1 whereas a location unoccupied by any atoms may have an atomic density of 0. In some cases, molecular occupancy fields may be approximated (or parameterized) using a computation model. For example, in some cases, the molecular occupancy fields (or atomic density fields) may be approximated (or parameterized) by a neural network called a neural field. Alternatively, molecular occupancy fields (or atomic density fields) may also be approximated (or parameterized) using conditional neural field, which is a neural network whose outputs are subject to one or more conditions (e g., a target structure). As described in more details below, the computation model (e.g., the neural field or conditional neural field) may include a convolutional neural network (e.g., a three-dimensional convolutional neural network) trained to encode the molecular occupancy field (or atomic density field) of a molecule into aAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 corresponding latent code. Furthermore, in some cases, the computation model (e.g., the neural field or conditional neural field) may include a multiplicative filter network (or another machine learning architecture) trained to decode latent codes (or modulation codes), which encode the molecular occupancy fields of individual molecules. In some cases, the neural field (or conditional neural field) may decode the latent code (or modulation code) of a molecule into the corresponding molecular occupancy field by at least providing a mapping between points in three-dimensional space (e g., (x,y, z) coordinates) to the atomic densities at the corresponding locations. In some cases, the computation model (e.g., the neural field or the conditional neural field) may be an autoencoder (e.g., a variational autoencoder) including the encoder (e.g., implemented with the three-dimensional convolutional neural network) and the decoder (e.g., implemented as the multiplicative filter network).

[0102] Representing the three-dimensional structures (or conformations) of molecules as a neural molecular field (or a conditional neural molecular field) offers a number of advantages over conventional three-dimensional representations of molecules. For example, the neural molecular field may represent the three-dimensional structures (or conformations) of molecules, which constitute complex high-dimensional data, in a relatively compact and lowdimensional space. Accordingly, neural molecular field representations may be compatible with expressive machine learning architectures, such as convolutional neural networks, transformers, and / or the like. When implemented using an expressive machine learning architecture, a molecular design computation model operating on the latent codes generated by the neural field (or conditional neural field) may achieve better performance, including more accurate predictions and better generalization. Moreover, the computational resources required for the molecule design computation model to operate on latent codes scale well with the size ofAttomey Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 molecules, thus allowing the molecule design computation model to be more efficient and perform better with larger sized molecules. The latent code of a molecule may also be decoded, for example, by the neural field (or conditional neural field), to recover the molecular occupancy field (or atomic density field) of the molecule at any arbitrary resolution (or scale). The neural molecular field representations of molecules also do not make any assumptions about molecular structure or geometry. As such, the molecule design computation model may operate on the latent code of an input molecule to generate the latent code of an output molecule without any a priori knowledge of the number of atoms present in the output molecule. The neural molecular field representations of molecules may be domain-agnostic and adapted for use in a variety of molecular design tasks capable of being expressed as fields mapping locations in three- dimensional space to a variety of physical quantities including, for example, atomic densities, surfaces, pharmacophores, molecular orbitals, electron densities, and / or the like.

[0103] In some example embodiments, the neural field (or a conditional neural field) may be trained to encode, into a low dimensional latent code, the molecular occupancy field representative of the three-dimensional structure (or conformation) of a molecule. For example, in some cases, the neural field (or conditional neural field) may be trained to model common molecular structural features, such as bonds, angles, valencies, symmetries, and / or the like. Accordingly, the neural field (or conditional neural field) may serve as a decoder that is shared amongst multiple molecules, each of which being associated with a unique latent code (or modulation code). In some cases, an encoder (e.g., a three-dimensional convolutional neural network (CNN) may be trained to generate, for each molecule, a latent code (or modulation code) that captures the structural features that are unique to individual molecules. In some cases, as a decoder, the neural field (or conditional field) may determine, based on a point in three-Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 dimensional space (e.g., (x, y, z) coordinates) and the latent code (or modulation code) of a molecule, the atomic density exhibited by the molecule at the location corresponding to the point in three-dimensional space. As noted, the latent code of a molecule is a low-dimensional, compact, and scalable (e.g., in molecule size and resolution) representation that does not make any assumptions on the three-dimensional structure of the molecule. Moreover, as described in more details below, the molecule design computation model may operate on a latent code representative of the three-dimensional structure of the input molecule instead of a conventional three-dimensional representation (e.g., point cloud, voxel grid, and / or the like) of an input molecule to achieve better performance and computational efficiency. For instance, in some cases, the latent code of the input molecule may be modified before the resulting modified latent code is decoded by the neural field to recover the molecular occupancy field (or atomic density field) of the corresponding output molecule. Doing so may enable the molecule design computation model to scale up to larger sized molecules with available computational resources while preserving the benefits operating on a three-dimensional representation of the input molecule.

[0104] In some example embodiments, the molecule design computation model may generate an output molecule by at least modifying the latent code encoding a molecular occupancy field representative of the three-dimensional structure (or conformation) of the input molecule. In some cases, the molecule design computation model may be trained to approximate a noisy data distribution of noisy latent codes of molecules, which exhibits smoother transitions between higher density regions populated by molecules with desired properties and lower density regions populated by molecules without the desired properties. As such, instead of modifying the latent code of the input molecule directly, the molecule design computation model may beAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 trained to modify a noisy latent code generated by adding noise (e.g., Gaussian noise and / or the like) to the latent code of the input molecule. Doing so may be tantamount to sampling noisy latent codes from the noisy data distribution, whose smoother density transitions and more gradual gradients support a more efficient exploration of the data distribution.

[0105] Contrastingly, sampling from the true data distribution of clean latent codes may be thwarted by the “jaggedness” of the true data distribution, which refers to the steep gradient changes that are present between the low- and high-density regions of the true data distribution. When sampling from the true data distribution of clean latent code, the molecule design computation model may be confined to areas within the immediate vicinity of the sample molecules in the training dataset used to train the molecule design computation model. This limitation gives rise to the phenomenon of mode collapse where the molecule design computation model is less robust and capable of generating few output molecules (e.g., those within the immediate vicinity of the sample molecules in the training dataset). Sampling from the noisy data distribution may avoid mode collapse but the presence of an excessive quantity of noise in the noisy latent code sampled from the noisy data distribution may thwart the subsequent recovery of the molecule therefrom. For example, noise in the noisy latent code may result in one or more atoms being missing from the molecule that is recovered by the molecular occupancy field decoding the noisy latent code. As such, it should be appreciated that the noise level associated with the noisy data distribution may be adjusted to tolerate some jaggedness in order to increase the quality of the molecules recovered from the noisy latent codes.

[0106] In some example embodiments, multiple noisy latent codes may be sampled from the noisy data distribution over multiple timesteps, with each successive noisy latent code being sample from an incrementally higher density region of the noisy data distribution. UponAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 meeting one or more criteria, a noisy latent code sampled from the noisy data distribution may be denoised to generate a denoised (or “clean”) latent code before the denoised latent code is further decoded into the molecular occupancy field (or atomic density field) of the corresponding output molecule. As a continuous function, the molecular occupancy field (or atomic density field) of a molecule recovered from the latent code of the molecule may be capable of representing the three-dimensional structure of the molecule at any arbitrary resolution. The compactness of the latent codes compared to the corresponding molecular occupancy fields may also improve the scalability of the generative process to larger sized molecules that would overwhelm prevalent time and computational resource constraints if rendered in conventional three-dimensional representations such as point cloud and voxel grids.

[0107] In some example embodiments, the molecule design computation model may be trained to perform structure-conditioned three-dimensional molecule generation in which the one or more output molecules are generated while conditioned on a target structure such that the three-dimensional structure of the one or more output molecules are consistent with that of the target structure. For example, in some cases, the target structure may include the three- dimensional structure of at least a portion of another molecule, such as an epitope, a binding pocket, and / or the like. Alternatively and / or additionally, the target structure may be a substructure, such as a scaffold (of a chemical compound), a backbone (of a protein molecule), and / or the like. In some cases, the molecule design computation model may perform structure- conditioned three-dimensional molecule generation by at least modifying, based on the target structure, a latent code representative of the three-dimensional structure of an input molecule. In some cases, the latent code may be a low-dimensional representation that encodes the molecular occupancy field representative of the three-dimensional structure of the input molecule.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1Moreover, in some cases, the molecule design computation model may be trained to operate on a noisy latent code generated by adding noise (e.g., Gaussian noise and / or the like) to the latent code encoding the molecular occupancy field (or atomic density field) of the input molecule. For instance, in some cases, the molecule design computation model may modify the noisy latent code before the modified latent code (or a noisy latent code sampled from the aforementioned noisy data distribution) is denoised and decoded to generate the molecular occupancy field representative of three-dimensional structure of the output molecule. The compactness of the latent code may reduce the computational burden associated with the structure-conditioned generative process such that the molecule design computation model is able to perform faster and scale to larger molecules. Conditioning the modification of the latent code on the target structure may ensure that the three-dimensional structure of the resulting output molecule is consistent with the target structure. In instances where the target structure includes the three-dimensional structure of at least a portion of another molecule (e.g., epitope, binding pocket, and / or the like), for example, the three-dimensional structure of the output molecule may be complementary to that of the target structure. Where the target structure is a substructure (e.g., scaffold, backbone, and / or the like), the three-dimensional structure of the output molecule may conform to that of the target structure.

[0108] In some example embodiments, the molecule design computation model may be trained to perform structure-conditioned three-dimensional molecule generation with a training dataset in which each training sample includes a target structure and a corrupted latent code representative of the corrupted three-dimensional structure of a sample molecule whose uncorrupted (or original) three-dimensional structure is consistent with that of the target structure prior to the corruption. For example, in some cases, the uncorrupted three-dimensional structureAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 of the sample molecule may be complementary to that of the target structure (e.g., epitope, binding pocket, and / or the like) such that the sample molecule is able to bind to target structure. Alternatively, the uncorrupted three-dimensional structure of the sample molecule may conform to that of the target structure (e.g., scaffold, backbone, and / or the like). In some cases, the corrupted latent code may be generated by corrupting an uncorrupted latent code representative of the uncorrupted three-dimensional structure of the sample molecule with the addition of noise (e.g., Gaussian noise such as isotropic Gaussian noise). In some cases, the molecule design computation model may operate on an embedding that combines the corrupted latent code of the sample molecule and a voxelized representation of the three-dimensional structure of the target structure. For instance, in some cases, the molecule design computation model may be trained to denoise, based on the three-dimensional structure of the target structure, the corrupted latent code and recover an uncorrupted latent code corresponding to the uncorrupted three-dimensional structure of the sample molecule. In some cases, the uncorrupted latent code may be further decoded to generate a molecular occupancy field representative of the uncorrupted three- dimensional structure of the sample molecule. In some cases, the training of the molecule design computation model may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the molecule design computation model to reduce (or minimize) a difference (e.g., mean squared error (MSE)) between the recovered three-dimensional structure of the sample molecule and the ground-truth three-dimensional structure of the sample molecule.

[0109] In some example embodiments, to avoid overfitting the molecule design computation model to the sample molecules in the training dataset, the molecule design computation model may be trained to recover noisy versions of the uncorrupted latent codes instead of the original latent codes. That is, the latent codes of the sample molecules in theAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 training dataset may be adulterated with additional noise, which is not the same as the noise that the molecule design computation model is trained to remove from the corrupted latent code of each sample molecule in the training dataset. As described in more details below, in some cases, the noisy latent code of a sample molecule may be generated by adulterating the latent code of the sample molecule with a first quantity of noise (e.g., Gaussian noise such as isotropic Gaussian noise and / or the like) before the resulting noisy latent code is further corrupted with a second quantity of noise (e.g., Gaussian noise such as isotropic Gaussian noise and / or the like) to generate the corrupted latent code. It should be appreciated that the molecule design computation model may be trained to denoise the corrupted latent code of the sample molecule by removing the first quantity of noise (to recover the original latent code of the sample molecule) but not the second quantity of noise (in order to continue sampling from the noisy data distribution).

[0110] FIG. 1 depicts a system diagram illustrating an example of a molecule design system 100, in accordance with some example embodiments. Referring to FIG. 1, the molecule design system 100 may include a molecule design engine 110, a training engine 120, and a client device 130. In the example of the molecule design system 100 shown in FIG. 1, the molecule design engine 110, the training engine 120, and the client device 130 may be communicatively coupled via a network 140. The client device 130 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable apparatus, and / or the like. The network 140 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0111] Referring again to FIG. 1 , in some example embodiments, the molecule design engine 110 may include a molecule design computation model 115 trained to generate a latent code encoding a molecular occupancy field (or atomic density field) of an output molecule by at least modifying a latent code encoding a molecular occupancy field (or atomic density field) of an input molecule. For example, in FIG. 1, the molecule design computation model 115 may modify a latent code 151 encoding an input molecule molecular occupancy field 150 to generate a modified latent code 158 that can be decoded to recover an output molecule molecular occupancy field 158. In some cases, the molecule design computation model 115 may be trained to perform structure-conditioned three-dimensional molecule generation. In the example shown in FIG. 1, the molecule design computation model 115 may be applied to generate, based at least on the input molecule molecular occupancy field 150 and a target structure 152, the output molecule molecular occupancy field 158. In some cases, a molecular occupancy field may be a continuous function that represents the three-dimensional structure of a molecule, including the spatial distribution of the constituent atoms, by mapping three-dimensional coordinates to the atomic densities at the corresponding locations. The atomic densities may be centered around individual atoms, meaning that a location nearer the center of an atom may exhibit a higher atomic density than a location farther away from the center of the atom. In some cases, the input molecule molecular occupancy field 150 and the output molecule molecular occupancy field 158 may include multiple channels, each of which corresponding to an atom type (or chemical element) in the corresponding molecule. For example, in some cases, a first channel may map three-dimensional coordinates to the atomic densities of a first atom type (or first chemical element) while a second channel may map three-dimensional coordinates to the atomic densities of a second atom type (or second chemical element).Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0112] In some example embodiments, the molecular occupancy field (or atomic density field) of the molecule may be approximated (or parameterized) by a computation model. For example, in some cases, the molecular occupancy field (or atomic density field) of the molecule may be approximated (or parameterized) by a neural network called a neural field (or a conditional neural field). In some cases, the molecular occupancy field (or atomic density field) of the molecule may be encoded into a low dimensional latent code, which is more compact and improves the scalability of the generative process. For instance, in FIG. 1, the molecule design engine 110 as including an encoder 111 trained to generate a latent code 151. When the latent code 151 is used for training the molecule design computation model 115, the latent code 151 may encode the input molecule molecular occupancy field 150 of an input molecule that the molecule design computation model 115 is trained to recover. Alternatively, where the molecule design computation model 115 is being applied to generate the output molecule molecular field 158 at inference time, the latent code 151 may be a noisy latent code not corresponding to the molecular field of any particular input molecule. In some cases, the encoder 111 may be a three- dimensional convolutional neural network (CNN) that has been trained to generate the latent code 151 to capture the structural features that are unique to the input molecule. In some cases, instead of operating directly on the input molecule molecular occupancy field 150, which imposes higher computational burden and limits scalability to larger sized molecules, the molecule design computation model 115 may generate the output molecule molecular occupancy field 158 by modifying the latent code 151 instead. As described in more details below, the latent code 151 may be further adulterated with noise (e.g., Gaussian noise and / or the like) before the molecule design computation model 115 is applied to modify the latent code 151. Moreover, the modified latent code 156 may be decoded by the neural field (or a conditionalAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 neural field), which serves as a decoder 119 shared by multiple molecules. Tn some cases, the decoder 119 may be trained to model the structural features that are common to all molecules. In some cases, the decoder 119 may decode the modified latent code 156 by at least providing, for each point in three-dimensional space, the atomic density of the output molecule at the corresponding location.

[0113] In some example embodiments, the molecule design computation model 115 may be trained to approximate a noisy data distribution populated by noisy latent codes. The noisy data distribution may exhibit smoother density transitions between higher density regions populated by molecules with one or more desired properties and lower density regions populated by molecules without the one or more desired properties. For small molecule drugs (or chemical compounds, examples of the one or more desired properties may include potency (e.g., biochemical potency, cellular potency, in vivo potency), clearance, and permeability. In the context of large molecule therapeutics (or biologies), examples of the one or more desired properties may include expression, binding affinity towards a target molecule, binding specificity, stability, non-immunogenicity, human-ness, absence of self-association (or nonaggregation), lack of chemical liabilities (e.g., aspartate isomerization, oxidation, deamidation) , and / or the like. In some cases, noise (e.g., Gaussian noise and / or the like) may be added to the latent code 151 such that the molecule design computation model 115 samples from the noisy data distribution with smother density transitions when modifying the latent code 151. In some cases, the molecule design computation model 115 may modify the latent code 151 over multiple timesteps, with one or more noisy latent codes being sampled from the noisy data distribution at each timestep. In some cases, the molecule design computation model 115 may modify the latent code 151 while guided by the gradient of a function (e.g., score function) indicative of theAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 local density changes in the noisy data distribution. In some cases, the latent code 151 may be modified based on the gradient of the function (e.g., score function) such that the noisy latent codes resulting from the modifications are sampled from incrementally higher density regions of the noisy data distribution. For instance, in some cases, the molecule design computation model 115 may be applied to modify the latent code 151 to generate a first noisy latent code and a second noisy latent code. The first noisy latent code and the second noisy latent code may be sampled from the noisy data distribution. In some cases, the function (e.g., score function) may be applied to determine a first value indicative of a first local change in a density of the noisy data distribution at a first location occupied by the first noisy latent code. Furthermore, the function (e.g., score function) may be applied to determine a second value indicative of a second local change in the density of the noisy data distribution at a second location occupied by the second noisy latent code. In some cases, where the first value and the second value indicates that the first noisy latent code is sampled from a higher density region of the noisy data distribution than the second noisy latent code, the molecule design computation model may be applied to further modify the first noisy latent code instead of the second noisy latent code.

[0114] In some cases, the function (e.g., score function) may be parameterized by the parameters (e.g., weights, biases, and / or the like) of the molecule design computation model 115, meaning that the parameters of the function may be determined during the training of the molecule design computation model 115. Those parameters may be adjusted, during the training of the molecule design computation model 115, for the function to assign a higher value (e.g., higher score) to a noisy latent code sampled from a region in the noisy data distribution with a more positive local change (e.g., an increase or a smaller decrease) in density than a noisy latentAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 code sampled from a region with a less positive local change (e.g., a decrease or a smaller increase) in density.

[0115] In some example embodiments, the latent code 151, or a latent code embedding 153 of the latent code 151, may be combined with a target structure embedding 155 of the target structure 152 to form a joint embedding 154. In some cases, the molecule design computation model 115 may operate on the joint embedding 154 to generate the output molecule molecular occupancy field 158. For example, in some cases, the molecule design computation model 115 may modify the latent code 151 of the input molecule molecular occupancy field 150 over multiple timesteps, with one or more noisy latent codes being sampled from the noisy data distribution at each timestep. As shown in FIG. 1, in some cases, the resulting modified latent code 156 may be denoised and decoded, for example, by the decoder 119, to generate the output molecule molecular occupancy field 158. In some cases, the latent code 151 of the input molecule molecular occupancy field 150 may be modified based on the target structure embedding 155 such that the three-dimensional structure encoded by the output molecule molecular occupancy field 158 is consistent with that of the target structure 152. For instance, in cases where the target structure 152 includes the three-dimensional structure of at least a portion of another molecule, such as an epitope, a binding pocket, and / or the like, the three-dimensional structure encoded by the output molecule molecular occupancy field 158 may be complementary to that of the target structure 152. Alternatively, where the target structure 152 is a substructure (e g., scaffold, backbone, and / or the like), the three-dimensional structure encoded by the output molecule molecular occupancy field 158 may conform to that of the target structure 152.

[0116] In some example embodiments, the molecule design computation model 115 may be trained based on a training dataset that includes, for each training sample, a targetAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 structure and a corrupted latent code representative of the corrupted three-dimensional structure of a sample molecule whose uncorrupted (or original) three-dimensional structure is consistent with that of the target structure. As shown in FIG. 1, in some cases, the training engine 120 may generate, for a sample molecule molecular occupancy field 182, a corrupted latent code 188 for inclusion in the training dataset. In some cases, the three-dimensional structure of the sample molecule may be consistent with that of the target molecule, which may be a substructure (e.g., scaffold, backbone, and / or the like) or at least a portion of another molecule (e.g., epitope, binding pocket, and / or the like). To generate the corrupted latent code 188, the training engine 120 may apply the encoder 111 to generate a latent code 186 encoding the sample molecule molecular field 182 before noise (e.g., Gaussian noise and / or the like) is added, for example, by a corruption engine 121, to generate the corrupted latent code 188. In some cases, the latent code 186 may be adulterated with a first quantity of noise (e.g., Gaussian noise and / or the like) before the corruption engine 121 adds a second quantity of noise to generate the corrupted latent code 188. The first quantity of noise projects the latent code 186 to the aforementioned noisy data distribution, which exhibits a smoother density transition than the true data distribution of the latent code 186. As such, the molecule design computation model 115 may be trained to remove the second quantity of noise but not the first quantity of noise from the corrupted latent code 188. For example, in some cases, the molecule design computation model 115 may be trained by at least adjusting one or more parameters (e.g., weights, biases, and / or the like) of the molecule design computation model 115 to increase the similarity between the latent code 186 and the latent code recovered by the molecule design computation model 115 from the corrupted latent code 188.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0117] FIG. 2 depicts a flowchart illustrating an example of a process 200 for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments. Referring to FIGS. 1-2, the process 200 may be performed by the molecule design engine 110 to generate the output molecule molecular occupancy field 158 by at least applying the molecule design computation model 115 to modify the input molecule molecular occupancy field 150. In some cases, the molecule design computation model 110 may be trained to operate on the latent code 151, which is a lower dimensional representation of the input molecule molecular occupancy field 150. One advantage of the latent code 151 is its compactness, meaning that the computational resource required for the molecule design computation model 115 to operate on the latent code 151 scales well with the size of the input molecule associated with the input molecule molecular occupancy field 150. The latent code 151 is also compatible with more expressive machine learning architectures. As such, the molecule design computation model 115 may be implemented with more expressive machine learning architectures to achieve better performance, including more accurate predictions and better generalization. Yet another advantage of the latent code 151 is that the latent code 151 does not make any assumptions about molecular structure or geometry, such that the molecule design computation model 115 is able to operate on the latent code 151 and generate the modified latent code 156 corresponding to the output molecule molecular occupancy field 158 without a priori knowledge of the number of atoms present in the output molecule.

[0118] At 202, an input molecule is received. In some example embodiments, the input molecule may be a small molecule (or a low-weight compound) having a molecular weight between approximately 100 Daltons and 1000 Daltons. Alternatively, the input molecule may be a large molecule, such as proteins, ranging between approximately 3000 Daltons and 150,000Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1Daltons in molecular weight. Tn some cases, it is also possible for the input molecule to be a modality between small and large molecules, such as macro cyclic peptides. In some cases, the three-dimensional structure (or conformation) of the input molecule may be represented by molecular occupancy. In some cases, the molecular occupancy field (or atomic density field) the input molecule may include atomic densities across a voxel grid (or a grid of voxels) in which each voxel is associated with a value indicative of the atomic density at the corresponding location.

[0119] At 204, a latent code of the input molecule is generated by at least applying a computation model to encode, into the latent code, a molecular occupancy field representative of a three-dimensional structure of the input molecule. In some example embodiments, the molecular occupancy field (or atomic density field) the input molecule may represent the three- dimensional strucutre (or conformation) of the input molecule by at least specifying the atomic density at one or more points in three-dimensional space. In some cases, these atomic densities may be centered around the individual atoms present in the input molecule. That is, in some cases, the atomic density of a point in a particular location in three-dimensional space may correspond to a distance between that location and the center of an atom in the input molecule. As such, a location nearer the center of an atom may exhibit a higher atomic density (e.g., closer to an atomic density value of 1) while a location farther away from the center of the atom may exhibit a lower atomic density (e.g., closer to an atomic density value of 0). In some cases, the molecular occupancy representative of the three-dimensional structure of the input molecule may be encoded by a computation model, such as a neural field, into a latent code (or modulation code). In some cases, the neural field may be a continuous function mapping every point in three-dimensional space to the atomic density of each atom type at the corresponding location inAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 three-dimensional space. In some cases, the neural field may be parameterized by a neural network. In some cases, the neural field may be trained to model common molecular structural features, such as bonds, angles, valencies, symmetries, and / or the like. In some cases, the neural field may encode the molecular occupancy field (or atomic density field) the input molecule such that the latent code resulting therefrom captures the structural features that are unique to the input molecule. In some cases, the latent code generated by the neural field may be unique to the input molecule. Moreover, in some cases, the neural field may encode the molecular occupancy field (or atomic density field) the input molecule such that the resulting latent code may be decoded, for example, by the neural field, to recover the three-dimensional structure of the input molecule. For instance, in some cases, the neural field may decode the latent code of the input molecule by at least mapping each point in three-dimensional space (e.g., three-dimensional coordinates (x, y, z)) to the atomic density at the corresponding location in three-dimensional space.

[0120] At 206, an output molecule is generated by at least modifying the latent code of the input molecule. In some example embodiments, the output molecule may be generated by at least applying a molecule design computation model to modify the latent code of the input molecule. In some cases, the molecule design computation model may generate the latent code of the output molecule by modifying the latent code of the input molecule over one or more successive iterations. Moreover, in some cases, the molecule design computation model may operate on a noisy version of the latent code of the input molecule such that each iteration of modification includes sampling a noisy latent code from a noisy data distribution. In some cases, the noisy data distribution, which are populated by noisy latent codes instead of clean latent codes, may exhibit smoother transitions between higher density regions and lower densityAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 regions. For example, in some cases, the addition of noise may be tantamount to projecting the latent code of the input molecule from the true data distribution of clean latent codes to the noisy data distribution. In some cases, the molecule design computation model may be applied to modify the noisy latent code of the input molecule to successively sample noisy latent codes from the noisy data distribution until one or more criteria are met. For instance, in some cases, the one or more criteria may include the noisy latent code being sampled from a sufficiently high density region of the noisy data distribution such that the molecule reconstructed from the noisy latent code has a sufficiently high likelihood of exhibiting the one or more desired properties. Upon meeting the one or more criteria, the noisy latent code sampled from the noisy data distribution may be denoised to generate a denoised (or “clean”) latent code. As described in more details below, the denoised latent code may be further decoded, for example, by a computation model (e.g., the neural field), into the molecular occupancy field (or atomic density field) the corresponding output molecule.

[0121] In some cases, the molecule design computation model may be trained to approximate the noisy data distribution using the noisy latent codes of sample molecules exhibiting one or more desired properties. Accordingly, the higher density regions of the noisy data distribution may be populated by the noisy latent codes of molecules with one or more desired properties while the lower density regions may be populated by the noisy latent codes of molecules without the one or more desired properties. Sampling from the noisy data distribution may be more efficient due to the noisy data distribution having a smoother density distribution (or less jagged density distribution) than the true data distribution of clean latent codes. However, the quantity of noise present in the noisy latent codes may affect the quality of the molecules reconstructed therefrom. For example, when an excessive quantity of noise is presentAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 in the noisy latent codes, the three-dimensional structure of the molecules recovered therefrom may exhibit artifacts such as missing atoms. As such, in some cases, the noise level associated with the noisy data distribution (e.g., the quantity of noise added to clean latent codes) may be adjusted to tolerate some jaggedness in order to increase the quality of the molecules recovered from the noisy latent codes.

[0122] At 208, the latent code of the output molecule is decoded to determine a molecular occupancy field of the output molecule. In some example embodiments, the neural field may be applied to decode the latent code of the output molecule. As noted, the molecule design computation model may be applied to modify the noisy latent code of the input molecule continuously until one or more criteria are met. In some cases, doing so may be tantamount to sampling, from a noisy data distribution populated by noisy latent codes, one or more successive noisy latent codes until the one or more criteria are met. The noisy latent code that is sampled from the noisy data distribution upon meeting the one or more criteria may be denoised before the clean latent code resulting therefrom is decoded by the neural field, serving as a decoder in this instance. As noted, in some cases, the neural field may decode the latent code by at least mapping points in three-dimensional space (e.g., three-dimensional coordinates (x, y, z)) to the atomic densities at the corresponding locations. For example, in some cases, the neural field may receive, as inputs, the latent code of the output molecule and one or more points in three- dimensional space (e.g., one or more three-dimensional coordinates (x,y, z)). In some cases, the neural field may output, for the latent code of the output molecule and a point in three- dimensional space (e.g., three-dimensional coordinatesa value indicative of the atomic density at the corresponding location. In some cases, atomic densities may be centered at each atom present in the output molecule. As such, the neural field may output a higher value for aAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 point in three-dimensional space near the center of an atom and a lower value for a point in three-dimensional space distant from the center of an atom. In some cases, the atomic densities at multiple different points in three-dimensional space may correspond to the molecular occupancy field (or atomic density field) the output molecule.

[0123] At 210, a three-dimensional structure of the output molecule is determined based at least on the molecular occupancy field of the output molecule. In some example embodiments, the three-dimensional structure (or conformation) of the output molecule may be further extracted from the molecular occupancy field (or atomic density field) the output molecule, as decoded by the neural field based on the latent code of the output molecule and one or more points in three-dimensional space (e.g., one or more three-dimensional coordinates (x, y, z)). As noted, in some cases, the molecular occupancy field (or atomic density field) the output molecule may be approximated by the neural field providing the atomic density at the one or more points in three-dimensional space (e.g., one or more three-dimensional coordinates (x, y, z)). For certain applications, such as those in chemistry and biology, requiring the three- dimensional structure of the output molecule instead of its molecular occupancy, the three- dimensional structure (or conformation) of the output molecule may be extracted from the molecular occupancy field (or atomic density field) of the output molecule. For example, in some cases, a discretized voxel grid may be rendered from the molecular occupancy field (or atomic density field) of the output molecule using a uniform discretization of space and the neural field. In some cases, the atomic density present at each voxel in the discretized voxel grid may be determined by applying the neural field to the corresponding point in three-dimensional space (e.g., three-dimensional coordinates (x,y, z)) and the latent code of the output molecule. In some cases, a peak finding algorithm may be applied to identify one or more peak (orAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 maximum) atomic density values across the discretized voxel grid. The number of atoms in the output molecule on each channel, each of which corresponding to a different atom type (or chemical element), as well as the three-dimensional coordinates of each atom may be inferred from the peak (or maximum) atomic density values. In some cases, a continuous refinement may be applied to determine the local maximum of the neural field such that the coordinates of each identified atom may be refined around the neighborhood of coordinates found with the peak detector. This continuous refinement scheme may enable the detection of atomic coordinates that lie beyond the initial coarse uniform discretization. Furthermore, as noted, the quality of the reconstruction of the output molecule may be improved by at least reducing the quantity of noise present in the noisy latent code from which the output molecule originates.

[0124] In some example embodiment, the molecule design computation model 115 may perform structure-conditioned generation of three-dimensional molecules in which the modifying of the latent code encoding the molecular occupancy field (or atomic density field) of the input molecule is conditioned on a target structure such that the three-dimensional structure (or conformation) of the output molecule generated therefrom is consistent with the target structure. In some cases, the molecule design computation model 115 may operate on the latent code 151 encoding the input molecule molecular occupancy field 150. For example, in FIG. 1, the molecule design computation model 115 may modify the latent code 151 encoding the input molecule molecular occupancy field 150 to generate the modified latent code 156, which may be decoded by the decoder 1 19 to recover the output molecule molecular occupancy field 158. In some cases, the encoder 111 encoding the input molecule molecular occupancy field 150 to generate the latent code 151 and the decoder 119 decoding the modified latent code 156 to generate the output molecule molecular occupancy field 158 may form an autoencoderAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 architecture implementing the neural field (or conditional neural field). As noted, the neural field (or conditional neural field) may be a continuous function mapping points in three- dimensional space (e.g., three-dimensional coordinates (x,y, z)) to atomic densities at the corresponding location. In some cases, the neural field (or conditional neural field) may be parameterized by a neural network. In some cases, the neural field may be parameterized by a three-dimensional convolutional neural network implementing an encoder and a multiplicative filter network implementing the decoder.

[0125] To further illustrate, FIG. 3A depicts a schematic diagram illustrating an example of a neural-field based autoencoder 300, in accordance with some example embodiments. In the example shown in FIG. 3 A, the three-dimensional structure (or conformation) of a molecule 310 may be represented as a function x GG IR3, wherein n is the number of atom types (or chemical elements) present in the molecule 310 and v is the atomic density field of the molecule 310. In some cases, the encoder 111 may extract a latent code z of the molecule 310, which can be decoded back into a density field by the decoder 119. In some cases, the encoder 111 may be implemented as a three-dimensional convolutional neural network (CNN) whose input may be a low-resolution voxel grid of the molecule 310 denoted, for example, as Q G JR>SXSXSX’1with grid dimension s. The latent code z G ]R>sxsxsxdoutput by the encoder 111 may share the same spatial dimension as the low-resolution voxel but exhibit a higher number of channels d. In some cases, a spatially arranged latent code may be used in order to leverage certain machine learning architectures, such as UNet-based architectures, for generative modeling. Doing so may avoid the loss of spatial information as part of the generative process.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0126] In some cases, the decoder 119 may be implemented as a conditional neural field based on a multiplicative filter network architecture. In some cases, the multiplicative filter network architecture uses Gabor filters, which are suitable for modelling sparse atomic density fields. The latent space may be regularized by imposing a slight Kullback-Leibler (KL) penalty towards a standard normal on the latent vectors. The loss over a set T> of sample molecules (in a training dataset) may be defined as Equation (1) as follows:wherein denotes the encoder 111,denotes the decoder 119,— J ' iCv), <7(v)), [ / i(v), cr(v)]and is small. In some cases, training may be focused on non-empty voxel spaces by at least upsampling coordinates x close to the center of each atom.

[0127] FIG. 3B depicts a schematic diagram illustrating an example of a process 325 for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments. In the example shown in FIG. 1, the input molecule corresponding to the input molecule molecular occupancy field 150 and the target structure 152 may form a proteinligand complex 301 in which the input molecule is the ligand and the target structure 152 is the binding pocket in a protein molecule in which the binding interaction between the ligand and the protein molecule takes place. In some cases, the encoder 111 may be a three-dimensional convolutional neural network (CNN) trained to generate the latent code 151 by at least encoding the voxelized input molecule molecular occupancy field 150. In some cases, the latent code 151 may be compact, low-dimensional representation of the input molecule molecular occupancy field 150.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0128] As shown in FIG. 3B, a noisy latent code 308 may be generated by adding noise 307, such as the Gaussian noise s shown, to the latent code 151. FIG. 3B shows the noise 307 having a noise level cr, which determines the quantity of the noise 307 (e.g., Gaussian noise E). Moreover, FIG. 3B shows the three-dimensional structure y encoded by the noisy latent code 308. In some cases, the addition of the noise 307 may project the latent code 151 from a true data distribution of clean (or noiseless) latent codes to a noisy data distribution with smoother density transitions between higher density regions populated by molecules having one or more desired properties and lower density regions populated by molecules without the one or more desired properties. Sampling from the noisy data distribution, which includes drawing noisy latent codes from the noisy data distribution by modifying the noisy latent code 308, may be more efficient than sampling from the true data distribution at least because the gradual gradient of the noisy data distribution enables noisy latent codes to be sampled from regions of the noisy data distribution that are farther away from the location of the noisy latent code 308.

[0129] Referring again to FIG. 3B, in some cases, a three-dimensional target structure representation 303 of the target structure 152 may be generated, for example, by a structure computation model 305. In some cases, the three-dimensional target structure representation 303 may be a voxel grid representation of the target structure 152 in which the atoms forming the target structure 152 are represented by discrete atomic density values across a voxel grid. For example, in some cases, the atomic density field of the target structure 152 may be discretized into low resolution voxels forming the three-dimensional target structure representation 303, which is then encoded with by an encoder 309 into the target structure embedding 155. In some cases, the target structure embedding 155 may correspond to the latent code of the target structure 152. In some cases, the three-dimensional target structure representation 303 and theAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 noisy latent code 308 may be combined to form the joint embedding 154. For instance, FIG. 3B shows the encoder 309 generating the latent code embedding 153 corresponding to the noisy latent code 308 and the target structure embedding 155 corresponding to the three-dimensional target structure representation 303. In some cases, the encoder 309 may be trained to encode the noisy latent code 308 and the three-dimensional target structure representation 303 to a common embedding space. That is, the encoder 309 may generate the latent code embedding 153 and the target structure embedding 155 to have the same spatial dimensions. In the example shown in FIG. 3B, the latent code embedding 153 and the target structure embedding 155 may be combined (e.g., concatenated) to generate the joint embedding 154.

[0130] In some example embodiments, the molecule design computation model 115 may operate on the joint embedding 154 in order to generate, based at least on the target structure 152, the output molecule molecular occupancy field 158. In some cases, the molecule design computation model 115 may generate the output molecule molecular occupancy field 158 by at least modifying the latent code embedding 153 of the noisy latent code 308. In some cases, the molecule design computation model 115 may modify the latent code embedding 153 while guided by the gradient of a function (e.g., score function) indicative of the local density changes in the noisy data distribution. For example, in some cases, the molecule design computation model 115 may sample the noisy data distribution via gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC)) such that the resulting modified latent code 156 is sampled from a sufficiently high density region of the noisy data distribution to ensure that the corresponding output molecule exhibits the one or more desired properties. In some cases, the molecule design computation model 115 may continue to modify the latent code embedding 153 and sample from the noisy data distribution until one or moreAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 criteria are met. For instance, in some cases, the molecule design computation model 115 may continue to sample from the noisy data distribution until a threshold quantity of noisy latent codes have been sampled from the noisy data distribution. Alternative and / or additionally, the molecule design computation model 115 may continue to sample from the noisy data distribution until the function outputs a value that indicates the noisy latent code is drawn from a sufficiently high density region of the noisy data distribution. In the example shown in FIG. 3B, once the one or more criteria are met, the modified latent code 156 may be decoded by the decoder 119 in order to recover the output molecule molecular field 158 of the corresponding output molecule. In some cases, the decoder 119 is a neural field (or a conditional neural field) trained to recover the output molecule molecular field 158 by at least mapping, based at least on the modified latent code 156, points in three-dimensional space to the atomic density of the output molecule at the corresponding locations. Furthermore, in some cases, the three-dimensional structure of the corresponding output molecule, including the types and coordinates of the constituent atoms, may be extracted from the output molecule molecular occupancy field 158.

[0131] FIG. 3C depicts a schematic diagram illustrating another example of a process 350 for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments. As in FIG. 3B, FIG. 3C shows the input molecule corresponding to the input molecule molecular occupancy field 150 and the target structure 152 also forming the protein-ligand complex 301 in which the input molecule is the ligand and the target structure 152 is the binding pocket in a protein molecule in which the binding interaction between the ligand and the protein molecule takes place. In some cases, the three-dimensional target structure representation 303 of the target structure 152 may be generated, for example, by discretizing the three-dimensional target structure representation 303 into low resolution voxels.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1Moreover, in some cases, FIG. 3C shows the latent code 151 of the input molecule molecular field 150 being generated, for example, by the encoder 111 encoding a voxelized molecular field 251 generated by discretizing the input molecule molecular field 150. In some cases, noise (e.g., Gaussian noise and / or the like) may be added to the latent code 151 in order to generate the noisy latent code 308. As noted, in some cases, the addition of noise to the latent code 151 projects the latent code 151 into a noisy data distribution from which the molecule design computation model 115 subsequently samples additional noisy latent codes by modifying the noisy latent code 308. For example, in FIG. 3C, the molecule design computation model 115 may operate on a combination (e.g., concatenation) of the three-dimensional target structure representation 303 (or the target structure embedding 155 generated by the encoder 309) and the noisy latent code 308. In some cases, the molecule design computation model 115 may modify the noisy latent code 308 while being conditioned on the target structure 152 specified by the three-dimensional target structure representation 303 (or the target structure embedding 155 generated by the encoder 309) such that the modified latent code 156 encodes an output molecule whose three- dimensional structure is consistent with that of the target structure 152.

[0132] As shown in FIG. 3C, the modifying of the noisy latent code 308 and the corresponding sampling of additional noisy latent codes from the noisy data distribution may continue until one or more criteria are satisfied. For example, in some cases, the molecule design computation model 115 may be applied to continue modifying the noisy latent code 308 and sampling additional noisy latent codes from the noisy data distribution until a threshold quantity of noisy latent codes have been sampled from the noisy data distribution. Alternative and / or additionally, the molecule design computation model 115 may continue to sample from the noisy data distribution until the function outputs a value that indicates the noisy latent code isAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 drawn from a sufficiently high density region of the noisy data distribution. In the example shown in FIG. 3C, once the one or more criteria are met, the modified latent code 156 may be decoded by the decoder 119 in order to recover the output molecule molecular field 158 of the corresponding output molecule. In some cases, the decoder 119 is a neural field (or a conditional neural field) trained to recover the output molecule molecular field 158 by at least mapping, based at least on the modified latent code 156, points in three-dimensional space to the atomic density of the output molecule at the corresponding locations. Furthermore, in some cases, the three-dimensional structure of the corresponding output molecule, including the types and coordinates of the constituent atoms, may be extracted from the output molecule molecular occupancy field 158.

[0133] FIG. 3D depicts a schematic diagram illustrating an example of a process 375 for training the molecule design computation model 115, in accordance with some example embodiments. In some example embodiments, the molecule design computation model 115 may be a denoiser. Given some noise e~J ' (0, <72) (e.g., the noise 307) and a clean latent z (e.g., the latent code 151 encoding the voxelized molecular field 251 of the input molecule molecular field 150), y = z + e (e.g., the noisy latent code 308) may be defined as a noisy version of z. The molecule design computation model 115, as a denoiser zg, may aim to remove the noise conditioned on the encoding of the target structure zrec(e.g., the target structure embedding 155 of the three-dimensional structure representation 303 of the target structure 152). That is, the denoiser zgmay be trained such that zg(y\zrec; cr) z. In practice, the denoiser zgmay include some additional pre- or post-processing, called preconditioning, as defined by Equation (2) below.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 wherein the data is assumed to be normalized to have unit variance and zero mean. Owing to the spatial structure of latent vectors, in some cases, UNet architectures may be used to parameterize the molecule design computation model (denoted as UQ). The UNet architecture may be efficient and scalable, rendering it suitable for complex generative tasks.

[0134] The training of the molecule design computation model 115 shown in FIG. 3D may aim to optimize the following loss given the noise level o.

[0135] When considering more noise levels, the following reweighing scheme may be applied:wherein u(cr) denotes a single layer multilayer perceptron (MLP) trained with the denoiser and o is sampled along some predetermined distribution p(cr).

[0136] FIG. 4A depicts a flowchart illustrating an example of a process 400 for structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments. Referring to FIGS. 1, 3A-B, and 4A, the process 400 may be performed by the molecule design engine 110 to generate the output molecule molecular field 158 by at least applying the molecule design computation model 115 to modify, based on the target structure 152, the input molecule molecular field 150. In some cases, the molecule design computation model 110 may be trained to operate on the latent code 151, which is a lower dimensional representation of the input molecule molecular occupancy field 150. One advantage of the latent code 151 is its compactness, which enables the molecule design computation model 151 to scale and accommodate larger sized molecules with the available computationalAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 resources. Moreover, in some cases, the modification of the latent code 151 encoding the input molecule molecular occupancy field 150 may be conditioned on the target structure 152. One advantage of structure-conditioned three-dimensional molecule generation is that the three- dimensional structure (or conformation) of the output molecule represented by the output molecule molecular field 158 may be consistent with the three-dimensional structure of the target structure 152. For example, where the target structure 152 includes at least a portion of another molecule (e.g., epitope, binding pocket, and / or the like), the three-dimensional structure of the output molecule may be complementary to that of the target structure 152. Where the target strucutre 152 is a substructure (e.g., scaffold, backbone, and / or the like), the three-dimensional structure of the output molecule may conform to that of the target structure 152.

[0137] At 402, the molecule design engine 110 may determine a three-dimensional target structure representation of a target structure. In some example embodiments, the molecule design engine 110 may apply the structure computation model 305 to generate the three- dimensional target structure representation 303 of the target structure 152. In some cases, the three-dimensional target structure representation 303 may be a voxel grid representation of the target structure 152. In some cases, the three-dimensional target structure representation 303 may include a voxel grid in which each voxel is associated with a value corresponding to an atomic density at a corresponding location. For example, a voxel in the three-dimensional target representation 303 that is closer to the center of an atom in the target structure 152 may have an atomic density value than a voxel that is farther away from the center of any atoms in the target structure 152.

[0138] At 404, the molecule design engine 110 may generate a noisy latent code. In some example embodiments, when the molecule design computation model 115 is applied toAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 generate an output molecule having, for example, the output molecule molecular field 158, the molecule design engine 110 may apply the molecule design computation model 115 to operate on the noisy latent code 308 generated to include random noise, such as the noise 307, which may be Gaussian noise and / or the like. That is, instead of a latent code encoding the molecular occupancy field (or atomic density field) of a particular input molecule, the noisy latent code 308 may include the noise 307 such that the noisy latent code 308 does not encode the molecular occupancy field (or atomic density field) of any particular input molecule. Instead, as described in more details below, the molecule design computation model 115 may modify the noisy latent code 308 to generate the modified latent code 156 encoding the output molecule molecular field 158 of the output molecule. In some cases, the noisy latent code 308 may be embedded, for example, by the encoder 309, to generate the latent code embedding 153 which may be, as described in more details below, combined with the embedding of the target structure 152.

[0139] At 406, the molecule design engine 110 may generate an embedding combining the three-dimensional target structure representation and the noisy latent code. In some example embodiments, the molecule design engine 110 may generate the joint embedding 154 by at least combining the latent code embedding 153 (of the noisy latent code 308) and the target structure embedding 155. As shown in FIG. 2A, in some cases, the latent code 151 may be adulterated with the noise 307 (e.g., Gaussian noise and / or the like) to generate the noisy latent code 308 before the latent code embedding 153 is generated by encoding the noisy latent code 308. In some cases, the addition of the noise 307 may project the latent code 151 into a noisy data distribution with smoother density transitions than the true (or noiseless) data distribution of the latent code 151. As described in more details below, the molecule design computation model 115 may be applied to modify the noisy latent code 308, which may beAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 tantamount to sampling noisy latent codes from the noisy distribution, until one or more criteria are satisfied. In some cases, the noisy latent code 308 may be modified based on the target structure 152. Accordingly, in some cases, the encoder 309 may further encode the three- dimensional target structure representation 303 and the noisy latent code 308 such that the resulting latent code embedding 153 and the target structure embedding 155 occupy a common embedding space with the same latent and spatial dimensions before being combined into the joint embedding 154.

[0140] At 408, the molecule design engine 110 may apply the molecule design computation model 115 to generate a modified latent code encoding an output molecule molecular occupancy field (or atomic density field) of an output molecule by at least modifying, based at least on the three-dimensional target structure representation of the target structure, the noisy latent code such that a three-dimensional structure of the output molecule is consistent with a three-dimensional structure of the target structure. In some example embodiments, the molecule design engine 110 may apply the molecule design computation model 115 to modify the latent code embedding 153 corresponding to the noisy latent code 308. In some cases, modifying the latent code embedding 153 may be tantamount to sampling from a noisy distribution in which higher density regions are populated by noisy latent codes of molecules exhibiting one or more desired properties and lower density regions are populated by noisy latent codes of molecules without the one or more desired properties. In some cases, the molecule design computation model 115 may modify the latent code embedding 153 over multiple iterations while guided by a function (e.g., score function) whose output is a value indicative of the local change in density at the location in the noisy distribution from which a noisy latent code is sampled. Guided by the output of the function (e.g., score function), the molecule designAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 computation model 115 may modify the latent code embedding 153 such that the noisy latent codes are sampled from incrementally higher density regions of the noisy distribution. Modifying the latent code embedding 153 may modify the three-dimensional structure of the corresponding molecule by at least modifying the atomic densities that may be decoded from the resulting modified latent code 156. In some cases, the molecule design computation model 110 may continue to be applied to modify the latent code embedding 153 until one or more criteria are satisfied. For instance, in some cases, the one or more criteria may be satisfied when a threshold quantity of iterations have modifications have been made to the latent code embedding 153. In some cases, the one or more criteria may also include the modified latent code 156 being sampled from a sufficiently high density region of the noisy distribution.

[0141] At 410, the molecule design engine 110 may decode the modified latent code to determine the output molecule molecular occupancy field (or atomic density field) of the output molecule. In some example embodiments, the molecule design engine 110 may apply the neural field serving as the decoder 119 to decode the modified latent code 156. For example, as noted, the decoder 119 may be applied to decode the modified latent code 156 generated by the molecule design computation model 115 once one or more criteria are satisfied. In some cases, the modified latent code 156 may be denoised before the decoder 119 is applied to decode the modified latent code 156. The denoising projects the modified latent code 156 from the noisy distribution from which the molecule design computation model 115 samples noisy latent codes to the true distribution of the modified latent code 156.

[0142] In some example embodiments, the modified latent code 156 may be unique to the output molecule having the output molecule molecular occupancy field 158. Meanwhile, the neural fieldmay be a continuous function mapping the three-dimensional coordinatesAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 associated with the modified latent code 156 to atomic densities at any arbitrary resolution. For example, in some cases, the neural fieldserving as the decoder 119 may approximate the output molecule molecular occupancy field 158 by at least providing the atomic density at any location in three-dimensional space at any arbitrary resolution. Accordingly, in some cases, the decoder 119 may be trained to decode the modified latent code 156 by at least determining, based at least on the modified latent code 156, an atomic density exhibited by the output molecule at one or more points in three-dimensional space (e.g., identified by the three- dimensional coordinates (x,y, z)). The atomic densities may be centered at the center of the atoms forming the output molecule, meaning that the atomic density at a location nearer the center of an atom may be higher than the atomic density at a location farther away from the center of an atom. Moreover, the output molecule molecular occupancy field 158 may include multiple channels, each of which having the atomic densities of a different atom type (or chemical element). As described in more details below, the three-dimensional structure of the output molecule may be determined based on the output molecule molecular occupancy field 158.

[0143] At 412, the molecule design engine 110 may determine, based at least on the modified latent code, a three-dimensional structure of the output molecule. In some example embodiments, the three-dimensional structure (or conformation) of the output molecule may be further extracted from the output molecule molecular occupancy field 158. As noted, the output molecule molecular occupancy field 158 may be approximated by the neural fieldServing as the decoder 119, the neural fieldmay decode the modified latent code 156 by at least providing a mapping between points in three-dimensional space and the atomic densities at the corresponding locations. However, some applications, such as those in chemistry and biology,Attomey Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 may require the three-dimensional structure of the output molecule instead of the corresponding output molecule molecular occupancy field 158. Accordingly, in some cases, the three- dimensional structure (or conformation) of the output molecule may be extracted using the approximation of the output molecule molecular occupancy field 158 provided by the neural field for the corresponding modified latent code 156. For example, a discretized voxel grid may be rendered from the output molecule molecular occupancy field 158 using a uniform discretization of space and the neural fieldThe molecular density present at each voxel in the discretized voxel grid may be determined by applying the neural fieldto the corresponding coordinates (x, y, z) and the modified latent code 156. A peak finding algorithm may then be applied to determining one or more peak (or maximum) atomic density values across the discretized voxel grid. The number of atoms in the output molecule on each channel, each of which corresponding to a different atom type (or chemical element), as well as the three- dimensional coordinates of each atom may be inferred from the peak (or maximum) atomic density values. In some cases, a continuous refinement may be applied to determine the local maximum of the neural fieldsuch that the coordinates of each identified atom may be refined around the neighborhood of coordinates found with the peak detector. This continuous refinement scheme may enable the detection of atomic coordinates that lie beyond the initial coarse uniform discretization.

[0144] FIG. 4B depicts a flowchart illustrating an example of a process 4B for training a molecule design computation model to perform structure-conditioned generation of three-dimensional molecules, in accordance with some example embodiments. Referring to FIGS. 1, 3A-B, and 4B, the process 450 may be performed by the training engine 110 to train the molecule design computation model 115 to perform structure-conditioned generation of three-Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 dimensional molecules. For example, as described in more details below, the molecule design computation model 115 may be trained to generate an output molecule by at least modifying, based at least on the target structure 152, the latent code 151 encoding the input molecule molecular occupancy field 150 of the input molecule. The latent code 151 is a compact representation of the input molecule molecular occupancy field 150 that imposes less computational overhead than operating directly on the input molecule molecular occupancy filed 150. As such, training the molecule design computation model 115 to operate on the latent code 151 may be advantageous at least because the compactness of the latent code 151 increases the scalability of the generative process to larger sized molecules. In some cases, the molecule design computation model 115 is also trained to modify the latent code 151 based on the target structure 152. Doing so is advantageous at least because the molecule design computation model 115 may generate the three-dimensional structure of the output molecule to be consistent with that of the target structure 152. Structural consistency, such as complementarity or conformity to the three-dimensional structure of the target structure 152, may increase the likelihood of the output molecule exhibiting one or more desired properties.

[0145] At 452, a training dataset may be generated to include, for each training sample in the training dataset, a sample target structure and a corrupted latent code encoding a corrupted molecular occupancy field (or atomic density field) of a sample molecule whose three- dimensional structure is consistent with a three-dimensional structure of the sample target structure. In some example embodiments, the training dataset for training the molecule design computation model 115 may include a plurality of training samples. In some cases, each training sample may include a corrupted latent code (or modulation code), such as the corrupted latent code 188 shown in FIG. 1. As shown in FIG. 1, in some cases, the corrupted latent code 188Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 may be generated by adding noise to the latent code 186. For example, in some cases, the latent code 186 may include a first quantity of noise (e.g., Gaussian noise and / or the like) for projecting the latent code 186 to a noisy distribution with smoother density transitions than the true distribution of the latent code 186. Accordingly, in some cases, the corrupted latent code 188 may be generated by the addition of a second quantity of noise. As described in more details below, the molecule design computation model 115 may be trained to remove the second quantity of noise but not the first quantity of noise.

[0146] In some example embodiments, each training sample in the training dataset may include a target structure whose three-dimensional structure (or conformation) is consistent with that of the sample molecule whose molecular occupancy field is encoded by the corrupted latent code (or modulation code). For example, in some cases, the target structure may include the three-dimensional structure of at least a portion of another molecule, such as an epitope, a binding pocket, and / or the like. In those instances, the uncorrupted three-dimensional structure of the sample molecule may be complementary to that of the target structure. Alternatively, the target structure may be a substructure, such as a backbone, a scaffold, and / or the like. Where the target structure is a substructure, the uncorrupted three-dimensional structure of the sample molecule may conform to that of the target structure.

[0147] At 454, the molecule design computation model 115 may be trained to recover an uncorrupted latent code encoding a molecular occupancy field (or atomic density field) of the sample molecule included with each training sample. In some example embodiments, the training of the molecule design computation model 115 may include applying the molecule design computation model 115 to modify the corrupted latent code included with each training sample. In some cases, the training of the molecule design computation model 115 may includeAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 adjusting one or more parameters (e.g., weights, biases, and / or the like) of the molecule design computation model 115 to increase the similarity between the modified latent code generated by the molecule design computation model 115 and the uncorrupted latent code (or ground truth latent code) of the sample molecule included with each training sample. The uncorrupted latent code of the sample molecule may be decoded, for example, by the conditional neural fieldimplementing the decoder 119, to the molecular occupancy field (or atomic density field) of the sample molecule. For example, in some cases, the decoder 119 may decode the uncorrupted latent code by at least mapping each point in three-dimensional space to an atomic density at the corresponding locations.

[0148] FIG. 5A depicts a schematic diagram illustrating an example of a neural field decoding a latent code, in accordance with some example embodiments. In some cases, atoms may be represented as continuous Gaussian-like shapes in three-dimensional space, centered around the atomic coordinates of each atom. A molecule may be defined as a field, or a continuous function, mapping every point in the three-dimensional space to the atomic density of each atom type, v. IR3-> ]Rn, wherein n is the number of atom types (or chemical elements) in the training dataset D. The molecular occupancy field vafor each atom type (or chemical element) a may be computed, based on Equation (3) below, by integrating the occupancy generated by all atoms of this type.In Equation (3), atis the ithatom of type (or chemical element) a, for a total of naatoms. Each atom’s radium may be set to r = .5A for all atom types (or chemical elements). Molecular fieldsAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 are smooth functions taking values between 0 (for a location far away from all atoms) and 1 (for a location at the center of an atom).

[0149] As shown in FIG. 5A, in some example embodiments, the molecular occupancy field v, such as the input molecule molecular occupancy field 150, may be approximated (or parameterized) by a conditional neural field (or neural network) / >: IR3x IRd-> IR". Moreover, the input molecule may be mapped to the latent code 151. In FIG. 5 A, the conditional neural fieldmay serve as the decoder 119 whose parameters < > are set during training such that the neural fieldis able to determine, based on the latent code (or modulation code) z of a molecule, a mapping between a point in three-dimensional space to the atomic densities of the molecule. In FIG. 5A, the conditional neural fieldmaps a location in three-dimensional space identified by the coordinates (x, y, z) to a value d indicative of the atomic density at the location. For example, the decoder 119 may be trained based on the training dataset T> in which each molecule is mapped to a latent code (or modulation code) z E IRd.

[0150] Referring again to FIG. 5A, the decoder 119 (e.g., the conditional neural field may map points in three-dimensional space, each of which being identified by the coordinates (x, y, z), to the atomic density value d at the corresponding location. The value d may have a value between 0 (if the location (x, y, z) is far away from all atoms) and 1 (if the location (x,y, z) is at the center of an atom). In the example shown in FIG. 5A, a voxelized molecule representation 500 of the input molecule molecular occupancy field 150 may be encoded by the encoder 111 to form the latent code 151. In some cases, the voxelized molecule representation 500 may be a low-resolution voxel grid that is generated by voxelizing the input molecule molecular occupancy field 150. Voxelization in this context may include determining,Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 for each voxel in a voxel grid (e g., low-resolution voxel grid), an atomic density value. As described in more details below, in some cases, the encoder 111 may be a three-dimensional convolutional neural network (CNN) that includes one or more residual blocks, each of which containing one or more convolution layers followed by BatchNorm, ReLU, and pooling layers (e g., max pooling layers, average pooling layers, and / or the like). In some cases, the latent code 151 may be unique to the three-dimensional structure of the input molecule. Moreover, as shown in FIG. 5A, the latent code 151 may be decoded by the decoder 119 (e.g., the conditional neural field denoted asin FIG. 5 A) to the atomic density d for any point (x, y, z) in space.

[0151] In some example embodiments, the neural field ftp may be implemented as a multiplicative filter network (MFN). In some cases, the objective for training the decoder 119 may include learning the parameters of the decoder 119 (e.g., the parameters of the neural field ftp implemented as a multiplicative filter network (MFN)) and the modulation code z such that for any molecular occupancy field v and coordinate x G IR3, the decoder 119 (e.g., the neural field ftp may be able to determine the atomic density at the location indicated by the coordinate x (e.g., f(p(x, z') = v(x)). In some cases, the molecular occupancy fields v may be approximated by the decoder 119 (e.g., the neural field fp) with a linear combination of an exponential large number of parameterized basis function K, such that amplitudes are modulated by the individual latent codes (or modulation codes) z as shown below.bias, wherein Akand skdenote the amplitudes and basis functions, respectively. This approximation (or parameterization) of the neural fieldmay be achieved by modeling the decoder 119 (e.g., the neural fieldwith the multiplicative filter network (MFN). In some cases, theAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 multiplicative filter networkmay be a type of coordinate-based network capable of performing this linear combination under some assumptions on the basis functions. The parameters of these functions are introduced in Equation (5) below.

[0152] FIG. 5B depicts a block diagram illustrating an example of a conditional multiplicative filter network 525 for implementing a neural field, such as the conditional neural field f(f> implementing the decoder 119 shown in FIG. 5 A. In some the example shown in FIG. 5B, the conditional multiplicative filter network 525 may include an L quantity of multiplicative blocks In this context, “conditioning” refers to adjustments made to the parameters of the conditional multiplicative filter network 525 based on additional information or context. In some cases, the conditioning of the parameters of the conditional multiplicative filter network 525 may be implemented with feature-wise linear modulation (FiLM) layers, although other conditioning strategies are also possible. An example of a multiplicative block 550 (or the Ithmultiplicative block is shown in FIG. 5C. As shown in FIG. 5C, the multiplicative block 550 may include a fully-connected layer (denoted FC), a feature-wise linear modulation (FiLM) layer, and an element wise product (denoted 0) with the spatial basis function S^D . The neural field that is modeled by the multiplicative filter network 525 with one or more of the multiplicative block 550 may be expressed by the following recursive expression. / lCo)(x) = S^o) (x),wherein the spatial basis function s is parameterized by to®, 0 denotes the Hadamard product, and= W^l~^z are the bias and scale modulation terms.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1The feature-wise linear modulation (FiLM) layers are not conditioned on the spatial basis functions s^t), which guarantees that the amplitudes Akdepend on the latent (or modulation) code z. As described in more details below, during training, the parameters of the neural field and the latent (or modulation) codes z (one per each moleculein the training dataset D) may be learned through either auto-decoding or auto-encoding.

[0153] To learn the parameters of the neural field (pand the latent (or modulation) codes z (one per each molecule in the training dataset 2)) through auto-decoding, each latent (or modulation code) z may be initialized randomly and learned directly (together with the parameters < >) through backpropagation. This may be achieved by solving the optimization problem expressed as Equation (4) below. argwherein the integral may be approximated by sampling finite sets of points X c IR3. While auto-decoding is typically applied in settings with relatively few samples, the training may be scaled to larger datasets (e.g., containing one million sample molecules).

[0154] With auto-encoding, the latent (or modulation) codes z may be generated by an encoder parameterized by parameters ip before being decoded by the corresponding neural field to the corresponding molecular occupancy fields. In some cases, the encodermay be a trainable three-dimensional convolutional network encoder that ingests voxel grids Q (low- resolution voxel grids) as inputs. The auto-encoding approach may be flexible and compatible with other encoder architectures operating on other molecule representations (e.g., graph neuralAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 networks (GNNs) and point clouds). The parameters i of the encoder nd the parameters (p of the neural fieldmay be learned with the following objective:Once the training of the encoderis complete, the latent (or modulation) codes z may be generated with the trained encoderInstead of learning the latent (or modulation) codes z individually, the auto-encoding approach learns the encoderwhich allows data augmentation to be leveraged more efficiently. As a result, the auto-encoding approach supports the learning of a more structured latent space.

[0155] In some example embodiments, the latent (or modulation) codes z and the neural field provide access to the (learned) continuous molecular occupancy field fp . z). In some cases, the three-dimensional conformation of the molecules may be extracted from the latent (or modulation) codes z by at least identifying the atoms in the corresponding molecular occupancy fieldstheir approximate locations, and type (chemical element). To do so, a discretized voxel grid may be determined from the molecular occupancy field v using a uniform discretization of space and the neural fieldz). A peak finding algorithm may be applied to infer the number of atoms in the molecule on each channel of the grid (each of which representing a different atom type (or chemical element)) and their (discretized) coordinates. A continuous refinement may be applied to find the local maximum of the neural fieldFor each identified atom a, its coordinates may be refined around the neighborhood of the coordinates found with the peak detector: xa= arg max [f^(x,z)] , xE IR3: || x— x° || <rAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 wherein [ / ^(x, z)]^ denotes the field restricted to the channel corresponding to the atom type (or chemical element). This continuous refine may be capable of identifying coordinates beyond the initial coarse uniform discretization. In practice, the refinement process may be batched across molecules while leveraging an optimization algorithm Boyden-Fletcher-Goldfarb-Shannon (BFGS) (e.g., limited memory BFGS (L-BFGS)) to estimate the parameters.

[0156] As noted, the conditional neural fieldapproximation of the continuous molecular occupancy field v may offer a number of advantages when used to represent the three- dimensional structure of large sized molecules. For example, conditioning the multiplicative filter network 525 shown in FIG. 5B provides the flexibility to choose any type of spatial basis that satisfies a multiplicative-sum property. In some cases, setting the spatial basis to filters (e.g., Gabor filters) that account for the sparse nature of molecular occupancy fields may give rise to better performance than other filters (e.g., Fourier filters). For instance, for each layer I of the multiplicative filter network 525, Gabor parameterization may be expressed as Equation (6) below:wherein denotes the mean of the Gabor filter,denotes the scale, fl® denotes the frequency, and (■,■) denotes the concatenation operator. Equation (6) combines real and imaginary parts of the complex Gabor filter. The use of the Gabor filters permits the removal of phase parameters, thus reducing the overall parameter count of the multiplicative filter network 525.

[0157] The overall conditional formulation of the multiplicative filter network 525 is parameter efficient and shares parameters across molecules and channels (for different atomAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 types (or chemical elements)). In some cases, the parameters m of the spatial basis functions may be excluded from the feature-wise linear modulation (FiLM) layers in each multiplicative block 550 to further reduce the parameter count. In some cases, the conditional neural field may be trained on any free-form discretization of the input field. Molecular occupancy values may be computed on-the-fly, thus allowing the conditional neural fieldto be trained with large batch sizes and on large three-dimensional molecules. In some cases, upsampling points near the center of each atom improved training time. The neural fields v accurately reconstructs the input data. Operating on the latent (or modulation codes) z makes the model extremely robust to noise in the latent space. As described in more details below, a generative process that includes sampling latent (or modulation) codes z followed by decoding them into molecules achieves at least one order magnitude faster molecule sampling time than other state-of-the-art methods.

[0158] FIG. 6 depicts a schematic diagram illustrating an example of a process 600 for sampling latent (or modulation) codes from a noisy data distribution populated by noisy latent codes, in accordance with some example embodiments. In FIG. 6, p(z) denotes the distribution of latent (or modulation) codes z (e.g., the latent code 151) while p(v) denotes the (unknown) distribution) of molecular occupancy fields v, defined more formally as the pushforward of p(z) via the mapping z i-> f z'). The function (e.g., score function) approximating the smoothed densities of the codes p(y), gg(y) ~ V logp(z) may be estimated via neural empirical Bayes (NEB). The function gg(e.g., score function) may output values indicative of the local change in the density of the distribution p(y). Sampling from the smoothed distribution p(y), which may be guided by the function ge(e.g., score function), benefits from faster mixing than sampling from the original density p(z). This smoothed densityAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 distribution p(y) may be defined by transforming the random variable Z with the addition of noise A (e.g., isotropic Gaussian noise) with a known noise levelY = Z + N, where N~N (0, The noise level o may present a tradeoff between the simplicity of the denoising objective and sampling quality.

[0159] Neural empirical Bayes (NEB) may be based on an empirical Bayes view of (denoising) score-based models that relates the estimator of clean data (e.g., denoiser) and the score function g9of the smoothed distribution p(y) at a fixed noise level. In some cases, the denoiser may be the least-square estimator of Z given Y = y, which is the Bayes estimator (e.g., z(y) = E[Z | Y = y]. Under Gaussian noise, the denoiser and the smoothed score function may be related by z(y) = y + o-2V logp(y). (7)

[0160] The denoiser may be parameterized by a neural network and learned by minimizing the following objective. The score function g0may be recovered from the learned denoiser via Equation (8) and used for sampling soothed codes. In practice, the empirical loss may be optimized based on the latent codes inferred from a set of molecular fields in the training dataset 2).

[0161] In some example embodiments, the score function g9may guide the sampling of noisy latent (or modulation) codes z from the noisy distribution p(z) using a walk-jump sampling (WJS) scheme. This approach samples molecules from the noisy distribution p(z) using the score function g9of noisy latent (or modulation) codes z instead of clean latent (or modulation) codes z. As shown in the process 600 in FIG. 6, in some cases, the walk-jump sampling (WJS) scheme includes “walking” the noisy distribution p(z) to sample noisy latent (orAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 modulation) codes z therefrom and “jumping” from the noisy distribution p(z) back to the true data distribution of three-dimensional molecules by decoding the noisy latent (or modulation) codes z into the noisy molecular occupancy fields vkrepresentative of the three-dimensional structure of the corresponding molecules. For example, the first noisy latent codein FIG. 6 corresponds to the molecular occupancy fieldof a molecule, including by encoding the type and location of the atoms in this molecule. For visual clarity, at least some of those atoms are annotated with circles. In some cases, the “walking” of the noisy distribution p(z) may include sampling the second noisy latent code z2therefrom. The second noisy latent code z2may encode the molecular occupancy field v2of a different molecule. In the example show in FIG. 6, this difference between the molecules associated with the molecular occupancy fieldsand v2include atoms of different types and at different locations. In other words, sampling the second noisy latent code z2from the noisy data distribution may be akin to generating a different molecule than sampling the first noisy latent code z from the noisy data distribution.

[0162] In some cases, mixing may be improved by initializing the sampling chains by adding uniform noise to Gaussian noise (with the same noise level a used when training the denoiser). In practice, the uniform noise may be defined over the range of code values, e.g., the trainingdataset of the latent (or modulation) codes. Noisy latent (or modulation) codes z may be sampled from the noisy distribution p(y) with gradient-based Markov Chain Monte Carlo (e.g., Langevin Markov Chain Monte Carlo (MCMC)) that discretize the underdamped Langevin diffusion starting from y0and u0= 0:Attomey Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 where Btis the standard Brownian motion in Mriand y is the friction (the “mass” is set to 1). This stochastic differential equation (SDE) using the ABOBA scheme, given a discretization step 8 and a fixed number of walk steps K. At a given time step K, clean samples may be estimated by denoising the noisy latent (or modulation) code z (e.g., zK= z0(yK)), which are then used to obtain the atomic coordinates of the corresponding molecules.

[0163] Referring again to FIG. 6, the noisy distribution p(y) may be “walked” while guided by the function gesuch that successive samples are drawn from incrementally higher density regions of the noisy distribution p(y). For example, in some cases, the function gemay assign, to each noisy latent (or modulation) code zk, a value indicative of the local change in the density of the region of the noisy distribution p(y) from which the noisy latent (or modulation) code zkis drawn. The higher density regions of the noisy distribution p(y) may be populated by molecules exhibiting one or more desired properties, meaning that the molecule corresponding to the noisy latent code ykmay be more likely to exhibit the one or more desired properties than the molecule corresponding to previous noisy latent code yk-^. This is the “walk” of the noisy distribution p(y) shown in FIG. 6 while the “jump” corresponds to the decoding of the noisy latent (or modulation) code zkto the corresponding molecular occupancy field vkby the neural field As noted, in some cases, the neural fieldmay be a neural network (with the parameters < >) that approximates (or parameterizes) the molecular occupancy field vkby at least mapping each point (%, y, z) to the atomic density d at the corresponding location in three- dimensional space.

[0164] FIGS. 7A-B depict a schematic diagrams illustrating two examples of an output molecule 700 whose generation is conditioned on the target structure 152, in accordance with some example embodiments. In FIG. 7A, the target structure 152 include the three-Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 dimensional structure of at least a portion of another molecule, such as an epitope, a binding pocket, and / or the like. Accordingly, the example of the output molecule 700 generated in FIG. 7A may exhibit a three-dimensional structure (or conformation) that is complementary to the three-dimensional structure (or conformation) of the target structure 152. Alternatively, the target structure 152 in FIG. 7B may be a substructure, such as a scaffold, a backbone, and / or the like. In this example, the output molecule 700 may exhibit a three-dimensional structure (or conformation) that conforms to the three-dimensional structure of the target structure 152.

[0165] Experimental Examples

[0166] Various examples of the molecule design computation model described herein, which operates on latent codes encoding the molecular occupancy (or density field) representative of the three-dimensional structure of molecules, may be trained to perform three different generation tasks: small molecule generation, antibody complementarity determining (CDR) loop redesign conditioned on an epitope, and macro cyclic peptide generation conditioned on a protein pocket.

[0167] Small Molecule Generation

[0168] The small molecule generation task leveraged benchmark data from a test set containing 100 protein pockets.

[0169] The performance of the molecule design computation model (MDCM) described herein is evaluated against other models designed for pocket-conditioned ligand generation, including autoregressive models operating on point clouds (AR and Pocket2Mol), diffusion models operating on point clouds (DiffSBDD, TargetDiff, and DecompDiff), and a model based on Bayesian flow networks (MolCraft). For each model, 100 ligands were sampled per pocket. Affinity was measured with three metrics using AutoDock Vina: VinaScoreAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 computes the binding affinity (docking) score of the molecule on its original generated pose, VinaMin performs a local energy minimization on the ligand followed by dock scoring, and VinaDock performs full redocking (search and scoring) between ligand and target. A druglikeness (QED) score and a synthesizability (SA) score were measured for each generated molecule. Diversity was computed by averaging the Tanimoto distance of every pair of generated ligand per pocket. The number of atoms per molecule is the average number of (heavy) atoms per molecule. Additional metrics include steric clash, which computes the number of clashes between generated ligands and the associated pockets, and strain energy (SE), which measures the difference between the internal energy of the generated molecule’s pose (without pocket) and a relaxed pose. The results are reported in Table 1 below. As shown in Table 1, the molecule design computation model described herein achieved competitive performance compared to other state-of-the-art models.

[0170] Table 1

[0171] Antibody CDR Redesign

[0172] The antibody complementarity determining region (CDR) redesign task used the SabDab dataset, which includes antibody structures in complex with proteins, as well as the training, validation, and test splits from DiffAb. This split ensured that antibodies similar toAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 those in the test set (e.g., >50% CDR H3 identity) were removed from the training set. The testing split includes 19 targets, for which each CDR loop was individually redesigned.

[0173] The performance of the molecule design computation model (MDCM) described herein is evaluated against representative baselines: Rosetta (RAbD) and some machine learning based generative models: DiffAb and AbDiffuser. Additional comparisons were made against AbX and dyMEAN in the setting of H3 deisgn, for which the DiffAb splits were considered. Table 2 presents the results.

[0174] For each target, 100 molecule with unique sequences were sampled with the molecule design computation model described herein. Performance metrics were computed for each of the three complementarity determining regions (CDRs) on the heavy chain (e.g., Hl, H2, and H3 in Table 2) and the three complementarity determining regions (CDRs) on the light chain (e.g., LI, L2, and L3 in Table 2). The metrics computed include amino acid recovery of the seed (AAR) measured by the sequence identify between the reference CDR sequences and the generated ones, the alpha carbon (Ca) root-mean-square deviation (RMSD) between the generated structure and the original structure with antibody frameworks aligned with IMP, and the percentage of designed CDRs with lower (better) binding energy (AG) than the original CDR. The binding energy was calculated by InterfaceAnalyzer in Rosetta. Existing models apply some Rosetta-based relaxations prior to computing IMP to improve energy scores: DiffAb refines the structure with OpenMM and AbX uses FastRelax. The moleculd design computation model’s IMP is reported for unrelaxed designs. Table 2 below present the results. As shown in Table 2, the molecule design computation model is competitive with respesct to AAR, outperforming other models on RMSD. In particular, the molecule design computation model described herein outperforms in loops with more variability: H3 and L3. Despite any sort of relaxation, theAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 molecule design computation model achieved competitive IMP scores to approaches that relax their structures.

[0175] Table 2Method AAR T RMSD 1 IMP THl RAbD 22.9 2.26 43.9DiffAb 65.8 1.19 53.6AbDiffuser 76.3 1.58MDCM 69.9 0.47 30.3H2 RAbD 25.5 1.64 53.5DiffAb 49.3 1.08 29.8AbDiffuser 65.7 1.45MDCM 52.0 0.61 32.3H3 RAbD 22.1 2.90 23.3DiffAb 26.8 3.60 23.6DyMEAN 29.3 4.80 5.26AbX 30.3 3.41 42.9AbDiffuser 34.1 3.35MDCM 42.4 1.89 38.6Method AAR T RMSD 1 IMP TLI RAbD 34.3 1.20 46.8DiffAb 55.7 1.39 45.6AbDiffuser 81.4 1.46MDCM 78.6 0.86 49.7L2 RAbD 26.3 1.77 56.9DiffAb 59.3 1.37 50.0AbDiffuser 83.2 1.40MDCM 60.5 0.91 34.7L3 RAbD 20.7 1.62 55.6DiffAb 46.5 1.63 47.3AbDiffuser 73.2 1.59MDCM 66.0 0.91 38.0

[0176] Macro-cyclic Peptide GenerationAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1

[0177] Due to the lack of macro-cyclic peptide (MCP) and protein complex structures, a dataset of 186,685 MCP-protein complexes were curated following a “mutate then relax” strategy on an original set of 641 protein-MCP complexes sourced from the RCSB PDB. The source dataset includes MCP lengths ranging from 4 to 25 amino acid with an average of 10. Moreover, 78% of the MCPs in the source dataset contain one or more non-canonical amino acid residues, e.g., any amino acid residue that is neither L-canonical nor D-canonical. The source dataset was split into training, validation, and testing using a clustering approach.

[0178] The performance of the molecule design computation model (MDCM) described herein in generating macro-cyclic peptides was evaluated by calculating Tanimoto similarity (TS), which assesses the resemblance between the seed MCP and the sampled structures generated by the molecule design computation model (MDCM) modifying the latent code of the seed MCP.

[0179] Table 3 presents the Tanimoto similarity between the seed MCP and 20 sampled MCP molecules for each pocket. Pockets from the validation set were either relaxed to MCP mutants (labeled as mut# in Table 3) or derived from crystal structures (labeled as CP in Table 3).

[0180] Table 3Seed Tanimoto Sim lvwi_mut308 0.35+0.05 lvwe-mut90 0.39+0.07 lwb0-mut339 0.40±0.07 lwb0-mut389 0.53+0.132x7k_mut245 0.32±0.064gw4_mutl62 0.34+0.044gw5_mutl04 0.50+0.064mnw_CP 0.37+0.03Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-14os2_mut430 0.41±0.045jme_CP 0.39±0.037k2e_mutl51 0.31+0.037k2e_mut348 0.32+0.017pml_mut384 0.33+0.09

[0181] Additonal performance metrics for the MCP generation task include ligand RMSD (L-RMSD), which is the RMSD between the sampled MCP to the seed MCP calculated based on the backbone atoms (N, Ca, C, O). The same backbone logic was applied to compute interface RMSD (I-RMSD), which is the RMSD in the pocket. This metric is generally lower due to the pockets being identical. Moreover, the sampled MCP and the seed MCP were not aligned for I-RMSD since this metric is based on where the MCPs are in the pocket. The average L-RMSD was 2.71 and the average I-RMSD was 1.84 with per target RMSD reported in Table 4, which shows a comparison of the L-RMSD and I-RMSD values of various PDB entries.

[0182] Table 4PDB L-RMSD I-RMSD lbm2 2.17 ± 0.89 1.22 ± 0.28Ibxo 1.10 + 0.40 0.81 + 0.28 ljd2 2.36 + 0.32 3.19 + 1.22Ikmh 1.91 + 0.37 2.37+ 0.552bdx 6.19 + 0.20 2.63+ 0.392gpl 2.14 + 0.31 3.09+ 0.663dv5 2.46 + 1.42 1.09+ 0.403egh 3.03 + 0.53 1.68 + 0.323rqd 1.79 + 0.04 3.95+ 0.903s04 4.46 ± 0.87 2.97+ 0.453sud 2.42+ 0.55 1.45 + 0.193v3b 5.09 + 1.32 2.99+ 0.904cli 0.53 + 0.14 0.65+ 0.20Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-14dri 1.31 ± 0.18 0.96± 0.234dru 0.86 ± 0.10 0.73 ± 0.174i51 3.24 ± 0.77 2.25± 0.374jna 1.43 ± 0.22 1.05±0.174n7y 4.07 ± 1.21 2.14 ± 0.474twt 4.49 ± 0.80 3.58± 0.834wvu 3.41 ± 1.04 2.64 ± 0.724,\v9 2.96 ± 0.55 2.13± 0.685agu 4.78 ± 1.07 2.22± 0.465bmm 2.64 ± 0.12 1.24 ± 0.125cs2 4.25 ± 1.78 2.57± 0.685j31 3.92± 0.39 2.00± 0.445jh7 0.38± 0.24 0.95± 0.195jm4 4.60 ± 0.51 2.60± 0.4551rg 2.21 ± 0.89 1.10 ± 0.405mev 0.47 ± 0.19 0.80± 0.325ooc 1.86 ± 0.13 1.22 ± 0.32Average 2.71 ± 1.66 1.84 ± 1.08

[0183] Sampling within test pockets exhibits a correlation with the MCP seeds. This phenomenon is shown in FIG. 8, which illustrates the close alignment of the backbone between the sample MCP structures and the seed MCP structure. Most of the sampled MCP molecules display consistent repeating peptide bonds, linking the Ci carbon of one a-amino acid residue to the Nz nitrogen of the next amino acid residue. Closure bonds (such as the disulfide in panels a and b of FIG. 8 and the TV to C cyclization in panel c) were also maintained in the sampled MCP molecules generated by the molecule design computation model.

[0184] The sampled MCP molecules generated by the molecule design computation model exhibited Tanimoto similarity (TS) scores from 0.31 to 0.53. These mid-range scores reflect a strong similarity to the peptide backbone, with variability occurring at the functionalAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 groups of the residues. Table 5 presents the per-residue Tanimoto similarity (TS) between the seed MCP (panel (a) of Table 5) and the crystal MCP (panel (b) of Table 5) of the 20 sampled MCP molecules in the pocket relaxed o Ivwe mutant 90. An example of per-residue Tanimoto similarity for MCP molecule sampled with the seed mutant (Table 5(a)) and the crystal MCP (Table 5(b)) demonstrated that the highest Tanimoto similarity (TS) occured at the disulfide closure bond. This elevated per-residue similarity resulted from the preservation of closure bond residues throughout the curated MCP dataset. Furthermore, the per-residue Tanimoto similarity (TS) was higher across crystal MCP residues than at the mutated residues of the seed MCP mutant. Because all mutants originate from the crystal MCP, the dataset was closely tied to the crystal sequence, and the sampled MCP structures similarly reflected this connection.

[0185] Table 5(a) CYS HIS A20 B67 PHE CYS lvwe_mut90 0.63±0.24 0.24±0.17 0.26±0.07 0.14±0.12 0.29±0.14 0.53±0.21(b) CYS HIS PRO GLU PHE CYSIvwe-CP 0.61+0.23 0.24+0.12 0.31±0.14 0.28+0.10 0.33+0.08 0.57±0.23

[0186] FIG. 9 depicts a block diagram illustrating an example of a computing system 900, in accordance with some example embodiments. Referring to FIGS. 1-9, the computing system 900 may be used to implement the molecule design engine 110, the training engine 120, the client device 130, and / or any components therein.

[0187] As shown in FIG. 9, the computing system 900 can include a processor 910, a memory 920, a storage device 930, and input / output devices 940. The processor 910, the memory 920, the storage device 930, and the input / output devices 940 can be interconnected via a system bus 950. The processor 910 is capable of processing instructions for execution within the computing system 900. Such executed instructions can implement one or more componentsAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 of, for example, the molecule design engine 1 10, the training engine 120, the client device 130, and / or the like. In some example embodiments, the processor 910 can be a single-threaded processor. Alternately, the processor 910 can be a multi -threaded processor. The processor 910 is capable of processing instructions stored in the memory 920 and / or on the storage device 930 to display graphical information for a user interface provided via the input / output device 940.

[0188] The memory 920 is a computer readable medium such as volatile or nonvolatile that stores information within the computing system 900. The memory 920 can store data structures representing configuration object databases, for example. The storage device 930 is capable of providing persistent storage for the computing system 900. The storage device 930 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 940 provides input / output operations for the computing system 900. In some example embodiments, the input / output device 940 includes a keyboard and / or pointing device. In various implementations, the input / output device 940 includes a display unit for displaying graphical user interfaces.

[0189] According to some example embodiments, the input / output device 940 can provide input / output operations for a network device. For example, the input / output device 940 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0190] In some example embodiments, the computing system 900 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 900 can be used to execute any type of software applications. These applications can be used to performAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 940. The user interface can be generated and presented to a user by the computing system 900 (e.g., on a computer screen monitor, etc.).

[0191] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0192] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the termAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1“machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid- state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0193] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED) monitor, or an organic light emitting diode (OLED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, opticalAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.

[0194] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.

[0195] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed featuresAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desired results. Other implementations may be within the scope of the following claims.

Claims

Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising: receiving an input molecule, the input molecule being associated with a molecular occupancy field representative of a three-dimensional structure of the input molecule, where the molecular occupancy field of the input molecule maps a point in three- dimensional space to an atomic density of the input molecule at a corresponding location in three-dimensional space; generating a latent code of the input molecule by at least applying a computation model to encode, into the latent code of the input molecule, the molecular occupancy field of the input molecule; and generating an output molecule by at least modifying the latent code of the input molecule.

2. The method of claim 1, wherein the atomic density of the input molecule at the corresponding location in three-dimensional space comprises a distance between that location and a center of an atom forming the input molecule.

3. The method of any of claims 1 to 2, wherein the molecular occupancy field comprises a continuous function mapping, based at least on the latent code of the input molecule, the point in three-dimensional space to the atomic density of the input molecule at the corresponding location in three-dimensional space.

4. The method of any of claims 1 to 3, wherein the computation model includes an encoder that generates the latent code of the input molecule by at least encoding the molecular occupancy field of the input molecule.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-15. The method of claim 4, wherein the computation model further comprises a decoder that decodes the latent code of the input molecule to recover the molecular occupancy field of the input molecule, and wherein the decoder recovers the molecular occupancy field of the input molecule by at least determining, based at least on the latent code of the input molecule and a coordinate of the point in three-dimensional space, the atomic density of the input molecule at the corresponding location in three-dimensional space.

6. The method of any of claims 1 to 5, further comprising: generating a noisy latent code by at least adding noise to the latent code of the input molecule; and applying a molecule design computation model to modify the noisy latent code of the input molecule.

7. The method of claim 6, further comprising: applying a molecule design computation model to modify the noisy latent code over one or more successive iterations.

8. The method of claim 7, wherein the molecule design computation model is trained to approximate a noisy data distribution of molecules exhibiting one or more desired properties, wherein the molecule design computation model is trained based at least on a plurality of noisy latent codes of sample molecules exhibiting one or more desired properties, and wherein the modifying of the noisy latent code comprises sampling one or more noisy latent codes from the noisy data distribution.

9. The method of claim 8, wherein the molecule design computation model is applied to modify the noisy latent code and sample one or more noisy latent codes from the noisy data distribution until one or more criteria are met.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-110. The method of any of claims 1 to 9, further comprising: determining a molecular occupancy field of the output molecule by at least applying the computation model to decode a modified latent code resulting from the modifying of the latent code of the input molecule.

11. The method of claim 10, further comprising: determining, based at least on the molecular occupancy field of the output molecule, a three-dimensional structure of the output molecule.

12. The method of claim 11, wherein the three-dimensional structure of the output molecule is determined by at least rendering, based at least on the molecular occupancy field of the output molecule, a voxel grid comprising a plurality of voxels, wherein each voxel of the plurality of voxels is associated with an atomic density value at a corresponding location in three-dimensional space; identifying one or more peak atomic density values in a plurality of atomic density values across the voxel grid; and determining, based at least on one or more peak atomic density values, a quantity and a location of one or more atoms present in the output molecule.

13. The method of claim 12, wherein the rendering the voxel grid includes applying the computation model to determine, for each voxel in the voxel grid, the atomic density value at the corresponding location in three-dimensional space.

14. The method of any of claims 1 to 13, wherein the computation model comprises a neural field or a conditional neural field.

15. The method of any of claims 1 to 14, wherein the computation model comprises an autoencoder, wherein the autoencoder includes an encoder that encodes the molecularAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 occupancy field of the input molecule, and wherein the autoencoder further includes a decoder that recovers a molecular occupancy field of the output molecule by at least decoding a modified latent code resulting from the modifying of the latent code of the input molecule.

16. The method of claim 15, wherein the encoder comprises a three-dimensional convolutional neural network, and wherein the decoder comprises a multiplicative filter network.

17. The method of any of claims 1 to 16, wherein the computation model encodes the molecular occupancy field of the input molecule such that the latent code of the input molecule is unique to the input molecule and captures one or more structural features unique to the input molecule.

18. The method of any of claims 1 to 17, wherein the computation model is trained to model one or more common molecular features including one or more of bonds, angles, valencies, and symmetries.

19. The method of any of claims 1 to 18, wherein the molecular occupancy field of the input molecule comprises a voxel grid having a plurality of voxels, and wherein each voxel of the plurality of voxels is associated with an atomic density value of the input molecule at a corresponding location in three-dimensional space.

20. The method of any of claims 1 to 19, wherein the molecular occupancy field of the input molecule includes a plurality of channels, and wherein each channel of the plurality of channels corresponds to a different type of atom present in the input molecule.21 . A computer-implemented method for generating an output molecule having one or more desired properties, the method comprising: determining a three-dimensional target structure representation of a target structure; generating a noisy latent code;Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 applying a molecule design computation model to generate a modified latent code, wherein the modified latent code encodes an output molecule molecular occupancy field of the output molecule, wherein the molecule design computation model generates the modified latent code by at least modifying, based at least on the three-dimensional target structure representation, the noisy latent code, and wherein the molecule design computation model modifies the noisy latent code such that a three-dimensional structure of the output molecule encoded by the modified noisy latent code is consistent with a three-dimensional structure of the target structure; and decoding the modified latent code to determine the output molecule molecular occupancy field of the output molecule, wherein the decoding includes determining, based at least on the modified latent code, an atomic density exhibited by the output molecule at one or more points in three- dimensional space.

22. The method of claim 21, wherein the target structure includes at least a portion of another molecule, and wherein the molecule design computation model modifies the noisy latent code such that a three-dimensional structure of the output molecule is complementary to a three- dimensional structure of the target structure.

23. The method of claim 22, wherein the target structure comprises an epitope or a binding pocket.

24. The method of any of claims 21 to 22, wherein the target structure includes a substructure, and wherein the molecule design computation model modifies the noisy latent codeAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 while conditioned on the target structure such that a three-dimensional structure of the output molecule encoded by the modified noisy code conforms to a three-dimensional structure of the target structure.

25. The method of claim 24, wherein the substructure comprises a scaffold of a chemical compound or a backbone of a protein molecule.

26. The method of any of claims 21 to 25, further comprising: applying a decoder to decode the modified latent code, wherein the decoder comprises a neural field trained to approximate the output molecule molecular occupancy field by at least mapping, based at least on the modified latent code, the one or more points in three-dimensional space to an atomic density of the input molecule at a corresponding location in three-dimensional space.

27. The method of claim 26, wherein the neural field comprises a neural network approximating a continuous function that maps the one or more points to the atomic density of the output molecule at one or more corresponding locations in three-dimensional space.

28. The method of any of claims 26 to 27, wherein the neural field maps the one or more points to the atomic density of the output molecule at any arbitrary resolution.

29. The method of any of claims 26 to 28, wherein a location near a center of an atom in the output molecule has a higher atomic density than a location far away from all atoms in the output molecule.

30. The method of any of claims 21 to 29, further comprising: encoding the three-dimensional target structure representation to generate a target structure embedding; encoding the noisy latent code to generate a latent code embedding;Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 generating a joint embedding by at least combining the target structure embedding and the latent code embedding; and applying the molecule design computation model to generate the modified latent code by at least operating on the joint embedding.

31. The method of claim 30, wherein the target structure embedding and the latent code embedding are encoded to occupy a common embedding space and have a same spatial dimensions.

32. The method of any of claims 21 to 31, wherein the noisy latent code is generated to include random noise.

33. The method of any of claims 21 to 32, wherein the molecule design computation model samples noisy latent codes from a noisy data distribution by modifying the noisy latent code.

34. The method of claim 33, wherein the molecule design computation model samples noisy latent codes from the noisy data distribution while guided by a function that outputs a value indicative of a local change in a density of the noisy data distribution at a location of each noisy latent code sampled by the molecule design computation model.

35. The method of claim 34, wherein the molecule design computation model is applied to sample, based on the value output by the function, noisy latent codes from incrementally higher density regions of the noisy data distribution.

36. The method of claim 35, wherein higher density regions of the noisy data distribution are populated with noisy latent codes of molecules more likely to exhibit one or more desired properties while lower density regions of the noisy data distribution are populated by noisy latent codes of molecules less likely to exhibit the one or more desired properties.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-137. The method of any of claims 34 to 36, further comprising: applying the molecule design computation model to modify the noisy latent code to generate a first noisy latent code, applying the molecule design computation model to modify the noisy latent code to generate a second noisy latent code, applying the function to determine a first value indicative of a first local change in the density of the noisy data distribution at a first location occupied by the first noisy latent code, applying the function to determine a second value indicative of a second local change in the density of the noisy data distribution at a second location occupied by the second noisy latent code, and applying the molecule design computation model to further modify, when the first value and the second value indicates the first noisy latent code is sampled from a higher density region of the noisy data distribution than the second noisy latent code, the first noisy latent code and not the second noisy latent code.

38. The method of any of claims 32 to 37, further comprising: denoising the modified latent code prior to decoding the modified latent code.

39. The method of any of claims 21 to 38, further comprising: determining, based at least on the output molecule molecular occupancy field of the output molecule, a three-dimensional structure of the output molecule, wherein the determining the three-dimensional structure of the output molecule includes determining a type and a location of each atom in the output molecule.

40. The method of claim 39, wherein the determining the three-dimensional structure of the output molecule includesAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 determining, based at least on the output molecule molecular occupancy field, a voxel grid in which each constituent voxel is associated with a value corresponding to an atomic density at a corresponding location; identifying one or more peak atomic density values across the voxel grid; and determining, based at least on the one or more peak atomic density values, one or more coordinates of each atom in the output molecule.

41. The method of claim 40, wherein the determining the three-dimensional structure of the output molecule further includes refining the one or more coordinates of each atom in the output molecule around a neighborhood of each atom.

42. The method of any of claims 40 to 41, wherein the determining the three- dimensional structure of the output molecule further includes inferring one or more of a bond between atoms in the output molecule and an identity of an amino acid residue comprising the output molecule.

43. The method of any of claims 40 to 42, wherein the output molecular occupancy field includes a plurality of channels, wherein each channel of the plurality of channel corresponds to a different atom type, and wherein the one or more peak atomic density values are identified for each channel.

44. The method of any of claims 21 to 43, further comprising: generating a training dataset to include, for each training sample in the training dataset, a sample target structure and a corrupted latent code encoding a corrupted molecular occupancy field of a sample molecule whose three-dimensional structure is consistent with a three- dimensional structure of the sample target molecule; andAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 training the molecule design computation model to recover an uncorrupted latent code encoding a molecular occupancy field of the sample molecule included with each training sample.

45. The method of claim 44, wherein the corrupted latent code is generated by the addition of a first quantity of noise to corrupt a molecular occupancy field of the sample molecule and a second quantity of noise projecting the corrupted latent code to a noisy data distribution, and wherein the molecule design computation model is trained to approximate the noisy data distribution by at least removing the first quantity of noise but not the second quantity of noise.

46. The method of claim 45, wherein the first quantity of noise and the second quantity of noise comprise Gaussian noise.

47. The method of any of claims 44 to 46, further comprising: applying an encoder to generate the corrupted latent code by at least encoding the corrupted molecular occupancy field of the sample molecule.

48. The method of claim 47, wherein the encoder is trained to generate the corrupted latent code by at least adjusting one or more parameters of the encoder such that the encoder generates, for each training molecular occupancy field, a latent code matching a ground truth latent code of the training molecular occupancy field.

49. The method of any of claims 47 to 48, wherein the encoder is trained to generate the corrupted latent code along with a decoder trained to decode the modified latent code, and wherein the training of the encoder and the decoder includes adjusting one or more parameters of the encoder and the decoder such that the encoder generates, for each training molecularAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 occupancy field, a latent code from which the decoder is able to recover a molecular occupancy field matching the training molecular occupancy field.

50. The method of any of claims 21 to 49, wherein the three-dimensional target structure representation of the target molecule comprises a voxel grid representation of the target structure in which one or more atoms forming the target structure are represented by discrete atomic density values across a voxel grid.

51. A computer-implemented method, comprising: receiving a voxel grid representative of a three-dimensional structure of a molecule, wherein the voxel grid includes a plurality of voxels, and wherein each voxel of the plurality of voxels is associated with an atomic density value of the molecule at a corresponding location in three-dimensional space; determining a molecular occupancy field representative of the three-dimensional structure of the molecule, wherein the molecular occupancy field of the molecule comprises a continuous function mapping each point in three-dimensional space to a corresponding atomic density value of the molecule at a location of the point in three-dimensional space; and generating a latent code of the molecule by at least encoding the molecular occupancy field of the molecule.

52. The method of claim 51, further comprising: encoding, using a computation model, the molecular occupancy field of the molecule to generate the latent code of the molecule.

53. The method of claim 52, further comprising: decoding, using the computation model, the latent code of the molecule to recover the molecular occupancy field of the molecule, wherein the computation model decodes the latentAttorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-1 code of the molecule by at least mapping, based at least on the latent code of the molecule, one or more point in three-dimensional space to atomic densities of the molecule at corresponding locations in three-dimensional space.

54. The method of claim 53, further comprising: determining, based at least on the molecular occupancy field of the molecule, the three- dimensional structure of the molecule.

55. The method of claim 54, wherein the three-dimensional structure of the molecule is determined by at least rendering, based at least on the molecular occupancy field of the molecule, the voxel grid, wherein the rendering of the voxel grid includes applying the computation model to determine, for each voxel in the voxel grid, the atomic density value at the corresponding location in three-dimensional space; identifying one or more peak atomic density values in a plurality of atomic density values across the voxel grid; and determining, based at least on one or more peak atomic density values, a quantity and a location of one or more atoms present in the molecule.

56. The method of any of claims 52 to 55, wherein the computation model is trained to generate the latent code of the molecule to capture one or more structural features unique to the molecule.

57. The method of any of claims 52 to 56, wherein the computation model is trained to is trained to model one or more common molecular features including one or more of bonds, angles, valencies, and symmetries.Attorney Ref.: 14786-083-228 (103963-228083) / P39770-WO-158. The method of any of claims 52 to 57, wherein the computation model comprises a neural field or a conditional neural field.

59. The method of any of claims 51 to 58, wherein the molecular occupancy field of the molecule includes a plurality of channels, and wherein each channel of the plurality of channels corresponds to a different type of atom present in the input molecule.

60. The method of any of claims 51 to 59, further comprising: generating one or more additional molecules by at least modifying the latent code of the molecule, wherein the generating of the one or more additional molecules includes determining, based at least on a modified latent code resulting from the modifying of the latent code of the molecule, a molecular occupancy field of each additonal molecule, and determining, based at least on the molecular occupancy field of each additional molecule, a three-dimensional structure of each additional molecule.

61. A system, comprising: at least one data processor; and at least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 60.

62. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 60.