Systems and methods for molecule structure design and synthesis
A computational framework using a trained diffusion model to position and link molecular fragments on target biomolecules addresses inefficiencies in drug discovery, producing high-scoring, synthesizable molecules with improved affinity and selectivity.
Patent Information
- Application Number
- PCT/US2025/038747
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-07-22
- Publication Date
- 2026-01-29
AI Technical Summary
Current computational drug discovery platforms are inefficient and costly due to limitations in compound libraries and optimization processes, resulting in long timelines and high costs for identifying pharmaceutical compounds that interact with target biomolecules.
A computational framework utilizing a trained diffusion model to position and link molecular fragments on a target biomolecule, considering chemical properties and geometric dimensions, and employing a neighbor relationship or edge prediction model to generate molecule structures that interact with biological targets.
The method efficiently generates synthesizable, drug-like molecules with predicted affinity and selectivity, outperforming traditional methods in generating high-scoring ligands and reducing the time and cost associated with drug discovery.
Smart Images

Figure US2025038747_29012026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR MOLECULE STRUCUTURE DESIGN AND SYNTHESISCROSS REFENCE TO RELATED DISCLOSURES
[0001] This application claims the benefit of U.S. Provisional Patent Appl. No. 63 / 674,178, entitled “Systems and Methods for Small Molecule Design and Synthesis,” filed July 22, 2024, the disclosure of which is hereby incorporated by reference in their entirety.TECHNOLOGICAL FIELD
[0002] The disclosure is generally directed to systems and methods to design ligand molecule structures via deep learning computation, including developing drug-like molecules that interact with target biomolecules.BACKGROUND
[0003] A common goal in drug discovery is to identify pharmaceutical compounds that have the potential to treat disease. Advancements in experimental and computational techniques have improved drug discovery by enabling the rapid determination of detailed structures for a wide variety of biomolecules. These structures not only illuminate the molecular mechanisms underlying the functions of these biomolecules but also serve as guides for designing therapeutics that target these biomolecules. For example, during the COVID-19 pandemic, structures of key coronavirus proteins were resolved within months of the outbreak, enabling development of new targeted therapeutics.
[0004] A typical computational drug discovery platform utilizes a digital library of molecules and computationally assesses the ability of each molecule to bind within a pocket of a target biomolecule, such as a protein or a nucleic acid. Hits are then biochemically assessed for their ability to bind the protein and / or modulate the protein’s functions. Top hits identified via biochemical assessment are then optimized chemically by medicinal chemists and then optimized compounds are reassessed via biochemical experimentation. This process, however, is inefficient due to the limitations of compound libraries and the optimization process, resulting in long timelines and high costs.SUMMARY
[0005] Several embodiments of systems and methods are for generating molecule structures that interact with a target molecule. In many embodiments, a computational framework is utilized to position functional groups upon a target biomolecule and link the functional groups to yield a molecule structure. In several embodiments, a trained diffusion model is utilized to position functional groups upon a target biomolecule. In many embodiments, the diffusion model utilizes embeddings of functional groups. In many embodiments, the diffusion model is trained using structure data of known target-ligand interactions. In several embodiments, neighbor relationships or an edge prediction model is utilized to link the positioned functional groups, yielding an intact molecule structure. In many embodiments, the linking method considers available attachment points and chemical rules to yield the molecule structure. In some embodiments, a computationally generated molecule structure is synthesized to yield a ligand compound. In some embodiments, a synthesized ligand compound is further assessed for characterization. In some embodiments, a synthesized ligand compound is for use a drug. In some embodiments, a synthesized ligand compound is administered to a recipient to treat a medical condition.
[0006] In some aspects, the techniques described herein relate to a computational method to design molecules for interacting with a biological target, the method including: providing a structure of a target biomolecule having a site to be targeted; positioning, utilizing a computational diffusion model and a library of molecular fragments, molecular fragments upon the target site; wherein the molecule fragments are numerically embedded such that it captures chemical properties of the molecule fragments; wherein the computational diffusion model is trained utilizing a cohort of structural data of ligand molecules bound to its biological target; wherein each molecule fragment of the set of molecular fragments are to be positioned simultaneously; and linking, utilizing a linking computational model, the set of molecular fragments to yield a small molecule.
[0007] In some aspects, the techniques described herein relate to a computational method, further including: acquiring the library of molecular fragments; and reducing, utilizing a dimensionality reduction technique, dimensionality of chemical properties of each fragment of the library of molecular fragments into a set of vectors.
[0008] In some aspects, the techniques described herein relate to a method, wherein the reduction technique is one of the following: principal component analysis (PCA), T- distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), singular value decomposition (SVD), or linear discriminant analysis (LDA).
[0009] In some aspects, the techniques described herein relate to a method, wherein the chemical properties include one or more of the following: charge, molecular weight, number of hydrogen bond donors / acceptors, or number of rings.
[0010] In some aspects, the techniques described herein relate to a method, wherein the dimensionality reduction technique further captures an extended three-dimensional fingerprint of the molecular fragments.
[0011] In some aspects, the techniques described herein relate to a method, wherein the computational diffusion model is a diffusion denoising probabilistic model.
[0012] In some aspects, the techniques described herein relate to a method further including: training the computational diffusion model by: acquiring the cohort of structural data of molecules bound to its binding site of a biological target; fragmenting each molecule of the cohort into molecular fragments, wherein each molecular fragment of the small molecule is represented within the library of molecular fragments; embedding each molecular fragment into numerically embedded molecular fragments that captures the chemical properties of the molecule fragment; and entering the structural data of the biological target and the embedded molecular fragments into the computational diffusion model to learn to reconstruct the structural data from Gaussian noise.
[0013] In some aspects, the techniques described herein relate to a method, wherein the computational diffusion model learns to reconstruct the structural data from Gaussian noise by: progressively adding noise to positions and features of the structural data of the biological target and the embedded molecular fragments until transformed into Gaussian noise; and iteratively removing the noise to reconstruct the positions and the features of the structural data of the biological target and the embedded molecular fragments.
[0014] In some aspects, the techniques described herein relate to a method, wherein the molecules of the cohort of structural data include, prioritize, or are limited to ligands that are a particular type or have a particular functionality.
[0015] In some aspects, the techniques described herein relate to a method, wherein the linking computational model is a neighbor-relationship-based computational model that links molecular fragments by identifying neighbor relationships of the molecular fragments based on their three-dimensional proximity.
[0016] In some aspects, the techniques described herein relate to a method, wherein the identifying neighbor relationships of the molecular fragments based on their three- dimensional proximity includes generating a neighbor graph including nodes and edges; wherein the nodes represent the molecular fragments and the edges represent potential connection between nearby molecular fragments.
[0017] In some aspects, the techniques described herein relate to a method, wherein linking the set of molecular fragments includes enumerating spanning trees from the neighbor graph, wherein a spanning is a subset of edges from the neighbor graph that connects all nodes, wherein the spanning trees represent possible ways to connect the fragments while ensuring that each fragment is connected to at least one other fragment.
[0018] In some aspects, the techniques described herein relate to a method, wherein a spanning tree is a subset of edges from the neighbor graph that connects all nodes without forming any cycles.
[0019] In some aspects, the techniques described herein relate to a method, wherein the linking computational model is an edge prediction model.
[0020] In some aspects, the techniques described herein relate to a method further including synthesizing the small molecule.
[0021] In some aspects, the techniques described herein relate to a method further including assessing the synthesized small molecule for binding to the biological target.
[0022] In some aspects, the techniques described herein relate to a method further including assessing the synthesized small molecule in a biological function assay to assess an ability of the small molecule to alter a biological function of the biological target.
[0023] In some aspects, the techniques described herein relate to a method, wherein the method is repeated a plurality of times to yield a plurality of small molecules.
[0024] In some aspects, the techniques described herein relate to a method further including identifying one or more candidate small molecules from the plurality of small molecules by: assessing the plurality of molecules for a set of desired characteristics;wherein the one or more candidate small molecules are assessed to have a desired characteristic greater than a threshold.
[0025] In some aspects, the techniques described herein relate to a method further including identifying one or more candidate small molecules from the plurality of small molecules by: filtering out one or more small molecules from the plurality small molecules based on a set of desired characteristics; wherein the one or more molecules filtered out are determined to be below a threshold for a desired characteristic.
[0026] In some aspects, the techniques described herein relate to a method, wherein the set of desired characteristics include one or more of: difficulty to synthesize, cost to synthesize, solubility, bioavailability, polarity, pKa, or stability.
[0027] In some aspects, the techniques described herein relate to a method further including synthesizing the one or more candidate small molecules.
[0028] In some aspects, the techniques described herein relate to a method further including assessing the synthesized one or more candidate small molecules for binding to the biological target.
[0029] In some aspects, the techniques described herein relate to a method further including assessing the synthesized one or more candidate small molecules in a biological function assay to assess an ability of the one or more candidate small molecules to alter a biological function of the biological target.
[0030] In some aspects, the techniques described herein relate to a computational method to design small molecules for binding to a biological target, the method including: providing a structure of biological target having a binding site to be targeted; positioning, utilizing a computational diffusion model and a library of molecular fragments, a set of molecular fragments within the target binding site; wherein the molecular fragments are embedded as a set of vectors via a dimensionality reduction technique that captures chemical properties of the molecular fragments; wherein the computational diffusion model is trained utilizing a cohort of structural data of small molecules bound to its binding site of a biological target; wherein each molecular fragment of the set of molecular fragments are to be positioned simultaneously; and linking, utilizing a neighborrelationship-based computational model, the set of molecular fragments to yield a small molecule.
[0031] In some aspects, the techniques described herein relate to a computational method, further including: acquiring the library of molecular fragments; and reducing, utilizing the dimensionality reduction technique, dimensionality of chemical properties of each fragment of the library of molecular fragments into its set of vectors.
[0032] In some aspects, the techniques described herein relate to a method, wherein the reduction technique is one of the following: principal component analysis (PCA), T- distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (LIMAP), singular value decomposition (SVD), or linear discriminant analysis (LDA).
[0033] In some aspects, the techniques described herein relate to a method, wherein the chemical properties include one or more of the following: charge, molecular weight, number of hydrogen bond donors / acceptors, or number of rings.
[0034] In some aspects, the techniques described herein relate to a method, wherein the dimensionality reduction technique further captures an extended three-dimensional fingerprint of the molecular fragments.
[0035] In some aspects, the techniques described herein relate to a method wherein the computational diffusion model is a diffusion denoising probabilistic model.
[0036] In some aspects, the techniques described herein relate to a method further including: training the computational diffusion model by: acquiring the cohort of structural data of small molecules bound to its binding site of a biological target; fragmenting each small molecule of the cohort into molecular fragments, wherein each molecular fragment of the small molecule is represented within the library of molecular fragments; reducing, utilizing the dimensionality reduction technique, each molecular fragment into embedded molecular fragments including a set of vectors; and entering the structural data of the biological target and the embedded molecular fragments into the computational diffusion model to learn to reconstruct the structural data from Gaussian noise.
[0037] In some aspects, the techniques described herein relate to a method, wherein the computational diffusion model learns to reconstruct the structural data from Gaussian noise by: progressively adding noise to positions and features of the structural data of the biological target and the embedded molecular fragments until transformed into Gaussiannoise; and iteratively removing the noise to reconstruct the positions and the features of the structural data of the biological target and the embedded molecular fragments.
[0038] In some aspects, the techniques described herein relate to a method, wherein the small molecules of the cohort of structural data of small molecules bound to its binding site are known drugs.
[0039] In some aspects, the techniques described herein relate to a method, wherein the neighbor-relationship-based computational model links molecular fragments by identifying neighbor relationships of the molecular fragments based on their three- dimensional proximity.
[0040] In some aspects, the techniques described herein relate to a method, wherein the identifying neighbor relationships of the molecular fragments based on their three- dimensional proximity includes generating a neighbor graph including nodes and edges; wherein the nodes represent the molecular fragments and the edges represent potential connection between nearby molecular fragments.
[0041] In some aspects, the techniques described herein relate to a method, wherein linking the set of molecular fragments includes enumerating spanning trees from the neighbor graph, wherein a spanning is a subset of edges from the neighbor graph that connects all nodes, wherein the spanning trees represent possible ways to connect the fragments while ensuring that each fragment is connected to at least one other fragment.
[0042] In some aspects, the techniques described herein relate to a method, wherein a spanning tree is a subset of edges from the neighbor graph that connects all nodes without forming any cycles.
[0043] In some aspects, the techniques described herein relate to a method further including synthesizing the small molecule.
[0044] In some aspects, the techniques described herein relate to a method further including assessing the synthesized small molecule for binding to the biological target.
[0045] In some aspects, the techniques described herein relate to a method further including assessing the synthesized small molecule in a biological function assay to assess an ability of the small molecule to alter a biological function of the biological target.
[0046] In some aspects, the techniques described herein relate to a method, wherein the method is repeated a plurality of times to yield a plurality of small molecules.
[0047] In some aspects, the techniques described herein relate to a method further including identifying one or more candidate small molecules from the plurality of small molecules by: assessing the plurality of small molecules for a set of desired characteristics; wherein the one or more candidate small molecules are assessed to have a desired characteristic greater than a threshold.
[0048] In some aspects, the techniques described herein relate to a method further including identifying one or more candidate small molecules from the plurality of small molecules by: filtering out one or more small molecules from the plurality small molecules based on a set of desired characteristics; wherein the one or more molecules filtered out are determined to be below a threshold for a desired characteristic.
[0049] In some aspects, the techniques described herein relate to a method, wherein the set of desired characteristics include one or more of: difficulty to synthesize, cost to synthesize, solubility, bioavailability, polarity, pKa, or stability.
[0050] In some aspects, the techniques described herein relate to a method further including synthesizing the one or more candidate small molecules.
[0051] In some aspects, the techniques described herein relate to a method further including assessing the synthesized one or more candidate small molecules for binding to the biological target.
[0052] In some aspects, the techniques described herein relate to a method further including assessing the synthesized one or more candidate small molecules in a biological function assay to assess an ability of the one or more candidate small molecules to alter a biological function of the biological target.BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The description and claims will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments of the invention and should not be construed as a complete recitation of the scope of the invention.
[0054] Figure 1 provides a flow diagram of a method to generate molecule structures that targets a biomolecule.
[0055] Figure 2 provides a schematic of a computational processing system for generating molecular structures that target a biomolecule.
[0056] Figures 3A-3E provide schematics of a computational framework MedSAGE and fragment-based molecule generation. Figure 3A: An illustration of de novo structure- guided drug design for molecules: the process of generating novel molecular structures, using 3D information about the target protein’s binding site to guide compound design. Figure 3B: Examples of small molecules generated by prior diffusion-based methods (left) highlight common issues, including poor drug-likeness and low synthetic feasibility, often due to overly complex ring systems and the presence of reactive functional groups (e.g., epoxides). MedSAGE addresses these challenges with a novel fragment-based diffusion framework. Representative MedSAGE-generated molecules, shown on the right, exhibit improved structural realism and synthetic accessibility. Figure 3C: A curated set of small functional groups and ring systems for medicinal chemistry applications. Figure 3D: A schematic of encoding and decoding of fragments. Molecules are represented as centroids of their constituent fragments, which are labeled by lowdimensional latent vectors for efficient use in diffusion models. The latent space embedding smoothly encodes chemical and geometric properties of the fragments. Figure 3E: A schematic of the MedSAGE process. The generative model operates in two phases. The first phase operates on fragments, which are positioned and assigned types simultaneously within the context of the target protein binding pocket using a trained denoising diffusion model. In the second phase, the coarse-grained fragment representation is converted back to atomic-level structures, followed by connecting nearby fragments.
[0057] Figures 4A-4D provide molecular structures of the seventy most frequently utilized fragments from the fragment library.
[0058] Figure 5 provides a data graph depicting docking score as a function of the number of connection candidates. For fixed sets of fragments, we sampled multiple chemically valid connections using our connection algorithm. Each resulting molecule was docked using either high-throughput virtual screening (HTVS) mode or standard precision (SP) mode. We report the best docking score as a function of the number of thesampled candidates. The results show that docking scores tend to plateau after sampling approximately 30-40 connection variants.
[0059] Figures 6A-6C provide a schematic summary of MedSAGE molecule generation. MedSAGE generates realistic, diverse molecules that recapitulate scaffolds and key interactions of high-affinity reference ligands. Each figure includes selected examples of MedSAGE-generated molecules for a drug targets. The protein pocket structure (light grey surface and sticks) provided to MedSAGE along with a generated molecule (dark grey sticks) is shown for each target. The generated molecule was compared to the “reference ligand” — a drug-like molecule with high affinity for the target protein (<100 nM). Dark-background images illustrate ligand-pocket interactions, including hydrogen bonds, salt bridges, and TT-TT interactions. For each example, an additional six MedSAGE-generated molecules are provided to highlight the diversity of generated chemical structures. Figure 6A: The target biomolecule is a crystal structure of Aurora kinase (PDB 3D14), a serine / threonine kinase essential for mitosis and implicated in various cancers. The reference ligand is a potent and selective inhibitor with a binding affinity of 22 nM and demonstrated in vivo antitumor activity. Figure 6B: The target biomolecule is a crystal structure of the dopamine transporter from D. melanogaster (PDB 4XNX), used to study the binding mechanisms of antidepressants to various transporters. The reference ligand is reboxetine, an approved antidepressant with 20 nM affinity for this target. Figure 6C: The target biomolecule is a crystal structure of the gp120 subunit of the HIV-1 envelope trimer (PDB 6MTN). The reference ligand is an analog of BMS-626529 (temsavir), an approved antiretroviral medication.
[0060] Figures 7A-7E provides data graphs summarizing ligand characteristics for molecule structures generated by MedSAGE, PMDM, DiffSBDD, IPDiff, and random sampling. MedSAGE generates molecules with superior affinity and selectivity compared to other diffusion-based methods for molecule generation. Benchmarking was conducted using a curated dataset of 25 diverse proteins, each of therapeutic importance. For each target, several hundred molecules were generated by each method, and their properties were evaluated. The Reference set consists of drug-like ligands with high-affinity for the target proteins, serving as the positive control. PMDM, DiffSBDD, and IPDiff are molecules generated using publicly available all-atom diffusion models, used forcomparison. The Random Sample set are molecules randomly sampled from the Enamine REALSpace chemical library, without using the target protein structure, serving as the negative control. Ligand characteristics assessed were Predicted Binding Affinity (Figure 7A), Selectivity Score (Figure 7B), Synthetic Complexity Score (Figure 7C), Interactin Similarity to Reference Ligand (Figure 7D), and Number of Aromatic Rings (Figure 7E). Data are presented as the median values, with 95% confidence intervals (see Methods for details on statistical analyses). P-values calculated using two-sided t- test (n.s. indicates p>0.05, * p<0.05, ** p<0.01 , *** p<0.001 ).
[0061] Figure 8 provides a data depicting a comparison of docking scores between reference ligands and top MedSAGE-generated ligands. For each target, the MedSAGE- generated molecule with the best Glide docking score was selected. Molecules were compared to the docking score of the corresponding high-affinity reference ligand.
[0062] Figure 9 provides a data graph depicting a consensus between Vina and Glide docking scores. Benchmarking was performed on a curated dataset of 25 protein targets. For each target and method, several hundred molecules were generated per method and evaluated using both Glide and Vina docking scores. To establish a consensus metric, only molecules where the Glide and Vina scores differed by no more than 3 kcal / mol were retained. The Vina wasused score for comparison. Even under this criterion, MedSAGE- generated molecules continued to outperform those from other diffusion-based methods.
[0063] Figure 10 provides data graphs summarizing characteristics for molecule structures generated by MedSAGE. MedSAGE Produces Synthesizable and Drug-Like Molecules. For each target, several hundred ligands were generated using SAGE, using the 25 benchmarking target protein structures from the test set. The property distributions of the generated molecules were compared to those of the known, high-affinity reference ligands, and ligands generated with all-atom diffusion method DiffSBDD. Distributions are plotted using kernel density estimations for continuous metrics and binned, normalized histograms for discrete metrics. Docking scores are Glide scores at the target protein. Synthetic complexity score, calculated using the standard implementation in RDKit, is a relative measure of the ease of synthesizing the ligand, with lower scores indicating easier synthesis. All other properties are standard molecular descriptors computed with RDKit.
[0064] Figures 11A-11C provide data graphs and schematics of molecular characteristics summarizing the comparison between MedSAGE and traditional virtual screening methodologies. Figure 11 A: MedSAGE efficiently identifies high-scoring ligands compared to traditional virtual screening (VS) of chemical libraries. Docking score distributions are shown for three drug targets, with molecules identified through VS and molecules generated by MedSAGE. The median Glide docking score is plotted as a function of increasing percentile thresholds. The chemical libraries docked for virtual screening contained 30 million molecules for targets 1 and 2 and 400 million for target 3. MedSAGE was used to generate 2000 molecules for each target. Figure 11 B: Chemical diversity of top-scoring MedSAGE-generated molecules were compared with molecules identified by virtual screening, measured by 1-TC (Tanimoto Coefficient) of pairs of molecular fingerprints. Higher values indicate greater diversity. This analysis used the top-scoring 1000 molecules from each method. Figure 11 C: Selected examples of high- scoring ligands (top 1000) from virtual screening and MedSAGE, highlighting cases where MedSAGE-generated molecules that share chemical similarities with commercially available compounds (Enamine REALSpace) that are also high scoring.DETAILED DESCRIPTION
[0065] Described herein are various embodiments of computational systems and methods for designing molecule structures using a fragment-based approach. In several embodiments, molecule fragments are positioned on a target biomolecule. In many embodiments, molecule fragments are chemical functional groups. In some embodiments, molecule fragments are computationally embedded such that their properties are numerically captured within the embedding. In several embodiments, a trained diffusion model positions molecule fragments on the target biomolecule, considering chemical interaction and geometric dimensions. In some embodiments, a trained diffusion model is utilized to geometrically transform molecule fragments. In many embodiments, molecule fragments are linked, yielding a molecule structure. In various embodiments, molecules are linked by a neighbor relationship model and / or an edge prediction model. In some embodiments, the computational framework generates a variety of molecule structures that target the biomolecule. In some embodiments, acomputationally designed molecule is chemically synthesized. Synthesized molecules can be assessed for their ability to interact with and / or modulate activity of its target biomolecule. Furthermore, synthesized molecules can be utilized in a variety of applications, such as (for example) within medicinal formulations, as biological research tools, and / or agricultural products.Molecule Structure Design
[0066] Several embodiments are directed to generating molecule structures via a computational framework. In many embodiments, a computational framework generates a molecule structure de novo. In several embodiments, the computational framework comprises a trained diffusion model for three-dimensional positioning of molecule fragments upon a target biomolecule. In many embodiments, a computational model is configured to link molecule fragments to yield a molecule structure.
[0067] Provided in Fig. 1 is a computational method to generate a molecule structure utilizing a computational framework in accordance with various embodiments. Computational Method 100 begins with providing (101 ) a structure of a target biomolecule. The target biomolecule is any biomolecule that a user wants to design a ligand that interacts with the biomolecule. The biomolecule can be a proteinaceous species, a nucleic acid, a polysaccharide, or another biomolecule, but especially a macromolecule. Examples of common biomolecules to be targeted are enzymes, cell receptors, transmembrane channels, transporters, cancer markers, cancer neoantigens, and pathogen proteins.
[0068] In many embodiments, a representation of the target biomolecule is provided. For example, the crystal structure of the target biomolecule can be utilized. In many embodiments, a site on the target biomolecule is selected to be targeted. Generally, any site on a biomolecule can be utilized. Various options for sites include a binding pocket, a substrate binding site, a cofactor binding site, a site associated with a target conformation, a pore of a channel or transporter, a site that is modified (e.g., secondary modification), a site that is unique (e.g., non-human like site of a pathogen), etc.
[0069] Computational method 100 positions (103) molecule fragments upon the target biomolecule at the site to be targeted. Molecule fragments are chemical functional groups,which are multiatom molecular structures. In some implementations, a library of chemical functional groups can be utilized for computationally selecting and positioning functional groups. Generally, a library can comprise any chemical functional groups. In some embodiments, the chemical functional groups to select from are limited to functional groups used within medicinal compounds. In some embodiments, the chemical functional groups to select from are limited to compounds with organic atoms. In some embodiments, functional groups to select from are limited to functional groups having the following atoms: C, 0, N, P, S, H, F, Cl, and Br. In some embodiments, differences in protonation are considered two different functional groups (e.g., carboxylate (COO-) is different from carboxylic acid (COOH)). In some embodiments, functional groups that are tautomers are considered different groups.
[0070] Various computational techniques can be utilized to position molecule fragments onto a target biomolecule. In several embodiments a diffusion model is utilized. A diffusion model, which is often used for computer vision tasks, generally is trained to denoise images by removal of Gaussian noise. In the context of molecule design, a diffusion model can be trained to learn chemical interactions and geometric positioning that occur between functional groups of the molecule fragments and the target biomolecule. The geometric positioning can be in three dimensions. Upon a new target site, the model can select functional groups and iteratively “denoise” unspecific chemical interactions and geometric positions of the selected functional groups to yield an optimized set of functional groups. The selected and positioned functional groups are optimized based on the global chemical interactions and geometric positions among all the functional groups and the target molecule. Accordingly, the diffusion model can simultaneously globally assess the chemical interactions and geometric positions between the selected functional groups and the target biomolecule at each iteration. A diffusion model may incorporate an equivariant graph neural network, tensor field network, geometric vector perceptron, SE(3)-transformer, or any architecture that works with 3D coordinates of the atoms or molecular fragments.
[0071] A diffusion model can be trained in a variety of ways. A diffusion model can be trained on a set of target biomolecules with an associated ligand in which the chemical interactions and geometric positioning can be learned. These interactions and positionscan be determined or inferred from a variety of molecular interaction data. Some of the best data generated for learning these interactions and positions are crystal structures of biomolecules with the ligand associated thereon. Accordingly, in some embodiments, crystallography data or images of crystallized biomolecules having an associated ligand are used to train the diffusion model.
[0072] In many embodiments, biomolecule-ligand data used for training have characteristics that are desired to be learned by a diffusion model. For example, if ligands with drug qualities are desired, the biomolecule-ligand data used for training can be drug ligands associated with their target biomolecule. Various characteristics of target biomolecules, ligands, and / or target-ligand pairs can be included, prioritized, or limited to for training, as desired by the task. For example, it may be desirable to include, prioritize, or limit the training with target biomolecules that are a particular protein type (e.g., kinase, transcription factor, receptor, transmembrane channel, transporter, dimer, pathogen- derived, etc.). For example, it may be desirable to include, prioritize, or limit the training with ligands that are a particular type (e.g., small molecules, naturally derived molecules, macrocycles, peptidomimetics, peptides, etc.) or have a particular functionality (e.g., solubility, tolerability, pharmacokinetics, bioavailability, biodistribution, metabolism rate, metabolites formed, ability to cross blood brain barrier, pKa, stability, polarity, synthesis difficulty, synthesis costs, etc.). It may also be desirable to include, prioritize, or limit the training with biomolecule-ligand pairs that have a particular property (e.g., affinity, specificity, dissociation rate, covalent attachment, etc.). Filters, thresholds, and other criteria can be utilized to select target biomolecules and ligands. Training data may include qualities that are desired and / or qualities that are undesired, such that the model can learn which combination of functional groups to select and / or to disregard.
[0073] Various types of diffusion models can be utilized. Examples of types of diffusion models that can be utilized include score-based generative models, latent diffusion models, denoising diffusion implicit models, and denoising diffusion probabilistic models. In some embodiments, the diffusion model is a denoising diffusion probabilistic model.
[0074] The functional group data can be used as features with the diffusion model. In several embodiments, each functional group is computationally embedded for use in the model. Various embedding techniques can be utilized including (but not limited to)principal component analysis (PCA), T-distributed stochastic neighbor embedding (t- SNE), uniform manifold approximation and projection (UMAP), singular value decomposition (SVD), or linear discriminant analysis (LDA) In some embodiments, embeddings are learned in an unsupervised manner. In some embodiments, the embedding technique comprises a reduction technique. In many embodiments, the embedding of a functional group numerically captures one or more of its properties. These properties can include one or more of: Extended Three-Dimensional Fingerprint as well as chemical properties of charge, molecular weight, number of hydrogen bond donors / acceptors, number of rings, etc. In several embodiments, each functional group embedding comprises a set of vectors capturing the properties of the functional group. In some embodiments, a reduction technique decomposes functional groups into their constituent fragments and assigns a labeled point at each fragment’s centroid.
[0075] Various methodologies can be utilized to train a diffusion model. Upon acquiring the structural data of a cohort of biomolecule-ligand pairs, the ligand can be fragmented until each fragment of the ligand is represented within the library of molecule fragments. Each fragment can be embedded to yield a set of fragment embeddings, each embedding capturing one or more properties of the fragment. The fragment embeddings can be entered into the model such that it can learn to reconstruct the structural data from Gaussian noise. This can be achieved by progressively adding noise to positions and features of the structural data of the target biomolecule and the embedded molecule fragments until transformed into Gaussian noise and iteratively removing the noise to reconstruct the positions and the features of the structural data of the target biomolecule and the embedded molecule fragments.
[0076] Using a trained diffusion model, a molecule structure can be designed de novo to a target biomolecule of interest. The model can use structure data of the target biomolecule and denoise functional group embeddings upon the target structure data. The model outputs a global positioning of the functional group embeddings.
[0077] Computational method 100 links (105) molecule fragments to yield a molecule structure that targets the biomolecule of interest. Molecule fragments can be linked by a variety of computational methods. Generally, a computational model can comprehend relationships between functional groups and build chemically sound connectors toconnect all the groups. In one methodology, a neighbor graph of the positioned molecule fragments is generated and then utilized to identify potential connections, and selecting connections that meet a set of rules. In another methodology, a trained edge prediction model predicts bonding and interactions between molecule fragments. Other methodologies can also be used, and the various methodologies can be combined and / or applied sequentially.
[0078] In embodiments that generate a neighbor graph, the nodes of the graph can be represented by the molecule fragments with an assigned position. Edges can be represented by potential connections between the molecule fragments. In some embodiments, a potential connection between the molecule fragments is within a threshold distance. Using the potential connections, a set of spanning trees are generated in which each spanning tree comprises a unique set of connections and each molecule fragment is connected to at least one other molecule fragment. The generation of spanning trees can be configured such that the spanning tree does not comprise connections that form a cycle. Within each spanning tree, the model can identify potential attachment points within each molecule fragment and then determine viable connections between the molecule fragments based on attachment points and a basic set of chemical bonding rules, such as valence requirements, bond order, and atom connectivity. The model then outputs candidate molecule structures.
[0079] In embodiments that utilize a trained edge prediction model, the model can be trained on a set of known and viable molecule structures. The model can learn to predict whether a bond would be formed between two molecule fragments. In some embodiments, the model is a binary model, and the output is the presence or absence of a bond between two molecule fragments. In some embodiments, the model is a probabilistic model, and the output is the probability of a bond forming between two molecule fragments. The model can be instructed to satisfy chemical bonding rules, such as valence requirements, bond order, and atom connectivity. The prediction process can also utilize a diffusion model (e.g., denoising diffusion probabilistic model), which can predict molecular structures within connectors. An edge prediction model or diffusion model may incorporate an equivariant graph neural network, tensor field network,geometric vector perceptron, SE(3)-transformer, or any architecture that works with 3D coordinates of the atoms or molecular fragments.
[0080] Notably, molecule fragments (or atoms therein) can optionally be geometrically transformed during the linking process, which may yield a more optimal molecule structure. Examples of geometric transformations that can be performed include (but are not limited to) rotation and translation. A diffusion model (e.g., denoising diffusion probabilistic model) can be trained to apply geometric transformations, yielding an output of updated coordinates that reflect the transformations of the molecule fragments (or atoms therein). The model may incorporate an equivariant graph neural network, tensor field network, geometric vector perceptron, SE(3)-transformer, or any architecture that works with 3D coordinates of the atoms or molecular fragments.
[0081] Once a molecule structure is generated, in accordance with various embodiments, the compound is chemically synthesized. The compound to be synthesized can be a final molecule structure or can be further iterated upon. Generally, standard chemistry synthesis methods can be utilized. In some situations, synthesis protocols can be generated to produce the generated molecule structure. For more on compound synthesis, see, e.g., S. L. Schreiber, Proc Natl Acad Sci U S A. 2011 Apr 26;108(17):6699-702; J. Li, et al., Science. 2015 Mar 13; 347(6227 ):1221-6; and J. W. Lehmann, et al., Nat Rev Chem. 2018 Feb;2(2):0115; the disclosures of which are each incorporated by reference.
[0082] Once the compound is synthesized, it can be utilized in a variety of applications. For instance, a synthesized ligand compound can be utilized in a medicinal formula for treatment of a medical disorder or disease. In some situations, a synthesized ligand compound can be utilized as modulator in biochemical experimentation. In some situations, a synthesized ligand compound can be utilized in an agricultural product, such as (for example) an herbicide or a pesticide.
[0083] In some embodiments, a synthesized compound is assessed to determine association with the target biomolecule. Various biochemical experiments can be conducted to determine affinity, specificity, dissociation rate, covalent attachment, or other biomolecule-ligand characteristics. In some embodiments, a synthesized compound is assessed via biochemical, cellular, and / or vivo experimentation for drug-likecharacteristics and other molecule characteristics, including (but not limited to) solubility, tolerability, pharmacokinetics, bioavailability, biodistribution, metabolism rate, metabolites formed, ability to cross blood brain barrier, pKa, stability, and polarity. In some embodiments, a synthesized compound is administered to a pre-clinical model or to human recipient within a clinical trial to determine safety and / or efficacy.
[0084] Compounds can be generated for a variety of purposes and may be designed accordingly. For example, a compounds may be designed to agonize or antagonize an enzyme, provide steric interference, deliver a payload, mark a target for degradation, induce biomolecule-biomolecule interactions, provide a means of visualization, covalently link with its target, etc. In some embodiments, a synthesized compound is assessed to determine an effect on the target molecule. Biochemical, cellular, and in vivo experiments can be conducted to determine whether the compound agonizes an enzyme, antagonizes an enzyme, sterically interferes with an interaction, delivers a payload, marks a target for degradation, induces biomolecule-biomolecule interactions, yields biomolecule visualization, and / or covalently links with its target.
[0085] Altering enzyme activity is a common strategy in drug molecules to alter a medical condition. For example, statins used for treating cholesterol inhibit the enzyme HMG-CoA, reducing the production of LDL-cholesterol.
[0086] Steric hindrance can be useful to block interactions, such as blocking a protein on the surface of a pathogen to prevent the pathogen from infecting host cells. Various payloads can be delivered to specifically target biomolecules. For example, a cancer marker or neoantigen can be targeted with a ligand with an attached payload that provides visualization of the cancer or attacks to kill the cancer.
[0087] When a protein is marked with the ubiquitin peptide, it is signaled to be degraded by the cell’s degradation system. This attachment of ubiquitin onto a protein is performed by the enzyme E3 ubiquitin ligase. Using an E3 ubiquitin ligase attached to a ligand, target proteins can be marked for degradation, reducing their activity. These compounds can be referred to as proteolysis targeting chimera (PROTAC).
[0088] A molecular glue can induce interactions between two biomolecules by bringing them into proximity. Generally, a molecular glue can comprise two different ligands with a linker between the ligands. One ligand binds with a first biomolecule and a secondligand binds with a second biomolecule, increasing the interaction between the two biomolecules. For example, some proteins require dimerization (i.e. , two proteins coming together) to perform their function. A molecular glue can increase the dimerization of target proteins, increasing their functional activity. In another example, in alternative to attaching a E3 ligase to a ligand, a molecular glue can be utilized to target a protein to be degraded with one ligand and an E3 ubiquitin ligase with the second ligand, resulting in the target protein to be signaled for degradation.
[0089] A ligand can be linked to a fluorophore, a radioisotope, or other compound that enables visualization. Various scope systems or imaging modalities can be utilized to visualize biomolecules targeted with visualization-enabled ligand. This is a common technique to visualize cancer cells, which can be used during a surgical resection procedure to visualize the cancerous tissue to be removed.
[0090] A ligand can be configured to covalently bond with a target biomolecule. Covalent linkage establishes a long-term (and often permanent) interaction between the ligand and its target. To yield a covalent linkage, a ligand will have functional groups that can spontaneously or enzymatically form a covalent bond between the ligand and the target. Covalent binding is useful potent inhibition, long-term steric blockage, and maintaining the target in a particular conformation. For example, the COVID-19 drug Paxlovid covalently binds with a coronavirus protease to disrupt the virus lifecycle.Computational processing system
[0091] A computational processing system to generate molecule structures in accordance with various embodiments of the disclosure typically utilizes a processing system including one or more of a CPU, GPU and / or other processing engine. A GPU may be of benefit for performing various computational tasks described herein. In some embodiments, the computational processing system is housed within a computing device or a set of connected computing devices such as (but not limited to) a computer, mobile phone, a tablet computer, and / or portable computer. A set of computing devices can be connected in any manner that allows data communication, such as (for example) a wired connection, Bluetooth, Wi-Fi, a cellular system, a cloud system, or an internet modem connection.
[0092] An example of a computational processing system is provided in Fig. 2. Computational processing system 200 includes a processor system 202, an I / O interface 204, and a memory system 206. As can readily be appreciated, the processor system 202, I / O interface 204, and memory system 206 can be implemented using any of a variety of components appropriate to the requirements of specific applications including (but not limited to) CPUs, GPUs, ISPs, DSPs, wireless modems (e.g., Wi-Fi, Bluetooth modems), serial interfaces, depth sensors, IMUs, pressure sensors, ultrasonic sensors, volatile memory (e.g., DRAM) and / or non-volatile memory (e.g., SRAM, and / or NAND Flash). The memory system can store data applications, models and various other code for performing computational tasks, such as those described herein. The memory can store a molecule fragment library 208 which can be used to computationally generate molecule structures de novo. The memory can include a molecule fragment positioning model 210 to globally position functional groups of a putative ligand in relation to a target biomolecule. The memory can include a molecule fragment connection model 212 that link molecule fragments to yield a molecule structure that is predicted to interact with the target biomolecule. Generated molecule structures 214 can also be stored on the memory.EXAMPLES AND EXPERIMENTAL DATA
[0093] The various embodiments of the disclosure will be better understood with the various examples and corresponding experimental data provided within. A computational framework was developed that adapts diffusion models for molecule design. The framework generates molecules using a unique latent-space representation of chemical fragments while integrating physics-based guidance to optimize connections between fragments. In a benchmark of multiple methods across 25 therapeutically relevant protein targets, the framework achieved state-of-the-art performance, producing synthesizable, drug-like molecules with predicted affinity and selectivity closely matching known high- affinity ligands. Compared to large-scale virtual screening, the framework produced high- scoring molecules up to 100 times more efficientlyMed SAGE: Bridging Generative Al and Medicinal Chemistry for Structure-Based Design of Small Molecule Drugs
[0094] Structural biology is advancing at an extraordinary pace; advancements in experimental and computational techniques have enabled the rapid determination of detailed structures for a wide variety of biomolecules. These structures not only illuminate the molecular mechanisms underlying biological functions but also serve as guides for therapeutic design. A striking example of this progress occurred during the COVID-19 pandemic, where the structures of key coronavirus proteins were resolved within months of the outbreak, enabling development of new therapeutics.
[0095] Despite advances in structure determination, using protein structures to design effective therapeutics remains a complex and costly challenge.
[0096] In recent years, generative artificial intelligence (Al) has emerged as a promising approach for structure-based drug discovery across multiple therapeutic modalities, including biologies, peptides, and small molecules. Diffusion models, a class of generative Al methods, have delivered particularly impressive results recently for the de novo design of proteins and antibodies. By learning subtle patterns from large structural datasets, generative Al models can efficiently address complex design tasks that are challenging for traditional methods. Generative methods are also inherently very efficient, producing optimized solutions directly in contrast to the exhaustive search and evaluation of billions of candidates. Finally, generative Al models are highly adaptable and can be fine-tuned for specialized therapeutic applications.
[0097] However, current generative Al techniques — including diffusion models — have yet to provide robust solutions for small-molecule drug discovery, the modality encompassing the majority of drugs and new drug approvals. Despite notable theoretical advances, these methods continue to fall short in practical utility. Several recent de novo molecule generation methods we surveyed produced molecules that did not satisfy basic medicinal chemistry criteria, much less advanced properties like potency or selectivity.
[0098] Diffusion models have demonstrated significant success across domains where data can be represented as grids or sequences, including image generation, protein design, and natural language. However, complex molecular structures are not naturally represented by simple sequences or grids. 3D point clouds of atoms are a widelyused representation, but medicinal chemists typically conceptualize small molecules in terms of functional groups or “fragments” rather than individual atoms and elements. Fragment-based approaches entail unique challenges as well: thousands of chemical fragments exist, and these fragments can connect in highly complex ways.
[0099] We successfully overcome these obstacles enabling diffusion models to handle small molecule representations effectively via a framework referred to as MedSAGE (Fig. 3A). For example, we constructed mathematical embeddings of functional groups via a two-step generative process, and a carefully curated and prepared dataset. Additionally, we integrate the strengths of physics-based methods to guide the optimal connectivity of generated fragments. Our approach achieves state-of-the-art performance in de novo structure-guided small-molecule design, reliably producing synthesizable molecules precisely tuned to bind their targets while maintaining favorable drug-like attributes, improving upon conventional diffusion models (Fig. 3B). We also conduct a thorough comparison with real-world virtual screening campaigns. MedSAGE achieves unprecedented performance, making diffusion-based methods practically viable for structure-based small-molecule design, and laying the groundwork for a new paradigm in drug discovery.A Fragment-Based Molecular Representation for Diffusion Models
[0100] To simplify the learning process and eliminate the generation of invalid or unrealistic molecules, we developed a novel molecular diffusion approach that encodes both chemical and geometric information into a lower-dimensional latent space, guided by strong chemical intuition. Drawing inspiration from coarse-graining techniques used in molecular dynamics simulations, we represent molecules — including both ligands and protein targets — as fragments. In this context, fragments refer to small functional groups and ring systems that are relevant for medicinal chemistry applications. We then designed and trained diffusion models to operate directly on this fragment-based representation.
[0101] We first curated a comprehensive library of several thousand fragments relevant to drug discovery (Figs. 3C and 4A-4D), extracted from drug-like small molecule ligands (see Methods, Fragment Preparation). Each fragment was then assigned a lowdimensional vector label that encodes a smooth embedding of its Extended Three-Dimensional Fingerprint along with key chemical properties such as charge, polarity, hydrogen bond donors / acceptors, and atom count (Fig. 3D). These embeddings were learned in an unsupervised manner using t-SNE, enabling a unified embedding of all fragments in the library.
[0102] Compared to one-hot encoded fragment features or atomic element labels, these smooth and normalized vector labels are better suited for diffusion models and improve computational efficiency. The fragment-based representation significantly reduces degrees of freedom — by more than tenfold compared to all-atom representations commonly used in diffusion models for small molecules. For example, generating a benzene ring requires only a single fragment point and its associated label (six degrees of freedom in total), rather than precisely positioning and labeling all six carbon atoms (6x3 spatial coordinates + 6x8 element properties = 66 degrees of freedom). To encode molecular structures in this fragment representation, we developed a custom algorithm that decomposes molecules into their constituent fragments and assigns a labeled point at each fragment’s centroid (see Methods, Fragment Preparation).
[0103] MedSAGE designs molecules targeting a protein binding pocket structure through a two-phase process (Fig. 3E). In the first phase, we apply a denoising diffusion probabilistic model (DDPM) to generate the labels and positions of ligand fragments within the 3D pocket environment. DDPMs are a class of generative models that learn to reverse a noise-adding (diffusion) process by gradually refining noisy inputs back into structured outputs. The target pocket structure remains fixed throughout the noising and denoising process and is encoded using the same fragment representation. To distinguish between ligand and pocket points, we incorporate additional context features that indicate whether each point belongs to the ligand or the pocket. Our diffusion model employs an equivariant graph neural network (EGNN) for denoising predictions. The EGNN architecture effectively leverages geometric information by considering pairwise distances within the fragment point cloud.
[0104] To train our diffusion model, we constructed a custom dataset of 3D structures containing small molecule ligands with explicitly drug-like properties bound to protein targets. We began by downloading and processing all ligand-annotated structures from the Protein Data Bank (PDB). The dataset was then filtered to exclude commonbiomolecules such as lipids, peptides, carbohydrates, and nucleotides, as well as ligands that fell outside established drug-like property ranges. This curation process resulted in a dataset of approximately 35,000 protein-ligand complexes. These complexes were further processed to add hydrogens, fill missing sidechains, and determine protonation states. Pockets were selected using the known ligands. Special care was taken to account for protonation states in both the dataset and our fragment library. Protonation state helped improve accurate capturing of specific chemical features and interactions. See Methods, Dataset Preparation and Splitting, for further details.
[0105] In the second phase of MedSAGE, we convert the coarse-grained fragment representation back into an atomic structure to produce a final designed molecule. First, the generated fragment labels are decoded into specific fragments structures from our library using the learned latent space map (Fig. 3D). However, this step alone does not specify how nearby fragments should be chemically connected. To address this, we developed a search algorithm that identifies the optimal bond connections between fragments. The algorithm generates candidate molecules by adding bonds between adjacent fragments while enforcing chemical validity rules — for example, ensuring that an sp3carbon forms no more than four bonds (see Methods, Fragment Connecting). We then apply a physics-based scoring function (Glide) to evaluate the candidates and select the highest-scoring structure. While this connection step could eventually be replaced by a second Al model operating on atomic coordinates, we hypothesized that this physics approach would enforce strong physical constraints and leverage the robust scoring capabilities of existing physics-based docking methods. Importantly, this method remains computationally efficient because the fragments are already fixed and roughly positioned, significantly reducing the search space for potential connections. In practice, evaluating 40 connectivity candidates is often sufficient, after which scores tend to plateau (Fig. 5).Benchmarking Overview
[0106] To evaluate the performance of MedSAGE in structure-guided drug design, we constructed a rigorous benchmark of protein targets with therapeutic relevance. Specifically, we selected 25 diverse protein-ligand complexes from the test set, each containing a protein target bound to a high-affinity (<1 pM), drug-like small molecule —termed the "reference ligand" — which is either under development as a therapeutic or already an approved drug (Table 1 ). These targets (and proteins with similar sequences) were not used during model training. The average binding affinity of the reference ligands was 102 nM. For each protein target binding pocket, we generated 400 novel molecules using MedSAGE (Figs. 3A-3C). Each set of generated molecules was then evaluated across a comprehensive panel of 31 physicochemical and pharmacological properties, including logP, molecular charge, number of rotatable bonds, synthetic complexity, and predicted binding affinity (estimated via physics-based scoring functions) (Table 2). To place MedSAGE's performance into context, we conducted parallel evaluations of several leading generative Al methods on the same benchmarking targets, Diffusion Structurebased drug design (DiffSBDD), Interaction Prior-guided Diffusion model (IPDiff), and Probabilistic Molecular Design Model (PMDM). These methods utilize all-atom diffusion models to generate small molecules, allowing for a direct comparison with MedSAGE’s distinct latent fragment-based approach. For more on these models, see A. Schneuing, et al., Nat Comput. Sci. 2024 Dec;4(12):899-909; Z. Huang, et al., The Twelfth International Conference on Learning Representations 2024; and L. Huang, et al., Nat Commun. 2024 Mar 26;15(1 ):2657; the disclosures of which are hereby incorporated by reference.
[0107] Overall, MedSAGE demonstrates substantial improvements over existing leading methods by generating molecules that are highly realistic from a medicinal chemistry perspective, closely matching the structural and physicochemical properties of the drug-like reference ligands and adhering to established drug design rules (see Figs. 6A-6C and 7A-7E; Tables 1 and 1 ). This is despite not explicitly incorporating physicochemical property predictions or objectives into the generative process. As further illustrated in the Case Studies below, MedSAGE was able to identify scaffolds strikingly similar to the original reference ligands despite having no prior exposure to them during training. Additionally, molecules generated by MedSAGE frequently recapitulated key protein-ligand interactions, such as hydrogen bonds observed in the original high-affinity complexes (Fig. 6A). This indicates that MedSAGE effectively captures fundamental principles underlying ligand-protein binding.Med SAGE generates molecules with high predicted affinity and selectivity
[0108] The first metric we assessed was the predicted binding affinity of the generated molecules, approximated using docking scores from Schrodinger’s Glide (Fig. 7A). The median docking score of molecules generated by MedSAGE closely matched that of the high-affinity reference ligands (-9.2 vs. -9.5; difference not significant, p > 0.05, t-test; Table 2). Furthermore, for 23 out of the 25 targets, the top-ranked MedSAGE-generated molecule achieved a docking score notably better — typically 1-3 kcal / mol lower — than that of the reference ligand (Fig. 8). In contrast, molecules generated by other diffusionbased methods had weaker median docking scores, ranging from -5.5 to -6.3 (Fig. 7A, Table 3).
[0109] To further contextualize these results, we introduced a negative control by randomly sampling molecules from Enamine REALSpace and evaluating their docking scores and other properties. These control molecules, not designed or optimized for the protein targets in any way, served as a baseline and negative control. Their median docking score was -5.7 (Fig. yA), significantly worse than MedSAGE-generated molecules (p < 0.001 , t-test) but not statistically different from molecules generated by IPDiff and DiffSBDD (p > 0.05, t-test). This finding suggests that these alternative methods struggled to generate molecules specifically tailored to the target protein pockets.
[0110] We further evaluated predicted binding affinity using an alternative scoring method, AutoDock Vina. MedSAGE-generated molecules exhibited significantly better median Vina scores (-8.0) compared to the negative control set (-6.5, p < 0.001 , t-test) and were within 1 .2 units of the reference ligands (-9.2). Additionally, we developed a consistency scoring approach integrating both Glide and Vina scores, finding that MedSAGE continued to outperform other methods using this combined metric (Fig. 9).
[0111] Beyond assessing docking scores at the target protein, we also introduced a custom selectivity metric defined as the difference in docking scores between the intended target and a set of off-target proteins randomly selected from the benchmark set. Selectivity is crucial in drug design, as molecules optimized solely for affinity (or docking scores) can become overly bulky or hydrophobic, leading to non-specific binding to off-target proteins. MedSAGE-generated molecules displayed selectivity comparableto the reference ligands (Fig. 3b; 3.6 vs. 4.2, p > 0.05, t-test; Supp. Table 2). In contrast, molecules produced by the all-atom diffusion methods showed substantially lower selectivity (1 .3-1 .5), only slightly above that of the negative control molecules.
[0112] We hypothesized that the better selectivity of MedSAGE-generated molecules could stem from their ability to form specific key interactions — such as hydrogen bonds, salt bridges, and pi-pi stacking interactions — with the target proteins. To evaluate this, we quantified what fraction of interactions made by reference ligands were reproduced by each set of generated ligands. On average, MedSAGE-generated molecules reproduced 35% of the interactions observed in the high-affinity reference ligands versus 16% for the negative control (p<0.001 , t-test). In comparison, molecules generated by other diffusion-based methods reproduced only 12-23% of reference interactions.MedSAGE Produces Structurally Diverse, Synthesizable, and Drug-Like Molecules
[0113] We next evaluated broader physicochemical properties relevant to synthesizability and medicinal properties for developing drug candidates with favorable absorption, distribution, metabolism, and excretion (ADME) profiles. Across the majority of these metrics, MedSAGE-generated molecules closely matched the properties of the reference ligands (Fig. 10). MedSAGE also demonstrated the ability to generate structurally diverse ligands, even within individual protein targets. The average pairwise Tanimoto fingerprint similarity was less than 0.1 between pairs of generated ligands for a given target. Additionally, out of 400 ligands generated per target, there was an average of 200 unique scaffolds.
[0114] We evaluated ease of synthesis using a well-established synthetic complexity algorithm, that characterizes a molecule’s synthetic complexity as a score between 1 (easy to make) and 10 (very difficult to make). MedSAGE-generated molecules showed low average complexity scores (3.5), comparable to the reference ligands (3.2; p > 0.05, t-test). In contrast, molecules generated by DiffSBDD, IPDiff, and PMDM exhibited significantly higher synthetic complexity scores (4.9 to 5.3), consistent with the complex scaffolds and rare functional groups observed in these generated molecules. For the allatom diffusion methods, a deeper analysis revealed a high number of stereocenters (>3 per molecule) and a greater proportion of complex ring structures, such as bridged orspiro atoms (Table 3). In contrast, MedSAGE-generated molecules closely matched the average number of stereocenters observed in commercially available REALSpace molecules (1 .3 vs. 1 .4 stereocenters).
[0115] MedSAGE generated molecules with distributions of ring counts, rotatable bonds, and fraction of sp3carbons closely resembling those of the reference ligands (Fig. 10). However, we observed that MedSAGE-generated molecules contained slightly fewer heterocyclic rings compared to the reference ligands, perhaps reflecting a tendency to select simpler and more common ring fragments. In contrast, other diffusion-based methods (DiffSBDD and IPDiff) tended to generate molecules with fewer aromatic rings and a far higher prevalence of aliphatic rings (>3 aliphatic rings) than reference ligands (<1 aliphatic ring on average). This discrepancy likely arises from challenges in precisely positioning atoms during the diffusion generation process to form planar aromatic rings. MedSAGE avoids this issue because it generates molecules directly from fragment space, where different ring systems have distinct properties and are explicitly differentiated.
[0116] MedSAGE-generated molecules exhibited physicochemical properties closely aligned with medicinal chemistry guidelines, such as the Lipinski’s Rule of 5; these properties include molecular weight (340 Da, recommended <500 Da), number of hydrogen bond donors (3.0, recommended <5), hydrogen bond acceptors (4.2, recommended <10), and calculated logP (2.0, within the ideal range of 0-5). The average logP was 2.0, placing it within the ideal range (0-5) for oral bioavailability and CNS penetration, which is often desirable for therapeutics. In contrast, molecules generated by IPDiff and DiffSBDD tended to be more polar, with lower logP values of 0.87 and 0.7, respectively. Overall, these findings suggest that MedSAGE effectively generates molecules with physicochemical properties aligned with established principles of drug design.Case Studies
[0117] We examined the abilities of MedSAGE in several detailed case studies. In these examples, MedSAGE demonstrated a remarkable ability to “rediscover” scaffolds similar to the reference ligands (Figs. 6A-6C). Notably, the model was not trained on theseknown active molecules. This is very unlikely to occur by chance, considering the vast chemical space that MedSAGE can explore, with the potential to sample from trillions of possible molecular structures.
[0118] Our first case study is Aurora kinase, a serine / threonine kinase that plays a critical role in mitosis and is implicated in various cancers due to its involvement in chromosomal instability. The reference ligand, AK1 , is a potent and selective Aurora kinase inhibitor with a binding affinity of 22 nM, and significant activity in in vivo tumor models. With the Aurora kinase pocket as input, MedSAGE generated a molecule with a scaffold similar — but not identical — to AK1 (Fig. 6A). The generated compound preserves key structural features, including a core urea group flanked by two aromatic rings. This urea group forms critical hydrogen bonds with aspartate and lysine residues in the binding pocket. Additionally, both the reference and MedSAGE-generated molecules feature a heterocyclic aromatic ring at the end of the pocket that forms hydrogen bonds with a backbone amine. The MedSAGE-generated compound also forms additional hydrogen bonds with backbone carbonyl groups. The docking score of the MedSAGE-generated compound is -12, reflecting strong binding affinity, with an off-target average docking score of -5, indicating high selectivity for this target over off-target pockets.
[0119] In the next case study, the target is the crystal structure of the dopamine transporter (DAT) from D. melanogaster, which was used to study the binding mechanisms of antidepressants to the norepinephrine transporter (NET). The D. melanogaster DAT closely mimics the pharmacological profile of mammalian NETs, which are the targets of many antidepressants but are difficult to crystallize. The reference ligand is reboxetine, an approved antidepressant with 20 nM affinity for this target protein. MedSAGE generated a molecule with a scaffold closely resembling that of reboxetine, and a remarkably similar structural overlay (Fig. 6B). Two aromatic rings insert into a hydrophobic cleft, and an amino group within an aliphatic ring forms a salt bridge with a conserved aspartate residue. Notably, the generated molecule forms two additional hydrogen bonds compared to reboxetine and has a lower docking score (-9.7 vs. -8.4 for reboxetine).
[0120] While this particular molecule shares a similar scaffold with reboxetine, other MedSAGE-generated compounds display a wide range of diverse chemical structures(Fig. 6B). Nevertheless, all retain a key amine functionality that forms the characteristic ionic interaction with the aspartate — a hallmark of nearly all NET and DAT inhibitors. In other words, MedSAGE consistently captures the essential features required for binding while exploring diverse chemotypes that embody those features.
[0121] In our final case study, the target is the crystal structure of the gp120 subunit of the HIV-1 envelope (Env) trimer (Fig. 6C). The reference ligand, compound 484, is an analog of Temsavir (BMS-626529), an approved antiretroviral medication. Like cmp 484 and Temsavir, the MedSAGE-generated molecule contains a central piperazine ring, a carbonyl group that forms a key hydrogen bond, and a phenyl ring positioned in a narrow pocket — key structural features shared across the class of known inhibitors. Notably, the generated molecule forms a hydrogen bond with an aspartate residue that is not made by cmp 484, but is made by temsavir, which is known to be more potent. Interestingly, the piperazine ring adopts a different geometry in the generated molecule, enabling an additional hydrogen bond interaction.
[0122] Interestingly, across these case studies, we observe that MedSAGE-generated molecules tend to form more hydrogen bonds than the reference ligands. This is also reflected in the higher average number of hydrogen bond donors (Fig. 10).MedSAGE identifies high-scoring ligands more efficiently than traditional virtual screening
[0123] Virtual screening (VS) is a widely used computational approach for identifying drug candidates by evaluating large chemical libraries with physics-based or Al scoring functions. However, the exponential increase in the size of these libraries — now reaching hundreds of billions of compounds — has dramatically increased computational demands. Generative methods, such as MedSage, offer a potential approach to complement or even replace traditional virtual screening methods, which can design chemically diverse molecules with high predicted affinity.
[0124] To assess the potential of generative methods for molecule design, we compared MedSAGE to virtual screening across three case studies using retrospective datasets. These datasets correspond to real world VS projects for three distinct drug targets: a transcription factor and two G protein-coupled receptors (GPCRs). The selectedtargets span a range of binding pocket characteristics, from a shallow interface to a deep, buried pocket. The virtual screening datasets were collected by docking and scoring tens to hundreds of millions of ligands from the Enamine REALSpace using Schrodinger’s Glide (>50 weeks of compute time per target). MedSAGE was used to generate 2000 molecules for each target, which were then scored with Glide (<1 day of compute time per target).
[0125] MedSAGE produced diverse, high-scoring molecules more efficiently than virtual screening across all three case studies (Fig. 11A). The median docking score of the top 10% of MedSAGE-generated molecules was better than that of the top 0.1 % of compounds identified through virtual screening. This trend increased at higher percentiles, where the median docking score of the top 1 % of MedSAGE-generated molecules was comparable to that of the top 0.001 % of VS hits (equivalent to selecting the top 10 out of every 1 million docked compounds). MedSAGE-generated molecules were also chemically diverse, with an average Tanimoto coefficient of less than 0.1 between pairs of molecular fingerprints, indicating low structural similarity (Fig. 11 B, depicting scores of (1 -Tanimoto coefficient)). The top 1 k molecules from virtual screening showed lower diversity, with an average Tanimoto coefficient above 0.1. The top 1 ,000 ligands yielded by MedSage averaged 595 unique scaffolds, whereas the top 1 ,000 ligands yielded by virtual screening averaged 560 unique scaffolds. The findings were consistent across all three drug targets, which reflected a wide range of docking scores.
[0126] In this context, a potential limitation of MedSAGE is that its generated molecules may be less readily synthesizable than those from combinatorial libraries, which are typically built from simple, predefined building blocks and reactions. To assess this, we qualitatively compared MedSAGE-generated molecules to top molecules from virtual screening (Fig. 11 C). Notably, many MedSAGE-generated molecules had high similarity to top library compounds, sharing same common chemical fragments. Therefore, MedSAGE can be used to rapidly screen chemical libraries of commercially available compounds to identify commercially and readily available high-affinity candidates. To do so, MedSAGE can design high-affinity candidates, which can be used to search for compounds within chemical libraries to identify analogs with high similarity.This approach can be used to quickly identify readily available compounds to streamline drug screening, saving time and resources.Conclusions
[0127] In this study, we developed MedSAGE, a state-of-the-art generative model for the de novo design of small molecules targeting protein structures, enabling rapid design of drug-like candidates with high predicted affinity and supportive of the various embodiments described herein. While diffusion models have achieved remarkable success in generating images and even biologies, their application to molecule design has often resulted in molecules that are often unfeasible from a medicinal chemistry perspective. We reimagine diffusion-based methods for medicinal chemistry, using functional groups — rather than individual atoms — as building blocks and mathematically encoding them through embeddings of chemical properties. MedSAGE drastically reduces the degrees of freedom required to represent and generate chemically sophisticated, drug-like molecules. This framework is particularly advantageous given the relatively small amount of high-quality training data. Our two-phase workflow — first placing fragments in 3D via a diffusion model and then establishing chemical bonds under physics-based and chemical validity constraints — merges the strengths of Al-driven generation with robust physical principles.
[0128] We conducted a rigorous and comprehensive evaluation of MedSAGE alongside other leading methods, benchmarking performance across a diverse set of protein targets and metrics. MedSAGE consistently achieved state-of-the-art results, generating molecules with docking scores on par with high-affinity reference ligands. Beyond binding affinity, the method produces ligands with structural and physicochemical properties that align with medicinal chemistry best practices. Notably, these outcomes emerge without explicitly optimizing for such properties during generation — suggesting that generative Al, when structured correctly, can implicitly learn and apply medicinal chemistry principles that are otherwise challenging to model.MethodsFragment Preparation
[0129] We created a custom fragment library by combining curated fragments relevant to drug discovery with automatically detected fragments from the prepared ligand dataset (Figs. 4A-4D). We also explicitly accounted for differing protonation and tautomeric states of each fragments (a protonated version of a fragment is a separate fragment in our library). The dataset consists of 3700 unique fragments.
[0130] To obtain the fragment library, we began by defining a minimal set of functional groups, {F}, that are important for drug-like ligands (methyl, ethyl, phenyl, carboxy, nitro, cyclohexyl, sulfonamide etc.). We then iterate over each molecule in our dataset to find additional fragments not in {F}. First each molecule is converted to a graph representation. Then we remove subgraphs that match a graph in {F} as long as these subgraphs are not part of rings and only connected to other subgraphs by single bonds. Then we split apart any ring structures that are connected by single bonds. The remaining connected components of the molecule graph are considered the non-library fragments and added to {F}. We do this for each ligand in our dataset to obtain the complete library. In many cases, fragments can be matched in multiple way. We start with fragments that are terminal, they are connected to the rest of the molecule by a single bond. If there are multiple terminal fragment matches, we remove the largest one first.
[0131] To create the fragment latent space, we computed a vector of properties for each fragment. This consisted of the Extended Three-Dimensional Fingerprint as well as chemical properties of charge, molecular weight, number of hydrogen bond donors / acceptors, and number of rings (for more on the Extended Three-Dimensional Fingerprint, see S.D. Axen, et. al., J Med Chem. 2017 Sep 14;60(17):7393-7409, the disclosure of which is hereby incorporated by reference). We then used t-SNE (sci-kit learn python library), to embed these vectors into a latent space with three dimensions, which was then normalized to be centered at 0 and have a standard deviation of 1 along each dimension. This approach yielded a smooth and interpretable latent space, and fragments with similar properties were found to cluster together.Architecture and Training
[0132] To position and assign fragments, we used Diffusion Denoising Probabilistic Models (DDPMs). DDPMs are a class of generative models that learn to reverse a diffusion process by gradually removing noise from a sample. The diffusion process starts with the fragment point cloud and progressively adds noise to the positions and features, eventually transforming them into Gaussian noise. The generative denoising process aims to learn the reverse mapping, starting from the Gaussian noise and iteratively removing the noise to reconstruct the original fragment latent points. During the noising and denoising processes, the pocket fragments remain fixed, and the example is centered at the centroid of the pocket fragments. We add additional context features to each point to specify whether the point belongs to the pocket or to the ligand. The learnable function that models the dynamics of the diffusion process is implemented as a E(3)-equivariant graph neural network (EGNN) following Hoogeboom et. al. where the fragment point cloud is converted to a fully connected graph, with edges labeled with distances between the points.Fragment Connecting
[0133] Given the 3D arrangement of fragments, we developed a custom algorithm to find fragment connections that result in a single fully connected molecule. The algorithm starts by identifying the "neighbor" relationships between the chemical fragments based on their 3D proximity. It creates a "neighbor graph" where nodes represent fragments, and edges represent potential connections between nearby fragments within a threshold distance. The algorithm then enumerates spanning trees from the neighbor graph. A spanning tree is a subset of edges from the neighbor graph that connects all nodes without forming any cycles. These spanning trees represent possible ways to connect the fragments while ensuring that each fragment is connected to at least one other fragment.
[0134] Given a proposed connectivity between the fragments, the algorithm identifies the unique attachment points (positions where fragments can be connected) on each fragment. The algorithm then enumerates all valid connections available between the fragments by considering the available attachment points and simple chemical rules for new bonds (e.g. don’t connect fluorine and oxygen).
[0135] The algorithm outputs a list of candidate molecules composed of the set of initial fragments linked in various ways that fit the 3D arrangement. The candidate molecules are then scored using a physics-based scoring function (Glide) in the context of the protein binding pocket. The best-scoring candidate from each unique set of fragments is selected based on the docking score. For large molecules, the connectivity algorithm can generate hundreds of candidates based on the numerous possibilities to link the fragments together. To maintain efficiency, these candidates can be randomly subsampled before proceeding to the scoring step. In practice, 30-40 were candidates was sufficient to optimize the connectivity, and scores subsequently plateau, but the actual number of candidates can vary dependent on desired outcome (Fig. 5).Benchmarking
[0136] Number of fragments: number of fragments for benchmarking was selected based on the number of fragments in the reference ligand. This same information was provided to all tested diffusion methods. This controlled for size variability. A second model could be used to predict the number of fragments or a distribution of the number of fragments.
[0137] Statistical analyses: In bar plots, data are reported as median values with 95% confidence intervals. For each target protein, metrics were first averaged across the generated molecules, and the median of these averages was then computed across all target proteins. This ensures that molecules generated for a particular target protein are not treated as statistically independent samples.Vina Docketing and Docking Score Cross Validation
[0138] We further evaluated predicted binding affinity using an alternative scoring method, AutoDock Vina. MedSAGE uses Glide docking scores to guide how to connect fragments — which could introduce a bias favoring Glide scores. Therefore, we wanted to assess whether MedSAGE-generated molecules maintained favorable affinity predictions with a different scoring function. Indeed, MedSAGE-generated molecules exhibited significantly better median Vina scores (-8.0) compared to the negative control set (-6.5, p < 0.001 , t-test) and were within 1 .2 units of the reference ligands (-9.2).
[0139] Interestingly, some molecules generated by other diffusion-based methods showed considerably better median Vina scores than Glide scores at first glance; however, upon inspection, we found many of these scores were unrealistically favorable (often < -20), well below even the best scores obtained by high-affinity reference ligands (lowest observed -13). Such extremely low scores typically suggest artifacts from the docking method.
[0140] To systematically address discrepancies arising from scoring-function biases and potential docking artifacts, we introduced a stringent scoring approach combining both Vina and Glide scores. Specifically, we retained only molecules for which Glide and Vina scores were within a tolerance threshold of 3 kcal / mol, then used the Vina scores. This threshold was informed by examining discrepancies between Glide and Vina scores for reference ligands, where the average difference in score was 0.5 kcal / mol, with a maximum difference of 2.5 kcal / mol. Applying this consistency metric revealed that MedSAGE-generated molecules continued to outperform those from other diffusionbased methods, though by a narrower margin than previously observed for Glide scores alone (Fig. 9). This result highlights MedSAGE’s robust performance across multiple affinity metrics and emphasizes the importance of cross-validation with different scoring functions to identify reliably high-affinity molecules.Table 1: Proteins with Therapeutic Relevance Selected for BenchmarkingTable 2: Benchmarking Metrics and Statistics for MedSAGE vs. Reference Ligands vs. REALSpaceMetrics are reported as the median with a 95% confidence interval (CI). The Reference vs. MedSAGE column shows p-values from a two-sided Welch’s t-test comparing Reference Ligands to MedSAGE-generated ligands. Similarly, the REALSpace vs. MedSAGE column presents p-values from a two-sided Welch’s t-test comparing REALSpace Ligands to MedSAGE-generated ligands. All statistical calculations were performed by first averaging results for each independent benchmarking protein, then calculating the statistic. Therefore, N = 25 for each metric and comparison.Supplementary Table 2: Benchmarking Metrics and Statistics for PMDM, DIffSBDD, IPDiff vs. Reference Ligands and REALSpaceMetrics are reported as the median with a 95% confidence interval (CI). The vs. REAL and vs. reference rows presents p-values from a two-sided Welch’s t-test comparing ligands to REALSpace ligands or reference ligands for the metric in the row directly above. All statistical calculations were performed by first aggregating results for each independent benchmarking protein, then calculating the statistic. Therefore, N = 25 for each metric and comparison.
Claims
WHAT IS CLAIMED IS:1 . A computational method to design molecules for interacting with a biological target, the method comprising: providing a structure of a target biomolecule having a site to be targeted; positioning, utilizing a computational diffusion model and a library of molecular fragments, molecular fragments upon the target site; wherein the molecule fragments are numerically embedded such that it captures chemical properties of the molecule fragments; wherein the computational diffusion model is trained utilizing a cohort of structural data of ligand molecules bound to its biological target; and wherein each molecule fragment of the set of molecular fragments are to be positioned simultaneously; and linking, utilizing a linking computational model, the set of molecular fragments to yield a small molecule.
2. The computational method of claim 1 , further comprising: acquiring the library of molecular fragments; and reducing, utilizing a dimensionality reduction technique, dimensionality of chemical properties of each fragment of the library of molecular fragments into a set of vectors.
3. The method of claim 2, wherein the reduction technique is one of the following: principal component analysis (PCA), T-distributed stochastic neighbor embedding (t- SNE), uniform manifold approximation and projection (UMAP), singular value decomposition (SVD), or linear discriminant analysis (LDA).
4. The method of any one of claims 1 -3, wherein the chemical properties comprise one or more of the following: charge, molecular weight, number of hydrogen bond donors / acceptors, or number of rings.
5. The method of any one of claims 2-4, wherein the dimensionality reduction technique further captures an extended three-dimensional fingerprint of the molecular fragments.
6. The method of any one of claims 1 -5, wherein the computational diffusion model is a diffusion denoising probabilistic model.
7. The method of any one of claims 1 -6 further comprising: training the computational diffusion model by: acquiring the cohort of structural data of molecules bound to its binding site of a biological target; fragmenting each molecule of the cohort into molecular fragments, wherein each molecular fragment of the small molecule is represented within the library of molecular fragments; embedding each molecular fragment into numerically embedded molecular fragments that captures the chemical properties of the molecule fragment; and entering the structural data of the biological target and the embedded molecular fragments into the computational diffusion model to learn to reconstruct the structural data from Gaussian noise.
8. The method of claim 8, wherein the computational diffusion model learns to reconstruct the structural data from Gaussian noise by: progressively adding noise to positions and features of the structural data of the biological target and the embedded molecular fragments until transformed into Gaussian noise; and iteratively removing the noise to reconstruct the positions and the features of the structural data of the biological target and the embedded molecular fragments.
9. The method of any one of claims 1 -8, wherein the molecules of the cohort of structural data include, prioritize, or are limited to ligands that are a particular type or have a particular functionality.
10. The method of any one of claims 1-9, wherein the linking computational model is a neighbor-relationship-based computational model that links molecular fragments by identifying neighbor relationships of the molecular fragments based on their three- dimensional proximity.
11. The method of claim 10, wherein the identifying neighbor relationships of the molecular fragments based on their three-dimensional proximity comprises generating a neighbor graph comprising nodes and edges; wherein the nodes represent the molecular fragments and the edges represent potential connection between nearby molecular fragments.
12. The method of claim 11 , wherein linking the set of molecular fragments comprises enumerating spanning trees from the neighbor graph, wherein a spanning is a subset of edges from the neighbor graph that connects all nodes, wherein the spanning trees represent possible ways to connect the fragments while ensuring that each fragment is connected to at least one other fragment.
13. The method of claim 12, wherein a spanning tree is a subset of edges from the neighbor graph that connects all nodes without forming any cycles.
14. The method of any one of claims 1 -13, wherein the linking computational model is an edge prediction model.
15. The method of any one of claims 1 -14 further comprising synthesizing the small molecule.
16. The method of claim 15 further comprising assessing the synthesized small molecule for binding to the biological target.
17. The method of claim 15 or 16 further comprising assessing the synthesized small molecule in a biological function assay to assess an ability of the small molecule to alter a biological function of the biological target.
18. The method of any one of claims 1-14, wherein the method is repeated a plurality of times to yield a plurality of small molecules.
19. The method of claim 18 further comprising identifying one or more candidate small molecules from the plurality of small molecules by: assessing the plurality of molecules for a set of desired characteristics; wherein the one or more candidate small molecules are assessed to have a desired characteristic greater than a threshold.
20. The method of claim 18 or 19 further comprising identifying one or more candidate small molecules from the plurality of small molecules by: filtering out one or more small molecules from the plurality small molecules based on a set of desired characteristics; wherein the one or more molecules filtered out are determined to be below a threshold for a desired characteristic.
21. The method of claim 19 or 20, wherein the set of desired characteristics include one or more of: difficulty to synthesize, cost to synthesize, solubility, bioavailability, polarity, pKa, or stability.
22. The method of any one of claims 19-21 further comprising synthesizing the one or more candidate small molecules.
23. The method of claim 22 further comprising assessing the synthesized one or more candidate small molecules for binding to the biological target.
24. The method of claim 22 or 23 further comprising assessing the synthesized one or more candidate small molecules in a biological function assay to assess an ability of the one or more candidate small molecules to alter a biological function of the biological target.
25. A computational method to design small molecules for binding to a biological target, the method comprising: providing a structure of biological target having a binding site to be targeted; positioning, utilizing a computational diffusion model and a library of molecular fragments, a set of molecular fragments within the target binding site; wherein the molecular fragments are embedded as a set of vectors via a dimensionality reduction technique that captures chemical properties of the molecular fragments; wherein the computational diffusion model is trained utilizing a cohort of structural data of small molecules bound to its binding site of a biological target; and wherein each molecular fragment of the set of molecular fragments are to be positioned simultaneously; and linking, utilizing a neighbor-relationship-based computational model, the set of molecular fragments to yield a small molecule.
26. The computational method of claim 25, further comprising: acquiring the library of molecular fragments; and reducing, utilizing the dimensionality reduction technique, dimensionality of chemical properties of each fragment of the library of molecular fragments into its set of vectors.
27. The method of claim 25 or 26, wherein the reduction technique is one of the following: principal component analysis (PCA), T-distributed stochastic neighbor embedding (t-SNE), uniform manifold approximation and projection (UMAP), singular value decomposition (SVD), or linear discriminant analysis (LDA).
28. The method of any one of claims 25-27, wherein the chemical properties comprise one or more of the following: charge, molecular weight, number of hydrogen bond donors / acceptors, or number of rings.
29. The method of any one of claims 25-28, wherein the dimensionality reduction technique further captures an extended three-dimensional fingerprint of the molecular fragments.
30. The method of any one of claims 25-29 wherein the computational diffusion model is a diffusion denoising probabilistic model.31 . The method of any one of claims 25-30 further comprising: training the computational diffusion model by: acquiring the cohort of structural data of small molecules bound to its binding site of a biological target; fragmenting each small molecule of the cohort into molecular fragments, wherein each molecular fragment of the small molecule is represented within the library of molecular fragments; reducing, utilizing the dimensionality reduction technique, each molecular fragment into embedded molecular fragments comprising a set of vectors; and entering the structural data of the biological target and the embedded molecular fragments into the computational diffusion model to learn to reconstruct the structural data from Gaussian noise.
32. The method of claim 31 , wherein the computational diffusion model learns to reconstruct the structural data from Gaussian noise by: progressively adding noise to positions and features of the structural data of the biological target and the embedded molecular fragments until transformed into Gaussian noise; and iteratively removing the noise to reconstruct the positions and the features of the structural data of the biological target and the embedded molecular fragments.
33. The method of any one of claims 25-32, wherein the small molecules of the cohort of structural data of small molecules bound to its binding site are known drugs.
34. The method of any one of claims 25-33, wherein the neighbor-relationship-based computational model links molecular fragments by identifying neighbor relationships of the molecular fragments based on their three-dimensional proximity.
35. The method of claim 34, wherein the identifying neighbor relationships of the molecular fragments based on their three-dimensional proximity comprises generating a neighbor graph comprising nodes and edges; wherein the nodes represent the molecular fragments and the edges represent potential connection between nearby molecular fragments.
36. The method of claim 35, wherein linking the set of molecular fragments comprises enumerating spanning trees from the neighbor graph, wherein a spanning is a subset of edges from the neighbor graph that connects all nodes, wherein the spanning trees represent possible ways to connect the fragments while ensuring that each fragment is connected to at least one other fragment.
37. The method of claim 36, wherein a spanning tree is a subset of edges from the neighbor graph that connects all nodes without forming any cycles.
38. The method of any one of claims 25-37 further comprising synthesizing the small molecule.
39. The method of claim 38 further comprising assessing the synthesized small molecule for binding to the biological target.
40. The method of claim 38 or 39 further comprising assessing the synthesized small molecule in a biological function assay to assess an ability of the small molecule to alter a biological function of the biological target.41 . The method of any one of claims 25-37, wherein the method is repeated a plurality of times to yield a plurality of small molecules.
42. The method of claim 41 further comprising identifying one or more candidate small molecules from the plurality of small molecules by: assessing the plurality of small molecules for a set of desired characteristics; wherein the one or more candidate small molecules are assessed to have a desired characteristic greater than a threshold.
43. The method of claim 41 or 42 further comprising identifying one or more candidate small molecules from the plurality of small molecules by: filtering out one or more small molecules from the plurality small molecules based on a set of desired characteristics; wherein the one or more molecules filtered out are determined to be below a threshold for a desired characteristic.
44. The method of claim 42 or 43, wherein the set of desired characteristics include one or more of: difficulty to synthesize, cost to synthesize, solubility, bioavailability, polarity, pKa, or stability.
45. The method of any one of claims 42-44 further comprising synthesizing the one or more candidate small molecules.
46. The method of claim 45 further comprising assessing the synthesized one or more candidate small molecules for binding to the biological target.
47. The method of claim 45 or 46 further comprising assessing the synthesized one or more candidate small molecules in a biological function assay to assess an ability of the one or more candidate small molecules to alter a biological function of the biological target.
Citation Information
Patent Citations
Systems and methods for artificial intelligence-guided biomolecule design and assessment
US11450407B1
Systems and Methods for Generating Ligand Compounds
US20230317212A1