Generating torsion profiles using circular mixture distributions

The use of a circular mixture distribution and machine learning models to generate torsion profiles of small molecules addresses inefficiencies in existing methods, providing high-precision torsion angle predictions for drug development applications.

WO2026017861A1PCT designated stage Publication Date: 2026-01-22MONTE ROSA THERAPEUTICS AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/070655
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2025-07-18
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing methods for generating three-dimensional molecular conformations are inefficient and lack accuracy, particularly in identifying relevant conformations for drug development applications, as they rely on classical force fields or computationally expensive ab initio methods.

Method used

A method using a circular mixture distribution to generate torsion profiles of rotatable bonds in small molecules, employing a graph neural network and a distribution prediction machine learning model to process molecular structures and predict torsion profiles with high precision.

Benefits of technology

Enables high-precision generation of torsion profiles, facilitating the selection of relevant conformations for drug development by accurately predicting torsion angles and strain energy, thereby improving the efficiency and accuracy of molecular docking and virtual screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025070655_22012026_PF_FP_ABST
    Figure EP2025070655_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Described herein are techniques for generating the torsion profile of rotatable bonds of a small molecule using a circular distribution, and methods of using the same.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GENERATING TORSION PROFILES USING CIRCULAR MIXTURE DISTRIBUTIONS

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 673,347, filed on July 19, 2024, the contents of which is herein incorporated by reference.

[0004] TECHNICAL FIELD

[0005] Described herein are methods for generating torsion profiles for a small molecule using machine learning.

[0006] BACKGROUND

[0007] This specification relates to torsion profiles of a molecule. A torsion profile provides a measure of the potential energy variation of a molecule as a function of its torsional angle. A torsion angle is a dihedral angle between two planes formed by four atoms in the molecule.

[0008] This specification also relates, in part, to processing data using machine learning models. Machine learning models receive an input and generate an output, e.g., a prediction, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.

[0009] SUMMARY

[0010] Molecular conformation determines the chemical and physical properties of a molecule. Conformation generation is widely used in different applications in drug development such as quantitative structure-activity relationships (QSAR), molecular docking and virtual screening. However, generating three-dimensional (3D) structure is not trivial since the produced ensemble of conformations includes identifying relevant diverse conformations that are representative of all possible conformations in a conformational space.

[0011] For example, low-energy conformation or conformation-having torsion angles that are observed in experimental structures are usually treated as relevant conformations. To handle 3D conformation generation at large-scale, the conformation energy is often calculated from classical force field methods, while torsion likelihood is extracted from a torsion angle library that stores torsion knowledge derived from one or more experimental structure databases, e.g., the Cambridge Structural Database (CSD). Both methods are efficient but with limited accuracy. Other methods to calculate the energy with greater accuracy exist, e.g., ab initio methods, but can be computationally inefficient.

[0012] Described herein are high-precision techniques for generating the torsion profile of the rotatable bonds of a small molecule using a circular mixture distribution. In this context, the rotatable bond angle refers to a dihedral angle between two planes formed by four atoms in the molecule. Generating the torsion profile of a small molecule can inform the selection of relevant conformations, e.g., based on the strain of the ligand.

[0013] Thus, provided herein are methods for generating the torsion profile of rotatable bonds of a small molecule using a circular distribution, and methods of using the same.

[0014] In one aspect, there is provided a method for receiving data representing a structure of a molecule, processing the data using a graph neural network to generate an embedding of a graph representation of at least one portion of the molecule, wherein the graph representation includes: a set of node embeddings, each node embedding representing an atom in the at least one portion of the molecule; a set of edge embeddings, each edge embedding representing a bond between a pair of atoms in the at least one portion of the molecule; one or more sets of torsional features, each including an identification of a torsion atom in the at least one portion of the molecule, wherein a torsion atom comprises a node at a terminus of a rotatable bond, and a corresponding phase angle of the rotatable bond, and processing the embedding of the graph representation using a distribution prediction machine learning model to generate one or more parameters of a circular mixture distribution characteristic of a torsion profile of the at least one portion of the molecule.

[0015] In some implementations, the method further includes generating the torsion profile of the at least one portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution.

[0016] In some implementations, receiving data representing the structure of a molecule includes receiving data representing a two-dimensional structure of a molecule.

[0017] In some implementations, receiving data representing the two-dimensional structure of the molecule includes receiving a Simplified Molecular Input Line Entry System (SMILES) string. In some implementations, receiving data representing the structure of the molecule further includes generating one or more fragments of the molecule, wherein each fragment includes a portion of the molecule.

[0018] In some implementations, generating the one or more fragments includes identifying one or more rotatable bonds in the molecule, at least one torsion atom associated with each rotatable bond, and one or more local atoms to the at least one torsion atom.

[0019] In some implementations, identifying the one or more local atoms to the at least one torsion atom includes, for each torsion atom, identifying one or more covalently bonded atoms to the torsion atom, evaluating the one or more covalently bonded atoms to identify partial structures including the one or more covalently bonded atoms, wherein the partial structures comprise one or more of functional groups and ring systems, and including atoms of the partial structures as local atoms in the one or more fragments.

[0020] In some implementations, the method further includes, in response to an identification of a ring system, identifying possible ortho substitutions of the ring system as a partial structure.

[0021] In some implementations, the graph representation of at least one portion of the molecule includes a graph representation of one or more fragments of the molecule.

[0022] In some implementations, processing the data using the graph neural network to generate the embedding of a graph representation of at least one portion of the molecule includes processing the data representing the at least one portion of the molecule to generate a set of node embeddings and a set of edge embeddings, and to identify one or more torsion features, using one or more augmentation neural networks to update the set of edge embeddings in accordance with the one or more torsion features, and updating the set of node embeddings and the set of updated edge embeddings by performing message passing operations conditioned on a topology of the graph representation.

[0023] In some implementations, using one or more augmentation neural networks to update the set of edge embeddings in accordance with the torsion features includes, for each edge embedding updating an edge embedding of a non-rotatable bond by processing one or more node embeddings of atoms corresponding with the non-rotatable bond using an augmentation neural network, or updating an edge embedding of a rotatable bond by generating and aggregating respective outputs of at least two augmentation neural networks.

[0024] In some implementations, updating an edge embedding of the rotatable bond includes generating a respective output with each augmentation neural network by processing one or more node embeddings of atoms corresponding with the rotatable bond using the augmentation neural network, and aggregating the respective outputs of each augmentation neural network by summing a first output, a second output weighted by a sine of the phase angle of the torsion features, and a third output weighted by a cosine of the phase angle.

[0025] In some implementations, the method further includes generating the embedding of the graph representation using graph pooling.

[0026] In some implementations, generating the embedding of the graph representation using graph pooling includes using an attention mechanism to determine a set of attention weights corresponding with the set of node embeddings, and aggregating the node embeddings in accordance with the attention-weights to generate the embedding of the graph representation.

[0027] In some implementations, the distribution prediction machine learning model includes a mixture density network configured to generate a set of parameters for N circular distributions in the circular mixture distribution.

[0028] In some implementations, the one or more parameters include, for each circular distribution in the circular mixture distribution, a set of parameters including a mean value, a standard deviation value, and a mixing coefficient value.

[0029] In some implementations, the graph neural network, one or more augmentation neural networks, and distribution prediction machine learning model have been jointly trained by operations including, for each of a number of training iterations, receiving a corresponding ground truth torsion profile for each rotatable bond in the at least one portion of the molecule, identifying a set of allowable torsion angle values for the at least one portion of the molecule, sampling one or more torsion angle values from the identified set of allowable values, for each sampled torsion angle values, processing the at least one portion of the molecule with the sampled torsion angle value to generate the one or more parameters of the circular mixture distribution characteristic of the torsion profile of the at least one portion of the molecule, and updating values of respective sets of parameters including a set of graph neural network parameters, one or more sets of augmentation neural network parameters, and a set of distribution prediction machine learning model parameters in accordance with minimizing a discrepancy between the ground truth torsion profile of the at least one portion of the molecule and the generated torsion profile of the at least one portion of the molecule.

[0030] In some implementations, the method further includes, for each molecule in a set comprising a plurality of molecules, generating the one or more parameters of the circular mixture distribution characteristic of the torsion profile of the at least one portion of the molecule, generating the torsion profile of the at least one portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution, and using the torsion profile to characterize a likelihood of observing a conformer of the molecule in a binding site of a second molecule.

[0031] In some implementations, the method further includes selecting one or more of the molecules for physical synthesis based at least on the respective likelihood of observing the conformer of the molecule in the binding site of the second molecule.

[0032] In some implementations, the method further includes physically synthesizing one or more of the selected molecules.

[0033] In some implementations, the method further includes, for each portion of the molecule comprising a rotatable bond, generating the torsion profile of the portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution generated for the portion of the molecule, and performing conformational sampling using the torsion profiles corresponding with each portion of the molecule.

[0034] In some implementations, performing conformational sampling using the torsion profiles corresponding with each portion of the molecule includes sampling a torsion angle from the torsion profile for each portion of the molecule, and generating a molecular conformation defined by the sampled torsion angles.

[0035] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0036] As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a sample” includes a plurality of samples, including mixtures thereof.

[0037] The terms “determining,” “measuring,” “evaluating,” “assessing,” “assaying,” and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of’ can include determining the amount of something present in addition to determining whether it is present or absent depending on the context.

[0038] As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value.

[0039] Throughout this specification, an “embedding” refers to an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values.

[0040] A “set” of entities refers to a collection of one or more of entities

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0042] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.

[0043] DESCRIPTION OF DRAWINGS

[0044] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0045] FIG. l is a diagram showing an example of a system for predicting the torsion profile of a molecule using machine learning.

[0046] FIG. 2 depicts an example of a torsion profile for a torsional fragment of a small molecule.

[0047] FIGS. 3 A and 3B depict an example of torsion fragment generation.

[0048] FIG. 4 shows an example of multiple torsion atom sets around a single rotatable bond.

[0049] FIG. 5 shows an example of a method for torsion profile prediction directly from small molecule two-dimensional (2D) structures.

[0050] FIG. 6 depicts an example machine learning model architecture for predicting torsion profile(s) directly from small molecule 2D structure(s). FIG. 7 shows an example circular distribution and demonstrates the advantages of predicting a torsion profile using a circular distribution.

[0051] FIG. 8A depicts example results of torsion profiles, e.g., using the example system of FIG. 1.

[0052] FIG. 8B illustrates the accuracy of the energy strain computed from a torsion profile generated with the techniques of this specification with respect to the energy strain computed from a torsion profile generated using a molecular mechanics force-field function for each a number of torsion angles.

[0053] FIG. 9 is a flow diagram of an example process for generating a torsion profile using parameters of a circular mixture distribution.

[0054] FIG. 10 is a flow diagram of an example process for generating an embedding of a graph representation of a molecule.

[0055] DETAILED DESCRIPTION

[0056] Provided herein are high-precision techniques for generating the torsion profile of rotatable bonds of a small molecule. A rotatable bond refers to a chemical bond between at least two atoms of a molecule with one or more degrees of freedom, e.g., the bond can freely rotate around an axis without breaking the bond. In this context, the rotatable bond angle refers to a dihedral angle between two planes formed by four atoms in the molecule. The torsion profile provides a measure of the potential energy variation of a molecule as a function of its torsional angle.

[0057] A key property of torsion itself is that it is periodic. If the bond is rotated through 360 degrees, the energy should stay the same. In this specification, there are two approaches described that fulfill this property.

[0058] Small molecule ligands can have multiple conformations — the conformation in solution may or may not be the conformation that the small molecule takes when in complex, e.g., as part of a ternary complex. This so-called bioactive conformation interacts with the environment that can accordingly change the energy hypersurface of the molecule. Therefore, some conformations may have higher energy than others in bound state which they might have to composite with interactions to the environment. Thus, provided herein are methods for calculating strain energy from a molecule that can be used to select more relevant conformations, e.g., more relevant bioactive conformations. TERNARY COMPLEX MODELS

[0059] The methods and compositions described herein are useful, for example, in scoring model(s) of ternary complex(es) (e.g., two proteins and a small molecule such as a molecular glue). In some cases, the model(s) are of region(s) of the ternary complex(es), e.g., the interface region(s) of the ternary complex. While ternary complexes are contemplated specifically, as will be understood by a person of skill in the art, the methods are applicable and can be extended to complexes of one or more proteins in combination with one or more small molecule ligands.

[0060] In some cases, the model(s) are obtained from a database. For example, the Protein Data Bank (PDB, wwPDB consortium. (2019) Protein Data Bank: the single global archive for 3D macromolecular structure data. Nucleic Acids Res 47: D520-D528) or the AlphaFold Protein Structure Database (Varadi, M. et al. “AlphaFold Protein Structure Database in 2024: Providing structure coverage for over 214 million protein sequences.” Nucleic Acids Research, gkadlOl 1 (2023)). In some cases, the model(s) are generated by determining a three-dimensional structure experimentally, e.g., using X-ray crystallyography, nuclear magnetic resonance (NMR spectroscopy), cry-electron microscropy (cryoEM), small-angle X-ray scattering (SAXS), small-angle neutron scattering (SANS), or combinations thereof. In some cases, the model(s) are generated using computer modeling. In some cases, the model(s) of ternary complexes are generated, e.g., as described in PCT / US2023 / 082629. In some cases, the computer modeling is carried out using an artificial intelligence program, e.g., according to the methods described in Jumper et al., “Highly Accurate Protein Structure Prediction with AlphaFold,” Nature 596:583-89 (2021) or Evans et al., “Protein Complex Prediction with AlphaFold-Multimer,” bioRxiv (2021).

[0061] Ternary complex(es) can be represented by a set of data parameters in digital format, e.g., in a PDB file, for use in the methods described herein. Thus, in some cases, the methods described herein comprise providing data defining the three-dimensional structure of the ternary complex(es). Various software is available for the generation and manipulation of data files defining three-dimensional structures. For example, PyMOL, ROSETTA, MAESTRO, ChimeraX, AlphaFold2, etc.

[0062] In some cases, the ternary complex comprises an E3 ligase substrate receptor protein, a target protein (e.g., neosubstrate), and a small molecule (e.g., molecular glue).

[0063] In some cases, the ligase substrate receptor protein is an E3 ligase substrate receptor protein identified in the following table. Table 1. Examples of E3 ligase substrate receptor proteins

[0064] In some cases, the E3 ligase substrate receptor protein is at least 80%, e.g., at least

[0065] 90%, at least 95%, or at least 99% identical to an E3 ligase substrate receptor protein from

[0066] Table 1.

[0067] In some cases, the E3 ligase is an enzymatically active portion of an E3 ligase substrate receptor protein from Table 1.

[0068] The compositions and methods describe herein are useful, for example, in identification and / or prediction of programmable E3 ligase substrate receptor proteins and, optionally, target proteins, e.g., as described in WO2023091567, which is hereby incorporated by reference in its entirety.

[0069] Molecular Glues

[0070] In some cases, the E3 ligase binding modulator is a molecular glue.

[0071] A molecular glue is a small molecule that stabilizes the interaction of two or more biomolecules (e.g., proteins) at a protein-protein interaction (PPI) interface, e.g., by chemically inducing or strengthening surface interactions between the proteins. In some cases, the molecular glue stabilizes the interaction of an E3 ligase substrate receptor protein and one or more target protein(s).

[0072] In some cases, the molecular glue functions as a molecular glue drug by modulating (e.g., increasing or promoting) one or more of: the stability of protein-protein interact! on(s), degradation of protein(s), sequestration of protein(s) (e.g., into specific regions of a cell), phosphorylation of protein(s), de-phosphorylation of protein(s), and stabilization of protein(s).

[0073] In some cases, the modulation is directly of the target protein (the “glued” target). In some cases, the modulation is indirect (e.g., of a target downstream of the “glued” target). Molecular Glue Degraders

[0074] Thalidomide and immunomodulatory imide drugs (IMiDs), such as lenalidomide, and pomalidomide, are examples of molecular glue drugs that induce degradation of normally unrecognized target proteins (sometimes referred to as “neosubstrates") by generating an interaction between an E3 ligase substrate receptor (e.g., cereblon) and a target protein (e.g., IKZF1 / 3).

[0075] Molecular glue drugs, such as these, that induce the degradation of protein(s) are sometimes referred to as a molecular glue degraders. Molecular glue degraders are believed to create neosubstrate recognition interfaces on the surface of the E3 ligase substrate receptor protein that engage in induced protein-protein interactions with neosubstrates.

[0076] TORSION PROFILING FOR SMALL MOLECULES

[0077] In some cases, the methods described herein comprise generating torsional profile(s) for a small molecule or fragment(s) thereof.

[0078] In some cases, generating a torsional profile comprises generating torsion fragment(s). In some cases, generating torsional profiles comprises torsional scan(s), e.g., of torsion fragment(s). In some cases, the torsional scan(s) comprise Molecular Mechanical (MM) methods and / or Quantum Mechanical (QM) methods. In some cases, generating torsional profiles comprises predicting torsional profiles, e.g., of torsion fragment(s), e.g., with a machine learning algorithm. In some cases, generating torsional profiles comprises predicting torsional profiles directly from a small molecule 2D structure, without directly generating torsion fragment(s), e.g., with a machine learning algorithm.

[0079] In some cases, the torsion profile is a function, such that f(x) = y, where x is the conformation (e.g., torsional angle) around a rotatable bond of a torsion fragment and y is the energy or predicted energy of the conformation (e.g., from MM and / or QM calculations or predictive models). In some cases, the torsion profile is a continuous function interpolated from a set of discrete conformations (x) and relative energies (y) for a plurality of conformations of a small molecule or fragment(s) thereof. An example of a torsion profile for a torsion fragment of a small molecule is depicted in FIG. 2.

[0080] In some cases, the method comprises generating an energy score (e.g., a relative energy score) for the small molecule or fragment(s) thereof, e.g., based on the torsion profile(s). In some cases, the energy score is calculated by subtracting the global minimum (e.g., identified using a continuous function as described above) from the energy score (e.g., relative energy score) of the small molecule or fragment thereof (e.g., an energy score generated using any of the methods described herein).

[0081] An example system for generating a torsion profile of a small molecule or fragment is depicted in FIG. 1. The torsion profile prediction system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0082] In particular, the torsion profile prediction system 100 can receive molecular structure data 110 characterizing a small molecule or fragment of a small molecule, e.g., a portion of the small molecule. As an example, the data can represent the two-dimensional structure of a molecule, e.g., as a Simplified Molecular Input Line Entry System (SMILES) string. In some cases, the system 100 can receive the data and generate one or more fragments 125 of the molecule, e.g., using a fragmentation engine 120, as is described in detail below.

[0083] Fragmentation Based Torsion Profiling

[0084] In some cases, a torsional profile is generated based on conformer generation and torsional scanning of torsion fragments (e.g., torsional fragments generated as described above). An example of fragmentation based torsion profiling is depicted in FIG. 2.

[0085] Torsion Fragments

[0086] In some cases, generating a torsional profile comprises generating torsion fragments for a small molecule. In some cases, the torsion fragment(s) each comprise a rotatable bond and flanking regions representing the essential chemical environment of the small molecule surrounding the rotatable bond. In particular, the torsion profile prediction system 100 can use the fragmentation engine 120 to process the molecular structure data 110 to generate one or more fragments 125.

[0087] In some cases, generating torsion fragments comprises generating substructures of the small molecule, each comprising a set of atoms comprising a rotatable bond and flanking regions representing the essential chemical environment of the small molecule surrounding the rotatable bond. In one example, generating a substructure of the small molecule comprises: identifying starting torsion atom(s) for the small molecule as belonging to the set of atoms (e.g., 4 starting torsion atom(s)) and incrementally expanding the set of atoms as follows: first, atoms covalently bonded to the starting torsion atoms (i.e., neighboring atoms which are one bond removed from the initial torsion atoms) are added to the set, with partially selected ring systems being added completely to prevent ring breaking; then, functional groups for which some atoms are already included in the set are added to the set in their entirety; finally, substitutions at ortho positions are added, if they are present.

[0088] The substructure comprising the set of atoms is then extracted from the parent molecule to form the torsion fragment. In the case in which there are invalid valences in the torsion fragment, the system can add hydrogen to carbons or methyl groups to non-carbons to fix the valences. This procedure is iterated for additional starting torsion atom(s) to yield a set of torsion fragments representing the small molecule.

[0089] Functional groups are sets of connected atoms that determine properties and reactivity of a parent molecule. Functional groups are defined through their heteroatoms and its environments.

[0090] Identification of functional groups is described, for example, in Patai’s Chemistry of Functional Groups. Wiley. doi: 10.1002 / 9780470682531; Feldman et al., (2005) CO: a chemical ontology for identification of functional groups and semantic comparison of small molecules. FEBS Lett 579:4685-4691; Bobach Cet al., (2012) Automated compound classification using a chemical ontology. J Cheminform 4:40; Djoumbou et al., (2016) ClassyFire: automated chemical classification with a comprehensive, computable taxonomy. J Cheminform 8:61; Lewis et al., (2015) Modern 2D QSAR for drug discovery. WIREs Comput Mol Sci 4:505-522; Salmina et al., (2016) Extended functional groups (EFG): an efficient set for chemical characterization and structure-activity relationship studies of chemical compounds. Molecules 21 : 1-8; Sterling T et al., (2015) ZINC 15 - Ligand discovery for everyone. J Chem Inf Model 55:2324-2337; Baell JB et al., (2010) New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays J Med. Chem 53:2719-2740; and Bruns RF (2012) Watson IA rules for identifying potentially reactive or promiscuous compounds. J MedChem 55:9736-9772.

[0091] In some cases, the functional groups are a set of manually curated set of substructure features. In some cases, the functional groups are automatically identified in the small molecule, e.g., as described in Ertl, P., “An algorithm to identify functional groups in organic molecules,” J Cheminform 9, 36 (2017).

[0092] In some cases, the functional groups are identified in the small molecule as follows, iterating through non-aromatic atoms: 1. mark all heteroatoms in a molecule, including halogens; 2. mark also the following carbon atoms: atoms connected by non-aromatic double or triple bond to any heteroatom, atoms in nonaromatic carbon-carbon double or triple bonds, acetal carbons, i.e. sp3 carbons connected to two or more oxygens, nitrogens or sulfurs; these O, N or S atoms must have only single bonds, all atoms in oxirane, aziridine and thiirane rings (such rings are traditionally considered to be functional groups due to their high reactivity); 3. merge all connected marked atoms to a single functional group; 4. extract function groups also with connected unmarked carbon atoms, these carbon atoms are not part of the functional group itself, but form its environment; 5. optionally, sp3 carbons connected to urea are considered as atoms of the functional group.

[0093] In some cases, the functional groups are identified in the small molecule as follows: 1. environments on carbon atoms are deleted, the only exception are substituents on carbonyl that are retained (to distinguish between aldehydes and ketones); 2. all free valences on heteroatoms are filled by the “R atoms” (this atom may represent hydrogen or carbon) with exception of: hydrogens on the -OH groups, and hydrogens on the simple amines and thiols (i.e. FGs with just single central N or S atom) are not replaced, this allows to distinguish secondary and tertiary amines, and thiols and sulfides; 3. all remaining environment carbons (on heteroatoms and carbonyls) are replaced by the “R atoms”; exceptions are environments on single atomic N or O FGs with one carbon connected, where this carbon is retained also with its type (aliphatic or aromatic), this allows to distinguish between amines and anilines, and alcohols and phenols; 4. optionally, sp3 carbons connected to urea are considered as atoms of the functional group.

[0094] In some cases, the ring system definition supersedes a functional group definition, e.g., of a traditionally considered functional group such as oxirane, aziridine, or thiirane.

[0095] An example of generating torsion fragments is depicted in FIGS. 3A-3B. For example, the torsion profile prediction system 100 can use the fragmentation engine 120 to generate the torsion fragments in FIGS. 3A-3B.

[0096] In particular, FIG. 3 A shows a fragmentation procedure with modified CC-885 molecular glue. For the sake of clarity of the procedure, the CC-885 was modified by substituting one of the hydrogens at ortho position of the phenyl ring with a methyl group. The highlighted atoms in the numbered five steps show the incremental generation of a fragment for a specific torsion. The four torsion atoms severed as initial atom set of the fragment. The atom set was further expanded by adding their neighboring atoms (atoms within one bond away). To avoid ring breaking, the remaining atoms overlapping with the ring were added. Functional groups were added as defined in Ertl, “An Algorithm to Identify Functional Groups in Organic Molecules,” Journal of Cheminformatics 9, 36 (2017), with the addition that sp3 carbons connected to urea were considered as atoms of the functional group as well. The missing carbon atom was added to prevent any truncation of the defined functional group. Finally, the ortho substitution of the phenyl ring was added.

[0097] FIG. 3B shows minimal torsion fragments of the modified CC-885. The first fragment was generated as described in Figure 2. The rest of the fragments were generated by considering other torsion atoms of modified CC-885 as the starting atom set for the fragmentation algorithm. Torsion atoms are highlighted in red.

[0098] Conformer Generation

[0099] In some cases, the methods described herein comprise providing or generating a plurality of 3D conformers of the small molecules or torsion fragments. In some cases, the 3D conformer data is provided directly. In some cases, the 3D conformer is generated, e.g., from structural data for small molecule(s) and / or scaffolds (e.g., ID and / or 2D data). In some cases, generating 3D conformers comprises generating 1 3D conformer per stereoisomer, e.g., using RDKit, OMEGA, Corina, etc. In some cases, stereoisomers are 1 per rotatable bond flipped on all stereo centers. In some cases, energy (e.g., in kcal / mol) is generated for the 3D conformers. 3D conformers are generated using any suitable means, including, as one example, as described in Hawkins et al., “Conformer Generation with OMEGA: Algorithm and Validation Using High Quality Structures from the Protein Databank and the Cambridge Structural Database,” J. Chem. Inf. Model. 2010, 50, 572-584.

[0100] In the methods described herein, generating a plurality of conformers per torsion fragment functions, for example, to increase the probability that the true global minimum conformation is identified at each angle during torsion scanning.

[0101] In some cases, from 10 to 500 conformers, e.g., from 10 to 450, 10 to 400, 10 to 350, 10 to 300, 10 to 250, 10 to 200, 10 to 150, 10 to 100, 10 to 50, 50 to 500, 50 to 450, 50 to 400, 50 to 350, 50 to 300, 50 to 250, 50 to 200, 50 to 150, 50 to 100, 100 to 500, 100 to 450, 100 to 400, 100 to 350, 100 to 300, 100 to 250, 100 to 200, 100 to 150, 150 to 500, 150 to 450, 150 to 400, 150 to 350, 150 to 300, 150 to 250, 150 to 200, 200 to 500, 200 to 450, 200 to 400, 200 to 350, 200 to 300, 200 to 250, 250 to 500, 250 to 450, 250 to 400, 250 to 350, 250 to 300, 300 to 500, 300 to 450, 300 to 400, 300 to 350, 350 to 500, 350 to 450, 350 to 400, 400 to 500, 400 to 450, or 450 to 500 conformers are generated for each torsion fragment. In some cases, at least 10, e.g., at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 conformers are generated for each torsion fragment. In some cases, the plurality of conformers is pruned. In some cases, pruning the conformers comprises removing conformers with an RMSD difference of less than 0.5 A among conformers.

[0102] In some cases, the conformers in the plurality are minimized. In some cases, the conformers are minimized using a force field, e.g., a molecular mechanics (MM) force field, e.g., MM2, MM3, MMFF94, MMFF94s, etc., or a universal force field, e.g., UFF, UFF4MOF, UFF-TIC, UFF-4m, etc.

[0103] Energy Score Calculations

[0104] Varying the torsion in a molecule not only affects the electronic character of the rotatable bond, but also nonbonded and steric interactions between chemical groups of the opposite sides of the bond. Thus, in some cases, the methods described herein comprise generating energy score(s) for different 3D conformations of a small molecule or fragment(s) thereof, e.g., for the plurality of conformers of the torsion fragments described above. In some cases, generating an energy score comprises a molecular mechanics (MM) and / or quantum mechanics (QM) calculation.

[0105] In some cases, method comprises rotating a reference angle of a rotatable bond (e.g., a conformer of the plurality of conformers of the torsion fragments described above) in degree increments to generate a set of test angles, and calculating an energy score (e.g., a relative energy score, e.g., relative to the reference angle) for each of the reference and test angles. In some cases, the reference angle is rotated in 1-359° increments, e.g., 1-30° increments, preferably 15° increments or thereabouts to generate a set of test conformations, e.g., for a full rotation around the bond (from 0° to 360°). Rotating in 15° increments, for example, results in 24 different test angles.

[0106] In some cases, the energy score comprises a molecular mechanics (MM) force-field calculation. Suitable MM force-field calculations include, but are not limited to: universal force-field (UFF), as described in doi: 10.1021 / ja00051a040, Merck molecular force-field (MMFF94), as described in doi: 10.1002 / (SICI)1096-987X(199604)17:5 / 6<490::AID- JCC1>3.0.CO;2-P, doi: 10.1002 / (SICI)1096-987X(199604)17:5 / 6<520::AID- JCC2>3.0.CO;2-W, doi : 10.1002 / (SICI) 1096-987X( 199604) 17 : 5 / 6<553 : : AID- JCC3>3.0.CO;2-T, doi: 10.1002 / (SICI)1096-987X(199604)17:5 / 6<587::AID- JCC4>3.0.CO;2-Q, doi : 10.1002 / (SICI) 1096-987X( 199604) 17 : 5 / 6<616 : : AID- JCC5>3.0.CO;2-X, doi: 10.1002 / (SICI)1096-987X(199905)20:7<720::AID-JCC7>3.0.CO;2- X, doi: 0.1002 / (SICI)1096-987X(199905)20:7<730::AID-JCC8>3.0.CO;2-T, optimized potentials for liquid simulations force-field (OPLS), as described in doi:10.1021 / ja9621760 and doi:10.1021 / jp003919d, general Amber force-field (GAFF), as described in doi: 10.1002 / jcc.20035, and CHARMM general force-field (CGENFF), as described in doi: 10.1002 / jcc.21367.

[0107] In some cases, the energy score comprises a quantum mechanics (QM) force-field calculation. Suitable QM calculations include, but are not limited to, pole basis sets, as described in doi: 10.1063 / 1.1674902, correlation-consistent basis sets, as described in doi: 10.1063 / 1.456153, polarization-consistent basis sets, as described in doi: 10.1063 / 1.1413524, completeness-optimized basis sets, as described in doi: 10.1002 / jcc.20358, and Karlsruhe basis sets, as described in doi: 10.1039 / C4CP04286G.

[0108] In some cases, the energy score comprises a combined MM / QM force-field calculation. In this case, one or more of the MM calculations and one or more of the QM calculations can be combined.

[0109] In some cases, for each angle (e.g., reference angle or test angle), the lowest energy conformer is selected, and, in some cases, combined to yield a set of conformers equal in number to the number of angles for further calculation.

[0110] Thus, in some cases, a first energy score is calculated for each of the reference and test angles of each of the plurality of conformers, the lowest energy conformer for each angle is selected, and a second energy score is calculated for the lowest energy conformer. For example, when 100 conformers are generated for each torsion fragment, the reference conformation is rotated at 15° increments for each of the conformers, and the lowest energy of the 100 conformers is selected for each of the 100 conformers (e.g., the lowest energy conformation by MM calculation), a total of 24 conformers for each minimal torsion fragments are selected for further calculation (e.g., QM calculation, e.g., single-point QM calculation).

[0111] In some cases, MM calculations are carried out in gas phase using a force field, e.g., an MM force field, e.g., MM2, MM3, MMFF94 or MMFF94s, preferably MMFF94s. In some cases, QM calculations are carried out in gas phase. In one example, single-point QM calculation comprises a procedure comprising a first optimization and, optionally a second optimization. In some cases, the first optimization is performed at the hf3c level, and the second optimization is performed at the B3LYP / 6-31G* or B3LYP / 6-31+G* baseline level. In some cases, the latter is carried out only if the fragment contains sulfur. In some cases, the reference torsion is frozen at a set angle for both steps. Machine Learning Based Torsional Profile Prediction Using a Set of Torsion Features

[0112] In some cases, torsion profiling comprises machine learning-based torsion profile prediction. QM calculations are computationally expensive. To increase efficiency, in some cases, a machine learning algorithm (e.g., a neural network) is trained to predict torsion profiles around rotatable bonds (e.g., for different 3D conformations of a small molecule or fragment(s) thereof, e.g., for the plurality of conformers of the torsion fragments described above.

[0113] Such prediction is useful, for example, to utilize torsion profiles (and, therefore, entropy penalty and torsional strain, e.g., as described below) in virtual screening campaigns, e.g., for molecular glue drugs (e.g., molecular glue degraders). As another example, the predicted torsion profile can be used to filter out conformers which are unlikely in any ligand-based approaches, e.g., based on an unacceptable level of strain, thereby decreasing the false-positive rate of selected conformers that are under high strain when docked. As yet another example, the predicted torsion profile can be used to decrease the number of conformers in an ensemble to determine a more relevant conformational space, e.g., for further processing.

[0114] In this case, an example system can use extracted torsion features to predict the torsional profile. For example, the local atomic environment of each torsion atom can be represented through a set of torsion feature(s). The example system can process the extracted torsion features to generate a torsion profile.

[0115] In some cases, the feature(s) comprise symmetry function(s), e.g., functions used to convert the chemical environment around a torsion definition to a float vector, e.g., for further processing, geometric characteristic(s) and chemical characteristic(s). In some cases, the feature(s) comprise dihedral angle, length of the central rotatable bond, distance between the two terminal atoms, product of the atomic numbers of the two central and two terminal atoms, and combinations thereof.

[0116] In some cases, the feature(s) are combined to produce a N-dimensional descriptor vector. In some cases, the descriptor vector is an atomic environment vector that encodes the chemical environment around a given rotatable bond. In some cases, the machine learning algorithm (e.g., the neural network) makes its predictions based on the descriptor vector.

[0117] In some cases, the symmetry function(s) comprise an atom-centered symmetry function(s) (ACFs). See, e.g., Behler and Parrinello, “Generalized Neural -Network Representation of High-Dimensional Potential -Energy Surfaces,” Phys. Rev. Lett. 98: 14601 (2007); see also Rai et al., “TorsionNet: A Deep Neural Network to Rapidly Predict Small Molecule Torsion Energy Profiles with the Accuracy of Quantum Mechanics,” J. Chem. Inf. Model. 64(4):785-800 (2022), which also describes how ACSF can be implemented for torsion profile prediction. ACSFs model the local chemical environment of an atom z via radial and angular distributions of the surrounding nuclei. Thus, in some cases, the symmetry function(s) comprise radial ACSF(s), angular ACSF(s), and / or torsion ACSF(s).

[0118] In some cases, the radial ACSF takes the form: where N is the number of atoms and nj is the distance between atoms z and j. / / and « are parameters modulating the width and position of the Gaussian function.

[0119] In some cases, the angular ACSF takes the form: where Qijt is the angle spanned by the atoms z, j and k. The term in brackets characterizes the distribution of angles, k is a parameter which takes the values = ±1 and shifts the maximum of the angular term between 0° and 180°, while C, controls its width. ;, k, and fk are cutoff functions, using the shorthand notation fc(rf) =f . The introduction of terms based on rjk introduces asymmetric behavior into the angular functions, leading to a smaller spatial extent for angles close to 180°.

[0120] In some cases, the torsion ACSF takes the form: where dihedral angles (pijki includes atoms j and k of the specified rotatable bond. Pairs i and I are any atoms of type a and ?, respectively, but located on opposite side of the rotatable bond (Jk). In some cases, ACSFs are defined for a set of elements (e.g., pairs (radial) and triples (angular)). In one example, terms corresponding to these specific combinations of elements are counted in the summations of the radial and / or angular ACSF equation(s). Thus, for example, if a molecular system contains the elements H, C and O, in some cases, the environment of a hydrogen atom would be described by a set of radial functions for the pairs H-H, H-C and H-0 and angular functions for the triples H-H-H, H-H-C, H-H-0, H-C-C, H- C-0 and H-O-O. For each of these elemental combinations, several ACSFs can be introduced respectively, varying in the parameters , p, X and C, (using e.g. a set of ACSFs with Gaussians of different widths q to describe H-C distances). In some cases, all the sets of radial and angular ACSFs are collected into a descriptor vector encoding the chemical environment of the central nucleus.

[0121] In some cases, the symmetry function(s) comprise weighted atom-centered symmetry function(s). See, e.g., Gasteggar, “wACSF — Weighted atom-centered symmetry functions as descriptors in machine learning potentials,” J. Chem. Phys. 148:241709 (2018). Unlike ACSFs, which use separate functions to describe different combinations of elements, wACSFs account for the composition of the chemical environment in an implicit manner by introducing element-dependent weighting functions into the radial and angular symmetry functions. In this way, the number of wACSFs required to describe a system does not depend on the number of different elements present (unlike, for example, ACSFs). Thus, in some cases, the symmetry function(s) comprise a radial wACSF(s) and / or an angular wACSF(s).

[0122] In some cases, the radial wACSF takes the form: and / or the angular wACSF takes the form: fijfikfjk

[0123] Where Zj and Z are the atomic numbers of the nuclei j and k respectively. g(Zz) and h(Zj, Zk) are weighting functions, which modify the contribution of each radial and angular term based on the chemical elements of the atoms involved. While g and h can in principle use a wide variety of different definitions, in some cases g(Zz) = Zj and / or A(Z, Zj) = ZiZj. In some cases, In some cases, the symmetry function(s) comprise a weighted torsion symmetry function. In some cases, the weighted torsion symmetry function takes the form:

[0124] Wldlh= 2'-<ZyZlS»it, S ij ZA(1 + ^coSeljk,y x defined by torsion angles Oijki, where atoms j and k are the ones that form the specified rotatable bond. The atoms i and I are any neighboring atoms of the bond atoms but are respectively located on the opposite side of the rotatable bond (jk).

[0125] In some cases, a cutoff function fcis used to ensure that only the energetically relevant regions close to the central nucleus are encoded in the symmetry function(s), e.g., the ACSF or wACSF.

[0126] In some cases, the cutoff function is defined as: where rcis the cutoff radius specifying the size of the region surrounding the central atom.

[0127] In some cases, parameter(s) of the descriptor(s), e.g., the symmetry function(s) are pre-set. In some cases, param eter(s) of the descriptor(s), e.g., the symmetry function(s) are directly incorporated into the machine learning algorithm. The skilled artisan will appreciate that suitable parameters (e.g., pre-set parameters and / or initial parameters) can be chosen and optimized based on the data set(s), and it is within the skill of the artisan to select those parameters.

[0128] In some cases, the methods described herein comprise encoding the chemical environment of torsion atom(s) in an N-dimensional vector (e.g., as described herein). In some cases, the vector is used to train a machine learning model to predict a torsion profile (e.g., as described herein), around a rotatable bond (e.g., a rotatable bond described herein).

[0129] Thus, described herein are training sets comprising torsion profile(s) around rotatable bond(s), e.g., defined by four atoms, i,j, k, and / , where atoms j and & make the rotatable bond and atoms i and j are the neighboring atoms. In some cases, the training data is generated by fragmentation based torsion profiling (e.g., as described above). In some cases, the training data is augmented by enumeration. For example, for a given rotatable bond defined by atoms j and k, there may be one or more possible neighboring atom z and neighboring atom I. For a given rotatable bond, in a given 3D configuration, with a given relative energy (e.g., calculated as described above), the there are possibly multiple torsion atom sets ijkl from which description vector(s) can be produced, which differ in the planes from which the torsion angle is calculated from, but not in the potential energy of the geometry. This is depicted, for example, in FIG. 4. Thus, the training data may be augmented, for example, by including vectors of multiple torsion sets around rotatable bond(s), but, for example, keeping the relative potential energy the same during training.

[0130] Also described herein are torsion profile prediction system(s) comprising an embedding neural network and a comparison engine. The embedding neural network is configured to process data defining the chemical environment of torsion atom(s), in accordance with values of a set of embedding neural network parameters, to generate an embedding of the chemical environment in an embedding space (i.e., a space of possible embeddings). The embedding space can be any appropriate space, e.g., a 10 dimensional, 50 dimensional, or 100 dimensional Euclidian space. In some cases, the machine learning model is an embedding neural network, e.g., a deep neural network (DNN). In some cases, an embedding encodes the features in a fixed-dimensional representation.

[0131] In some cases, the embedding neural network is trained using an embedding training system to generate a target representation comprising a relative energy score. In some cases, the embedding training system is trained on a data set, e.g., a training set comprising torsion profile(s) around rotatable bond(s), e.g., as described herein, e.g., an augmented data set described herein.

[0132] The embedding neural network can have any appropriate neural network architecture that enables it to perform its described functions, e.g., processing surface patches to generate embeddings. In particular, the embedding neural network can include any appropriate types of neural network layers in any appropriate numbers and connected in any appropriate configuration.

[0133] In some cases, the methods described herein comprise active learning, e.g., dropoutbased active learning. See, e.g., Tsymbalov, E., Panov, M., Shapeev, A. (2018). Dropout- Based Active Learning for Regression. In: , et al. Analysis of Images, Social Networks and Texts. AIST 2018. Lecture Notes in Computer Science(), vol 11179. Springer, Cham. In some cases, active learning comprises calculating the uncertainty of a minimal torsion fragment, e.g., as the sum of the standard deviation of the potential energy prediction of its conformer(s), e.g., as calculated using dropout-based active learning. The dropout rate can be any suitable rate, e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, or 0.8. In some cases, the N most uncertain predicted minimal torsion fragments generated using the predictive model are selected for torsion profiling (e.g., fragmentation based torsion profiling, e.g., as described above), the calculated set is added to the training set, and then the neural network is re-trained on this training set. In some cases, the N most uncertain prediction minimal torsion fragments are a certain number of predictions or less than a certain percent of the predictions (e.g., 90% or less, e.g., 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% or less). In some cases, this procedure is repeated one or more times, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times.

[0134] Torsion Profile Prediction from Sampling Multiple Conformers

[0135] In some cases, the system can represent a conformation as a feature vector input by applying a symmetry function, e.g., wACSF, to capture the local chemical environment around a torsion definition, e.g., as described above. The system can then generate a predicted energy value for each feature vector input, e.g., by processing the feature vector input using a regression model that is configured to predict a single energy value for each feature vector input. More specifically, the system can generate the torsion profile by predicting energy values from conformations sampled by rotating through the rotatable bond and plotting the torsion angles against the predicted energy values from the regression model. In this case, to fulfill the periodicity property of a torsion profile, the system can ensure feature vector inputs generated at a degree of rotation of 0 degrees and 360 degrees are the same.

[0136] Torsion Profile Prediction Directly from a Single Conformer

[0137] This section describes techniques for predicting a torsion profile from a single conformer of a torsion fragment. In some cases, the methods described herein comprise predicting torsion profile(s) directly from small molecule 2D structure(s), i.e., with minimal fragmentation, conformer generation, and energy score calculation as described above. An example of this method is depicted in FIG. 5.

[0138] In particular, the system can generate an embedding of a graph representation of the small molecule or fragments of the molecule and use the embedding of the graph representation to generate the torsion profile. In the particular example depicted in FIG. 1, the torsion profile prediction system 100 can generate the graph embedding 185 by processing the fragments 125 using a graph representation engine 130 and can process the graph embedding 185 using a distribution prediction machine learning model 190 to generate the torsion profile 195. In some cases, predicting torsion profile(s) directly from small molecule 2D structure(s) comprises sampling from a probability distribution of a torsion retrieved by its energy torsion profile using the function below describing the Boltzmann distribution: where the probability Pxof angle x is calculated by its energy state x, molecular Boltzmann constant KB and temperature T.

[0139] For example, sampling from the Boltzmann distribution can be used to combine data from experimental small molecule structures, e.g., which can represent the probability of how often a certain torsion angle appears in experimental data, and torsion profiles derived from quantum mechanical calculations, e.g., which represents the calculated energy values of different angles. More specifically, the system 100 can leverage the Boltzmann distribution to retrieve the probability of observing a certain angle for the quantum mechanical energy values in order to combine data for training the models 160 and 190 of the torsion profile prediction system.

[0140] Furthermore, in some cases, the distribution prediction machine learning model 190 can predict the torsion profile 195 as a torsion energy profile. In other cases, the machine learning model 190 can predict the torsion profile 190 as a likelihood distribution of the torsion angle. In this case, the likelihood distribution can be converted to an energy scale by computing the negative logarithm of the likelihood, e.g., -log(likelihood).

[0141] The system 100 can process the fragments 125 to generate a molecule graph representation 175 defined by G=(V,E) for each fragment, e.g., the molecule graph G is parameterized by a set of node embeddings (V) and a set of edge embeddings (E). In this case, the set of node embeddings can represent the atoms of the torsion fragment and the set of edge embeddings can represent the chemical bonds between a pair of atoms in the torsion fragment.

[0142] The system can process each of the torsion fragments 125 to generate initial embeddings 140, e.g., the initial node embeddings 142 and initial edge embeddings 144. In particular, the initial embeddings 140 can be generated using one or more node and edge features. For example, the system 100 can process the torsion fragments 125 using an open- source cheminformatics tool, e.g., RDKit, to generate a number of features 150. As another example, the system 100 can receive the features 150. An example of selected node 152 and edge 154 features, respectively, are depicted in the tables below. Node features

[0143] Edge features

[0144] In this case, the initial embeddings 140 can be generated by processing the node and edge features using an embedding neural network or an embedding layer configured to process a set of node or edge features to generate a corresponding node or edge embedding that represents the respective atom or chemical bond in an embedding space. For example, the system 100 can process each fragment using an encoder block, e.g., of the graph neural network 160, to embed the node and edge features into an N-dimensional vector embedding space. In other cases, the initial embeddings can be generated by setting the initial embeddings 140 to a default embedding, e.g., an embedding where all values are zero, or randomly sampling values, e.g., from a probability distribution such as a Gaussian distribution, to generate the initial embeddings 140.

[0145] In addition to the node 142 and edge 144 embeddings, additional features defined hereafter as torsion features 150 can be used to make the graph G unique for each torsion fragment. In particular, the system 100 can process the molecular structure data 110, e.g., the SMILES string, corresponding to each fragment to identify one or more torsion features.

[0146] As an example, a torsion feature 156 can include a list of the atoms making the rotatable bond of the torsion, e.g., as defined by the corresponding numbers of the nodes representing the atoms. As another example, a torsion feature 156 can be the “phase angle”, e.g., the angle of an improper torsion defined by the atoms involved in the rotatable bond and two atoms bonded to one of them. In this case, one of the branched atoms can belong to the torsion definition, and therefore the phase angle can also be viewed as the difference between the dihedral angles to the main and any other branch, e.g., any other atom, at the same side of the rotatable bond. This phase angle has been described previously as crucial to make the model invariant to permutations such as cis-trans, e.g., as described Raush, E., et. al “Graph- Convolutional Neural Net Model of the Statistical Torsion Profiles for Small Organic Molecules (doi:10.1021 / acs.jcim.2c00790).

[0147] In particular, the phase angle is useful to predict any asymmetric torsion profile where the profile is different between 0 and 180 to 180 to 360. In some cases, the system can extract the phase angle of a conformation by setting the torsion angle to zero and calculating the other torsions defined by the branched atoms, e.g., as described in Adams, K., et. al. “Learning 3D representations of Molecular Chirality with Invariance to Bond Rotations” (doi:10.48550 / arXiv.2110.04383).

[0148] Torsion features

[0149] The system 100 can process the initial node embeddings 142, initial edge embeddings 144, and torsion features 150 using a graph neural network (GNN) 160. A graph neural network 160 is a machine learning model that can process graph- structured data in an efficient way. In particular, GNNs are capable of processing graph- structured data in efficient ways, such as through message passing, a method for updating node embeddings based on the aggregation of information from nearby nodes into messages, to update the relationships between the nodes and edges.

[0150] In more detail, the GNN 160 can include an encoder block, a sequence of one or more message passing neural network layers, and a decoder block. The encoder block can, for each atom in the molecule, process a set of atom features of the atom to generate a node embedding for the graph node that represents the atom. Further, the encoder block can, for each bond in the molecule, process a set of bond features of the bond to generate an edge embedding for the graph edge that represents the edge. Each message passing layer is configured to process the set of node embeddings and the set of edge embeddings, by neural network operations that are parametrized by a set of neural network parameters of the message passing layer and that are conditioned on the topology of the graph representing the molecule, to update the node embeddings and the edge embeddings. The decoder block can process the node and edge embeddings generated by the final message passing layer to generate the GNN output, in this case, the graph embedding 185. In this case, the GNN 160 can include one or more augmentation neural network(s) 170 to generate updated embeddings, e.g., node and edge embeddings that incorporate the information from the torsion features 150, as an intermediate processing output in the GNN 160. The augmentation neural network(s) 170 can have any appropriate machine learning architecture, e.g., a neural network, that can be configured to process one or more of the embeddings 142, 144 and the torsion features 150. In particular, each augmentation neural network 170 can have any appropriate number of neural network layers (e.g., 1 layer, 5 layers, or 10 layers) of any appropriate type (e.g., fully-connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).

[0151] More specifically, the system 100 can update the set of edge embeddings using the torsion features 150 (and can update the set of node embeddings by performing message passing with the updated edge embeddings, which will be described in more detail below). In particular, the system 100 can update the edge embeddings of a rotatable bond using a different technique than is used to update the edge embeddings of a non-rotatable bond. For example, the system 100 can update the edge embeddings of a rotatable bond using a weighted combination of outputs of multiple augmentation neural networks 170 and can update the embeddings of a non-rotatable bond using a single augmentation neural network 170.

[0152] In an example in which the one or more augmentation neural networks 170 are implemented as multi-layer perceptrons (MLPs), the system 100 can apply the MLP on the edge embeddings and their connected nodes as defined in the equation below: e-y = MLP([v--1, v]-1, e-J1]) where Vt and Vj are the node features and etj is the edge connecting them. As specified above, the edges connected to one of the nodes representing an atom involved in the rotatable bond of the torsion definition can be treated differently. For example, the edges can be updated with three MLPs by which two are weighted by either sines or cosines of the “phase” angles a as defined in the following equation.

[0153] The system can then use the updated edge embeddings e ■ to update the node embeddings, e.g., using message passing: v‘ = MLP I [?'-*, - MLPdv '.e' )] I \ lull A-1 / where Vj are the neighboring nodes and the node that is updated. More specifically, message passing is a method for updating node and edge embeddings based on the aggregation of information from nearby node and edge embeddings into messages. As an example, a message can contain information from the nodes in the one-hop away neighborhood of the node. In particular, the message used to update each node embedding can be a summation of the neighboring one-hop nodes multiplied by the relational weights of the edges multiplied by the trainable weights of the graph neural network. The current embedding for each node can then depend on the relational weights defined in the edge embeddings of the respective nodes’ neighbors.

[0154] As an example, the GNN 160 can have three message passing layers, e.g., three graph convolutional layers. The GNN 160 can update the node and edge embeddings using the message passing layers to generate the graph representation 175 and can perform graph pooling 180 to process the node and edge embeddings generated by the final message passing layer to generate the graph representation engine 130 output, in this case, the graph embedding 185.

[0155] In particular, the system 100 can further apply softmax-based attention pooling to the final, e.g., updated, node embeddings in the graph representation 175 as part of graph pooling 180, to generate a set of attention weights corresponding with the node embeddings. In more details, the system 100 can compute an attention score as defined by: where v represents each of the node embeddings. The graph representation engine 130 can then apply the attention-weights to generate the graph embedding 185 through graph pooling 180: where N refers to the number of nodes in the graph.

[0156] While described above in terms of ensuring the graph G is unique for each torsion fragment by incorporating the torsion features 150 into the edge embeddings 144 using the augmentation neural network(s) 170, in another example, the system 100 can process the node 142 and edge 144 embeddings, without the torsion features 150, using a graph neural network that is configured to generate representations that are invariant to bond rotations about internal molecular bonds to generate the graph embedding 185. For example, the system 100 can process the node 142 and edge 144 embeddings to generate the graph embedding 185 using the graph neural network described in detail in Adams, K., et. al. “Learning 3D Representations of Molecular Chirality with Invariance to Bond Rotations” (arXiv:2110.04383).

[0157] In more detail, rather than modifying the graph convolution layers as described above, an alternative approach to encoding torsional information in molecular graphs can involve integrating torsion-specific operations within the graph pooling mechanism disclosed in “Learning 3D Representations of Molecular Chirality with Invariance to Bond Rotations”. More specifically, given a rotatable bond between atoms j and k, the system 100 can extract the node embeddings Vj and vkfrom the final message-passing layer of a graph convolution network.

[0158] The system 100 can then identify respective neighboring atoms, denoted as i and / , to construct the set of all possible dihedral quadruplets (I- Vj, vk, vt), where vtand vtare the corresponding node embedding of atoms i and / , respectively. For each dihedral quadruplet, the system can determine the corresponding phase angle (Pi -^,1and incorporate this information through trigonometric encoding:

[0159] In this case, the torsion-specific embeddings can be generated using an MLP and pooled together to obtain a comprehensive representation for the rotatable bond. The generated torsion-aware representation cj kcan then be utilized either directly or in combination with the graph representation vector to parameterize the wrapped Gaussian mixture model for torsion profile prediction.

[0160] In either case, the system 100 can then process the graph embedding 185 using a distribution prediction machine learning (ML) model 190 to generate one or more parameters 192, e.g., one or more summary statistics, parameterizing a circular distribution characterizing the torsional profile for each torsion fragment 125. A circular distribution refers to a probability distribution that is wrapped around a circle to represent probability over a periodic domain. Circular distributions can be effectively used to model angular data, e.g., since the end and the beginning of the periodic domain are the same. As an example, the circular distribution can be a circular mixture distribution, e.g., a circular distribution including one or more circular distributions as component distributions. In this case, each component distribution contributes to the circular distribution based on a contribution weight, e.g., a mixture coefficient.

[0161] As an example, the param eter(s) 192 can include a mean, standard deviation, and mixture coefficients for each component distribution in a mixture of wrapped Gaussians. As another example, the param eter(s) 192 can include a mean direction and concentration parameter of a Von Mises distribution. As yet another example, the parameter(s) 192 can include a mean direction and spread of a wrapped Cauchy distribution. As a further example, the parameter(s) 192 can include a rate parameter of a wrapped exponential distribution.

[0162] The distribution prediction ML model 190 can have any appropriate machine learning architecture, e.g., a neural network, that can be configured to process a graph representation 185, e.g., a set of node and edge embeddings, to generate one or more values, e.g., the summary statistics of the circular distribution. In particular, the distribution prediction ML model 190 can have any appropriate number of neural network layers (e.g., 1 layer, 5 layers, or 10 layers) of any appropriate type (e.g., fully-connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).

[0163] The distribution prediction machine learning model 190 can have been trained on a set of training examples by a machine learning training technique to optimize an objective function. The objective function can measure, for each training example, a discrepancy between: (i) a ground truth torsion profile of the molecule and (ii) the generated torsion profile of the molecule. In particular, the objective function can be determined by quantifying the difference between the ground truth parameters, e.g., summary statistics, and the generated parameters. The objective function can measure a discrepancy between target and generated molecule properties in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss.

[0164] The machine learning training technique can be any technique appropriate for training the property prediction machine learning model. For instance, the machine learning training technique can be a stochastic gradient descent training technique, e.g., using RMSprop or Adam.

[0165] In some cases, the GNN 160, the one or more augmentation network(s) 170, and the distribution prediction machine learning model 190 can have been jointly trained, e.g., using a corresponding ground truth torsion profile for each rotatable bond in the at least one portion of the molecule. More specifically, the joint training can involve using all the different torsion fragments, e.g., generated by the fragmentation engine 130, to create a representative torsion angle set.

[0166] For example, the system 100 can observe torsion angles from each of the torsion fragments 125 to create the torsion angle set that is sampled during training. An example of generating the torsion angle set will be covered in more detail in an example embodiment described below. In some cases, the system 100 can remove fragments from the dataset that have less than a threshold number of torsion angles, e.g., the system can remove fragments with less than 10 torsion angles.

[0167] In contrast to representing each fragment with a single angle, which would result in an overrepresentation of some fragments, and oversampling the underrepresented fragment and their corresponding angles to create a balanced training set, the system can use a random sampler that randomly selects an angle for each fragment at each epoch from the torsion angle set generated for the fragment. This improves training efficiency by ensuring that corrupted angles are sampled less during training, e.g., thereby encouraging the GNN 160, the one or more augmentation network(s) 170, and the distribution prediction machine learning model 190 to train on conformers with real observed sample angles.

[0168] In this case, the system 100 can update values of respective sets of parameters including a set of graph neural network parameters, one or more sets of augmentation neural network parameters, and a set of distribution prediction machine learning model parameters, e.g., in accordance with an objective function. For example, the objective function can maximize a measure of likelihood of an observed torsion angle generated from the circular mixture distribution parameterized by the one or more generated parameters of the circular mixture distribution using the distribution prediction machine learning model 100.

[0169] For example, the distribution prediction ML model 190 can be implemented as a mixture density network (MDN), e.g., a Gaussian mixture model as described in Bishop, C. “Mixture Density Networks” (Feb. 1994). An example embodiment that involves using an MDN to predict a mean, standard deviation, and mixing coefficients of a mixture of six wrapped Gaussians will be described in more detail below.

[0170] In particular, generating summary statistics of a circular distribution can be advantageous relative to predicting summary statistics of a linear distribution, e.g., as depicted in FIG. 7.

[0171] FIG. 7 depicts an example histogram of observed angles of a torsion fragment with two corresponding modeled distributions: (i) a standard linear Gaussian model and (ii) a circular wrapped Gaussian model, e.g., from fitting the underlying data using the respective distribution types. The fits clearly demonstrate that the prediction of the circular wrapped Gaussian model represents the underlying data with high fidelity with respect to the linear approach.

[0172] Virtual Screening and Drug Discovery

[0173] The system 100 can use the param eter(s) 192 to generate a torsion profile 195 of the torsion fragment, e.g., using a circular mixture model parameterized by the one or more parameter(s) 192. For example, in the case that the parameter(s) 192 include respective means, standard deviations, and mixing coefficients of a mixture of wrapped Gaussians, the system 100 can model the torsion profile 195 using a Gaussian mixture model parameterized by the means, standard deviations, and mixing coefficients of each of the component distributions in the mixture of wrapped Gaussians.

[0174] For example, the system 100 can generate the torsion profile 195 using a circular mixture model for each molecule in a molecule library as part of a virtual screening process. The generated torsion profiles 195 can be used to vet candidate conformers of a molecule in a binding site of a second molecule, e.g., by providing a measure of likelihood of observing the conformer of the molecule in the binding site based on the observed torsion angle(s) of the conformer. In particular, conformers with torsion angle(s) that results in high strain energy, e.g., which are energetically unfavorable, will have a low probability of being observed, thereby making the conformers unlikely drug candidates.

[0175] As another example, the system 100 can perform conformation sampling for a molecule by generating the torsion profile of each rotatable bond in the molecule, as described above. In particular, the system 100 can determine conformations that are likely to be observed by sampling respective torsion angles from each of the torsion profiles corresponding with the rotatable bonds in the molecule, and can provide the molecule conformation defined by the sampled torsion angles. These conformers are suitable for various 3D ligand-based virtual screening methods, including shape comparison, pharmacophore matching, and other related techniques. In some cases, the system 100 can evaluate the molecules using the torsion profiles 195 to select one or more of the molecules in the library for physical synthesis. As an example, the selected molecules can be physically synthesized, e.g., such that further experiments can be performed to evaluate one or more molecular properties of the molecules as part of a drug discovery process. Torsional Strain and Torsional Entropy Estimation

[0176] In some cases, the methods described herein comprise generating entropy penalty term(s) and / or torsional strain term(s), e.g., for a 3D conformation of a small molecule or fragment thereof, e.g., as described herein.

[0177] Torsional Strain Estimation of a Conformer

[0178] In some cases, the torsional strain of a 3D conformation is estimated based on torsional profile(s) of its rotatable bond(s), calculated, e.g., as described above. In the case in which the prediction does not result in a continuous function, the torsion strain of a single rotatable bond can be calculated by calling the interpolation function of the corresponding minimal torsion fragment with the specified torsional angle in the given 3D conformation. In particular, the system can use an interpolation function to generate a continuous function from a set of predicted energy values, e.g., a set of predicted energy values generated by a regression model that processes a corresponding set of sampled torsion definitions. In some cases, the resulting strain term is corrected using the global minimum (e.g., as described above). In some cases, the torsional strain of a 3D conformation of an entire small molecule is defined by the sum of all strains of its rotatable bonds.

[0179] Torsional Entropy Estimation of a Molecule

[0180] Upon binding a protein receptor, a molecule adapts a binding pose that restricts its torsional states and thus reduces the torsional entropy of the small-molecule. The loss of entropy varies depending on the molecule and contributes to relative binding energies. Thus, entropic loss contributions can be used as a component in ranking small-molecules in terms of binding activity and, therefore, in some cases, the methods described herein comprise calculating torsional entropy, e.g., torsional entropy change or loss.

[0181] In some cases, torsional entropy is calculated from torsional profiles (e.g., generated as described herein). In some cases, a potential energy profile (in kcal / mol units) is estimated for each torsion in the molecule (e.g., as described herein). In some cases, estimation the entropy comprises calculating a partition function p tor for each torsion, which is defined as an integral of torsional energy profile over the full angle range: where R is the molar Boltzmann constant (also known as ideal gas constant) equal to 1.9872036e-3 kcal mol'1K'1. In some cases, using the partition function, the probability for a given angle state is estimated as:

[0182] Using statistical mechanics, the entropy contribution for each torsion is given by the equation below:

[0183] In a tight binding complex, it can be assumed that all the torsional states are fixed and, effectively, all the torsional entropy is lost. Hence, in some cases, the total torsional entropy change is defined as a sum of all torsional entropies of rotatable bonds within a molecule.

[0184] IMPLEMENTATION

[0185] FIG. 9 is a flow diagram of an example process for generating a torsion profile using parameters of a circular mixture distribution. For convenience, the process 900 will be described as being performed by a system of one or more computers located in one or more locations. For example, a torsion profile prediction system, e.g., the torsion profile prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 900.

[0186] The system can receive data representing a structure of a molecule (step 910). As an example, the data can represent the two-dimensional structure of a molecule, e.g., as a Simplified Molecular Input Line Entry System (SMILES) string. In some cases, the system can receive the data and generate one or more fragments of the molecule, e.g., one or more portions of the molecule.

[0187] In particular, the system can identify one or more rotatable bonds in the molecule, at least one torsion atom associated with each rotatable bond, and one or more local atoms to the at least one torsion atom. For example, the system can identify one or more covalently bonded atoms to the torsion atom as local atoms and can evaluate covalently bonded atoms to the local atoms to identify partial structures, e.g., functional groups and ring systems, that can be included as local atoms in the fragment. In some cases, the system can identify a ring system. In this case, the system can identify possible ortho substitutions of the ring system as partial structures.

[0188] The system can process the data using a graph neural network to generate an embedding of a graph representation of at least one portion of the molecule (step 920). For example, the system can generate the embedding of a graph representation of one or more fragments of the molecule. An example process for generating the embedding of a graph representation using a graph neural network that includes one or more augmentation neural networks will be covered in more detail in FIG. 10.

[0189] The system can then process the embedding of the graph representation using a distribution prediction machine learning model to generate one or more parameter(s) of a circular mixture distribution (step 930). The distribution prediction machine learning model can have any appropriate machine learning architecture, e.g., a neural network, that can be configured to process the embedding of the graph representation.

[0190] For example, the distribution prediction machine learning model can be a mixture density neural network configured to predict a set of one or more parameters for a number of circular distributions in the circular mixture distribution. In particular, the parameters can be summary statistics of the circular mixture distribution. As an example, the system can generate a mean value, a standard deviation value, and a mixing coefficient value for each of N circular distributions, e.g., N component distributions, in the circular mixture distribution.

[0191] The distribution prediction machine learning model can have been trained on a set of training examples by a machine learning training technique to optimize an objective function. The objective function can measure, for each training example, a discrepancy between: (i) a ground truth torsion profile of the molecule and (ii) the predicted torsion profile of the molecule. In particular, the objective function can be determined by quantifying the difference between the ground truth parameters, e.g., summary statistics, and the generated parameters. The objective function can measure a discrepancy between the ground truth and predicted torsion profile in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss. The machine learning training technique can be any technique appropriate for training the property prediction machine learning model. For instance, the machine learning training technique can be a stochastic gradient descent training technique, e.g., using RMSprop or Adam.

[0192] In particular, the graph neural network, one or more augmentation networks, and the distribution prediction machine learning can have been jointly trained, e.g., using a corresponding ground truth torsion profile for each rotatable bond in the at least one portion of the molecule. In this case, the system can update values of respective sets of parameters including a set of graph neural network parameters, one or more sets of augmentation neural network parameters, and a set of distribution prediction machine learning model parameters, e.g., in accordance with an objective function. In particular, the system can update the respective sets of parameters in accordance with minimizing a discrepancy between the ground truth torsion profile of the at least one portion of the molecule and the generated torsion profile of the at least one portion of the molecule.

[0193] For example, the system can receive a corresponding ground truth torsion profile for each rotatable bond in the at least one portion of the molecule, can identify a set of allowable torsion angle values for the at least one portion of the molecule, and can sample one or more torsion angle values from the identified set of allowable values to generate a number of examples for the at least one portion of the molecule. The system can then process the at least one portion of the molecule with the sampled torsion angle value to generate the one or more parameters of the circular mixture distribution characteristic of the torsion profile of the at least one portion of the molecule. The system can then determine the discrepancy between the ground truth torsion profile and the generate torsion profile of the at least one portion of the molecule and update the respective sets of parameters in accordance with minimizing the discrepancy.

[0194] The system can generate a torsion profile using the parameter(s) of the circular mixture distribution (step 940). In particular, the system can generate the torsion profile of the at least one portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution. As an example, in the case that the parameter(s) include respective means, standard deviations, and mixing coefficients of a mixture of wrapped Gaussians, the system can model the torsion profile using a Gaussian mixture model parameterized by the means, standard deviations, and mixing coefficients of each of the component distributions in the mixture of wrapped Gaussians.

[0195] The system can use the generated torsion profile to characterize a likelihood of observing a conformer of the molecule in a binding site of a second molecule. In particular, the system can generate a torsion profiles for each molecule in a molecule library and use the generated torsion profiles to evaluate candidate conformers, e.g., by providing a measure of likelihood of observing the conformer of the molecule in the binding site based on the observed torsion angle(s) of the conformer. In particular, conformers with torsion angle(s) that result in high strain energy, e.g., which are energetically unfavorable, will have a low probability of being observed, thereby making the conformers unlikely drug candidates. In some cases, the system can generate the torsion profile of each portion of a molecule including a rotatable bond, e.g., to generate a torsion profile of each rotatable bond in the molecule. More specifically, the system can generate the one or more parameters of the circular mixture distribution for each portion of the molecule, and can generate the torsion profile for each portion of the molecule, e.g., using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution generated for the portion of the molecule. As an example, the circular mixture model can be a mixture of wrapped Gaussians, as described above.

[0196] In this case, the system can performing conformational sampling using the torsion profiles corresponding with each portion of the molecule. For example, the system can sample, e.g., select, a torsion angle from the torsion profile for each portion of the molecule, and can provide the molecule conformation specified by the sampled torsion angles as a conformation.

[0197] In some cases, the system can select one or more of the molecules for physical synthesis based at least on the respective likelihood of observing the conformer of the molecule in the binding site of the second molecule. As an example, the selected molecules can be physically synthesized, e.g., such that further experiments can be performed to evaluate one or more molecular properties of the molecules as part of a drug discovery process.

[0198] FIG. 10 is a flow diagram of an example process for generating an embedding of a graph representation of at least one portion of the molecule using a graph neural network (GNN). For convenience, the process 1000 will be described as being performed by a system of one or more computers located in one or more locations. For example, a torsion profile prediction system, e.g., the torsion profile prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1000, e.g., using a graph neural network.

[0199] The system can generate a set of node embeddings and a set of edge embeddings (step 1010). In particular, the system can process the data representing the at least one portion of the molecule, e.g., one or more fragments, and can generate an initial set of node and edge embeddings for each fragment corresponding to the atoms and chemical bonds connecting a pair of atoms, respectively. For example, the system can obtain a set of node and edge features and can embed the features into an N-dimensional vector embedding space using an encoder block, e.g., of the graph neural network. In other cases, the initial set of node and edge embeddings can be generated by setting the initial embeddings to a default embedding, e.g., an embedding where all values are zero, or randomly sampling values, e.g., from a probability distribution such as a Gaussian distribution, to generate the initial embeddings.

[0200] The system can identify one or more torsion features (step 1020). For example, the system can process the molecular structure data, e.g., the SMILES string, corresponding to each fragment to identify one or more torsion features. As an example, a torsion feature can include a list of the atoms making the rotatable bond of the torsion, e.g., as defined by the corresponding numbers of the nodes representing the atoms. As another example, a torsion feature can be the “phase angle”, e.g., the angle of an improper torsion defined by the atoms involved in the rotatable bond and two atoms bonded to one of them.

[0201] The system can use one or more augmentation neural network(s) to update the set of edge embeddings in accordance with the torsion features (step 1030), e.g., to incorporate the information from the torsion features in an updated set of node and edge embeddings. The one or more augmentation neural network(s) can have any appropriate machine learning architecture, e.g., a neural network, that can be configured to process the set of node embeddings, the set of edge embeddings, and the torsion features. In some cases, the augmentation neural network(s) can be included in the graph neural network.

[0202] In particular, the system can update the set of edge embeddings and then update the node embeddings in accordance with the updated edge embeddings. For example, the system can update an edge embedding of a non-rotatable bond by processing one or more node embeddings of atoms corresponding with the non-rotatable bond using an augmentation neural network. As another example, the system can update an edge embedding of a rotatable bond by generating and aggregating respective outputs of three augmentation neural networks. In particular, the system can generate a respective output with each augmentation neural network by processing one or more node embeddings of atoms corresponding with the rotatable bond using the augmentation neural network and can aggregate the respective outputs by summing a first output, a second output weighted by a sine of the phase angle of the torsion features, and a third output weighted by a cosine of the phase angle.

[0203] The system can then update the set of node embeddings and edge embeddings by performing message passing (step 1040). Message passing is a method for updating node and edge embeddings based on the aggregation of information from nearby node and edge embeddings into messages by conditioning the updates on the topology of the graph representation. For example, a message can contain information from the nodes in the one- hop away neighborhood of the node. In particular, the system can perform message passing using one or more message passing layers of the graph neural network. The system can apply graph pooling (step 1050) and can generate an embedding of the graph representation of at least one portion of the molecule (step 1060). In particular, the system can use an attention mechanism to determine a set of attention weights corresponding with the updated node embeddings and can apply a graph pooling operation, e.g., an aggregation, in accordance with the attend on-weights to generate the graph embedding. As an example, the system can apply an aggregation, e.g., a weighted average, in accordance with the attention-weights to generate the embedding of the graph representation. As an example, the system can process the embedding of the graph representation using a distribution prediction machine learning model to generate a torsion profile, e.g., as described in process 900.

[0204] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0205] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0206] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0207] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0208] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0209] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0210] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0211] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.

[0212] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a Pytorch framework.

[0213] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0214] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0215] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0216] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0217] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0218] EXAMPLES

[0219] The invention is further described in the following example, which does not limit the scope of the invention described in the claims.

[0220] Example 1: Predicting Torsion Profile Using a Mixture of Wrapped Gaussians The following is an example implementation of a torsion profile prediction system 100.

[0221] Dataset: The Crystallography Open Database (COD) (doi: 10.1093 / nar / gkr900) was used as resource in addition to quantum mechanically derived torsion profiles. Only organic molecules were extracted from COD. Metalorganic compounds were not considered because they might have distorted ligand conformation coming from the coordination to the metal that leads to unrealistic torsion angles. The rest of the molecules were fragmentized as described above, and for each fragment, the observed torsion angles w stored.

[0222] The data set was further processed to retrieve any possible proper torsions by enumerating the torsions. The final set from COD consists of 13,605 fragments having between 10 and 28,000 torsions. The median number of torsions for each fragment is 24. To increase the data set further, torsion profiles from QM calculations were converted to yield an additional set of 4,597 fragments, each having 2,000 torsion angles.

[0223] Model: The model was divided into three stages: feature extraction, graph pooling, and mixture density network. The model was implemented using Pytorch Geometric (Fey, M., and Lenssen, J. “Fast Graph Representation Learning with PyTorch Geometric” (arXiv: 1903.02428)), and Pytorch (Paszke, A., et. al. “PyTorch: An Imperative Style, High- Performance Deep Learning Library” (arXiv: 1912.01703)). The example architecture of a GNN and MDN for predicting torsion profile(s) directly from small molecule 2D structure(s) is depicted in FIG. 6.

[0224] In this case, the MDN generates different outputs like means (p), standard deviations (c) and mixing coefficients (a) that are used to parameterize a mixture of wrapped gaussians. The MDN is configured to generate the outputs as follows: p = EL U Line ar (6)) + 1 a = E LU (Line ar (G)) + 1.1 a = Softmax(Linear(Gy)

[0225] In particular, the MDN can be configured with different activation functions to ensure that the predicted values of p, c, a are meaningful. For example, the MDN can apply an exponential linear unit (ELU) activation function to a Linear activation function to ensure that the c remains positive, e.g., since a negative standard deviation is not meaningful. As another example, the MDN can apply a Softmax activation function to ensure that the mixing coefficient is normalized, e.g., since the sum of the mixing coefficients is necessarily one. As yet another example, the MDN can apply an ELU activation function to a Linear activation function to ensure the p for the angles is positive, e.g., since the angles are predicted in radians on either a full, e.g., 0 to 2TI, or half, 0 to 7t, period domain.

[0226] In this case, the MLP used for augmenting the node and edge embedding consists of a Linear layer followed by batch normalization, ELU activation function and dropout. The dropout was set to 0.2 in all experiments.

[0227] In the experiments performed, predicting the torsion profile from a fragment using the GNN and MDN model took approximately 10 ms on a computing device with 1 GPU (NVIDIA A10G).

[0228] Training: The model parameters of the model depicted in FIG. 6 were updated using the Adam optimizer (Kingma, D. and Ba, J. “Adam: A Method for Stochastic Optimization”

[0229] (arXiv: 1412.6980)) with a learning rate of 0.001. The model was trained to optimize the following loss function: where LMDN minimizes the negative log-likelihood of a torsion angle computed using mixture model formed by k= 6 wrapped gaussians (WN).

[0230] As demonstrated in FIG. 7, the choice of a circular mixture distribution was crucial to learn the torsion profiles since torsion angles are within angular space which has a circular property. In this case, the wrapped Gaussian distribution is suitable to model such distributions and was thus here implemented as the probability function of the loss function to learn the period of 2TI of the torsion range. The parameters p, c and a used in the wrapped gaussian mixture model (WGMM) were generated from the model. The additional parameter w is an integer and was choosing as w G {—2, —1,0, 1,2}, e.g., where -2 and 2 indicate a periodicity of 2TI for the torsion range and -1 and 1 indicate a periodicity of TI for the torsion range.

[0231] The model was trained over 15,000 epochs using a batch size of 2,048. For each fragment and each training iteration, a new torsion angle was randomly selected from the torsion angle set. In this example, every fragment is only optimized to a single torsion angle per iteration, but the techniques of this specification are not limiting, e.g., that could change depending on the likelihood of the torsion angle as determined by number of occurrence in the torsion angle set. In particular, randomly sampling a new torsion angle each iteration, prevents the model from learning any biases that would come from the heavily imbalanced data set. The learned parameters can be used to predict a torsion profile from a torsion fragment by plotting angles against their predicted likelihood derived from the WGMM.

[0232] Results: Example results are depicted in FIGS. 8 A and 8B. In particular, FIG. 8 A includes the generated torsion profiles using the aforementioned method of this example for ten molecules.

[0233] More specifically, the ten molecules depicted in FIG. 8A include anomalous torsion conformations that, in some cases, are not properly modeled with force-field parameterization approaches that rely on look-up tables which only include the potential energy of the most common torsion conformations.

[0234] As depicted in FIG. 8A, generating the predicted torsion profile using the GNN and MDN of the aforementioned example to generate predicted parameters of a Gaussian mixture distribution is able to reproduce the observed torsion profile with high fidelity.

[0235] Furthermore, FIG. 8B demonstrates the accuracy of using the GNN and MDN of the aforementioned method of this example to predict torsion profiles with respect to MMFF94, a commonly-used molecular mechanics force-field. In particular FIG. 8B illustrates a comparison of letter-value plots that show an aggregated measure of root mean square deviation (RMSD) between torsion energies computed using a particular method and the ground truth energy for a set of molecules with each of a number of rotatable bonds.

[0236] In this case, the Platinum dataset of bioactive conformations, e.g., as described in Friedrich, Nils-Ole, et. al. “High-Quality Dataset of Protein-Bound Ligand Conformations and Its Application to Benchmarking Conformer Ensemble Generators” J. Chem. Inf. Model. 2017, 57, 3, 529-539, was used to evaluate the accuracy of energy strain computed from the predicted torsion profiles for the MMFF94 method and method of the aforementioned example.

[0237] More specifically, the 3D bioactive conformations of the Platinum dataset were converted into SMILES to remove any 3D information. The SMILES were then used as inputs to generate 250 conformers for each using a conformer ensemble generator method, e.g., OMEGA, as described in Hawkins, Paul C.D., et. al “Conformer Generation with OMEGA: Algorithm and Validation Using High Quality Structures from the Protein Databank and Cambridge Structural Database” J. Chem. Inf. Model. 2010, 50, 4, 572-584.

[0238] In this case, each conformer was evaluated by retrieving either the strain energy from predicted torsion profiles or calculating the energy using the force field MMFF94s. For each compound, the top three solutions either from strain energy or MMFF94s energy were retrieved and their RMSD to the bioactive conformation was calculated by RDKit. The minimum RMSD of each compound was compared with the two methods as shown in the letter-value plot depicted in FIG. 8B.

[0239] The lower the RMSD is the closer to the bioactive conformation. Notably, the torsion profile prediction system achieves lower RMSD than the molecular force field function, especially with respect to an increasing number of rotating bonds, thereby indicating an advantage of the approach for more flexible molecules with a large number of rotatable bonds.

[0240] OTHER EMBODIMENTS

[0241] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method comprising: receiving data representing a structure of a molecule; processing the data using a graph neural network to generate an embedding of a graph representation of at least one portion of the molecule, wherein the graph representation comprises: a set of node embeddings, each node embedding representing an atom in the at least one portion of the molecule; a set of edge embeddings, each edge embedding representing a bond between a pair of atoms in the at least one portion of the molecule; one or more sets of torsional features, each comprising an identification of a torsion atom in the at least one portion of the molecule, wherein a torsion atom comprises a node at a terminus of a rotatable bond, and a corresponding phase angle of the rotatable bond; and processing the embedding of the graph representation using a distribution prediction machine learning model to generate one or more parameters of a circular mixture distribution characteristic of a torsion profile of the at least one portion of the molecule.

2. The method of claim 1, further comprising: generating the torsion profile of the at least one portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution.

3. The method of any one of claims 1-2, wherein receiving data representing the structure of a molecule comprises receiving data representing a two-dimensional structure of a molecule.

4. The method of claim 3, wherein receiving data representing the two-dimensional structure of the molecule comprises receiving a Simplified Molecular Input Line Entry System (SMILES) string.

5. The method of any one of claims 1-4, wherein receiving data representing the structure of the molecule further comprises generating one or more fragments of the molecule, wherein each fragment comprises a portion of the molecule.

6. The method of claim 5, wherein generating the one or more fragments comprises identifying one or more rotatable bonds in the molecule, at least one torsion atom associated with each rotatable bond, and one or more local atoms to the at least one torsion atom.

7. The method of claim 6, wherein identifying the one or more local atoms to the at least one torsion atom comprises, for each torsion atom: identifying one or more covalently bonded atoms to the torsion atom; evaluating the one or more covalently bonded atoms to identify partial structures comprising the one or more covalently bonded atoms, wherein the partial structures comprise one or more of functional groups and ring systems; and including atoms of the partial structures as local atoms in the one or more fragments.

8. The method of claim 7, further comprising: in response to an identification of a ring system, identifying possible ortho substitutions of the ring system as a partial structure.

9. The method of any one of claims 1-8, wherein the graph representation of at least one portion of the molecule comprises a graph representation of one or more fragments of the molecule.

10. The method of any one of claims 1-9, wherein processing the data using the graph neural network to generate the embedding of a graph representation of at least one portion of the molecule comprises: processing the data representing the at least one portion of the molecule to generate a set of node embeddings and a set of edge embeddings, and to identify one or more torsion features; using one or more augmentation neural networks to update the set of edge embeddings in accordance with the one or more torsion features; andupdating the set of node embeddings and the set of updated edge embeddings by performing message passing operations conditioned on a topology of the graph representation.

11. The method of claim 10, wherein using one or more augmentation neural networks to update the set of edge embeddings in accordance with the torsion features comprises, for each edge embedding: updating an edge embedding of a non-rotatable bond by processing one or more node embeddings of atoms corresponding with the non-rotatable bond using an augmentation neural network; or updating an edge embedding of a rotatable bond by generating and aggregating respective outputs of at least two augmentation neural networks.

12. The method of claim 11, wherein updating an edge embedding of the rotatable bond comprises: generating a respective output with each augmentation neural network by processing one or more node embeddings of atoms corresponding with the rotatable bond using the augmentation neural network; and aggregating the respective outputs of each augmentation neural network by summing a first output, a second output weighted by a sine of the phase angle of the torsion features, and a third output weighted by a cosine of the phase angle.

13. The method of any one of claims 1-12, further comprising generating the embedding of the graph representation using graph pooling.

14. The method of claim 13, wherein generating the embedding of the graph representation using graph pooling comprises: using an attention mechanism to determine a set of attention weights corresponding with the set of node embeddings; and aggregating the node embeddings in accordance with the attend on-weights to generate the embedding of the graph representation.

15. The method of any one of claims 1-14, wherein the distribution prediction machine learning model comprises a mixture density network configured to generate a set of parameters for N circular distributions in the circular mixture distribution.

16. The method of claim 15, wherein the one or more parameters comprise, for each circular distribution in the circular mixture distribution, a set of parameters comprising a mean value, a standard deviation value, and a mixing coefficient value.

17. The method of any one of claims 10-16, wherein the graph neural network, one or more augmentation neural networks, and distribution prediction machine learning model have been jointly trained by operations comprising, for each of a number of training iterations: receiving a corresponding ground truth torsion profile for each rotatable bond in the at least one portion of the molecule; identifying a set of allowable torsion angle values for the at least one portion of the molecule; sampling one or more torsion angle values from the identified set of allowable values; for each sampled torsion angle values, processing the at least one portion of the molecule with the sampled torsion angle value to generate the one or more parameters of the circular mixture distribution characteristic of the torsion profile of the at least one portion of the molecule; and updating values of respective sets of parameters comprising a set of graph neural network parameters, one or more sets of augmentation neural network parameters, and a set of distribution prediction machine learning model parameters in accordance with minimizing a discrepancy between the ground truth torsion profile of the at least one portion of the molecule and the generated torsion profile of the at least one portion of the molecule.

18. The method of any one of claims 1-17, further comprising, for each molecule in a set comprising a plurality of molecules: generating the one or more parameters of the circular mixture distribution characteristic of the torsion profile of the at least one portion of the molecule; generating the torsion profile of the at least one portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution; andusing the torsion profile to characterize a likelihood of observing a conformer of the molecule in a binding site of a second molecule.

19. The method of claim 18, further comprising: selecting one or more of the molecules for physical synthesis based at least on the respective likelihood of observing the conformer of the molecule in the binding site of the second molecule.

20. The method of claim 19, further comprising physically synthesizing one or more of the selected molecules.

21. The method of claim 1, further comprising: for each portion of the molecule comprising a rotatable bond, generating the torsion profile of the portion of the molecule using a circular mixture model parameterized by the one or more parameters of the circular mixture distribution generated for the portion of the molecule; and performing conformational sampling using the torsion profiles corresponding with each portion of the molecule.

22. The method of claim 21, wherein performing conformational sampling using the torsion profiles corresponding with each portion of the molecule comprises: sampling a torsion angle from the torsion profile for each portion of the molecule; and generating a molecular conformation defined by the sampled torsion angles.

Citation Information

Patent Citations

  • Degron and neosubstrate identification

    WO2023091567A1

  • Ternary complex modelling for molecular glues

    WO2024123853A1