Systems and methods for dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models

A graph neural network with chirality-aware representations and diffusion modeling effectively predicts protein-ligand binding complexes, addressing the limitations of existing methods by enhancing ligand pose accuracy and binding site recovery.

WO2025160309A1PCT designated stage Publication Date: 2025-07-31IAMBIC THERAPEUTICS INC +5

Patent Information

Application Number
PCT/US2025/012811
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-11
Filing Date
2025-01-23
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing deep generative models for protein structure prediction provide incomplete information about protein function and are insufficient for structure-based drug design, particularly in predicting protein-ligand interactions.

Method used

A graph neural network with chirality-aware pairwise representations is used to process molecular structures, incorporating atomic nodes, local coordinate frames, and stereospecific embeddings, and a diffusion model to generate accurate protein-ligand binding complex structures.

Benefits of technology

The method achieves high-accuracy prediction of protein-ligand binding complexes, improving ligand pose accuracy by up to 78% and binding site recovery by 59% compared to existing methods, while handling dynamic protein conformations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012811_31072025_PF_FP_ABST
    Figure US2025012811_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods described herein include embodiments for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate a geometry prior; sampling an initial geometrical structure of the binding complex based on the geometry prior; and processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.59215-726.601 SYSTEMS AND METHODS FOR DYNAMIC-BACKBONE PROTEIN-LIGAND STRUCTURE PREDICTION WITH MULTISCALE GENERATIVE DIFFUSION MODELS CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No.63 / 624,758, filed January 24, 2024, U.S. Provisional Application No.63 / 662,579, filed June 21, 2024, U.S. Provisional Application No.63 / 714,805, filed October 31, 2024, and U.S. Provisional Application No. 63 / 730,862, filed December 11, 2024, each of which are incorporated herein by reference in their entirety. BACKGROUND

[0002] Recent developments in deep generative learning have demonstrated substantial progress in the synthesis of complex vision and language data. Two such strategies for generative modeling are (a) auto-regressive models, and (b) diffusion-based generative models. Recent works on applying these strategies have demonstrated that deep generative models are capable of producing de novo-designed proteins with experimentally validated functions.

[0003] Recently, there was remarkable success in predicting protein crystal structures, however, while deep generative models demonstrably can be used to generate molecular structures of proteins, single- structure formulation of the protein folding problem provides incomplete information about protein function and is often insufficient for structure-based drug design. SUMMARY

[0004] In some aspects, the present disclosure provides a graph neural network for learning a chirality- aware pairwise representation of a molecule, comprising: a graph transformer comprising: i. a set of atomic nodes; ii. a set of local coordinate frame nodes; and iii. a set of stereospecific pairwise embeddings between the set of atomic nodes and the set of local coordinate frame nodes.

[0005] In some embodiments, the set of atomic nodes encodes a set of atoms in the molecule. In some embodiments, the set of atomic nodes encodes a group number for the set of atoms in the molecule. In some embodiments, the group number comprises an integer ranging from 1 to 18. In some embodiments, the set of atomic nodes encodes a period. In some embodiments, the period number comprises an integer ranging from 1 to 7. In some embodiments, the set of atoms comprises heavy- atoms of the molecule. In some embodiments, the heavy-atoms comprise atoms comprising at least two protons.

[0006] In some embodiments, the set of local coordinate frame nodes encodes a set of local coordinate frames of the molecule. In some embodiments, a local coordinate frame, in the set of local coordinate frames, comprises three atoms consecutively bonded to one another in the molecule. In some embodiments, the set of local coordinate frame nodes encodes a group number for a central atom in aAttorney Docket No.59215-726.601 local coordinate frame in the molecule. In some embodiments, the set of local coordinate frame nodes encodes a period number for a central atom in a local coordinate frame in the molecule.

[0007] In some embodiments, the set of stereospecific pairwise embeddings encodes a pair of local coordinate frames in the molecule. In some embodiments, the set of stereospecific pairwise embeddings encodes a molecular symmetry annotation for the pair of local coordinate frames in the molecule. In some embodiments, the molecular symmetry annotation encodes, when the pair of local coordinate frames shares the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other. In some embodiments, the molecular symmetry annotation encodes, when the pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or a different side of the common double or aromatic bond.

[0008] In some embodiments, the set of stereospecific pairwise embeddings comprises a set of edge embeddings connecting the set of atomic nodes with the set of local coordinate frame nodes. In some embodiments, the set of edge embeddings are based on encodings of the set of atomic nodes and the set of local coordinate frame nodes. In some embodiments, the set of edge embeddings comprises edges up to the fourth nearest neighbor atoms in the molecule.

[0009] In some embodiments, the graph transformer is configured to process the set of atomic nodes, the set of local coordinate frame nodes, and the set of stereospecific pairwise embeddings to generate the representation of the molecule. In some embodiments, the process is configured to recursively update a set of path representations for the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings. In some embodiments, the process is configured to recursively update the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings based on the set of path representations.

[0010] In some aspects, the present disclosure provides a graph neural network for processing a representation of a molecule, comprising: a graph transformer comprising: i. a set of heavy-atom nodes ^^ encoding, for each heavy-atom in the molecule, the group and the period of the heavy-atom in cdimensions such that ^^ ∈ ^^Nheavy−atoms × c; ii. a set of local coordinate frame nodes ^^ encoding, foreach local coordinate frame in the molecule, (i) the group and the period of the central atom and (ii)types of bonds in the local coordinate frame in c dimensions, such that ^^ ∈ ^^Nframes × c;a set ofstereochemistry encodings ^^ encoding, for each pair of local coordinate frames in the molecule,molecular symmetry annotations in cs dimensions, such that ^^ ∈molecularsymmetry annotations comprising: 1. a first molecular symmetry annotation encoding, when a pair of local coordinate frames share the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other local coordinate frame; and 2. a second molecular symmetry annotation encoding, when the pair of local coordinate frames shares aAttorney Docket No.59215-726.601 common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or the different side of the common double or aromatic bond; and iv. a set of pair representations G encoding, for each pair between heavy-atoms and local coordinate frames in themolecule, ^^ and ^^ in cp dimensions, such that ^^ ∈ ^^Nframes × Nheavy−atoms × cp;the graphtransformer is configured to process the representation of a molecule, by: 1. processing H, F, G, and S for n iterations to generate featureswherein the features^^^^^^^^^^,the same dimensions as H, F, G, and S, respectively, and n is an integer greater than zero, wherein the processing comprises: a. recursively updating path representations ^^^^, ^^^^+1, ^^outbased on ^^^^, ^^^^, ^^^^, and ^^^^, in the n iterations, wherein l is an iteration index ranging from zero to n; and b. recursively updating ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, based on ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^..

[0011] In some embodiments, the graph neural network is trained based on at least one of: a negative log likelihood evaluated based on an output 3-dimensional mixture probability distribution of a geometry of the molecule and an observed geometry of the training sample; a SE(3)-invariant denoising score matching loss based on a frame aligned point error determined between a denoised molecular geometry and a ground-truth molecular geometry of the molecule; a mean squared loss for fitting level-1 chemical checker (CC) embeddings representing harmonized and integrated bioactivity data; a binary cross entropy loss for classifying whether a specific CC entry is available for any molecule in the chemical checker dataset; and a cross-entropy loss for predicting the masked tokens which is added to encourage learning on molecular graph topology distributions.

[0012] In some aspects, the present disclosure provides a method for learning a chirality-aware pairwise representation of a molecule, comprising processing: (a) a set of atomic nodes encoding a set of atoms in the molecule; (b) a set of local coordinate frame nodes encoding a set of local coordinate frames in the molecule; and (c) a set of stereospecific pairwise embeddings, between the set of atomic nodes and the set of local coordinate frame nodes; to generate the representation of the molecule.

[0013] In some embodiments, the processing comprises using a graph transformer. In some embodiments, the processing comprises processing a noisy geometry to generate a denoised molecular geometry. In some embodiments, the processing comprises processing the set of atomic nodes, the set of local coordinate frame nodes, and the set of stereospecific pairwise embeddings to generate the representation of the molecule. In some embodiments, the processing comprises recursively updating a set of path representations for the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings. In some embodiments, the processing comprises recursively updating the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings based on the set of path representations. In some embodiments, the processing comprises denoising the noisy molecular geometry based on the recursive updating to generate the denoised molecular geometry. In some embodiments, the processing comprises denoising with SE(3)-invariance.Attorney Docket No.59215-726.601

[0014] In some aspects, the present disclosure provides a method for processing a representation of a molecule, comprising: (a) generating a set of heavy-atom nodes ^^ encoding, for each heavy-atom inthe molecule, the group and the period of the heavy-atom in c dimensions such that ^^ ∈^^Nheavy−atoms × c; (b) generating a set of local coordinate frame nodes ^^ encoding, for each local coordinate frame in the molecule, (i) the group and the period of the central atom and (ii) types ofbonds in the local coordinate frame in c dimensions, such that ^^ ∈ ^^Nframes × c;generating a set ofstereochemistry encodings ^^ encoding, for each pair of local coordinate frames in the molecule,molecular symmetry annotations in cs dimensions, such that ^^ ∈molecularsymmetry annotations comprising: i. a first molecular symmetry annotation encoding, when a pair of local coordinate frames share the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other local coordinate frame; and ii. a second molecular symmetry annotation encoding, when the pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or the different side of the common double or aromatic bond; (d) generating a set of pair representations G encoding, for each pair between heavy-atoms and local coordinate frames in themolecule, ^^ and ^^ in cp dimensions, such that ^^ ∈ ^^Nframes × Nheavy−atoms × cp;processing a noisygeometry to generate a denoised molecular geometry, by: i. processing H, F, G, and S for n iterations to generate featureswherein the featuresand ^^^^^^^^^^comprise the same dimensions as H, F, G, and S, respectively, and n is an integer greater than zero, wherein the processing comprises: ii. recursively updating path representations ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^, based on ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, in the n iterations, wherein l is an iteration index ranging from zero to n; and iii. recursively updating ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, based on ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^.

[0015] In some aspects, the present disclosure provides a method of training a neural network, comprising: (a) providing a training sample of a representation a molecule; (b) processing, using a neural network, the representation of the molecule; (c) updating parameters of the neural network based on at least one of: i. a negative log likelihood evaluated based on an output 3-dimensional mixture probability distribution of a geometry of the molecule and an observed geometry of the training sample; ii. a SE(3)-invariant denoising score matching loss based on a frame aligned point error determined between the denoised molecular geometry and a ground-truth molecular geometry of the molecule; iii. a mean squared loss for fitting level-1 chemical checker (CC) embeddings representing harmonized and integrated bioactivity data; iv. a binary cross entropy loss for classifying whether a specific CC entry is available for any molecule in the chemical checker dataset; and a cross-entropy loss for predicting the masked tokens which is added to encourage learning on molecular graph topology distributions.Attorney Docket No.59215-726.601

[0016] In some aspects, the present disclosure provides a graph neural network for processing a representation of a protein, comprising: a graph transformer comprising: i. a sparse graph of the protein, comprising: 1. a set of amino acid sequence nodes; 2. a set of atomic nodes; 3. a perturbed protein geometry; and 4. a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes; and ii. one or more stacks of invariant point attention blocks; wherein the graph transformer is configured to process an input amino acid sequence, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0017] In some embodiments, the set of atomic nodes comprises a set of backbone atom nodes. In some embodiments, the set of edges connect the set of backbone atom nodes to the set of amino acid sequence nodes. In some embodiments, the set of edges are generated based on an inclusion probability, the inclusion probability being based on a distance between a backbone atom and a residue in the perturbed protein geometry. In some embodiments, the graph transformer is configured to process a diffusion time. In some embodiments, the diffusion time comprises a random Fourier encoding. In some embodiments, the set of edges are initialized with a randomly Fourier encoded signed sequence distance between two connected nodes if the two connected nodes are located on the same chain, and zeros if the two connected nodes are located on different chains. In some embodiments, the perturbed coordinates are sampled from a learned reverse-time SDE.

[0018] In some embodiments, the one or more stacks of invariant point attention blocks are configured to compute attention scores on the graph. In some embodiments, the one or more stacks of invariant point attention blocks are configured to associate each node of the sparsely-connected graph with a plurality of replica coordinate frames. In some embodiments, the one or more stacks of invariant point attention blocks are configured to output a plurality translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames. In some embodiments, a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation.

[0019] In some aspects, the present disclosure provides a graph neural network for processing a representation of a protein, comprising: a graph transformer comprising: i. a sparse graph of the protein, comprising: 1. a set of amino acid sequence nodes; 2. a set of atomic nodes comprising a set of backbone atom nodes; 3. a perturbed protein geometry sampled from a learned reverse-time SDE; and 4. a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes, wherein the set of edges are generated based on an inclusion probability function; and ii. one or more stacks of invariant point attention blocks configured to: 1. compute attention scores on the graph; 2. associate each node of the sparsely-connected graph with a plurality of replica coordinate frames, wherein a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation; and 3. output a plurality translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames; wherein the graph transformer is configured to process an input amino acidAttorney Docket No.59215-726.601 sequence, a diffusion time, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0020] In some aspects, the present disclosure provides a method for processing a representation of a protein, comprising: (a) providing a graph transformer comprising: i. a sparse graph of the protein, comprising: 1. a set of amino acid sequence nodes; 2. a set of atomic nodes; 3. a perturbed protein geometry; and 4. a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes; and ii. one or more stacks of invariant point attention blocks; and (b) processing, using the graph transformer, an input amino acid sequence, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0021] In some aspects, the present disclosure provides a method for processing a representation of a protein, comprising: (a) providing a graph transformer comprising: i. a sparse graph of the protein, comprising: 1. a set of amino acid sequence nodes; 2. a set of atomic nodes comprising a set of backbone atom nodes; 3. a perturbed protein geometry sampled from a learned reverse-time SDE; and 4. a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes, wherein the set of edges are generated based on an inclusion probability function; and ii. one or more stacks of invariant point attention blocks configured to: 1. compute attention scores on the graph; 2. associate each node of the sparsely-connected graph with a plurality of replica coordinate frames, wherein a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation; and 3. output a plurality translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames; (b) processing, using the graph transformer, an input amino acid sequence, a diffusion time, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0022] In some aspects, the present disclosure provides a method for predicting contacts between a protein and a ligand, comprising processing (i) a protein graph, (ii) a ligand graph, and (iii) a set of intermolecular edges connecting the protein graph and the ligand graph, to autoregressively sample the contacts based on a probability distribution output by a neural network that permits a plurality of contact modes. In some embodiments, the method further comprises generating the ligand graph. In some embodiments, the method further comprises generating the protein graph.

[0023] In some embodiments, the neural network outputs the probability distribution based on the (i) the protein graph, (ii) the ligand graph, and (iii) the set of intermolecular edges. In some embodiments, the neural network comprises one or more neural network blocks. In some embodiments, the neural network comprises an invariant point attention neural network that processes the protein graph to generate a first set of features. In some embodiments, the neural network comprises a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features, while using edge embeddings of the ligand graph as a relative positional encoding term. In some embodiments, the neural network comprises a second multi-head cross-attention neural network thatAttorney Docket No.59215-726.601 processes the first set of features, the second set of features, and the set of intermolecular edges to generate a third set of features. In some embodiments, the neural network comprises a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges, wherein the updated features comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0024] In some aspects, the present disclosure provides a method for predicting contacts between a protein and a ligand, comprising: (a) enumerating a set of intermolecular edges connecting a sparsely- connected protein graph and a ligand graph, wherein the set of intermolecular edges connects residues of the sparsely-connected protein graph and heavy-atoms of the ligand graph; (b) processing, using a neural network, (i) the sparsely-connected protein graph, (ii) the ligand graph, and (iii) the set of intermolecular edges, to predict the contacts, wherein: i. the sparsely-connected protein graph encodes an amino-acid sequence of the protein, a perturbed geometry of the protein, and a diffusion time; ii. the ligand graph encodes (A) a set of atoms of the ligand, (B) a set of local coordinate frames of the ligand, and (C) a set of stereospecific pairwise embeddings between the set of atoms and the set of local coordinate frames; and iii. the neural network comprises: 1. one or more neural network blocks each comprising: a. an invariant point attention neural network that processes the sparsely-connected protein graph to generate a first set of features; b. a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features, while using edge embeddings of the ligand graph as a relative positional encoding term; c. a second multi-head cross-attention neural network that processes the first set of features, the second set of features, and the pairwise intermolecular edges to generate a third set of features; and d. a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges, wherein the updated features comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0025] In some embodiments, the method further comprises training the neural network to reduce a difference between (i) the posterior distribution of observed contact maps in training data, and (ii) the predicted contacts. In some embodiments, the method further comprises training the neural network to reduce an element-wise difference between (i) the observed contact maps in the training data, and (ii) softmax-transformation of predicted contacts. In some embodiments, the probability distribution comprises a categorical posterior distribution over a sequence of the protein.

[0026] In some embodiments, the processing comprises segmenting and tokenizing the protein graph to generate a first set of patches representing the protein. In some embodiments, the processing comprises frame sampling the ligand graph to generate a second set of patches representing the ligand. In some embodiments, the neural network comprises an intra-patch self-attention mechanism for generating a cross-attention map based on the first set of patches or the second set of patches. In some embodiments, the neural network comprises an inter-patch triangular-gated self-attention mechanism for processing a self-attention map based on the first set of patches and the second set of patches. InAttorney Docket No.59215-726.601 some embodiments, the neural network comprises a graph-attention mechanism for processing a graph attention map based on the set of intermolecular edges.

[0027] In some embodiments, the autoregressive sampling comprises generating (i) a plurality of intermediate contact probabilities, (ii) a plurality of intermediate distance histograms, or (iii) both. In some embodiments, the autoregressive sampling comprises updating the set of intermolecular edges based on (i) the plurality of intermediate contact probabilities, (ii) the plurality of intermediate distance histograms, or (iii) both.

[0028] In some aspects, the present disclosure provides a method for predicting contacts between a protein and a ligand, comprising: (a) segment tokenizing a protein representation to generate a set of protein representation patches; (b) frame subsampling a ligand representation to generate a set of ligand representation patches; (c) generating an intermolecular representation comprising a plurality of edges that connect a first subset of nodes in the protein representation and a second subset of nodes in the ligand representation; and (d) sampling the contacts from a posterior probability distribution that permits a plurality of modes based on (i) the set of protein representation patches, (ii) the set of ligand representation patches, and (iii) the intermolecular representation, wherein the sampling comprises: i. using a neural network that comprises (i) an intra-patch attention mechanism between the sparse edges and the dense edges to process the set of protein representation patches and the set of ligand representation patches, (ii) an inter-patch self-attention mechanism to process the set of protein representation patches and the set of ligand representation patches, and (iii) a graph attention mechanism to process the intermolecular representation; ii. autoregressively sampling a plurality of intermediate contact maps for a plurality of iterations, while updating the plurality of edges of the intermolecular representation based on the plurality of intermediate contact maps; and iii. outputting the contacts based on a final contact map of the plurality of intermediate contact maps.

[0029] In some aspects, the present disclosure provides an autoregressive neural network for predicting contacts between a protein and a ligand, comprising: (a) an input comprising (i) a protein graph, (ii) a ligand graph, (iii) a set of intermolecular edges connecting the protein graph and the ligand graph; and (b) an output comprising parameters of a probability distribution that permits a plurality of contact modes.

[0030] In some embodiments, the autoregressive neural network further comprises a graph neural network for providing the ligand graph. In some embodiments, the autoregressive neural network further comprises a graph neural network for providing the protein graph. In some embodiments, the autoregressive neural network further comprises one or more neural network blocks. In some embodiments, the autoregressive neural network further comprises an invariant point attention neural network that processes the protein graph to generate a first set of features. In some embodiments, the autoregressive neural network further comprises a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features, while using edge embeddings of the ligand graph as a relative positional encoding term. In some embodiments, the autoregressive neuralAttorney Docket No.59215-726.601 network further comprises a second multi-head cross-attention neural network that processes the first set of features, the second set of features, and the set of intermolecular edges to generate a third set of features. In some embodiments, the autoregressive neural network further comprises a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges, wherein the updated features comprise the contacts for the last neural network block in the plurality of neural network blocks. In some embodiments, the protein graph is segmented and tokenized into a first set of patches representing the protein. In some embodiments, the ligand graph is frame-sampled to generate a second set of patches representing the ligand. In some embodiments, the autoregressive neural network further comprises an intra-patch self-attention mechanism for generating a cross-attention map based on the first set of patches or the second set of patches. In some embodiments, the autoregressive neural network further comprises an inter-patch triangular-gated self-attention mechanism for processing a self-attention map based on the first set of patches and the second set of patches. In some embodiments, the autoregressive neural network further comprises a graph-attention mechanism for processing a graph attention map based on the set of intermolecular edges. In some embodiments, the output further comprises (i) a plurality of intermediate contact probabilities, (ii) a plurality of intermediate distance histograms, or (iii) both. In some embodiments, the autoregressive neural network is configured to autoregressively update the set of intermolecular edges based on (i) the plurality of intermediate contact probabilities, (ii) the plurality of intermediate distance histograms, or (iii) both.

[0031] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising: (a) sampling an initial geometrical structure of the binding complex from a geometry prior; and (b) denoising, using a machine-learned stochastic differential equation (SDE), the initial geometrical structure to generate the geometrical structure of the binding complex.

[0032] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising predicting the geometrical structure of the binding complex based on a geometry prior comprising (i) noise structured on a template geometrical structure of the protein and (ii) predicted contacts between the protein and the ligand.

[0033] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising predicting the geometrical structure of the binding complex based on a sequence representation of the protein and a graph representation of the ligand, and optionally, on a geometry prior comprising predicted contacts between the protein and the ligand.

[0034] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising: (a) processing (i) a first representation comprising a protein representation and (ii) a second representation comprising aAttorney Docket No.59215-726.601 ligand representation to generate a geometry prior comprising (i) noise structured on the geometrical structure of the binding complex and (ii) predicted contacts between the protein and the ligand; and (b) denoising the geometrical structure of the binding complex sampled from the geometry prior.

[0035] In some embodiments, the geometry prior is based on a first representation comprising a protein representation. In some embodiments, the geometry prior is based on a second representation comprising a ligand representation. In some embodiments, the first representation comprises a protein complex representation of a plurality of proteins. In some embodiments, the second representation comprises a plurality of ligand representations. In some embodiments, the geometry prior comprises contacts between the protein and the ligand in the binding complex. In some embodiments, the geometry prior comprises an initial geometry of the protein in the binding complex. In some embodiments, the geometry prior comprises an initial geometry of the ligand in the binding complex. In some embodiments, the geometry prior comprises an initial geometry of a macromolecule in the binding complex.

[0036] In some embodiments, the method further comprises generating the contacts by iteratively sampling binding interface spatial proximity distributions of the binding complex. In some embodiments, the method further comprises generating the initial geometry of the protein. In some embodiments, the method further comprises generating the initial geometry of the ligand.

[0037] In some embodiments, the template geometrical structure comprises an experimentally determined geometrical structure. In some embodiments, the experimentally determined geometrical structure is determined using X-ray crystallography, nuclear magnetic resonance spectroscopy, or cryo-electron microscopy. In some embodiments, the template geometrical structure comprises a computationally determined geometrical structure. In some embodiments, the computationally determined geometrical structure is determined using an electronic structure calculation, a molecular dynamics simulation, a Monte Carlo simulation, a machine learning model, or any combination thereof. For example, a hybrid QM / MM calculation can be performed to determine a geometrical structure. In some embodiments, the template geometrical structure comprises coordinates of the backbone atoms of the protein. In some embodiments, the template geometrical structure comprises coordinates of the atoms of the protein that are distal from the ligand in the binding complex, wherein an atom is distal if it is further than a predetermined distance from the ligand.

[0038] In some embodiments, the predicting the geometrical structure of the binding complex comprises fixing the geometrical structure of the protein to the template geometrical structure of the protein. In some embodiments, the predicting the geometrical structure of the binding complex comprises allowing the geometrical structure of the protein to depart from the template geometrical structure of the protein. In some embodiments, the predicting the geometrical structure of the binding complex comprises fixing the geometrical structure of the ligand to a template geometrical structure of the ligand. In some embodiments, the predicting the geometrical structure of the binding complexAttorney Docket No.59215-726.601 comprises allowing the geometrical structure of the ligand to depart from a template geometrical structure of the ligand.

[0039] In some embodiments, the method further comprises generating a second ligand that is configured to bind to the protein. In some embodiments, the generating the second ligand comprises performing a gradient-based design using a differentiable protein sequence and / or a molecular graph generator. In some embodiments, the geometry prior comprises a finite-time marginal of a SDE configured to inject structured noise into the data distribution.

[0040] In some embodiments, the binding complex comprises an apoprotein. In some embodiments, the binding complex comprises a holoprotein.

[0041] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising: (a) performing an E(3)-equivariant forward-time-noising process of a truncated stochastic differential equation (SDE) to^^ = ^^∗ < ∞ based on a contact map to generate a geometry prior, such that the geometry priorcomprises a partially-diffused structured distribution of the geometrical structure of the binding complex, wherein the SDE is configured to: i. retain information about domain packing of the protein; ii. retain information about ligand binding interfaces of the protein; and iii. erase information about residue-scale local details of the protein; (b) sampling a noisy geometrical structure from the geometry prior; and (c) performing, using a machine-learned reverse-time SDE, a SE(3)-equivariant reverse- time-denoising process on the noisy geometrical structure of the binding complex sampled from the geometry prior, by attracting coordinates of the ligand and the protein based on the contact map, the machine-learned reverse-time SDE comprising: i. inputs comprising: 1. a protein representation comprising: a. a set of protein heavy-atom nodes: b. a set of protein heavy-atom coordinates; c. a protein residue-wise representation from a protein encoder; d. a protein atom type; and e. a first random Fourier encoding of diffusion time step t; 2. a ligand representation comprising: a. a set of ligand coordinates; b. a ligand representation from a path-integral graph transformer; and c. a second random Fourier encoding of the diffusion time step t; 3. a protein-ligand representation comprising: a. a protein- ligand graph; ii. outputs comprising a set of displacement vectors for the set of protein heavy-atom coordinates and the set of ligand coordinates.

[0042] In some embodiments, the neural network of the machine-learned reverse-time SDE is trained with a loss function configured to reduce (1) a difference between an observed ground-truth geometrical structure and the denoised geometrical structure of the binding complex. In some embodiments, the SDE retains information about domain packing of the protein in the forward-noising process by diffusing the protein’s backbone atoms around a template protein backbone structure. In some embodiments, the SDE retains information about ligand binding interfaces of the protein in the forward-noising process by diffusing the ligand atoms around the protein’s residues that are predicted to contact the ligand atoms, the diffusing based on a drift term which is based on a contact map, wherein the contact map comprises predicted contacts between the ligand atoms and the protein’s residues, andAttorney Docket No.59215-726.601 wherein the drift term is defined relative to the protein’s residues. In some embodiments, the SDE erases information about residue-scale local details by diffusing non-backbone atoms of the protein, excluding the protein backbone, around the protein backbone of the binding complex, wherein the drift term of the protein residue coordinates in the SDE is relative to the protein backbone of the binding complex. In some embodiments, the loss function is time-dependent such that the loss function is normalized by a factor that decreases with reverse-time.

[0043] In some embodiments, the protein-ligand representation further comprises: (a) a first set of edges connecting protein heavy-atom nodes and the residue node that the protein atom belongs to; (b) a second set of edges connecting pairs of protein heavy-atom nodes that are within the same residue; (c) a third set of edges connecting pairs of protein heavy-atom nodes that are within a first predetermined distance; and (d) a fourth set of edges connecting protein heavy-atom nodes and ligand atom nodes that are within a second first predetermined distance; wherein the edges in the first, second, third, and fourth set of edges that comprise a protein heavy-atom node are initialized with features that encode (i) whether two nodes of an edge belong to the same residue or the same ligand molecule, and (ii) whether there is a covalent bond between two nodes of the edge, wherein the two nodes are considered be covalently bonded if a distance between the two nodes are less than about an average Van der Waals (VdW) radius of atoms constituting the two nodes.

[0044] In some embodiments, the sampling the initial geometrical structure of the binding complex from the geometry prior is based on an inverse temperature parameter. In some embodiments, the inverse temperature parameter increases during the sampling process. In some embodiments, the denoising the initial geometrical structure to generate the geometrical structure of the binding complex is based on an inverse temperature parameter. In some embodiments, the denoising comprises annealing the initial geometrical structure.

[0045] In some embodiments, the method further comprises generating a report indicating the geometrical structure of the binding complex. In some embodiments, the method further comprises generating a report indicating a representation of the ligand.

[0046] In some aspects, the present disclosure provides a method of generating an identification of a ligand predicted to bind with a protein to form a binding complex. The method can comprise predicting contacts between the ligand and the protein in the binding complex. The method can comprise generating a geometry prior based on the contacts. The method can comprise denoising the geometry prior to generate a structure of the binding complex. The method can comprise generating a report indicating the identification of the ligand based on the structure of the binding complex.

[0047] In some embodiments, the contacts comprise pairwise proximities among residues of the protein and frames of the graph of the ligand.

[0048] In some embodiments, the predicting the contacts is based on a sequence of the protein, an embedding of the protein, a graph of the ligand, or any combination thereof.Attorney Docket No.59215-726.601

[0049] In some embodiments, the denoising is performed with SE(3) equivariance, reducing temperature, or both.

[0050] In some embodiments, the method further comprises generating a report indicating the geometrical structure of the binding complex.

[0051] In some embodiments, the method further comprises generating a report indicating a representation of the ligand.

[0052] In some embodiments, the method further comprises generating a report indicating an uncertainty of the geometrical structure of the binding complex. The uncertainty can be provided at the resolution of each atom in the geometrical structure.

[0053] In some aspects, the present disclosure provides a computer-implemented method for identifying a ligand that binds to a protein, comprising: sending instructions to generate an identification of the ligand that binds to the protein to one or more computers, wherein the one or more computers are configured to implement any one of the methods or the neural networks disclosed herein to generate a report indicating the identification; receiving the report from the one or more computers.

[0054] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate a geometry prior; sampling an initial geometrical structure of the binding complex based on the geometry prior; and processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules.

[0055] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate an initial geometrical structure of the binding complex formed by the plurality of macromolecules based on plurality of representations, wherein the plurality of macromolecules comprise at least 32 amino acids and / or nucleotides; and processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of representations.

[0056] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules and a plurality of small molecules, comprising: processing an input representation comprising a plurality of representations of the plurality of macromolecules and the plurality of small molecules to generate a geometry prior; sampling an initial geometrical structure of the binding complex based on the geometry prior; and processing, using a neural network, the initial geometrical structure to generateAttorney Docket No.59215-726.601 the geometrical structure of the binding complex formed by the plurality of macromolecules and the plurality of small molecules.

[0057] In some embodiments, the plurality of macromolecules comprise at least 2, 3, 4, 5, or 6 macromolecules. In some embodiments, the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations. In some embodiments, the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.

[0058] In some embodiments, the input representation further comprises a representation of a cofactor of the binding complex. In some embodiments, the plurality of macromolecules comprises a protein, a nucleic acid, or both. In some embodiments, the input representation further comprises a representation of a ligand. In some embodiments, the input representation comprises a plurality of representations of a plurality of ligands.

[0059] In some embodiments, the input representation comprises at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides. In some embodiments, the input representation comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides.

[0060] In some embodiments, the processing the input representation comprises segmenting the protein representation into a plurality of segments. In some embodiments, the plurality of segments comprises wherein each segment of the plurality of segments comprises a uniform size. In some embodiments, each segment in the plurality of segments comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids. In some embodiments, each segment in the plurality of segments comprises 8 amino acids.

[0061] In some embodiments, the processing the input representation comprises predicting a plurality of contacts between a first macromolecule and a second macromolecule of the plurality of macromolecules. In some embodiments, the predicting the plurality of contacts between the first macromolecule and the second macromolecule is performed by processing the plurality of segments. In some embodiments, the processing the input representation comprises predicting a plurality of contacts between (1) a first macromolecule and a first small molecule, and / or (2) a second macromolecule and the first small molecule or a second small molecule.

[0062] In some embodiments, the predicting the plurality of contacts is based on a predetermined set of contacts. In some embodiments, the plurality of contacts comprises all or a subset of the predetermined set of contacts. In some embodiments, the predetermined set of contacts is provided by a user. In some embodiments, the predetermined set of contacts is generated from a template structure of the binding complex. In some embodiments, the template structure of the binding complex is provided by a user. In some embodiments, the predetermined set of contacts specifies the binding site. In some embodiments, the predetermined set of contacts specifies the binding site of aAttorney Docket No.59215-726.601 ligand, a cofactor, or a macromolecule. In some embodiments, the template structure of the binding complex is generated by molecular mechanics or homology modeling.

[0063] In some embodiments, the processing the initial geometrical structure to generate the geometrical structure of the binding complex comprises dynamically generating connections between pairs of atoms of the binding complex. In some embodiments, the dynamically generating the connections comprises using a mixture of probability distributions to assign the connections between the pairs of atoms. In some embodiments, the total number of connections in the ESDM graph is bounded by a hardware-specific threshold.

[0064] In some embodiments, the cofactor comprises an organic cofactor or an inorganic cofactor. In some embodiments, the organic cofactor comprises flavine, heme, NAD+, thiamin pyrophosphate, pyridoxal phosphate, methylcobalamin, cobalamine, biotin, coenzyme A, tetrahydrofolic acid, menaquinone, ascorbic acid, flavin mononucleotide, flavine adenine dinucleotide, coenzyme F420, Adenosine triphosphate, S-Adenosyl methionine, Coenzyme B, Coenzyme M, Coenzyme Q, Cytidine triphosphate, Glutathione, Lipoamide, Methanofuran, Molybdopterin, a nucleotide sugar, 3'- Phosphoadenosine-5'-phosphosulfate, Tetrahydrobiopterin, Tetrahydromethanopterin, or any combination thereof. In some embodiments, the inorganic cofactor comprises a metal ion, metal iron, an iron-sulfur complex, or another transition-metal organometallic complex. In some embodiments, the plurality of macromolecules comprises a post-translation modification or a plurality of amino acids with post-translation modifications. In some embodiments, the post-translation modification comprises acylation, alkylation, prenylation, flavination, amination, deamination, carboxylation, decarboxylation, nitrosylation, halogenation, sulfurylation, glutathionylation, oxidation, oxygenation, reduction, ubiquitination, SUMOylation, neddylation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glypiation, glycosylphosphatidylinositol anchor formation, lipoylation, heme functionalization, phosphorylation, phosphopantetheinylation, retinylidene Schiff base formation, diphthamide formation, ethanolamine phosphoglycerol functionalization, hypusine formation, beta-Lysine addition, acetylation, formylation, methylation, amidation, amide bond formation, butyrylation, gamma-carboxylation, glycosylation, polysialylation, malonylation, hydroxylation, iodination, nucleotide addition, phosphate ester formation, phosphoramidate formation, adenylation, uridylylation, propionylation, pyroglutamate formation, gluthathionylation, sulfenylation, sulfinylation, sulfonylation, succinylation, sulfation, glycation, carbonylation, isopeptide bond formation, biotinylation, carbamylation, oxidation, pegylation, citrullination, deamidation, eliminylation, disulfide bond formation, proteolytic cleavage, isoaspartate formation, racemization, protein splicing, chaperone-assisted folding, or any combination thereof.

[0065] In some embodiments, the nucleic acid comprises a helix, a bulge, a loop, a junction, a stem- loop, a hairpin-loop, a tetraloop, a pseudoknot, a single strand, a double strand, or any combination thereof. In some embodiments, the nucleic acid comprises DNA, RNA, 5mC, 5hmC, 5fC, 5mC, 5hmC, 5fC, 5CaC, 6mA, 4mC, 8-oxoG, Tg, an AP site, DNA ps, or any combination thereof.Attorney Docket No.59215-726.601

[0066] In some embodiments, the neural network is trained with a training dataset comprising predicted geometrical structures of binding complexes of macromolecules. In some embodiments, the predicted geometrical structures of binding complexes comprise a plurality of monomers aligned to form the geometric structures. In some embodiments, the neural network is trained using a loss function configured to reduce achiral and steric clashes, wherein the loss function is adjusted with a weighting function that stabilizes the training. In some embodiments, the weighting function increases in value as the number of training iterations increases. In some embodiments, the neural network is trained by distributing batches of training data across a plurality of processors based on sizes of the training data to reduce load imbalance between the plurality of processors. In some embodiments, the training data comprises at least 100k training samples of geometrical structures of binding complexes. In some embodiments, the neural network comprises at least 100M parameters.

[0067] In some embodiments, the neural network is further trained with a second training dataset comprising geometrical structures of macromolecules or binding complexes generated from experiments, simulations, or electronic structure calculations. In some embodiments, the geometrical structures generated from experiments comprise cryo-EM generated structures, X-ray crystallography generated structures, and / or NMR generated structures. In some embodiments, the geometrical structures generated from simulations comprise molecular dynamics simulation or Monte Carlo simulation generated structures. In some embodiments, the geometrical structures are generated with homology modeling. In some embodiments, the geometrical structures are generated without homology modeling. In some embodiments, the geometrical structures generated from electronic structure calculations comprise DFT generated structures.

[0068] In some embodiments, a computational cost of processing the initial geometrical structure to generate the geometrical structure of the binding complex scales sub-linearly as a function of the number of amino acid residues. In some embodiments, the neural network is configured to provide a mean TM-score accuracy of 0.9 for the geometrical structure of the binding complex. In some embodiments, a ligand root-mean-square deviation of the geometrical structure is below two Angstrom. In some embodiments, the neural network uses site-specific docking. In some embodiments, the method further comprises generating a confidence score for each amino acid in the geometrical structure. In some embodiments, the confidence score is above a threshold confidence value. In some embodiments, the threshold confidence value is 0.8.

[0069] In some embodiments, the method further comprises processing the geometrical structure of the binding complex using molecular mechanics. In some embodiments, the molecular mechanics comprises molecular dynamics. In some embodiments, the molecular mechanics comprises electronic structure calculations.

[0070] In some embodiments, the molecular mechanics is performed with restraints or without restraints. In some molecule mechanics calculations (e.g., molecular dynamics), a frequency of an interaction between two or more atoms may be similar or higher than a desired integration time stepAttorney Docket No.59215-726.601 (e.g., in numerical integration methods). Performing the molecular mechanics calculation with the desired integration time step can cause instabilities in the numerical integration that may lead the molecular mechanics calculation to yield inaccurate results. Thus, in some cases, a restraint can be used to treat the high frequency interactions as fixed interaction. This can allow the numerical integration to proceed with improved stability that yields accurate results, with the advantage of performing the integration at a larger integration time step. Larger integration time step can be advantageous for achieving better computational speed, especially when imposing the restraints do not substantially affect the physical accuracy of the molecular mechanics calculations.

[0071] For example, covalent bonds between carbon atoms and hydrogen atoms have a much higher frequency than bonds between carbon atoms. In these cases, a restraint can be imposed on the covalent bonds between the carbon atoms and the hydrogen atoms such that the covalent bond is treated as a rigid body (compared to say, a harmonic or an inharmonic spring), which can allow integration time steps to be much larger. However, in instances where a high frequency interaction is important to the physical accuracy of the molecular mechanics calculation, a restraint might not be used. A simple example of this may include, e.g., proton transfer between water molecules (Grotthuss mechanism).

[0072] In some embodiments, the restraints are imposed based on geometrical structure of the binding complex predicted from the neural network. In some embodiments, the restraints are imposed based on initial coordinates of the binding complex predicted from the neural network. For example, restraints on certain coordinates of the binding complex can be generated based on a confidence level prediction of the neural network. In parts of the geometrical structure where the confidence of prediction is very high, restraints can be imposed such that those sections / portions of the geometry do not move significantly or at all during further molecular mechanics calculations. The restraints can be defined in various coordinate systems. One example may be via internal coordinates of one or more molecules.

[0073] In some embodiments, the restraints are imposed on covalent bonds formed with a hydrogen atom. In some embodiments, the restraints are imposed on torsional modes of the binding complex.

[0074] In some embodiments, the method further comprises predicting experimental binding affinity between at least two molecules, wherein the at least two molecules comprises (a) at least two macromolecules, or (b) at least one macromolecule and at least one small molecule. In some embodiments, the input representation comprises a sequence of a macromolecule. In some embodiments, the sequence is an amino acid sequence of a protein. In some embodiments, the sequence is a nucleic acid sequence of a nucleic acid. In some embodiments, the input representation comprises a 3-dimensional structure of a molecule. In some embodiments, the 3-dimensional structure is of a protein, nucleic acid, or a small molecule. In some embodiments, the input representation comprises a sequence and a 3-dimensional structure of one or more macromolecules.Attorney Docket No.59215-726.601

[0075] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a macromolecule and one or more of (1) a cofactor, or (2) a ligand of a post translational modification, comprising: processing an input representation comprising a representation of the macromolecule and a representation of the cofactor or the ligand of the post translational modification to generate a geometry prior; sampling an initial geometrical structure of the binding complex based on the geometry prior; and processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the macromolecule and the cofactor or the ligand of the post translational modification.

[0076] In some embodiments, the input representation further comprises a representation of a second macromolecule. In some embodiments, a plurality of macromolecules comprises the macromolecule and the second macromolecule. In some embodiments, , a plurality of representations comprises a representation of each macromolecule of the plurality of macromolecules. In some embodiments, the plurality of macromolecules comprise at least 2, 3, 4, 5, or 6 macromolecules. In some embodiments, the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.

[0077] In some embodiments, the input representation comprises the representation of a cofactor of the binding complex. In some embodiments, the macromolecule comprises a protein, a nucleic acid, or both. In some embodiments, the input representation comprises the representation of a ligand of a post translational modification. In some embodiments, the input representation comprises a plurality of representations of a plurality of ligands of post translational modifications.

[0078] In some embodiments, the input representation comprises at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides. In some embodiments, the input representation comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides.

[0079] In some embodiments, the processing the input representation comprises segmenting the representation of the macromolecule into a plurality of segments. In some embodiments, the plurality of segments comprises wherein each segment of the plurality of segments comprises a uniform size. In some embodiments, each segment in the plurality of segments comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids. In some embodiments, each segment in the plurality of segments comprises 8 amino acids.

[0080] In some embodiments, the processing the input representation comprises predicting a plurality of contacts between the first macromolecule and the second macromolecule of the plurality of macromolecules. In some embodiments, the predicting the plurality of contacts between the first macromolecule and the second macromolecule is performed by processing the plurality of segments. In some embodiments, the processing the input representation comprises predicting a plurality of contacts between the macromolecule and the cofactor or ligand of the post translational modification.Attorney Docket No.59215-726.601 In some embodiments, the predicting the plurality of contacts is based on a predetermined set of contacts. In some embodiments, the plurality of contacts comprises all or a subset of the predetermined set of contacts. In some embodiments, the predetermined set of contacts is provided by a user. In some embodiments, the predetermined set of contacts is generated from a template structure of the binding complex. In some embodiments, the predetermined set of contacts specifies the binding site of a ligand, a cofactor, or a macromolecule. In some embodiments, the template structure of the binding complex is provided by a user. In some embodiments, the predetermined set of contacts specifies the binding site. In some embodiments, the template structure of the binding complex is generated by molecular mechanics or homology modeling.

[0081] In some embodiments, the processing the initial geometrical structure to generate the geometrical structure of the binding complex comprises dynamically generating connections between pairs of atoms of the binding complex. In some embodiments, the dynamically generating the connections comprises using a mixture of probability distributions to assign the connections between the pairs of atoms. In some embodiments, the total number of connections in an ESDM graph associated with the binding complex is bounded by a hardware-specific threshold.

[0082] In some embodiments, the cofactor comprises an organic cofactor or an inorganic cofactor. In some embodiments, the organic cofactor comprises flavine, heme, NAD+, thiamin pyrophosphate, pyridoxal phosphate, methylcobalamin, cobalamine, biotin, coenzyme A, tetrahydrofolic acid, menaquinone, ascorbic acid, flavin mononucleotide, flavine adenine dinucleotide, coenzyme F420, Adenosine triphosphate, S-Adenosyl methionine, Coenzyme B, Coenzyme M, Coenzyme Q, Cytidine triphosphate, Glutathione, Lipoamide, Methanofuran, Molybdopterin, a nucleotide sugar, 3'- Phosphoadenosine-5'-phosphosulfate, Tetrahydrobiopterin, Tetrahydromethanopterin, or any combination thereof. In some embodiments, the inorganic cofactor comprises a metal ion, metal iron, an iron-sulfur complex, or another transition-metal organometallic complex. In some embodiments, the post-translation modification comprises acylation, alkylation, prenylation, flavination, amination, deamination, carboxylation, decarboxylation, nitrosylation, halogenation, sulfurylation, glutathionylation, oxidation, oxygenation, reduction, ubiquitination, SUMOylation, neddylation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glypiation, glycosylphosphatidylinositol anchor formation, lipoylation, heme functionalization, phosphorylation, phosphopantetheinylation, retinylidene Schiff base formation, diphthamide formation, ethanolamine phosphoglycerol functionalization, hypusine formation, beta-Lysine addition, acetylation, formylation, methylation, amidation, amide bond formation, butyrylation, gamma-carboxylation, glycosylation, polysialylation, malonylation, hydroxylation, iodination, nucleotide addition, phosphate ester formation, phosphoramidate formation, adenylation, uridylylation, propionylation, pyroglutamate formation, gluthathionylation, sulfenylation, sulfinylation, sulfonylation, succinylation, sulfation, glycation, carbonylation, isopeptide bond formation, biotinylation, carbamylation, oxidation, pegylation, citrullination, deamidation,Attorney Docket No.59215-726.601 eliminylation, disulfide bond formation, proteolytic cleavage, isoaspartate formation, racemization, protein splicing, chaperone-assisted folding, or any combination thereof.

[0083] In some embodiments, the nucleic acid comprises a helix, a bulge, a loop, a junction, a stem- loop, a hairpin-loop, a tetraloop, a pseudoknot, a single strand, a double strand, or any combination thereof. In some embodiments, the nucleic acid comprises DNA, RNA, 5mC, 5hmC, 5fC, 5mC, 5hmC, 5fC, 5CaC, 6mA, 4mC, 8-oxoG, Tg, an AP site, DNA ps, or any combination thereof. In some embodiments, the neural network is trained with a training dataset comprising predicted geometrical structures of binding complexes of macromolecules. In some embodiments, the predicted geometrical structures of binding complexes comprise a plurality of monomers aligned to form the geometric structures.

[0084] In some embodiments, the neural network is trained using a loss function configured to reduce achiral and steric clashes, wherein the loss function is adjusted with a weighting function that stabilizes the training. In some embodiments, the weighting function increases in value as the number of training iterations increases. In some embodiments, the neural network is trained by distributing batches of training data across a plurality of processors based on sizes of the training data to reduce load imbalance between the plurality of processors. In some embodiments, the training data comprises at least 100k training samples of geometrical structures of binding complexes. In some embodiments, the neural network comprises at least 100M parameters. In some embodiments, the neural network is further trained with a second training dataset comprising geometrical structures of macromolecules or binding complexes generated from experiments, simulations, or electronic structure calculations. In some embodiments, the geometrical structures generated from experiments comprise cryo-EM generated structures, X-ray crystallography generated structures, and / or NMR generated structures. In some embodiments, the geometrical structures generated from simulations comprise molecular dynamics simulation or Monte Carlo simulation generated structures. In some embodiments, the geometrical structures are generated with homology modeling. In some embodiments, the geometrical structures are generated without homology modeling. In some embodiments, the geometrical structures generated from electronic structure calculations comprise DFT generated structures.

[0085] In some embodiments, a computational cost of processing the initial geometrical structure to generate the geometrical structure of the binding complex scales sub-linearly as a function of the number of amino acid residues. In some embodiments, the neural network is configured to provide a mean TM-score accuracy of 0.9 for the geometrical structure of the binding complex. In some embodiments, a ligand root-mean-square deviation of the geometrical structure is below two Angstrom. In some embodiments, the neural network uses site-specific docking.

[0086] In some embodiments, the method further comprises generating a confidence score for each amino acid in the geometrical structure. In some embodiments, the confidence score is above a threshold confidence value. In some embodiments, the threshold confidence value is 0.8.Attorney Docket No.59215-726.601

[0087] In some embodiments, the method further comprises processing the geometrical structure of the binding complex using molecular mechanics. In some embodiments, the molecular mechanics comprises molecular dynamics. In some embodiments, the molecular mechanics comprises electronic structure calculations. In some embodiments, the molecular mechanics is performed with restraints or without restraints. In some embodiments, the restraints are imposed on covalent bonds formed with a hydrogen atom. In some embodiments, the restraints are imposed on torsional modes of the binding complex.

[0088] In some embodiments, the method further comprises predicting experimental binding affinity between at least two molecules, wherein the at least two molecules comprises (a) at least two macromolecules, or (b) at least one macromolecule and at least one small molecule. In some embodiments, the input representation comprises a sequence of a macromolecule. In some embodiments, the sequence is an amino acid sequence of a protein. In some embodiments, the sequence is a nucleic acid sequence of a nucleic acid. In some embodiments, the input representation comprises a 3-dimensional structure of a molecule. In some embodiments, the 3-dimensional structure is of a protein, nucleic acid, or a small molecule. In some embodiments, the input representation comprises a sequence and a 3-dimensional structure of one or more macromolecules.

[0089] In some aspects, the present disclosure provides a method of computing a tensor operation between tensors of different ranks, comprising: (a) extracting a first slice of a first tensor and a second slice of a second tensor of a plurality of tensors; (b) broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof; (c) performing a tensor operation between the first slice and the second slice; and (d) repeating (a)-(c) to perform a plurality of tensor operations with other slices of the first tensor and the second tensor to compute the tensor operation between the first tensor and the second tensor.

[0090] In some embodiments, the method further comprises obtaining gradients of a result of the tensor operation with respect to the first slice and the second slice. In some embodiments, the method further comprises obtaining gradients of a result of the tensor operation with respect to the first tensor and the second tensor.

[0091] In some embodiments, the method further comprises storing the first tensor, the second tensor, or both in a first memory. In some embodiments, the method further comprises storing the first slice, the second slice, or both in a second memory.

[0092] In some embodiments, the second memory is faster than the first memory. In some embodiments, the second memory has smaller size than the first memory. In some embodiments, the second memory is more physically localized to transistors of a processing unit than the first memory.

[0093] In some embodiments, the processing unit is a graphical processing unit (GPU). In some embodiments, the method is performed on a graphical processing unit (GPU).

[0094] In some embodiments, the first memory has a speed of at least 1.5 TB / s. In some embodiments, the first memory has a size of at least 40 GB. In some embodiments, the first memoryAttorney Docket No.59215-726.601 is a high bandwidth memory (HBM). In some embodiments, the second memory has a speed of at least 19 TB / s. In some embodiments, the second memory has a size of at least 20 MB. In some embodiments, the second memory is a static random access memory (SRAM).

[0095] In some aspects, the present disclosure provides a method of training a neural network by: (a) computing tensor operations using the method to obtain a gradient; and (b) updating a parameter of the neural network based on the gradient.

[0096] In some aspects, the present disclosure provides a method of performing inference using a neural network by: (a) computing tensor operations to obtain a result; and (b) outputting a prediction based on the result.

[0097] In some embodiments, the method further comprises predicting or learning to predict a geometrical structure of a molecule using the neural network. In some embodiments, the method further comprises predicting or learning to predict a geometrical structure of a binding complex formed between a plurality of molecules using the neural network.

[0098] In some embodiment, the molecule or the plurality of molecules comprises small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof.

[0099] In some embodiments, the tensor operation comprises addition, subtraction, inner product, dot product, outer product, cross product, matrix multiplication, Cartesian product, Hadamard product, Kronecker product, Cracovian product, Frobenius inner product, Khatri-Rao product, face- splitting product, or any combination thereof.

[0100] In some embodiments, the plurality of tensors comprise at least 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 tensors. In some embodiments, the plurality of tensors comprise at most 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 tensors. In some embodiments, the first tensor has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second tensor has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the first slice has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second slice has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the first tensor has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second tensor has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the first slice has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second slice has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second tensor is greater than the rank of the first tensor by at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second slice is greater than the rank of the firstAttorney Docket No.59215-726.601 slice by at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second tensor is greater than the rank of the first tensor by at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second slice is greater than the rank of the first slice by at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the broadcasting the first slice matches the shape of the second slice or the transpose thereof. In some embodiments, the first tensor is a bias tensor in an attention mechanism, and the second tensor is a query tensor, a key tensor, or a value tensor in the attention mechanism.

[0101] In some aspects, the present disclosure provides a method of computing a tensor operation between tensors of different ranks, comprising: (a) storing a plurality of tensors comprising a first tensor and a second tensor in a first memory of a processing unit; and (b) computing a tensor operation between slices of the first tensor and the second tensor by: (i) dynamically allocating the slices of the tensors by: (1) extracting a first slice of a first tensor and a second slice of a second tensor; (2) broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof; and (3) storing the first slice and the second slice in a second memory of the processing unit, wherein the second memory is faster than the first memory, and wherein the second memory is more physically localized to transistors of a processing unit than the first memory; (ii) performing the tensor operation between the first slice and the second slice; and (c) repeating (i) and (ii) with other slices of the first tensor and the second tensor to complete the tensor operation between the first tensor and the second tensor.

[0102] In some aspects, the present disclosure provides a method of computing a triangular attention tensor, comprising: (a) extracting a plurality of slices from a plurality of tensors, wherein the plurality of slices comprises a first slice of a bias tensor of the plurality of tensors, a second slice of a query tensor of the plurality of tensors, a third slice of a key tensor of the plurality of tensors, and a fourth slice of a value tensor of the plurality of tensors; (b) broadcasting the first slice to match a rank and a shape of the second slice, the third slice, the fourth slice, or a transpose of the second slice, the third slice, or the fourth slice; (c) performing tensor operations using the plurality of slices; and (d) repeating (a)-(c) with other slices of the plurality of tensors to compute the triangular attention tensor.

[0103] In some embodiments, a token dimension of the triangular attention tensor is at least 128, 256, 384, 512, 768, 1024, or 1280. In some embodiments, a token dimension of the triangular attention tensor is at most 128, 256, 384, 512, 768, 1024, 1280, or 2560. In some embodiments, the token dimension represents atoms in a molecule, amino acids in a protein or peptide, a nucleotide in a nucleic acid, or any combination thereof.

[0104] In some aspects, the present disclosure provides a method for generating a geometrical structure of a molecule, comprising: (a) processing a representation of a molecule to obtain anAttorney Docket No.59215-726.601 attention tensor from particles of the molecule; and (b) updating the representation based on the attention tensor to obtain the geometrical structure of the molecule.

[0105] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of molecules, comprising: (a) processing a representation of a plurality of molecules to obtain an attention tensor from particles of the plurality of molecules; and (b) updating the representation based on the attention tensor to obtain the geometrical structure of the binding complex.

[0106] In some aspects, the present disclosure provides an active programming interface comprising a callable function configured to program a processor to perform operations comprising: (a) extracting a first slice of a first tensor and a second slice of a second tensor; (b) broadcasting the first slice to match the rank and at least one dimension of the second slice or a transpose thereof; (c) performing the tensor operation between the first slice and the second slice; and (d) repeating (a)-(c) with different slices to compute the tensor operation between the first tensor and the second tensor.

[0107] In some aspects, the present disclosure provides a method of computing a tensor operation between tensors of different ranks by dynamically allocating slices of the tensors and computing the tensor operation between the slices.

[0108] In some aspects, the present disclosure provides a method for generating a geometrical structure of a molecule, comprising: (a) processing a geometry prior of the molecule to obtain a prediction of the geometrical structure; (b) processing the prediction and the geometry prior to obtain a vector field; and (c) generating the geometrical structure of the molecule based on the vector field.

[0109] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of molecules, comprising: (a) processing a geometry prior of a plurality of molecules to obtain a prediction of the geometrical structure of the binding complex; (b) processing the prediction and the geometry prior to obtain a vector field; and (c) generating the geometrical structure of the binding complex based on the vector field.

[0110] In some embodiments, the vector field is time independent or time dependent.

[0111] In some embodiments, the molecule or the plurality of molecules comprises small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof.

[0112] In some embodiments, the generating the geometrical structure is performed while maintaining optimal transport. In some embodiments, the generating the geometrical structure is performed while maintaining SE(3) superposition. In some embodiments, the SE(3) superposition is maintained by applying a Kabsch transform to the geometrical structure. In some embodiments, the generating the geometrical structure is performed while optimizing transport cost. In some embodiments, the generating the geometrical structure is performed while maintaining identities of the particles. In some embodiments, the identities of the particles are maintained by obtaining a graph permutation of a representation of the molecule or the plurality of molecules.Attorney Docket No.59215-726.601

[0113] In some embodiments, the method further comprises generating the geometry prior. In some embodiments, the geometry prior is physics-based.

[0114] In some aspects, the present disclosure provides a method of generating a geometrical structure of a molecule, comprising: (a) generating a current geometrical structure of the molecule based on a conditioning signal; (b) generating a latent variable by encoding the conditioning signal; (c) updating the current geometrical structure of the molecule by: (i) obtaining a predicted geometrical structure of the molecule based on the latent variable, a time variable, and the current geometrical structure; (ii) applying a global SE(3) superposition to the prediction by aligning the predicted geometrical structure to the current geometrical structure, and wherein the aligning is performed by applying a rotation to the predicted geometrical structure; (iii) obtaining an optimal graph permutation of the predicted geometrical structure to find a consistent ordering of nodes between the predicted geometrical structure and the current geometrical structure; and (iv) updating the current geometrical structure by interpolating between the current geometrical structure and the predicted geometrical structure; (d) repeating (c) until an output geometrical structure is obtained.

[0115] In some aspects, the present disclosure provides a computer system comprising: a memory comprising executable instructions; and at least one processor configured to execute the instructions, wherein when the at least one processor executes the instructions, the at least one processor causes the system to generate a geometrical structure of a binding complex formed between a plurality of molecules using a flow matching generative model.

[0116] In some embodiments, the flow matching generative model is configured to generate the geometrical structure of the binding complex formed between small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof.

[0117] In some aspects, the present disclosure provides a computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computer- executable code adapted to be executed to implement any one of the methods or the neural networks disclosed herein.

[0118] In some aspects, the present disclosure provides a non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to implement any one of the methods or the neural networks disclosed herein.

[0119] In some aspects, the present disclosure provides a computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to perform any one of the methods or the neural networks disclosed herein.Attorney Docket No.59215-726.601 INCORPORATION BY REFERENCE

[0120] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0122] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:

[0123] FIGS. 1A-1F show illustrations and schematics for systems and methods a neural framework for protein-ligand complex structure prediction. In some embodiments, the framework can perform protein-ligand complex structure prediction with full receptor flexibility. FIG.1A shows a high-level schematic of the framework, in accordance with some embodiments. FIG. 1B shows a forward and reverse process of the framework, in accordance with some embodiments. A denoising diffusion process can generate the binding complex 3D atomistic structure. The top arrow shows a method for sampling from a reverse stochastic process, in accordance with some embodiments. The bottom arrow shows a forward stochastic process for training the framework, in accordance with some embodiments. The protein (colored as red-blue from N- to C-terminus) and ligand (colored as grey) 3D structures can be jointly generated from a learned stochastic differential equation (SDE), with an initial state ^^^^∗. The initial state ^^^^∗can be either template-independent or approximated by the protein backbone template. The initial state ^^^^∗can comprise predicted interface contact maps. FIGS. 1C-1E show certain technical design elements of the framework, in accordance with some embodiments. FIG. 1C shows ligand molecules and amino acids encoded as the collection of atoms, local coordinate frames (depicted as semi-transparent triangles), and stereospecific pairwise embeddings (depicted as dashed lines) representing their interactions. FIG.1D shows a forward-time SDE that introduces relative drift terms among protein ^^^^atoms, non-^^^^atoms, and ligand atoms, such that the SDE erases data in acontrollable manner at ^^ = ^^∗ to be sampled from a noise distribution. FIGS. 1E-1F showinformation flow in the equivariant structure denoiser (ESD), in accordance with some embodiments. ESD can operate on a heterogeneous graph formed by protein atoms (P), ligand atoms (L), proteinAttorney Docket No.59215-726.601 backbone frames (B) and ligand local frames (F) to predict clean atomic coordinates ^̂^^^, ^̂^^^using thecoordinates at a finite diffusion time ^^ > 0.

[0124] FIGS. 2A-2H show model performance of certain embodiments of the framework on benchmarking problems. FIGS. 2A-2D show performance of fixed-backbone blind protein-ligand docking. FIG.2A shows success rates over the test dataset plotted against the number of conformations sampled per protein-ligand pair; a success was defined as the ligand root-mean squared difference (RMSD) being lower than given threshold for at least one of the sampled conformations. FIG. 2B shows distributions of the physical plausibility of sampled conformations measured by the ligand heavy-atom steric clash rate with receptor atoms. FIG. 2C shows distributions of the geometrical accuracy measured by the ligand RMSD plotted against the number of ligand rotatable bonds, an indicator of molecular flexibility. FIG. 2D shows an overlay of the framework-predicted ligand and side-chain structures on the ground-truth for a challenging example, Protein DataBank ID 6MJQ (PDB: 6MJQ). FIGS. 2E-2G show results from a ligand-coupled binding site repacking via diffusion-based inpainting experiment. FIG. 2E shows a selected example (PDB:6TEL) where the framework accurately inpainted the binding site protein-ligand structure, while directly aligning AlphaFold2TM(AF2TM) prediction to the ground-truth complex resulted in steric clashes between the ligand and binding site residues. FIG.2F shows binding site accuracies (measured by the all-atom local distance different test score, “lDDT-BS”) and ligand clash rates over the test dataset. 32 conformations were sampled for each protein-ligand pair. The dots indicate the median value and the error bars indicate 25% and 75% percentiles. FIG. 2G shows success rates compared to reference methods. A success was defined as: lDDT-BS > 0.7, ligand RMSD < 2.0 Å, and clash rate = 0.0. The pink "true contact map" curves were obtained by initializing the geometry prior ^^^^∗using the true protein-ligand contact map, while the gold curves were obtained by generating both protein and ligand conformations end- to-end. FIG.2H shows evaluation of the framework compared to reference methods in blind docking experiments. The framework incorporated protein residue embeddings from a protein language model and embeddings from a pretrained small molecule encoder. Using triangular attention and Swin- Transformer tokenization scheme, substantial improvement in docking atomic accuracy & binding site prediction accuracy were observed for the framework, in contrast to the reference methods. The left panel and the right panel respectively show ligand docking accuracy as measured by RMSD below 2 Angstroms and 5 Angstroms. The framework only needed to sample 1 structure per protein-ligand pair for the experiments. The advantage of the framework may be attributed to the residual scale embeddings acquired from the contact predictor.

[0125] FIGS. 3A-3C show framework assessments on systems with large binding-induced protein conformational transitions. Apo protein structures were used as the input backbone template. FIG.3A shows statistics of the relative protein folding similarity with respect to apoprotein (“apo”) and holoprotein (“holo”) PDB structures (measured by ΔTM-Score, the difference between TM-ScoresAttorney Docket No.59215-726.601 computed against holo and apo structures) and binding site similarity with respect to holo (measured by lDDT-BS) for sampled structures. Purple dots were obtained with protein-only inputs and gold dots were obtained using protein+ligand inputs. Ligand-conditioning increased average ΔTM-Score from - 9.0% to -7.7% (p=0.03), and average lDDT-BS from 0.59 to 0.63 (p<0.001). FIGS. 3B-3C show two examples for which neither their holo nor apo reference structures were observed during training. A marginal improvement in ΔTM-Score or lDDT-BS may indicate substantial protein conformational differences, while the framework can qualitatively capture the correct protein state transitions.

[0126] FIG. 4 shows a neural network architecture for at least a portion of the contact prediction network, in accordance with some embodiments.

[0127] FIG.5 shows a neural network architecture for at least a portion of the contact predictor (CP) block, in accordance with some embodiments. Arrows indicate information flow directions, and "+" indicates an element-wise tensor summation. kNN denotes the neighbor size of local residue-residue edges.

[0128] FIGS. 6A-6B illustrate information flow in contact prediction, in accordance with some embodiments. Contact prediction can sample adjacency matrices among protein and ligand nodes using an autoregressive decoding scheme, where the adjacency matrices sampled on the last step (^^^^) are passed to the network to update the predicted histograms of pairwise distances ("distograms") and contact maps ^̂^. The protein-ligand graph can be tokenized into N patches. The distance distributions can be inferred from hybrid graph & dense matrix representations.

[0129] FIG. 7 shows a network architecture of a single block in the equivariant structure denoiser (ESD), in accordance with some embodiments. Arrows indicate information flow directions, and "+" indicates an element-wise tensor summation.

[0130] FIG.8 shows a computer system, in accordance with some embodiments.

[0131] FIG.9 shows some prediction problems that the framework can be used to solve, in accordance with some embodiments. The framework can be used in various ligand structure prediction problems, such as solving blind docking problems between a rigid-receptor and a ligand. The framework can incorporate an assumed ground-truth protein structure. The framework can operate without incorporating prior knowledge of the binding site of the protein or the ligand. The framework can be used in binding site structure prediction. The framework can incorporate knowledge of a known binding site location. The framework can predict the structure in a cropped portions of the protein along with the ligand structure, while accounting for binding site “pocket” flexibility. The framework can be used for ligand-guided protein structure refinement. The framework can incorporate an assumed ground-truth ligand geometry, aligned to an apo protein template. The framework can incorporate domain packing information by truncating the forward-time SDE at T* < 1.0. The framework can perform diffusion purification from the perturbed template by performing T* → 0.Attorney Docket No.59215-726.601

[0132] FIG. 10 shows results for a temperature-adjusted diffusion model, in accordance with some embodiments. Langevin-simulated annealing steps were mixed into the probability flow of ordinary differential equations. The steps were numerically stable, and there was no prior mismatching issue at t=T, which can occur in naïve score norm scaling.

[0133] FIG. 11 shows results for a temperature-adjusted diffusion model compared to reference methods, in accordance with some embodiments. State-selective protein structure prediction was performed using the PocketMiner Dataset. Performance was improved using the Langevin simulated annealing compared to without, as measured by the TM-Score.

[0134] FIGS.12A-12G show the framework performance on benchmarking problems. FIGS.12A-D show results of rigid backbone blind protein-ligand docking experiments. The cumulative fraction of predictions with ligand RMSD below 2 Å (FIG. 12A) and 5 Å (FIG. 12B) over the test dataset are plotted against the number of ligand poses sampled per protein-ligand pair. FIG. 12C show fidelity assessments of model-assigned confidence estimations. The ligand RMSDs are plotted against the model-assigned pLDDT score averaged over all ligand atoms (ligand pLDDT). FIG. 12D shows a precision-recall curve evaluation, based on the sampled structures ranked by the ligand pLDDT score, with the binary value of ligand RMSD < 2 Å being treated as the class label. FIGS. 12E-G show flexible binding site structure recovery. FIG. 12E shows visualization of prediction results on a test set example (PDB:6P8Y) near the structure recovery accuracy cutoff. The framework generates the binding site protein-ligand structure consistent with experimental data, while directly aligning the AF2TMprediction to ground-truth complex results in steric clashes (dashed circle). The red arrow indicates the qualitative conformational change from AF2TMto the experimental bound-state structure. FIG. 12F summarizes binding site accuracy (measured by the lDDT-BS score) and ligand clash rate over the test dataset. 32 conformations were sampled for each protein-ligand pair; dots indicate the median value and error bars indicate 25th and 75th percentiles. FIG. 12G shows structure recovery accuracy compared to baseline methods. Solid lines correspond to recovery rates evaluated based on strict cutoffs (lDDT-BS > 0.7, ligand RMSD < 2 Å, clash rate = 0.0), while the dashed lines correspond to relaxed cutoffs (lDDT-BS > 0.5, ligand RMSD < 2.25 Å clash rate < 0.05, with clash rate cutoff matching the 95% percentile of experimental structure statistics.

[0135] FIGS. 13A-13H shows framework predictions for 33 contrasting apo-holo pairs from the PocketMiner dataset. FIGS. 13A-C show TM-scores with respect to apo and holo experimental reference structures plotted for AF2TMstructures, framework-sampled apo structures, and framework- sampled holo structures, respectively. For framework-sampled structures, the model-assigned pLDDT scores averaged among all protein residues are indicated by the dot colors. FIGS. 13D-E show TM- scores for reference methods and the framework averaged against all samples and the subset of targets for which any framework-predicted structure is of protein pLDDT=0.8 or higher. Error bars indicate the standard error of the mean calculated from six sets of predictions using independent random seeds and different AF2TMmodel checkpoints to obtain the initial template. FIG. 13F shows the weightedAttorney Docket No.59215-726.601 Q-factor metric for conformational change prediction accuracy plotted between framework predictions and AF2TMpredictions on all ligand-bound holo structures. Grey dots correspond to framework predictions in the absence of ligand inputs. FIG. 13G shows the weighted Q-factors for predicted ligand-bound holo structures plotted against the maximum sequence similarity of each target to samples in the training dataset. FIG.13H shows fidelity assessment of framework internal confidence predictions. The ligand-binding-site accuracy was measured by lDDT-BS, plotted against the ligand RMSD for all framework predictions; the model-assigned pLDDT scores averaged among all ligand atoms are indicated by the dot colors.

[0136] FIGS.14A-14G show framework predictions on structures that were recently determined with confirmed ligand-induced conformational changes of backbone with RMSD > 2 Å. FIGS. 14A-14B show TM-scores for reference methods and the framework averaged against all samples and the subset of targets for which any framework-predicted structure is of protein pLDDT=0.8 or higher. Error bars indicate the standard error of the mean calculated from six sets of predictions using independent random seeds and different AF2TMmodel checkpoints to obtain the initial template. FIG. 14C show the weighted Q-factor metric for conformational change prediction accuracy plotted between framework predictions and AF2TMpredictions on all ligand-bound holo structures. Grey dots correspond to framework predictions in the absence of ligand inputs. FIG.14D show the weighted Q- factors for predicted ligand-bound holo structures plotted against the maximum sequence similarity of each target to samples in the training dataset. FIG.14E show fidelity assessment of framework internal confidence predictions. The ligand-binding-site accuracy as measured by lDDT-BS is plotted against the ligand RMSD for all framework predictions; the model-assigned pLDDT scores averaged among all ligand atoms are indicated by the dot colors. FIGS. 14F-14G show framework predictions on representative targets (PDB:6KPE, PDB:6PKH, PDB:6WQA) suggesting structural elements for protein self-assembly and plausible models for enzyme catalysis and the activation of membrane receptors.

[0137] FIGS. 15A-15B shows benchmark results for direct structure predictions of protein-ligand complexes inputting the protein sequence and ligand molecular graph into the framework.2652 kinase structures, composed of about 30% apo structures and 70% holo structures were used. FIG.15A shows data confirming that SE(3)-invariant denoising loss function, which is a modified version of the frame aligned point error (FAPE) loss with broken chiral symmetry, gave rise to better local accuracy as indicated by the lDDT-BS score. FIG.15B shows that incorporating ligand information improves the accuracy of the prediction. Further, conditioning the generation process with ligand information improves the local accuracy, as measured by lDDT-BS. In addition, the conditional generation with ligand information gave rise to improved global protein folding similarity measurement as measured by the TM-score.Attorney Docket No.59215-726.601

[0138] FIG. 16 shows models of E. Coli adenosine kinase (ADK) structures generated via unconditional generation and protein-ligand joint generation. Qualitative distinctions between sampled apo versus binding complex structures were noted.

[0139] FIG. 17 shows models of saccharopine dehydrogenase, 3UGK (apo) and 3UH1 (holo). The AlphaFold2TMmethod predicted only the apo structure. Meanwhile, the framework sampled both apo and holo states with high specificity (with TM-score > 0.95). The framework was able to distinguish between ligand bound and ligand free states, as shown in the plots on the right where the samples form distinct clusters.

[0140] FIGS. 18A-18B show relationships between the ability to resolve apo / holo structures and ligand structure modeling. The framework was tested on 118 structures with backbone RMSD greater than 2 Angstroms that were released after Jan. 2019. There was no overlap between these structures and the training data used to train the framework or AF2. The structures covered broad classes of ligand-dependent conformational changes, e.g., agonist, antagonist, etc. FIG. 18A shows testing accuracy as measured by TM-Score and RMSD on the structures resolved after Jan. 2019. FIG. 18B shows that predicted structure accuracy was correlated with the ability to resolve apo / holo structures , which suggests the framework’s ability to accurately identify binding sites. Ligand docking accuracy is suggested as a causal factor for performing accurate holo structure predictions as disclosed herein.

[0141] FIG.19 shows a schematic of an embodiment of the framework employing template structure retrieval / sampling. A template structure of a protein can be retrieved or sampled from a database or another model. The template structure can be analyzed to generate a contact map of the template structure. The template structure and / or the contact map can be used in the framework to generate protein-ligand complex structures.

[0142] FIG.20 shows results of template-conditioned structure generation experiments. The data was generated using 1238 structures of a benchmarking dataset. The dataset was temporally split: structures with deposit date before Apr. 30, 2018 were used as the training set; structures with deposit date between Apr. 30, 2018 and Jan. 1, 2019 were used as the validation set; and structures with deposit date after Jan.1, 2019 were used as the test set. The results showed that utilizing template embeddings can significantly improve global prediction accuracy (as measured by the TM-score) and local prediction accuracy (as measured by lDDT-BS).

[0143] FIGS.21A-21B show the validation loss against relative compute and interface LDDT for various models including the framework of the present disclosure.

[0144] FIG.22 show the percent root mean square deviation under 2 Angstrom for various models including AlphaFoldTM.

[0145] FIGS.23A-23B show various metrics for models. In particular, the framework, when compared to other models, shows an increased accuracy in prediction as well as prediction speed.

[0146] FIG.24 shows the receptor TM-score for the framework, indicating a large increase in accuracy from other models.Attorney Docket No.59215-726.601

[0147] FIGS.25A-25B shows the progression of linear regression of the framework with training iterations.

[0148] FIG.26 shows the accuracy of structure predictions by the framework as quantified against experimentally measured structures of binding complexes between macromolecules and ligands.

[0149] FIG.27 shows the performance of a flow matching framework on the PoseBustersTMbenchmark. The left panel shows the success rate for predicting the ligand-protein structures to within 2Å RMSD, with and without additionally requiring that the structures are physically reasonable (PB-valid). The center panel shows the percentage of structures where ligand stereochemistry is correct in the predicted structure. The right panel shows the timing for running a single inference on a 1024-residue protein. Comparisons are made to a diffusion-model framework and AlphaFold3TM.

[0150] FIG.28 shows improvements to the framework associated with various technical implementations. The left panel shows the performance on the PoseBustersTM%RMSD < 2Å metric increasing as improvements were made to the data, the model architecture, and the inference pipeline. The right panel shows reduced memory consumption and model inference speed in using an implementation of memory-efficient tensor operations.

[0151] FIG.29 shows the evaluation of the flow-matching framework on diverse categories of biomolecular interactions using recent structures from the Protein Databank. Representative structures are shown, each matching the median prediction accuracy for its category. Predicted structures are shown in darker shade and the experimental ground truth is shown in lighter shade. These structures were deposited after the training cutoff of the model, and have at most 40% homology to the training data. Protein-protein interaction (PPI) accuracy is reported using the DockQ metric and protein monomer accuracy is reported using the Local Distance Difference Test (LDDT).

[0152] FIG.30 shows implementation details of a memory-efficient tensor operation. The left panel shows the allocation of tensors in static random access memory (SRAM) and high bandwidth memory (HBM), which are different parts of a GPU memory. The right panel compares an implementation of a tensor operation (FlashAttention) with another (FlashTriangularAttention). DETAILED DESCRIPTION

[0153] In some aspects, provided herein are systems and methods for predicting structures of complexes formed between a protein and a ligand. Protein structures can be dynamically modulated by their interactions with small-molecule ligands or post-translational modifications that trigger downstream responses which can be crucial to the regulation of biological functions. Such dynamical protein conformational changes can occur on a vast variety of kinetic timescales (e.g., ranging from nanosecond-scale vibrations to millisecond-scale collective motions) which can be associated with induced-fit binding, pocket crypticity, and global folding changes resulting from activation- inactivation effects . Proposing ligands that selectively target protein conformations is an importantAttorney Docket No.59215-726.601 strategy for designing or discovering small-molecule-based therapeutics. However, solving computational problems for predicting protein-ligand structures is challenging when the problems are coupled to receptor conformational responses. The challenges can be attributed to (i) the prohibitive cost of simulating the physics of slow protein state transitions, and (ii) the static nature of existing protein folding prediction algorithms.

[0154] Some methods have advanced molecular modeling techniques to reduce simulation costs. Examples of those advances include enhanced sampling techniques, collective variable estimation, molecular docking guided by template-based modeling, iterative refinement protocols, and so on. However, these methods often require integrating expert knowledge and intervention or constraints from experimental data on a case-by-case basis. Case-by-case evaluation of modeling problems can hamper high-throughput evaluation of a plethora of target candidates and drug candidates. Meanwhile, existing deep-learning-based algorithms for ligand-binding proteins structure prediction may be limited to single structure regression-based formulations, especially for non-endogenous small molecule binding for which the ligand identity cannot be inferred from protein motifs. Recognized herein, is a need for a unified framework for predicting binding complex structures with high- throughput.

[0155] To create the unified framework, a framework for generative modeling was designed, which is capable of directly predicting binding complex structures at an atomistic resolution with an accuracy comparable to structure determination experiments. Predicting complex structures at various coarse- grained resolutions are possible too. Binding is driven at least partially by both the determination of the contextual information related to ligand function such as selectivity to orthosteric or allosteric sites, and the process of resolving energetically favorable inter-atomic structures based on sub-nanometer- scale physical interactions. The framework can be configured to jointly fold the 3D structure of the protein-ligand complex, and quantify the structural fluctuations. Thus, the framework was designed to incorporate biophysical inductive biases to accurately predict the binding complex structures.

[0156] The framework can operate using various details of input. In some embodiments, the framework can operate based on solely molecular graphs as ligand inputs, which can enable end-to- end gradient-based design for functional small-molecules and ligand-binding proteins when coupled to differentiable protein sequence and / or molecular graph generators. Incorporating of state-of-the-art protein representation learning techniques such as the use of sequence evolutionary signals, pretrained language models, higher-level attention mechanisms, and any combination thereof can further improves the methodology. In some embodiments, an initial (or a template) structure of the ligand, the protein, or both can be used.

[0157] Thus, in some aspects, the present disclosure provides a framework for generating a structure of a binding complex formed between a ligand and a protein. The framework can comprise predicting contacts between the ligand and the protein in the binding complex. The contacts can define which portions of the ligand are proximal to which portions of the protein. Predicting theAttorney Docket No.59215-726.601 contacts can be based on neural network processing of the sequence of the protein, an embedding of the protein, a graph of the ligand, or any combination thereof. The contacts can be, e.g., pairwise proximities among residues of the protein and atoms or frames of the ligand. The framework can comprise generating a geometry prior based on the contacts. The predicted contacts can be used as a basis for generating a probability distribution of potential geometrical structures of the binding complex. The framework can comprise denoising the geometry prior to generate a structure of the binding complex. The denoising can be performed, e.g., using a diffusion model, to extract high- likelihood geometrical structures from the geometry prior. The denoising can be performed with SE(3) equivariance, reducing temperature of the chemical system, or both. The resulting geometrical structure of the binding complex can be used to generate a report of the structure.

[0158] Accordingly, the present disclosure describes a framework for predicting the three-dimensional (3D) structures of proteins (e.g., druggable targets), ligands (e.g., small molecule drugs), and complexes formed between them. The framework can integrate multi-scale inductive bias in biomolecular complexes with diffusion models by designing a stochastic differential equation (SDE) with structured drift terms. The framework can leverage diffusion-based generative modeling to sample 3D structures from a statistical distribution learned from training data. Residue-level contact maps between a protein and a ligand can be sampled, and in some cases, iteratively in a hierarchical manner. The framework can generate an ensemble of structures of binding complexes given a protein sequence and ligand molecular graphs as inputs. The framework can be conditioned on features obtained from protein language models (PLMs) and / or template protein structures retrieved from experimentally-resolved homologs or computational models. The framework can be conditioned with evolutionary constraints, e.g., extracted from multiple sequence alignments (MSA) or PLMs. Both the prediction pipeline and the underlying neural network architecture can be designed to closely resemble the multi-scale hierarchical organization of biomolecular complexes.

[0159] The framework can comprise a neural network. The neural network can comprise a graph neural network configured to encode the atomic-scale chemical and / or geometrical features of individual small-molecules and amino-acid graphs into tensor representations. The graph neural network can be configured to learn long-range inter-atomic geometrical correlations for molecules based on their graph-topological properties. The neural network can comprise a contact predictor to generate residue-scale inter-molecular distance distributions, coarse-grained contact maps, and / or associated pair representations. The neural network can comprise an equivariant structure denoiser to generate the binding complex structure conditioned on outputs from the atomic-scale and residue-scale networks. The equivariant structure denoiser can perform a structured denoising diffusion process that is equivariant and preserves chirality constraints of the protein and ligand molecules. The neural network can be trained on millions of samples from molecular conformation and bioactivity databases. The neural network can comprise an attention-based network. The neural network can comprise a vision-language model.Attorney Docket No.59215-726.601

[0160] The framework can generalize to ligand-unbound or predicted protein structure inputs even when the neural network is trained solely on experimental protein-ligand complex structures that are not paired to alternative protein conformations. When applied to blind protein-ligand docking, results show that the framework can improve the ligand pose accuracy by up to 78% compared to the best performing reference method on the PDBBind2020 dataset. When applied to ligand binding site design, results show that the framework can perform inpainting to accurately recover up to 46% of failed AF2TMbinding sites with up to 59% success rate improvements compared to the method of RosettaTM. Results show that the framework can improve the ligand pose geometrical accuracy by up to 61% compared to certain reference methods (e.g., DiffDockTM). Results show that the framework outperform a reference protein structure prediction algorithm, AF2TMas demonstrated by the highest TM-score (0.906 on average), and by 11-13% improvements in terms of accuracy on domains that undergo substantial conformational changes upon ligand binding. The learning-based method for dynamic-backbone protein-ligand structure prediction can establish an accuracy and sampling efficiency advantage relative to reference baseline approaches. The framework can be used to predict a regulatory effect of a ligand on protein folding. The results show that the framework has advantages on selectively predicting protein structures that are subject to induced-fit binding or conformational selection.

[0161] In some aspects, the present disclosure provides a graph neural network for extracting chemical, physical, and / or geometrical features from molecular graph representations. The molecular graph representations can be for one or more input ligands and / or amino acid residues (standard and / or modified) of one or more input proteins. The molecular graph representations can contain the molecular topology and stereochemistry specifications. The graph neural network can be configured to learn long-range inter-atomic geometrical correlations for molecules based on their graph- topological properties. Some embodiments of the graph-based neural network can be referred to as a path-integral graph transformer.

[0162] In some aspects, the present disclosure provides a neural network for inferring the interactions and / or distance distributions among: (i) one or more input proteins and, optionally, their amino acid residues, and (ii) one or more ligand graphs and, optionally, their local coordinate frames and / or atoms. The neural network may use features from the graph-based neural network (e.g., path-integral graph transformer) and / or features obtained from (including but not limited to) protein sequence embeddings from Transformer-based protein language models (e.g., ESM-2), multiple sequence alignments (MSAs), and templates from sequence homologs and / or structure prediction networks (e.g. AF2TM, OpenFold, ESMFold, RosettaFold, TrRosetta, RFdiffusion). In some embodiments, the neural network can be implemented via a geometric vector perceptron (GVP)-like 3D graph neural network. In some embodiments, the neural network can be implemented without the features obtained from protein sequence embeddings from Transformer-based protein language models, MSAs, and templates from sequence homologs and structure prediction networks. In some embodiments, the neural network canAttorney Docket No.59215-726.601 be implemented with the features obtained from protein sequence embeddings from Transformer-based protein language models, MSAs, and templates from sequence homologs and structure prediction networks. In some embodiments, the neural network can comprise a patch-wise tokenizer of the proteins, a cross-attention block, a triangular attention block, a graph transformer block, or any combination thereof. In some embodiments, the neural network can be referred to as a contact predictor. In some embodiments, the contact predictor can be referred to as BindingFormerTM.

[0163] In some aspects, the present disclosure provides a neural network for generating the 3D structure of (a) the binding complex formed between a protein (or a protein complex) and a ligand (or multiple ligands), or (b) a ligand-free protein (or a protein complex). In some embodiments, the neural network is a denoising neural network. In some embodiments, the neural network may use features obtained from the graph-based neural network (e.g., path-integral graph transformer) and / or the neural network for inferring the interactions and / or distance distributions (e.g., contact predictor). In some embodiments, a confidence interval score can be predicted for the ligand, the protein, or both. The confidence interval score can be assigned to atoms of the ligand, the protein, or both. In some embodiments, the denoising neural network can be referred to as an equivariant structure denoiser.

[0164] Determining the structures and functional conformational changes of protein-ligand complexes at the proteome scale is a grand-challenge problem. As described herein, in some aspects, the present disclosure provides a framework having a generative deep learning approach that systematically incorporates small molecule information and biophysical inductive bias, thereby achieving improvement structure prediction capabilities compared to state-of-the-art molecular docking and protein structure prediction methods. The framework is capable of generating structures within seconds on a standard GPU. The framework can allow faster exploration of protein-ligand interactions in the vast space of sequence and chemical spaces by leveraging rapidly evolving databases of computationally modeled protein structures and chemical substances.

[0165] The framework can be differentiable end-to-end, and thus can be used for applications related to ligand and protein design. A closed-loop combination of the framework and recent advances in differentiable generative models for protein sequences, molecular graphs, and structure-based bioactivity models can accelerate the process of sequence and compound prioritization. The framework can be directly applied to accelerate various physical simulation studies on protein-ligand interactions, such as guiding the design and optimization of collective variables in enhanced-sampling molecular dynamics simulations.

[0166] As a data-driven approach, the framework is both generalizable and amenable to continuous improvement through the integration of better experimental and bioinformatic data. The framework can be applied to protein families with no experimentally determined homologs. The framework can be applied to proteins with post-translational modifications and multi-state large heteromeric protein complexes. Moreover, refining the framework on data such as high-resolution nuclear magneticAttorney Docket No.59215-726.601 resonance (NMR) and molecular dynamics data can enable it to capture protein structure under physiological conditions beyond the distribution of crystal-like structures.

[0167] The framework can be coupled with auxiliary data related to protein-compound interactions such as binding affinities and high-throughput mass spectrometry signals, and features from recent multi-modal biochemical language models, given that the incorporation of large-scale protein sequencing data in the form of MSAs has already been demonstrated a crucial component to achieve transferable protein structure prediction. The framework can be used as a general computational framework for rapid and accurate protein-ligand complex structure prediction to facilitate structure biology, drug discovery, and protein engineering. Neural Framework for Binding Complex Structure Prediction

[0168] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a protein and a ligand.

[0169] The method can comprise processing a first representation comprising a protein representation. The method can comprise processing a second representation comprising a ligand representation. The first representation can comprise a representation of one protein. The first representation can comprise a protein complex representation of a plurality of proteins. The second representation can comprise a representation of one ligand. The second representation can comprise a plurality of ligand representations.

[0170] The method can comprise generating a geometry prior. The geometry prior can refer to a set of assumptions or constraints based on an understanding of geometric relationships between atoms and molecules. A geometry prior can be incorporated into a learning process for a machine learning model, and it can be incorporated into an inference process for a machine learning model. Using a geometry prior can constrain the model to learning and outputting shapes and patterns based on geometric principles like symmetry, scale invariance, or topological constraints.

[0171] In some embodiments, the geometry prior can comprise noise structured on a template geometrical structure. The geometry prior can comprise predicted contacts between the protein and the ligand. The contacts can be generated by sampling binding interface spatial proximity distributions of the binding complex. The geometry prior can comprise an initial geometry of the protein. The geometry prior can comprise an initial geometry of the protein in the binding complex. The geometry prior can comprise an initial geometry of the ligand. The geometry prior can comprise an initial geometry of the ligand in the binding complex. The template geometrical structure can comprise an experimentally determined geometrical structure. The experimentally determined geometrical structure can be determined using X-ray crystallography, nuclear magnetic resonance spectroscopy, or cryo-electron microscopy. The experimentally determined geometrical structure can be retrieved from a protein structure database, e.g., Protein Data Bank. The template geometrical structure can comprise a computationally determined geometrical structure. The template geometrical structure can beAttorney Docket No.59215-726.601 determined using an electronic structure calculation, a molecular dynamics simulation, a Monte Carlo simulation, a machine learning model, or any combination thereof. The template geometrical structure can comprise coordinates of the backbone atoms of the protein. The geometry prior can comprise a finite-time marginal of a SDE configured to inject structured noise into the data distribution.

[0172] In some aspects, the present disclosure provides a system for implementing the method for generating a geometrical structure of a binding complex formed between a protein and a ligand. Equivariant Structure Denoiser

[0173] In some aspects, the present disclosure provides a method for denoising a geometrical structure of a binding complex formed between a protein and a ligand.

[0174] The method can comprise sampling an initial geometrical structure of the binding complex from the geometry prior. The initial geometrical structure can be a noisy structure. The method can comprise denoising, using a machine-learned stochastic differential equation (SDE), the initial geometrical structure to generate the geometrical structure of the binding complex.

[0175] The sampling the initial geometrical structure of the binding complex from the geometry prior can be based on a temperature parameter. The temperature parameter can be an inverse temperature parameter. The inverse temperature parameter can increase during the sampling process. The denoising the initial geometrical structure to generate the geometrical structure of the binding complex can be based on an inverse temperature parameter. The inverse temperature parameter can increase during the denoising process. The denoising can comprise annealing the initial geometrical structure.

[0176] The method can comprise predicting the geometrical structure of the binding complex based on the geometry prior. Predicting can comprise generating the geometrical structure of the binding complex based on the geometry prior. Predicting can comprise fixing the geometrical structure of the protein to the template geometrical structure of the protein. Predicting can be based on a template geometrical structure that comprises coordinates of the atoms of the protein that are distal from the ligand in the binding complex, wherein an atom is distal if it is further than a predetermined distance from the ligand. Predicting can comprise changing the geometrical structure of the protein from the template geometrical structure of the protein. Predicting can comprise fixing the geometrical structure of the ligand to a template geometrical structure of the ligand. Predicting can comprise changing the geometrical structure of the ligand from a template geometrical structure of the ligand.

[0177] The binding complex can comprise various proteins. For example, the binding complex can comprise an apoprotein or a holoprotein. The binding complex can comprise a protein comprising an allosteric site. The method can determine the geometrical structure the binding complex taking into account allosteric mechanisms that affect protein folding.

[0178] The method can comprise generating a second ligand, different from the first ligand, that is configured to bind to the protein. Generating the second ligand can be based on the structure of theAttorney Docket No.59215-726.601 first ligand. Generating the second ligand can comprise performing a gradient-based design using a differentiable protein sequence and / or a molecular graph generator.

[0179] The method can comprise performing an E(3)-equivariant forward-time-noising process. The E(3)-equivariant forward-time-noising process can be of a truncated stochastic differential equation(SDE) to ^^ = ^^∗ < ∞ based on a contact map to generate a geometry prior. The geometry prior cancomprise a partially-diffused structured distribution of the geometrical structure of the binding complex. The SDE can be configured to: (i) retain information about domain packing of the protein, (ii), retain information about ligand binding interfaces of the protein; and / or (iii) erase information about residue-scale local details of the protein.

[0180] The SDE can be configured to retain information about domain packing of the protein in the forward-noising process by diffusing the protein’s backbone atoms around a template protein backbone structure. The SDE can be configured to retain information about ligand binding interfaces of the protein in the forward-noising process by diffusing the ligand atoms around the protein’s residues that are predicted to contact the ligand atoms. The diffusing can be based on a drift term which is based on a contact map. The contact map can comprise predicted contacts between the ligand atoms and the protein’s residues. The drift term can be defined relative to the protein’s residues. The SDE can be configured to erase information about residue-scale local details by diffusing non-backbone atoms of the protein, excluding the protein backbone, around the protein backbone of the binding complex. The drift term of the protein residue coordinates in the SDE can be defined relative to the protein backbone of the binding complex.

[0181] The method can comprise performing, using a machine-learned reverse-time SDE. The reverse-time SDE can comprise a SE(3)-equivariant reverse-time-denoising process on the noisy geometrical structure of the binding complex sampled from the geometry prior. The denoising process can comprise attracting coordinates of the ligand and the protein towards each other based on the contact map. The machine-learned reverse-time SDE can comprise outputs comprising a set of displacement vectors for a set of protein heavy-atom coordinates, the set of ligand coordinates, or both.

[0182] The machine-learned reverse-time SDE can comprise inputs comprising a protein representation. The protein representation can comprise a set of protein heavy-atom nodes, a set of protein heavy-atom coordinates, a protein residue-wise representation from a protein encoder, a protein atom type, a first random Fourier encoding of diffusion time step t, or any combination thereof. The machine-learned reverse-time SDE can comprise inputs comprising a ligand representation. The ligand representation can comprise a set of ligand coordinates, a learned ligand representation, a second random Fourier encoding of the diffusion time step t, or any combination thereof. The machine-learned reverse-time SDE can comprise inputs comprising a protein-ligand representation. The protein-ligand representation can comprise a protein-ligand graph. The protein-ligand graph can comprise a first set of edges connecting protein heavy-atom nodes and the residue node that the protein atom belongs to. The protein-ligand graph can comprise a second set of edges connecting pairs of protein heavy-atomAttorney Docket No.59215-726.601 nodes that are within the same residue. The protein-ligand graph can comprise a third set of edges connecting pairs of protein heavy-atom nodes that are within a first predetermined distance. The protein-ligand graph can comprise a fourth set of edges connecting protein heavy-atom nodes and ligand atom nodes that are within a second first predetermined distance. The edges in the first, second, third, and fourth set of edges that comprise a protein heavy-atom node can be initialized with features that encode (i) whether two nodes of an edge belong to the same residue or the same ligand molecule, and / or (ii) whether there is a covalent bond between two nodes of the edge. The two nodes can be considered to be covalently bonded if a distance between the two nodes is less than about an average Van der Waals (VdW) radius of atoms constituting the two nodes.

[0183] The machine-learned reverse-time SDE can be trained with a loss function configured to reduce a difference between an observed ground-truth geometrical structure and the denoised geometrical structure of the binding complex. The loss function can be time-dependent such that the loss function is normalized by a factor that decreases with reverse-time.

[0184] In some aspects, the present disclosure provides a system implementing the method for denoising a geometrical structure of a binding complex formed between a protein and a ligand. Path-Integral Graph Transformer

[0185] In some aspects, the present disclosure provides a method for processing a representation of a molecule.

[0186] The representation can comprise a chirality-aware representation of a molecule. The representation can comprise a chirality-aware pairwise representation of a molecule. The method can be used to generate a representation of a molecule.

[0187] The method can comprise processing a set of atomic nodes. The set of atomic nodes can represent a set of atoms in the molecule. An atomic node can encode a group number (i.e., of the periodic table, ranging from 1 to 18) for an atom in the molecule. An atomic node can encode a period number (i.e., of the periodic table, ranging from 1 to 7) for an atom in the molecule. The set of atoms can comprise heavy-atoms of the molecule. The heavy-atoms can comprise atoms comprising at least two protons. The set of atoms can comprise all atoms in the molecule or a subset of atoms in the molecule. For example, the method can comprise generating a set of heavy-atom nodes ^^ encoding, for each heavy-atom in the molecule, the group and the period of the heavy-atom in c dimensions suchthat ^^ ∈ ^^Nheavy−atoms × c.

[0188] The method can comprise processing a set of local coordinate frame nodes. The set of local coordinate frame nodes can represent a set of local coordinate frames in the molecule. A local coordinate frame node can represent three atoms consecutively bonded to one another in the molecule. A local coordinate frame node can encode a group number (i.e., of the periodic table, ranging from 1 to 18) for a central atom in a corresponding local coordinate frame. A local coordinate frame node canAttorney Docket No.59215-726.601 encode a period number (i.e., of the periodic table, ranging from 1 to 7) for a central atom in a corresponding local coordinate frame. The set of local coordinate frames can comprise all local coordinate frames in the molecule or a subset of local coordinate frames in the molecule. For example, the method can comprise generating a set of local coordinate frame nodes ^^ encoding, for each local coordinate frame in the molecule, (i) the group and the period of the central atom and (ii) types ofbonds in the local coordinate frame in c dimensions, such that ^^ ∈ ^^Nframes × c.

[0189] The method can comprise processing a set of pairwise embeddings. A pairwise embedding can be stereospecific. A pairwise embedding can comprise an edge embedding. A pairwise embedding can comprise an edge embedding between an atomic node and a local coordinate frame node. A pairwise embedding can comprise an edge embedding between a pair of local coordinate frames nodes. A pairwise embedding can encode a molecular symmetry annotation for a pair of local coordinate frames in the molecule. A molecular symmetry annotation can encode, when a pair of local coordinate frames shares the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other local coordinate frame. A molecular symmetry annotation can encode, when a pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or the different side of the common double or aromatic bond. An edge embedding can comprise edges up to the fourth nearest neighbor atoms in the molecule. For example, the method can comprise generating a set of stereochemistry encodings ^^ encoding, for each pair of local coordinate frames in the molecule,molecular symmetry annotations in cs dimensions, such that ^^ ∈ ^^Nframesthe method can comprise generating a set of pair representations G encoding, for each pair betweenheavy-atoms and local coordinate frames in the molecule, ^^ and ^^ in cp dimensions, such that ^^ ∈^^Nframes × Nheavy−atoms × cp.

[0190] The set of atomic nodes, the set of local coordinate frame nodes, the set of pairwise embeddings, or any combination thereof can be processed to generate the representation of the molecule. The processing can comprise recursively updating a set of path representations for the set of atomic nodes, the set of local coordinate frame nodes, the set of pairwise embeddings, and any combination thereof. The processing can comprise recursively updating the set of atomic nodes, the set of local coordinate frame nodes, the set of pairwise embeddings, and any combination thereof based on the set of path representations. The processing can comprise denoising a noisy geometry to generate a denoised molecular geometry. The processing can comprise denoising a noisy molecular geometry based on a recursive update to generate a denoised molecular geometry. The denoising can be performed with SE(3)-invariance. For example, the processing can comprise processing H, F, G, and S for n iterations to generate features ^^^^^^^^^^, ^^^^^^^^^^, ^^^^^^^^^^, and ^^^^^^^^^^, wherein the features ^^^^^^^^^^, ^^^^^^^^^^,the same dimensions as H, F, G, and S, respectively, and n is an integer greater than zero. The processing can comprise recursively updating path representations ^^^^^^^^, ^^^^^^^^,Attorney Docket No.59215-726.601 ^^^^^^^^, and ^^^^^^^^, based on ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, in the n iterations, wherein l is an iteration index ranging from zero to n. The processing can comprise recursively updating ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, based on ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^.

[0191] The processing can comprise using a neural network. The neural network can be a graph neural network. The neural network can comprise an attention mechanism. The neural network can comprise a graph transformer. The neural network can be trained based on a negative log likelihood evaluated based on an output 3-dimensional mixture probability distribution of a geometry of a molecule and an observed or ground-truth geometry of the molecule in a training sample. The neural network can be trained based on a SE(3)-invariant denoising score matching loss based on a frame aligned point error determined between a denoised molecular geometry and an observed or ground-truth geometry of the molecule in a training sample. The neural network can be trained on bioactivity data. The bioactivity data can comprise (i) a physicochemical property, (ii) an absorption, distribution, metabolism, excretion, or toxicity (ADMET) property, (iii) a chemical reaction, (iv) a synthesizability, (v) a solubility, (vi) a chemical stability, (vii) a protease stability, (viii) a plasma protein binding property, (ix) a peptide-protein or protein-protein binding interaction, (x) an activity, (xi) a selectivity, (xii) a potency, (xiii) a pharmacokinetic property, (xiv) a pharmacodynamic property, (xv) an in vivo safety property, (xvi) a property reflecting a relationship to a biomarker, (xvii) a formulation property, or (xviii) any combination thereof. The neural network can be trained on conformers of molecules determined using quantum mechanical calculations. The neural network can be trained based on a mean squared loss for fitting level-1 chemical checker (CC) embeddings representing harmonized and integrated bioactivity data. The neural network can be trained based on a binary cross entropy loss for classifying whether a specific CC entry is available for a molecule in the chemical checker dataset. The neural network can be trained based on a cross-entropy loss for predicting masked tokens, which can encourage learning on molecular graph topology distributions.

[0192] In some aspects, the present disclosure provides a system for implementing the method for processing a representation of a molecule (e.g., as described with respect to FIG.8.) Protein Encoder

[0193] In some aspects, the present disclosure provides a method for processing a representation of a protein. The method can comprise generating an encoded representation of the protein.

[0194] The method can comprise using a neural network. The neural network can comprise a graph neural network. The neural network can comprise a graph transformer. The neural network can comprise a graph of the protein. The neural network can comprise a set of amino acid sequence nodes. The neural network can comprise a set of atomic nodes. The neural network can comprise one or more stacks of invariant point attention blocks. The one or more stacks of invariant point attention blocks can compute attention scores on the graph. The one or more stacks of invariant point attention blocks can associate each node of the graph with a plurality of replica coordinate frames. The one or moreAttorney Docket No.59215-726.601 stacks of invariant point attention blocks can output a plurality translation vectors and / or quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames. A first replica coordinate frame of the plurality of replica coordinate frames can be initialized as a copy of a protein backbone coordinates in frame representation.

[0195] The method can comprise processing an input amino acid sequence. The input amino acid sequence can comprise one or more indicators of post-translational modifications. The method can comprise processing a perturbed protein geometry. The method can comprise processing a diffusion time. The diffusion time can comprise a random Fourier encoding. The perturbed protein geometry can comprise coordinates sampled from a learned reverse-time SDE.

[0196] The graph can be a sparse graph. The sparse graph of the protein can comprise a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes. The set of atomic nodes can comprise a set of backbone atom nodes. The set of edges can connect the set of backbone atom nodes to the set of amino acid sequence nodes. The set of edges can be generated based on an inclusion probability. The inclusion probability can be based on a distance between a backbone atom and a residue in the perturbed protein geometry. The set of edges can be initialized with a randomly Fourier encoded signed sequence distance between two connected nodes if the two connected nodes are located on the same chain, and zeros if the two connected nodes are located on different chains.

[0197] In some aspects, the present disclosure provides a system for implementing method for processing a representation of a protein. Contact Predictor

[0198] In some aspects, the present disclosure provides a method for predicting contacts between a protein and a ligand. The method can comprise processing (i) a protein graph, (ii) a ligand graph, (iii) a set of intermolecular edges connecting the protein graph and the ligand graph, or any combination thereof.

[0199] The method can comprise sampling the contacts. The sampling can be performed autoregressively. The contacts can comprise pairwise proximities between protein residues of the protein and atoms or frames of the ligand. The sampling can be based on a probability distribution output by a neural network. The sampling can comprise generating (i) a plurality of intermediate contact probabilities, (ii) a plurality of intermediate distance histograms, or (iii) both. The sampling can comprise updating the set of intermolecular edges based on (i) the plurality of intermediate contact probabilities, (ii) the plurality of intermediate distance histograms, or (iii) both. The probability distribution can be configured to permits a plurality of contact modes. For instance, the probability distribution can be configured to permit modeling contacts between a ligand and a protein, the protein having more than one binding side for the ligand. The probability distribution can comprise a categorical posterior distribution over a sequence of the protein.Attorney Docket No.59215-726.601

[0200] The neural network can output a probability distribution based on the (i) the protein graph, (ii) the ligand graph, (iii) the set of intermolecular edges, or any combination thereof. The neural network can comprise one or more neural network blocks. The neural network can comprise an invariant point attention neural network. The invariant point attention neural network can process the protein graph to generate a first set of features. The neural network can comprise a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features. The first multi- head cross-attention neural network can use edge embeddings of the ligand graph as a relative positional encoding term. The neural network can comprise a second multi-head cross-attention neural network that processes the first set of features, the second set of features, the set of intermolecular edges, or any combination thereof to generate a third set of features. The neural network can comprise a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges. The updated features can comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0201] The method can comprise enumerating a set of intermolecular edges connecting a sparsely- connected protein graph and a ligand graph. The set of intermolecular edges can connect residues of the sparsely-connected protein graph and heavy-atoms of the ligand graph.

[0202] The method can comprise processing, using a neural network, (i) the sparsely-connected protein graph, (ii) the ligand graph, (iii) the set of intermolecular edges, or any combination thereof to predict the contacts. The sparsely-connected protein graph can encode an amino-acid sequence of the protein, a perturbed geometry of the protein, a diffusion time, or any combination thereof. The ligand graph can encode (A) a set of atoms of the ligand, (B) a set of local coordinate frames of the ligand, (C) a set of stereospecific pairwise embeddings between the set of atoms and the set of local coordinate frames, or (D) any combination thereof.

[0203] The processing can comprise segmenting and / or tokenizing the protein graph to generate a first set of patches representing the protein. The processing can comprise frame sampling the ligand graph to generate a second set of patches representing the ligand. The neural network can comprise an intra- patch self-attention mechanism for generating a cross-attention map based on the first set of patches or the second set of patches. The neural network can comprise an inter-patch self-attention mechanism for processing a self-attention map based on the first set of patches and the second set of patches. The inter-path self-attention mechanism can comprise a triangular-gated self-attention mechanism. The neural network can comprise a graph-attention mechanism for processing a graph attention map based on the set of intermolecular edges.

[0204] The method can comprise training the neural network to reduce a difference between (i) the posterior distribution of observed contact maps in training data, and (ii) the predicted contacts. The method can comprise training the neural network to reduce an element-wise difference between (i) the observed contact maps in the training data, and (ii) softmax-transformation of predicted contacts.Attorney Docket No.59215-726.601

[0205] In some aspects, the present disclosure provides an autoregressive neural network for predicting contacts between a protein and a ligand. The autoregressive neural network can comprise an input comprising (i) protein graph, (ii) a ligand graph, (iii) a set of intermolecular edges connecting the protein graph and the ligand graph, or (iv) any combination thereof. The protein graph can be segmented and / or tokenized into a first set of patches representing the protein. The ligand graph can be frame-sampled to generate a second set of patches representing the ligand. The autoregressive neural network can comprise an output comprising parameters of a probability distribution. The probability distribution can permit a plurality of contact modes. The output can comprise (i) a plurality of intermediate contact probabilities, (ii) a plurality of intermediate distance histograms, or (iii) both. The autoregressive neural network of claim can be configured to autoregressively update the set of intermolecular edges based on (i) the plurality of intermediate contact probabilities, (ii) the plurality of intermediate distance histograms, or (iii) both.

[0206] The autoregressive neural network can comprise one or more neural network blocks. The autoregressive neural network can comprise an invariant point attention neural network that processes the protein graph to generate a first set of features. The autoregressive neural network can comprise a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features. The second set of features can be generated while using edge embeddings of the ligand graph as a relative positional encoding term. The autoregressive neural network can comprise a second multi-head cross-attention neural network that processes the first set of features, the second set of features, the set of intermolecular edges, or any combination thereof to generate a third set of features. The autoregressive neural network can comprise a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges. The updated features can comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0207] The autoregressive neural network can comprise an intra-patch self-attention mechanism for generating a cross-attention map based on the first set of patches or the second set of patches. The autoregressive neural network can comprise an inter-patch triangular-gated self-attention mechanism for processing a self-attention map based on the first set of patches and the second set of patches. The autoregressive neural network can comprise a graph-attention mechanism for processing a graph attention map based on the set of intermolecular edges.

[0208] In some aspects, the present disclosure provides a system implementing the method for predicting contacts between a protein and a ligand. Neural Framework for Macromolecule-Macromolecule Binding Complex Structure Prediction

[0209] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules. In some embodiments, the method comprises processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate a geometry prior. In someAttorney Docket No.59215-726.601 embodiments, the method comprises sampling an initial geometrical structure of the binding complex based on the geometry prior. In some embodiments, the method comprises processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules.

[0210] In some embodiments, the method comprises processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate an initial geometrical structure of the binding complex formed by the plurality of macromolecules based on plurality of representations. In some embodiments, the method comprises the plurality of macromolecules comprise at least 32 amino acids and / or nucleotides. In some embodiments, the method comprises processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of representations.

[0211] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules and a plurality of small molecules. In some embodiments, the method comprises processing an input representation comprising a plurality of representations of the plurality of macromolecules and the plurality of small molecules to generate a geometry prior. In some embodiments, the method comprises sampling an initial geometrical structure of the binding complex based on the geometry prior. In some embodiments, the method comprises processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules and the plurality of small molecules.

[0212] In some embodiments, the plurality of macromolecules comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or 30 macromolecules. In some embodiments, the plurality of macromolecules comprises at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or 30 macromolecules. In some embodiments, the plurality of representations comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or 30 representations. In some embodiments, the plurality of representations comprises at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or 30 representations.

[0213] In some embodiments, the input representation further comprises a representation of a cofactor of the binding complex. In some embodiments, the plurality of macromolecules comprises a protein, a nucleic acid, or both. In some embodiments, the input representation further comprises a representation of a ligand. In some embodiments, the input representation comprises a plurality of representations of a plurality of ligands.

[0214] In some embodiments, the input representation comprises at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides. In some embodiments, the input representation comprises at most 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides. In some embodiments, the input representation comprises at most 1000, 1500, 2000,Attorney Docket No.59215-726.601 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides. In some embodiments, the input representation comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides. In some embodiments, the input representation comprises at most 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides.

[0215] In some embodiments, the processing the input representation comprises segmenting the protein representation into a plurality of segments. In some embodiments, each segment of the plurality of segments comprises a uniform size. In some embodiments, each segment in the plurality of segments comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids. In some embodiments, each segment in the plurality of segments comprises at most 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids. In some embodiments, each segment in the plurality of segments comprises 8 amino acids.

[0216] In some embodiments, the processing the input representation comprises predicting a plurality of contacts between a first macromolecule and a second macromolecule of the plurality of macromolecules. In some embodiments, the predicting the plurality of contacts between the first macromolecule and the second macromolecule is performed by processing the plurality of segments. In some embodiments, the processing the input representation comprises predicting a plurality of contacts between (1) a first macromolecule and a first small molecule, and / or (2) a second macromolecule and the first small molecule or a second small molecule.

[0217] In some embodiments, the predicting the plurality of contacts is based on a predetermined set of contacts. In some embodiments, the plurality of contacts comprises all or a subset of the predetermined set of contacts. In some embodiments, the predetermined set of contacts is provided by a user. In some embodiments, the predetermined set of contacts is generated from a template structure of the binding complex. In some embodiments, the template structure of the binding complex is provided by a user. In some embodiments, the template structure of the binding complex is generated by molecular mechanics or homology modeling.

[0218] In some embodiments, the processing the initial geometrical structure to generate the geometrical structure of the binding complex comprises dynamically generating connections between pairs of atoms of the binding complex. In some embodiments, the dynamically generating the connections comprises using a mixture of probability distributions to assign the connections between the pairs of atoms. In some embodiments, the total number of connections in the ESDM graph is bounded by a hardware-specific threshold.

[0219] In some embodiments, the cofactor comprises an organic cofactor or an inorganic cofactor. In some embodiments, the organic cofactor comprises flavine, heme, NAD+, thiamin pyrophosphate, pyridoxal phosphate, methylcobalamin, cobalamine, biotin, coenzyme A, tetrahydrofolic acid, menaquinone, ascorbic acid, flavin mononucleotide, flavine adenine dinucleotide, coenzyme F420,Attorney Docket No.59215-726.601 Adenosine triphosphate, S-Adenosyl methionine, Coenzyme B, Coenzyme M, Coenzyme Q, Cytidine triphosphate, Glutathione, Lipoamide, Methanofuran, Molybdopterin, a nucleotide sugar, 3'- Phosphoadenosine-5'-phosphosulfate, Tetrahydrobiopterin, Tetrahydromethanopterin, or any combination thereof. In some embodiments, the inorganic cofactor comprises a metal ion, metal iron, an iron-sulfur complex, or another transition-metal organometallic complex. In some embodiments, the plurality of macromolecules comprises a post-translation modification or a plurality of amino acids with post-translation modifications. In some embodiments, the post-translation modification comprises acylation, alkylation, prenylation, flavination, amination, deamination, carboxylation, decarboxylation, nitrosylation, halogenation, sulfurylation, glutathionylation, oxidation, oxygenation, reduction, ubiquitination, SUMOylation, neddylation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glypiation, glycosylphosphatidylinositol anchor formation, lipoylation, heme functionalization, phosphorylation, phosphopantetheinylation, retinylidene Schiff base formation, diphthamide formation, ethanolamine phosphoglycerol functionalization, hypusine formation, beta-Lysine addition, acetylation, formylation, methylation, amidation, amide bond formation, butyrylation, gamma-carboxylation, glycosylation, polysialylation, malonylation, hydroxylation, iodination, nucleotide addition, phosphate ester formation, phosphoramidate formation, adenylation, uridylylation, propionylation, pyroglutamate formation, gluthathionylation, sulfenylation, sulfinylation, sulfonylation, succinylation, sulfation, glycation, carbonylation, isopeptide bond formation, biotinylation, carbamylation, oxidation, pegylation, citrullination, deamidation, eliminylation, disulfide bond formation, proteolytic cleavage, isoaspartate formation, racemization, protein splicing, chaperone-assisted folding, or any combination thereof.

[0220] In some embodiments, the nucleic acid comprises a helix, a bulge, a loop, a junction, a stem- loop, a hairpin-loop, a tetraloop, a pseudoknot, a single strand, a double strand, or any combination thereof. In some embodiments, the nucleic acid comprises DNA, RNA, 5mC, 5hmC, 5fC, 5mC, 5hmC, 5fC, 5CaC, 6mA, 4mC, 8-oxoG, Tg, an AP site, DNA ps, or any combination thereof.

[0221] In some embodiments, the neural network is trained with a training dataset comprising predicted geometrical structures of binding complexes of macromolecules. In some embodiments, the predicted geometrical structures of binding complexes comprise a plurality of monomers aligned to form the geometric structures. In some embodiments, the neural network is trained using a loss function configured to reduce achiral and steric clashes. In some embodiments, the loss function is adjusted with a weighting function that stabilizes the training. In some embodiments, the weighting function increases in value as the number of training iterations increases. In some embodiments, the neural network is trained by distributing batches of training data across a plurality of processors based on sizes of the training data to reduce load imbalance between the plurality of processors. In some embodiments, the training data comprises at least 100k training samples of geometrical structures of binding complexes. In some embodiments, the neural network comprises at least 100M parameters.Attorney Docket No.59215-726.601

[0222] In some embodiments, the neural network is further trained with a second training dataset comprising geometrical structures of macromolecules or binding complexes generated from experiments, simulations, or electronic structure calculations. In some embodiments, the geometrical structures generated from experiments comprise cryo-EM generated structures, X-ray crystallography generated structures, and / or NMR generated structures. In some embodiments, the geometrical structures generated from simulations comprise molecular dynamics simulation or Monte Carlo simulation generated structures. In some embodiments, the geometrical structures are generated with homology modeling. In some embodiments, the geometrical structures are generated without homology modeling. In some embodiments, the geometrical structures generated from electronic structure calculations comprise DFT generated structures.

[0223] In some embodiments, a computational cost of processing the initial geometrical structure to generate the geometrical structure of the binding complex scales sub-linearly as a function of the number of amino acid residues. In some embodiments, the neural network is configured to provide a mean TM-score accuracy of 0.9 for the geometrical structure of the binding complex. In some embodiments, the neural network is configured to provide a mean TM-score accuracy of at least 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99 for the geometrical structure of the binding complex. In some embodiments, the neural network is configured to provide a mean TM-score accuracy of at most 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99 for the geometrical structure of the binding complex. In some embodiments, a ligand root-mean-square deviation of the geometrical structure is below two Angstroms. In some embodiments, a ligand root-mean-square deviation of the geometrical structure is below 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 Angstrom. In some embodiments, a ligand root-mean- square deviation of the geometrical structure is above 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 Angstrom. In some embodiments, the neural network uses site-specific docking. In some embodiments, the site- specific docking comprises using one or more residue numbers, one or more contacts, or both of one or more binding sites. In some embodiments, the neural network uses the site-specific docking as inputs. In some embodiments, the method further comprises generating a confidence score for each amino acid in the geometrical structure. In some embodiments, the confidence score is above a threshold confidence value. In some embodiments, the threshold confidence value is 0.8. In some embodiments, the threshold confidence value is at least 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. In some embodiments, the threshold confidence value is at most 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99.

[0224] In some embodiments, the method further comprises processing the geometrical structure of the binding complex using molecular mechanics. In some embodiments, the molecular mechanics comprises molecular dynamics. In some embodiments, the molecular mechanics comprises electronic structure calculations. In some embodiments, the molecular mechanics is performed with restraints or without restraints. In some embodiments, the restraints are imposed on covalent bonds formed with a hydrogen atom. In some embodiments, the restraints are imposed on torsional modes of the bindingAttorney Docket No.59215-726.601 complex. In some embodiments, the restraints are imposed based on geometrical structure of the binding complex predicted from the neural network. In some embodiments, the restraints are imposed based on initial coordinates of the binding complex predicted from the neural network.

[0225] In some embodiments, the method further comprises predicting experimental binding affinity between at least two molecules. In some embodiments, the at least two molecules comprises (a) at least two macromolecules, or (b) at least one macromolecule and at least one small molecule. In some embodiments, the input representation comprises a sequence of a macromolecule. In some embodiments, the sequence is an amino acid sequence of a protein. In some embodiments, the sequence is a nucleic acid sequence of a nucleic acid. In some embodiments, the input representation comprises a 3-dimensional structure of a molecule. In some embodiments, the 3-dimensional structure is of a protein, nucleic acid, or a small molecule. In some embodiments, the input representation comprises a sequence and a 3-dimensional structure of one or more macromolecules.

[0226] In some aspects, the present disclosure provides a system implementing the method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules. In some aspects, the present disclosure provides a system implementing the method for geometrical structure of a binding complex formed between a plurality of macromolecules and a plurality of small molecules (e.g., as described with respect to FIG.8.) Memory Efficient Tensor Operations

[0227] In some aspects, the present disclosure provides a method of computing a tensor operation between tensors of different ranks. In some embodiments, the method comprises extracting a first slice of a first tensor and a second slice of a second tensor of a plurality of tensors. In some embodiments, the method comprises broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof. In some embodiments, the method comprises performing a tensor operation between the first slice and the second slice. In some embodiments, the method comprises repeating to perform a plurality of tensor operations with other slices of the first tensor and the second tensor to compute the tensor operation between the first tensor and the second tensor.

[0228] In some embodiments, the method further comprises obtaining gradients of a result of the tensor operation with respect to the first slice and the second slice. In some embodiments, the method further comprises obtaining gradients of a result of the tensor operation with respect to the first tensor and the second tensor.

[0229] In some embodiments, the method further comprises storing the first tensor, the second tensor, or both in a first memory. In some embodiments, the method further comprises storing the first slice, the second slice, or both in a second memory.Attorney Docket No.59215-726.601

[0230] In some embodiments, the second memory is faster than the first memory. In some embodiments, the second memory has smaller size than the first memory. In some embodiments, the second memory is more physically localized to transistors of a processing unit than the first memory.

[0231] In some embodiments, the processing unit is a graphical processing unit (GPU). In some embodiments, the method is performed on a graphical processing unit (GPU).

[0232] In some embodiments, the first memory has a speed of at least 0.5, 1.0, 1.5, 2, 4, 8, 16, 32, or 64 TB / s. In some embodiments, the first memory has a speed of at most 1.5, 2, 4, 8, 16, 32, or 64 TB / s. In some embodiments, the first memory has a size of at least 1, 20, 40, 80, 160, or 320 GB. In some embodiments, the first memory has a size of at most 1, 20, 40, 80, 160, or 320 GB. In some embodiments, the first memory is a high bandwidth memory (HBM). In some embodiments, the second memory has a speed of at least 4, 8, 16, 19, 32, 64, or 128 TB / s. In some embodiments, the second memory has a speed of at most 4, 8, 16, 19, 32, 64, or 128 TB / s. In some embodiments, the second memory has a size of at least 4, 8, 16, 20, 32, 64, or 128 MB. In some embodiments, the second memory has a size of at most4, 8, 16, 20, 32, 64, or 128 MB. In some embodiments, the second memory is a static random access memory (SRAM).

[0233] In some aspects, the present disclosure provides a method of training a neural network. In some embodiments, the method comprises computing tensor operations to obtain a gradient. In some embodiments, the method comprises updating a parameter of the neural network based on the gradient.

[0234] In some aspects, the present disclosure provides a method of performing inference using a neural network. In some embodiments, the method comprises computing tensor operations to obtain a result. In some embodiments, the method comprises outputting a prediction based on the result.

[0235] In some embodiments, the method further comprises predicting or learning to predict a geometrical structure of a molecule using the neural network. In some embodiments, the method further comprises predicting or learning to predict a geometrical structure of a binding complex formed between a plurality of molecules using the neural network.

[0236] In some embodiment, the molecule or the plurality of molecules comprises small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof.

[0237] In some embodiments, the tensor operation comprises addition, subtraction, inner product, dot product, outer product, cross product, matrix multiplication, Cartesian product, Hadamard product, Kronecker product, Cracovian product, Frobenius inner product, Khatri-Rao product, face- splitting product, or any combination thereof.

[0238] In some embodiments, the plurality of tensors comprise at least 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 tensors. In some embodiments, the plurality of tensors comprise at most 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 tensors. In some embodiments, the first tensor has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200,Attorney Docket No.59215-726.601 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second tensor has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the first slice has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second slice has a rank of at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the first tensor has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second tensor has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the first slice has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the second slice has a rank of at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second tensor is greater than the rank of the first tensor by at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second slice is greater than the rank of the first slice by at least 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second tensor is greater than the rank of the first tensor by at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the rank of the second slice is greater than the rank of the first slice by at most 1, 2, 3, 4, 5, 6, 7, 9, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. In some embodiments, the broadcasting the first slice matches the shape of the second slice or the transpose thereof. In some embodiments, the first tensor is a bias tensor in an attention mechanism, and the second tensor is a query tensor, a key tensor, or a value tensor in the attention mechanism.

[0239] In some aspects, the present disclosure provides a method of computing a tensor operation between tensors of different ranks. In some embodiments, the method comprises storing a plurality of tensors comprising a first tensor and a second tensor in a first memory of a processing unit. In some embodiments, the method comprises computing a tensor operation between slices of the first tensor and the second tensor. In some embodiments, the tensor operation is computed by dynamically allocating the slices of the tensors. In some embodiments, the dynamically allocating is performed by extracting a first slice of a first tensor and a second slice of a second tensor. In some embodiments, the dynamically allocating is performed by broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof. In some embodiments, the dynamically allocating is performed by storing the first slice and the second slice in a second memory of the processing unit, wherein the second memory is faster than the first memory, and wherein the second memory is more physically localized to transistors of a processing unit than the first memory. In some embodiments, the tensor operation is computed by performing the tensor operation between the first slice and the second slice. In some embodiments, the method comprises repeating the tensor operation with other slices of the first tensor and the second tensor to complete the tensor operation between the first tensor and the second tensor.Attorney Docket No.59215-726.601

[0240] In some aspects, the present disclosure provides a method of computing a triangular attention tensor. In some embodiments, the method comprises extracting a plurality of slices from a plurality of tensors, wherein the plurality of slices comprises a first slice of a bias tensor of the plurality of tensors, a second slice of a query tensor of the plurality of tensors, a third slice of a key tensor of the plurality of tensors, and a fourth slice of a value tensor of the plurality of tensors. In some embodiments, the method comprises broadcasting the first slice to match a rank and a shape of the second slice, the third slice, the fourth slice, or a transpose of the second slice, the third slice, or the fourth slice. In some embodiments, the method comprises performing tensor operations using the plurality of slices. In some embodiments, the method comprises repeating the broadcasting and / or performing the tensor operation with other slices of the plurality of tensors to compute the triangular attention tensor.

[0241] In some embodiments, a token dimension of the triangular attention tensor is at least 128, 256, 384, 512, 768, 1024, or 1280. In some embodiments, a token dimension of the triangular attention tensor is at most 128, 256, 384, 512, 768, 1024, 1280, or 2560. In some embodiments, the token dimension represents atoms in a molecule, amino acids in a protein or peptide, a nucleotide in a nucleic acid, or any combination thereof.

[0242] In some aspects, the present disclosure provides a method for generating a geometrical structure of a molecule. In some embodiments, the method comprises processing a representation of a molecule to obtain an attention tensor from particles of the molecule. In some embodiments, the method comprises updating the representation based on the attention tensor to obtain the geometrical structure of the molecule.

[0243] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of molecules. In some embodiments, the method comprises processing a representation of a plurality of molecules to obtain an attention tensor from particles of the plurality of molecules. In some embodiments, the method comprises updating the representation based on the attention tensor to obtain the geometrical structure of the binding complex.

[0244] In some aspects, the present disclosure provides an active programming interface comprising a callable function configured to program a processor to perform operations. In some embodiments, the operations comprise extracting a first slice of a first tensor and a second slice of a second tensor. In some embodiments, the operations comprise broadcasting the first slice to match the rank and at least one dimension of the second slice or a transpose thereof. In some embodiments, the operations comprise performing the tensor operation between the first slice and the second slice. In some embodiments, the operations comprise repeating the extracting, broadcasting, and / or the tensor operation with different slices to compute the tensor operation between the first tensor and the second tensor.Attorney Docket No.59215-726.601

[0245] In some aspects, the present disclosure provides a method of computing a tensor operation between tensors of different ranks by dynamically allocating slices of the tensors and computing the tensor operation between the slices. Flow Matching Framework for Structure Prediction

[0246] In some aspects, the present disclosure provides a method for generating a geometrical structure of a molecule. In some embodiments, the method comprises processing a geometry prior of the molecule to obtain a prediction of the geometrical structure. In some embodiments, the method comprises processing the prediction and the geometry prior to obtain a vector field. In some embodiments, the method comprises generating the geometrical structure of the molecule based on the vector field.

[0247] In some aspects, the present disclosure provides a method for generating a geometrical structure of a binding complex formed between a plurality of molecules. In some embodiments, the method comprises processing a geometry prior of a plurality of molecules to obtain a prediction of the geometrical structure of the binding complex. In some embodiments, the method comprises processing the prediction and the geometry prior to obtain a vector field. In some embodiments, the method comprises generating the geometrical structure of the binding complex based on the vector field.

[0248] In some embodiments, the vector field is time independent or time dependent.

[0249] In some embodiments, the molecule or the plurality of molecules comprises small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof.

[0250] In some embodiments, the generating the geometrical structure is performed while maintaining optimal transport. In some embodiments, the generating the geometrical structure is performed while maintaining SE(3) superposition. In some embodiments, the SE(3) superposition is maintained by applying a Kabsch transform to the geometrical structure. In some embodiments, the generating the geometrical structure is performed while optimizing transport cost. In some embodiments, the generating the geometrical structure is performed while maintaining identities of the particles. In some embodiments, the identities of the particles are maintained by obtaining a graph permutation of a representation of the molecule or the plurality of molecules.

[0251] In some embodiments, the method further comprises generating the geometry prior. In some embodiments, the geometry prior is physics-based.

[0252] In some aspects, the present disclosure provides a method of generating a geometrical structure of a molecule. In some embodiments, the method comprises generating a current geometrical structure of the molecule based on a conditioning signal. In some embodiments, the method comprises generating a latent variable by encoding the conditioning signal. In some embodiments, the method comprises updating the current geometrical structure of the molecule. InAttorney Docket No.59215-726.601 some embodiments, updating the current geometrical structure comprises obtaining a predicted geometrical structure of the molecule based on the latent variable, a time variable, and the current geometrical structure. In some embodiments, updating the current geometrical structure comprises applying a global SE(3) superposition to the prediction by aligning the predicted geometrical structure to the current geometrical structure, and wherein the aligning is performed by applying a rotation to the predicted geometrical structure. In some embodiments, updating the current geometrical structure comprises obtaining an optimal graph permutation of the predicted geometrical structure to find a consistent ordering of nodes between the predicted geometrical structure and the current geometrical structure. In some embodiments, the method comprises updating the current geometrical structure by interpolating between the current geometrical structure and the predicted geometrical structure. In some embodiments, the method comprises repeating the updating until an output geometrical structure is obtained.

[0253] In some aspects, the present disclosure provides a computer system comprising: a memory comprising executable instructions; and at least one processor configured to execute the instructions, wherein when the at least one processor executes the instructions, the at least one processor causes the system to generate a geometrical structure of a binding complex formed between a plurality of molecules using a flow matching generative model.

[0254] In some embodiments, the flow matching generative model is configured to generate the geometrical structure of the binding complex formed between small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof. In some embodiments, the flow matching generative model is configured to parameterize a flow, wherein the flow describes an evolution of the probability density path. Machine Learning Techniques

[0255] Various aspects of the present disclosure may be used to generate protein structures and / or representations. The protein structures and / or representations may include ligand structures or representations. Deep generative learning can refer to a method of training a machine learning model to approximate a distribution of a given dataset, while enabling coherent samples to be generated from the data distribution. Machine learning models trained in this way may be referred to as deep generative models. Deep generative learning can be applied on chemical datasets to create deep generative models for chemistry that allows generation of coherent molecular structures that are not present in a given training dataset of chemical structures.

[0256] Various inputs may be provided to the generative machine learning model. An input of a protein may be provided. The input of the protein can comprise: a primary structure of the protein, a secondary structure of the protein, a tertiary structure of the protein, a quaternary structure of the protein, a physical property of the protein, a binding site of the protein, a relative accessible surface area of the protein, or any combination thereof. The input can be embedding in a vector, e.g., using aAttorney Docket No.59215-726.601 tokenization scheme or a neural network. An input of a ligand can comprise a chemical structure of the ligand. The chemical structure can be a human-readable representation of the ligand, e.g., SMILES. The chemical structure can be an atomistic structure or a coarse-grained structure of the ligand.

[0257] Various machine learning models may be used. In some cases, the machine learning model comprises a language model, a graph model, a flow model, a generative adversarial network, a variational autoencoder, an autoregressive model, an autoencoder, a diffusion model, or any combination thereof.

[0258] In some cases, a “representation”, a “descriptor”, a “latent vector”, and the like, may refer to a description of a molecular entity. For example, commonly used descriptors include SMILES, primary sequences, or nucleic acid sequences. In some cases, a representation may be used as features or feature values to a neural network. In some cases, a representation be features or feature values from a neural network.

[0259] In some cases, the features or feature values may comprise an identifier of an atom, a functional group, an amino acid, a secondary structure, a tertiary structure, or a motif. In some cases, the features or feature values may comprise an electronic configuration, a charge, a size, a bond angle, a dihedral angle, or any combination thereof. In some cases, the motif may comprise a collection of atoms. In some cases, the features or feature values may comprise a geometric descriptor of an atom, an amino acid, a secondary structure, a tertiary structure, or a motif. In some cases, the features or feature values may comprise a physical property.

[0260] Some examples of secondary structures include alpha helices and beta sheets of proteins, where each can be formed when hydrogen bonding between residues of a protein is stabilized. Some examples of a tertiary structure can include larger geometrical features within a protein such as pockets, hairpins, concatenations, loop regions, globular regions, etc. The secondary structure of proteins may include the hydrogen bonding networks of subsections in a primary sequence. Alpha helices and beta sheets can be discerned, for example, within a Ramachandran plot of a protein due to the arrangement of the backbone amide groups hydrogen bonding. Tertiary structure may be seen as more interaction types contribute to the conformational landscape of a protein. Tertiary structures may depend on non-polar / hydrophobic van der Waal interactions, electrostatic interactions, and other various interactions described herein.

[0261] In some cases, a representation is in MOL2, PDB, MOL, PDBQ / PDBQT, SDF, CIF, CML, XML, ASN1, PARM, CRD, or TRJ. In some cases, a representation comprise SMILES, SELFIES, or InChI.

[0262] In some cases, a representation comprises one or more electron configurations. In some cases, an electron configuration may comprise one or more atomic orbitals, one or more molecular orbitals, or both. In some cases, an electron configuration may comprise valence electrons of an atom. In some cases, an electron configuration may comprise a character of an electron (e.g., s, p, d,Attorney Docket No.59215-726.601 f, and any mixtures thereof). In some cases, an electron configuration may comprise an electron spin. In some cases, an electron configuration may comprise electron density. In some cases, an electron configuration may be represented in various basis functions, including but not limited to, atomic orbitals, molecular orbitals, or plane waves.

[0263] Within various chemoinformatic formats can be differently encoded information. In some cases, an atomistic representation may comprise the relative cartesian coordinates of atoms to each other. In some cases, an atomistic representation may comprise the relative cartesian coordinates of atoms to an arbitrary point. In some cases, an atomistic representation may comprise thermodynamic estimations of values such as solvation energy, potential energy of bond lengths, bond angles, dihedral angles, 1-4 intramolecular interaction energies, intramolecular energies among adjacent bond angles, hydrogen bonding energies, and non-bonded interaction energies. In some cases, an atomistic representation may comprise atom type definitions and generalizations. In some cases, an atomistic representation may comprise polarizability parameters. In some cases, an atomistic representation may comprise Lennard-Jones van der Waal parameters. In some cases, an atomistic representation may comprise electrostatic charge parameters. In some cases, an atomistic representation may comprise bond length, bond angle, and dihedral force constants. In some cases, an atomistic representation may comprise bond length, bond angle, and dihedral equilibrium values. In some cases, an atomistic representation may comprise dihedral phase and periodicity force constants.

[0264] A graph, graph model, and graphical model can refer to a method of conceptualizing or organizing information into a graphical representation comprising nodes and edges. In some cases, a graph can refer to the principle of conceptualizing or organizing data, wherein the data may be stored in a various and alternative forms such as linked lists, dictionaries, spreadsheets, arrays, in permanent storage, in transient storage, and so on, and is not limited to specific cases disclosed herein. In some cases, the machine learning model can comprise a graph model.

[0265] The machine learning model can comprise a neural network comprising various architectures, loss functions, optimization algorithms, priors, and various other neural network design choices. In some cases, the machine learning model can comprise a neural network. In some cases, the machine learning model can comprise an autoencoder. In some cases, the machine learning model can comprise a generative model. In some cases, the machine learning model can comprise a variational autoencoder. In some cases, the machine learning model can comprise a generative adversarial network. In some cases, the machine learning model can comprise a flow model. In some cases, the machine learning model can comprise an autoregressive model. In some cases, the machine learning model can comprise a diffusion model. In some cases, the machine learning model can comprise a neural network with one or more layers. In some cases, the machine learning model can comprise a neural network with one or more fully connected layers. In some cases, the machine learning model can comprise a neural network with one or more convolutional layers. In some cases,Attorney Docket No.59215-726.601 the machine learning model can comprise a neural network with one or more message-passing layers. In some cases, the machine learning model can comprise a neural network with a bottleneck layer. In some cases, a layer may comprise an attention mechanism, a generalized message-passing graph neural network, or both. In some cases, a generalized message-passing graph neural network comprises a graph convolutional neural network.

[0266] In some cases, the machine learning model can comprise a neural network with residual blocks. In some cases, the machine learning model can comprise a neural network with attention. In some cases, the machine learning model can comprise a neural network with one or more non- linearities. In some cases, the machine learning model can comprise a neural network with one or more dropout layers. In some cases, the machine learning model can comprise a neural network with one or more batch normalization layers. In some cases, the machine learning model can comprise a regression loss function. In some cases, the machine learning model can comprise a logistic loss function. In some cases, the machine learning model can comprise a variational loss. In some cases, the machine learning model can comprise a prior. In some cases, the machine learning model can comprise a Gaussian prior. In some cases, the machine learning model can comprise a non-Gaussian prior. In some cases, the machine learning model can comprise an adversarial loss. In some cases, the machine learning model can comprise a reconstruction loss. In some cases, the machine learning model is trained with the Adam optimizer. In some cases, the machine learning model is trained with the stochastic gradient descent optimizer. In some cases, the model learning model hyperparameters are optimized with Gaussian Processes. In some cases, the machine learning model is trained with train / validation / test data splits. In some cases, the machine learning model is trained with k-fold data splits, with any positive integer for k.

[0267] The machine learning model can comprise a variety of manifold learning algorithms. In some cases, the machine learning model can comprise a manifold learning algorithm. In some cases, the manifold learning algorithm comprises principal component analysis. In some cases, the manifold learning algorithm comprises a uniform manifold approximation algorithm. In some cases, the manifold learning algorithm comprises an isomap algorithm. In some cases, the manifold learning algorithm comprises a locally linear embedding algorithm. In some cases, the manifold learning algorithm comprises a modified locally linear embedding algorithm. In some cases, the manifold learning algorithm comprises a Hessian eigen mapping algorithm. In some cases, the manifold learning algorithm comprises a spectral embedding algorithm. In some cases, the manifold learning algorithm comprises a local tangent space alignment algorithm. In some cases, the manifold learning algorithm comprises a multi-dimensional scaling algorithm. In some cases, the manifold learning algorithm comprises a t-distributed stochastic neighbor embedding algorithm (t-SNE). In some cases, the manifold learning algorithm comprises a Barnes-Hut t-SNE algorithm.

[0268] In some cases, the methods of the disclosure further comprise reducing one or more representations using a machine learning model. The terms “reducing”, “dimensionality reduction”,Attorney Docket No.59215-726.601 “projection”, “component analysis”, “feature space reduction”, “latent space engineering”, “feature space engineering”, “representation engineering”, or “latent space embedding”, as used herein, generally refer to a method of transforming a given input data with an initial number of dimensions to another form of data that has fewer dimensions than the initial number of dimensions. In some cases, the terms can refer to the principle of reducing a set of input dimensions to a smaller set of output dimensions. In some cases, the terms can refer to the principle of reducing a set of input dimensions to a set of output dimensions of a same or larger size.

[0269] The term “normalizing”, as used herein, generally refers to a collection of methods for adjusting a dataset to align the dataset to a common scale. In some cases, a normalizing method can comprise multiplying a portion or the entirety of a dataset by a factor. In some cases, a normalizing method can comprise adding or subtracting a constant from a portion or the entirety of a dataset. In some cases, a normalizing method can comprise adjusting a portion or the entirety of a dataset to a known statistical distribution. In some cases, a normalizing method can comprise adjusting a portion or the entirety of a dataset to a normal distribution. In some cases, a normalizing method can comprise adjusting the dataset so that the signal strength of a portion or the entirety of a dataset is about the same.

[0270] Converting can comprise one or more steps of various conversions of data. In some cases, converting can comprise normalizing data. In some cases, converting can comprise performing a mathematical operation that computes a score based on a distance between 2 points in the data. The 2 points can be, e.g., particles in a protein, a ligand, or a binding complex. The particles can be atoms or coarse-grained particles (e.g. protein residues). In some cases, the distance can comprise a distance between two edges in a graph. In some cases, the distance can comprise a distance between two nodes in a graph. In some cases, the distance can comprise a distance between a node and an edge in a graph. In some cases, the distance can comprise a Euclidean distance. In some cases, the distance can comprise a non-Euclidean distance. In some cases, the distance can be computed in a frequency space. In some cases, the distance can be computed in Fourier space. In some cases, the distance can be computed in Laplacian space. In some cases, the distance can be computed in spectral space. In some cases, the mathematical operation can be a monotonic function based on the distance. In some cases, the mathematical operation can be a non-monotonic function based on the distance. In some cases, the mathematical operation can be an exponential decay function. In some cases, the mathematical operation can be a learned function.

[0271] In some cases, converting can comprise transforming data in one representation to another representation. In some cases, converting can comprise transforming data into another form of data with less dimensions. In some cases, converting can comprise linearizing one or more curved paths in the data. In some cases, converting can be performed on data comprising data in Euclidean space. In some cases, converting can be performed on data comprising data in graph space. In some cases, converting can be performed on data in a discrete space. In some cases, converting can be performedAttorney Docket No.59215-726.601 on data comprising data in frequency space. In some cases, converting can transform data in discrete space to continuous space, continuous space to discrete space, graph space to continuous space, continuous space to graph space, graph space to discrete space, discrete space to graph space, or any combination thereof. In some cases, converting can comprise transforming data in discrete space into a frequency domain. In some cases, converting can comprise transforming data in continuous space into a frequency domain. In some cases, converting can comprise transforming data in graph space into a frequency domain.

[0272] In some cases, reducing can comprise transforming a given input data with any initial number of dimensions to another form of data that has any number of dimensions fewer than the initial number of dimensions. In some cases, reducing can comprise transforming input data into another form of data with fewer dimensions. In some cases, reducing can comprise linearizing one or more curved paths in the input data to the output data. In some cases, reducing can be performed on data comprising data in Euclidean space. In some cases, reducing can be performed on data comprising data in graph space. In some cases, reducing can be performed on data in a discrete space. In some cases, reducing can transform data in discrete space to continuous space, continuous space to discrete space, graph space to continuous space, continuous space to graph space, graph space to discrete space, discrete space to graph space, or any combination thereof.

[0273] The terms “clustering”, “cluster analysis”, or “generating modules”, as used herein, generally refer to a method of grouping samples in a dataset by some measure of similarity. Samples can be grouped in a set space, for example, element ‘a’ is in set ‘A’. Samples can be grouped in a continuous space, for example, element ‘a’ is a point in Euclidean space with distance ‘l’ away from the centroid of elements comprising cluster ‘A’. Samples can be grouped in a graph space, for example, element ‘a’ is highly connected to elements comprising cluster ‘A’. These terms can refer to the principle of organizing a plurality of elements into groups in some mathematical space based on some measure of similarity. Computer Systems

[0274] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 8 shows a computer system 801 that is programmed or otherwise configured to, for example, process a representation of a molecule, generate a structure of a molecule, or both.

[0275] The computer system can implement a framework of the present disclosure. The framework can generate a structure of a binding complex formed between a plurality of macromolecules. The framework can generate a structure of a binding complex formed between a plurality of macromolecules and a plurality of ligands. The framework can generate a structure of a binding complex in the presence of cofactors. The framework can generate a structure of a binding complex formed between a ligand and a protein. The framework can predict contacts between the ligand and the protein in the binding complex. The contacts can define which portions of the ligand are proximalAttorney Docket No.59215-726.601 to which portions of the protein. Predicting the contacts can be based on neural network processing of the sequence of the protein, an embedding of the protein, a graph of the ligand, or any combination thereof. The contacts can be, e.g., pairwise proximities among residues of the protein and atoms or frames of the ligand. The framework can generate a geometry prior based on the contacts. The predicted contacts can be used as a basis for generating a probability distribution of potential geometrical structures of the binding complex. The framework can denoise the geometry prior to generate a structure of the binding complex. The denoising can be performed, e.g., using a diffusion model, to extract high-likelihood geometrical structures from the geometry prior. The denoising can be performed with SE(3) equivariance, reducing temperature of the chemical system, or both. The resulting geometrical structure of the binding complex can be used to generate a report of the structure. The method can generate a report of the uncertainty of the geometrical structure of the binding complex. The uncertainty can be provided at the resolution of each atom in the geometrical structure. The uncertainty can be provided as an average (e.g., mean) over each atom in the geometrical structure.

[0276] The computer system 801 may regulate various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, processing a representation of a molecule, generating a structure of a molecule, or both. The computer system 801 may be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device may be a mobile electronic device.

[0277] The computer system 801 includes a central processing unit (CPU, also “processor” and “computer processor” herein) or a graphical processing unit (GPU) 805, which may be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 801 also includes memory or memory location 810 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 815 (e.g., hard disk), communication interface 820 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 825, such as cache, other memory, data storage and / or electronic display adapters. The memory 810, storage unit 815, interface 820 and peripheral devices 825 are in communication with the CPU 805 through a communication bus (solid lines), such as a motherboard. The storage unit 815 may be a data storage unit (or data repository) for storing data. The computer system 801 may be operatively coupled to a computer network (“network”) 830 with the aid of the communication interface 820. The network 830 may be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet.

[0278] The network 830 in some cases is a telecommunication and / or data network. The network 830 may include one or more computer servers, which may enable distributed computing, such as cloud computing. For example, one or more computer servers may enable cloud computing over the network 830 (“the cloud”) to perform various aspects of analysis, calculation, and generation of the present disclosure, such as, for example, processing a representation of a molecule, generating a structure of a molecule, or both. Such cloud computing may be provided by cloud computing platforms such as, forAttorney Docket No.59215-726.601 example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM cloud. The network 830, in some cases with the aid of the computer system 801, may implement a peer-to- peer network, which may enable devices coupled to the computer system 801 to behave as a client or a server.

[0279] The CPU 805 may comprise one or more computer processors and / or one or more graphics processing units (GPUs). The CPU 805 may execute a sequence of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 810. The instructions may be directed to the CPU 805, which may subsequently program or otherwise configure the CPU 805 to implement methods of the present disclosure. Examples of operations performed by the CPU 805 may include fetch, decode, execute, and writeback.

[0280] The CPU 805 may be part of a circuit, such as an integrated circuit. One or more other components of the system 801 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0281] The storage unit 815 may store files, such as drivers, libraries and saved programs. The storage unit 815 may store user data, e.g., user preferences and user programs. The computer system 801 in some cases may include one or more additional data storage units that are external to the computer system 801, such as located on a remote server that is in communication with the computer system 801 through an intranet or the Internet.

[0282] The computer system 801 may communicate with one or more remote computer systems through the network 830. For instance, the computer system 801 may communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user may access the computer system 801 via the network 830.

[0283] Methods as described herein may be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 801, such as, for example, on the memory 810 or electronic storage unit 815. The machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by the processor 805. In some cases, the code may be retrieved from the storage unit 815 and stored on the memory 810 for ready access by the processor 805. In some situations, the electronic storage unit 815 may be precluded, and machine-executable instructions are stored on memory 810.

[0284] The code may be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or may be compiled during runtime. The code may be supplied in a programming language that may be selected to enable the code to execute in a pre-compiled or as- compiled fashion.Attorney Docket No.59215-726.601

[0285] Aspects of the systems and methods provided herein, such as the computer system 801, may be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine- executable code may be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media may include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non- transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0286] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.Attorney Docket No.59215-726.601

[0287] The computer system 801 may include or be in communication with an electronic display 835 that comprises a user interface (UI) 840 for providing, for example, processing a representation of a molecule, generating a structure of a molecule, or both. Examples of UIs include, without limitation, a graphical user interface (GUI) and web-based user interface.

[0288] Methods and systems of the present disclosure may be implemented by way of one or more algorithms. An algorithm may be implemented by way of software upon execution by the central processing unit 805. The algorithm can, for example, process a representation of a molecule, generate a structure of a molecule, or both. List of Embodiments

[0289] The following list of embodiments of the invention are to be considered as disclosing various features of the invention, which features can be considered to be specific to the particular embodiment under which they are discussed, or which are combinable with the various other features as listed in other embodiments. Thus, simply because a feature is discussed under one particular embodiment does not necessarily limit the use of that feature to that embodiment.

[0290] Embodiment 1. A graph neural network for learning a chirality-aware pairwise representation of a molecule, comprising: a graph transformer comprising: a set of atomic nodes; a set of local coordinate frame nodes; and a set of stereospecific pairwise embeddings between the set of atomic nodes and the set of local coordinate frame nodes.

[0291] Embodiment 2. The graph neural network of Embodiment 1, wherein the set of atomic nodes encodes a set of atoms in the molecule.

[0292] Embodiment 3. The graph neural network of Embodiment 2, wherein the set of atomic nodes encodes a group number for the set of atoms in the molecule.

[0293] Embodiment 4. The graph neural network of Embodiment 3, wherein the group number comprises an integer ranging from 1 to 18.

[0294] Embodiment 5. The graph neural network of any one of Embodiments 2-4, wherein the set of atomic nodes encodes a period number for the set of atoms in the molecule.

[0295] Embodiment 6. The graph neural network of Embodiment 5, wherein the period number comprises an integer ranging from 1 to 7.

[0296] Embodiment 7. The graph neural network of any one of Embodiments 2-6, wherein the set of atoms comprises heavy-atoms of the molecule.

[0297] Embodiment 8. The graph neural network of Embodiment 7, wherein the heavy-atoms comprise atoms comprising at least two protons.

[0298] Embodiment 9. The graph neural network of any one of Embodiments 1-8, wherein the set of local coordinate frame nodes encodes a set of local coordinate frames of the molecule.Attorney Docket No.59215-726.601

[0299] Embodiment 10. The graph neural network of Embodiment 9, wherein a local coordinate frame, in the set of local coordinate frames, comprises three atoms consecutively bonded to one another in the molecule.

[0300] Embodiment 11. The graph neural network of Embodiment 9 or 10, wherein the set of local coordinate frame nodes encodes a group number for a central atom in a local coordinate frame in the molecule.

[0301] Embodiment 12. The graph neural network of any one of Embodiments 9-11, wherein the set of local coordinate frame nodes encodes a period number for a central atom in a local coordinate frame in the molecule.

[0302] Embodiment 13. The graph neural network of any one of Embodiments 1-12, wherein the set of stereospecific pairwise embeddings encodes a pair of local coordinate frames in the molecule.

[0303] Embodiment 14. The graph neural network of Embodiment 13, wherein the set of stereospecific pairwise embeddings encodes a molecular symmetry annotation for the pair of local coordinate frames in the molecule.

[0304] Embodiment 15. The graph neural network of Embodiment 14, wherein the molecular symmetry annotation encodes, when the pair of local coordinate frames shares the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other.

[0305] Embodiment 16. The graph neural network of Embodiment 14 or 15, wherein the molecular symmetry annotation encodes, when the pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or a different side of the common double or aromatic bond.

[0306] Embodiment 17. The graph neural network of any one of Embodiments 1-16, wherein the set of stereospecific pairwise embeddings comprises a set of edge embeddings connecting the set of atomic nodes with the set of local coordinate frame nodes.

[0307] Embodiment 18. The graph neural network of Embodiment 17, wherein the set of edge embeddings are based on encodings of the set of atomic nodes and the set of local coordinate frame nodes.

[0308] Embodiment 19. The graph neural network of Embodiment 17 or 18, wherein the set of edge embeddings comprises edges up to the fourth nearest neighbor atoms in the molecule.

[0309] Embodiment 20. The graph neural network of any one of Embodiments 1-19, wherein the graph transformer is configured to process the set of atomic nodes, the set of local coordinate frame nodes, and the set of stereospecific pairwise embeddings to generate the representation of the molecule.

[0310] Embodiment 21. The graph neural network of Embodiment 20, wherein the process is configured to recursively update a set of path representations for the set of atomic nodes, the set ofAttorney Docket No.59215-726.601 local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings.

[0311] Embodiment 22. The graph neural network of Embodiment 21, wherein the process is configured to recursively update the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings based on the set of path representations.

[0312] Embodiment 23. A graph neural network for processing a representation of a molecule, comprising: a graph transformer comprising: a set of heavy-atom nodes ^^ encoding, for each heavy-atom in the molecule, the group and the period of the heavy-atom in c dimensions such that ^^ ∈^^Nheavy−atoms × c; a set of local coordinate frame nodes ^^ encoding, for each local coordinate frame in the molecule, (i) the group and the period of the central atom and (ii) types of bonds in the localcoordinate frame in c dimensions, such that ^^ ∈ ^^Nframes × c; a set of stereochemistry encodings ^^encoding, for each pair of local coordinate frames in the molecule, molecular symmetry annotationsin cs dimensions, such that ^^ ∈molecular symmetry annotationscomprising: a first molecular symmetry annotation encoding, when a pair of local coordinate frames share the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other local coordinate frame; and a second molecular symmetry annotation encoding, when the pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or the different side of the common double or aromatic bond; and a set of pair representations G encoding, for each pair between heavy-atoms and local coordinate frames in the molecule, ^^ and ^^ in cpdimensions, such that ^^ ∈ ^^Nframes × Nheavy−atoms × cp;graph transformer is configured toprocess the representation of a molecule, by: processing H, F, G, and S for n iterations to generate featureswherein the featuresand ^^^^^^^^^^comprise the same dimensions as H, F, G, and S, respectively, and n is an integer greater than zero, wherein the processing comprises: recursively updating path representations ^^^^, ^^^^+1, ^^outbased on ^^^^, ^^^^, ^^^^, and ^^^^, in the n iterations, wherein l is an iteration index ranging from zero to n; and recursively updating ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, based on ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^.

[0313] Embodiment 24. The graph neural network of any one of Embodiments 1-23, wherein the graph neural network is trained based on at least one of: a negative log likelihood evaluated based on an output 3-dimensional mixture probability distribution of a geometry of the molecule and an observed geometry of the training sample; a SE(3)-invariant denoising score matching loss based on a frame aligned point error determined between a denoised molecular geometry and a ground-truth molecular geometry of the molecule; a mean squared loss for fitting level-1 chemical checker (CC) embeddings representing harmonized and integrated bioactivity data; a binary cross entropy loss for classifying whether a specific CC entry is available for any molecule in the chemical checker dataset;Attorney Docket No.59215-726.601 and a cross-entropy loss for predicting the masked tokens which is added to encourage learning on molecular graph topology distributions.

[0314] Embodiment 25. A method for learning a chirality-aware pairwise representation of a molecule, comprising processing: a set of atomic nodes encoding a set of atoms in the molecule; a set of local coordinate frame nodes encoding a set of local coordinate frames in the molecule; and a set of stereospecific pairwise embeddings, between the set of atomic nodes and the set of local coordinate frame nodes; to generate the representation of the molecule.

[0315] Embodiment 26. The method of Embodiment 25, wherein the set of atomic nodes encodes a set of atoms in the molecule.

[0316] Embodiment 27. The method of Embodiment 26, wherein the set of atomic nodes encodes a group number for the set of atoms in the molecule.

[0317] Embodiment 28. The method of Embodiment 27, wherein the group number comprises an integer ranging from 1 to 18.

[0318] Embodiment 29. The method of any one of Embodiments 26-28, wherein the set of atomic nodes encodes a period number for the set of atoms in the molecule.

[0319] Embodiment 30. The method of Embodiment 29, wherein the period number comprises an integer ranging from 1 to 7.

[0320] Embodiment 31. The method of any one of Embodiments 26-30, wherein the set of atoms comprises heavy-atoms of the molecule.

[0321] Embodiment 32. The method of Embodiment 31, wherein the heavy-atoms comprise atoms comprising at least two protons.

[0322] Embodiment 33. The method of any one of Embodiments 25-32, wherein the set of local coordinate frame nodes encodes a set of local coordinate frames of the molecule.

[0323] Embodiment 34. The method of Embodiment 33, wherein a local coordinate frame, in the set of local coordinate frames, comprises three atoms consecutively bonded to one another in the molecule.

[0324] Embodiment 35. The method of Embodiment 33 or 34, wherein the set of local coordinate frame nodes encodes a group number for a central atom in a local coordinate frame in the molecule.

[0325] Embodiment 36. The method of any one of Embodiments 33-35, wherein the set of local coordinate frame nodes encodes a period number for a central atom in a local coordinate frame in the molecule.

[0326] Embodiment 37. The method of any one of Embodiments 25-36, wherein the set of stereospecific pairwise embeddings encodes each pair of local coordinate frames in the molecule.

[0327] Embodiment 38. The method of Embodiment 37, wherein the set of stereospecific pairwise embeddings encodes a molecular symmetry annotation for a pair of local coordinate frames in the molecule.Attorney Docket No.59215-726.601

[0328] Embodiment 39. The method of Embodiment 38, wherein the molecular symmetry annotation encodes, when the pair of local coordinate frames shares the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other local coordinate frame.

[0329] Embodiment 40. The method of Embodiment 38 or 39, wherein the molecular symmetry annotation encodes, when the pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or the different side of the common double or aromatic bond.

[0330] Embodiment 41. The method of any one of Embodiments 25-40, wherein the set of stereospecific pairwise embeddings comprises a set of edge embeddings connecting the set of atomic nodes with the set of local coordinate frame nodes.

[0331] Embodiment 42. The method of Embodiment 41, wherein the set of edge embeddings are based on encodings of the set of atomic nodes and the set of local coordinate frame nodes.

[0332] Embodiment 43. The method of Embodiment 41 or 42, wherein the set of edge embeddings comprises edges up to the fourth nearest neighbor atoms in the molecule.

[0333] Embodiment 44. The method of any one of Embodiments 25-43, wherein the processing comprises using a graph transformer.

[0334] Embodiment 45. The method of any one of Embodiments 25-44, wherein the processing comprises processing a noisy geometry to generate a denoised molecular geometry.

[0335] Embodiment 46. The method of any one of Embodiments 25-45, wherein the processing comprises processing the set of atomic nodes, the set of local coordinate frame nodes, and the set of stereospecific pairwise embeddings to generate the representation of the molecule.

[0336] Embodiment 47. The method of any one of Embodiments 25-46, wherein the processing comprises recursively updating a set of path representations for the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings.

[0337] Embodiment 48. The method of Embodiment 47, wherein the processing comprises recursively updating the set of atomic nodes, the set of local coordinate frame nodes, the set of stereospecific pairwise embeddings, and the set of edge embeddings based on the set of path representations.

[0338] Embodiment 49. The method of Embodiment 48, wherein the processing comprises denoising the noisy molecular geometry based on the recursive updating to generate the denoised molecular geometry.

[0339] Embodiment 50. The method of Embodiment 48 or 49, wherein the processing comprises denoising with SE(3)-invariance.

[0340] Embodiment 51. A method for processing a representation of a molecule, comprising: generating a set of heavy-atom nodes ^^ encoding, for each heavy-atom in the molecule, the groupAttorney Docket No.59215-726.601and the period of the heavy-atom in c dimensions such that ^^ ∈ ^^Nheavy−atoms × c; generating a set oflocal coordinate frame nodes ^^ encoding, for each local coordinate frame in the molecule, (i) the group and the period of the central atom and (ii) types of bonds in the local coordinate frame in cdimensions, such that ^^ ∈ ^^Nframes × c; generating a set of stereochemistry encodings ^^ encoding, foreach pair of local coordinate frames in the molecule, molecular symmetry annotations in csdimensions, such that ^^ ∈ ^^Nframes × Nframes × cs, the molecular symmetry annotations comprising: afirst molecular symmetry annotation encoding, when a pair of local coordinate frames share the same central atom, whether an atom in one of the pair of local coordinate frames is above or below a plane formed by the other local coordinate frame; and a second molecular symmetry annotation encoding, when the pair of local coordinate frames shares a common double or aromatic bond, whether unshared bonds of the pair of local coordinate frames are on the same side or the different side of the common double or aromatic bond; generating a set of pair representations G encoding, for each pair between heavy-atoms and local coordinate frames in the molecule, ^^ and ^^ in cpdimensions, suchthat ^^ ∈ ^^Nframes × Nheavy−atoms × cp; processing a noisy geometry to generate a denoised moleculargeometry, by: processing H, F, G, and S for n iterations to generate features ^^^^^^^^^^, ^^^^^^^^^^, ^^^^^^^^^^, andand ^^^^^^^^^^comprise the same dimensions as H, F, G, and S, respectively, and n is an integer greater than zero, wherein the processing comprises: recursively updating path representations ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^, based on ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, in the n iterations, wherein l is an iteration index ranging from zero to n; and recursively updating ^^^^+^^, ^^^^+^^, ^^^^+^^, and ^^^^+^^, based on ^^^^^^^^, ^^^^^^^^, ^^^^^^^^, and ^^^^^^^^.

[0341] Embodiment 52. A method of training a neural network, comprising: providing a training sample of a representation a molecule; processing, using a neural network, the representation of the molecule according to the method of any one of Embodiments 25-51; updating parameters of the neural network based on at least one of: a negative log likelihood evaluated based on an output 3- dimensional mixture probability distribution of a geometry of the molecule and an observed geometry of the training sample; a SE(3)-invariant denoising score matching loss based on a frame aligned point error determined between the denoised molecular geometry and a ground-truth molecular geometry of the molecule; a mean squared loss for fitting level-1 chemical checker (CC) embeddings representing harmonized and integrated bioactivity data; a binary cross entropy loss for classifying whether a specific CC entry is available for any molecule in the chemical checker dataset; and a cross-entropy loss for predicting the masked tokens which is added to encourage learning on molecular graph topology distributions.

[0342] Embodiment 53. A graph neural network for processing a representation of a protein, comprising: a graph transformer comprising: a sparse graph of the protein, comprising: a set of amino acid sequence nodes; a set of atomic nodes; a perturbed protein geometry; and a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes; and one or moreAttorney Docket No.59215-726.601 stacks of invariant point attention blocks; and wherein the graph transformer is configured to process an input amino acid sequence, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0343] Embodiment 54. The graph neural network of Embodiment 53, wherein the set of atomic nodes comprises a set of backbone atom nodes.

[0344] Embodiment 55. The graph neural network of Embodiment 54, wherein the set of edges connect the set of backbone atom nodes to the set of amino acid sequence nodes.

[0345] Embodiment 56. The graph neural network of Embodiment 54 or 55, wherein the set of edges are generated based on an inclusion probability, the inclusion probability being based on a distance between a backbone atom and a residue in the perturbed protein geometry.

[0346] Embodiment 57. The graph neural network of any one of Embodiments 53-56, wherein the graph transformer is configured to process a diffusion time.

[0347] Embodiment 58. The graph neural network of Embodiment 57, wherein the diffusion time comprises a random Fourier encoding.

[0348] Embodiment 59. The graph neural network of any one of Embodiments 53-58, wherein the set of edges are initialized with a randomly Fourier encoded signed sequence distance between two connected nodes if the two connected nodes are located on the same chain, and zeros if the two connected nodes are located on different chains.

[0349] Embodiment 60. The graph neural network of any one of Embodiments 53-59, wherein the perturbed coordinates are sampled from a learned reverse-time SDE.

[0350] Embodiment 61. The graph neural network of any one of Embodiments 53-60, wherein the one or more stacks of invariant point attention blocks are configured to compute attention scores on the graph.

[0351] Embodiment 62. The graph neural network of any one of Embodiments 53-61, wherein the one or more stacks of invariant point attention blocks are configured to associate each node of the sparsely-connected graph with a plurality of replica coordinate frames.

[0352] Embodiment 63. The graph neural network of any one of Embodiments 53-62, wherein the one or more stacks of invariant point attention blocks are configured to output a plurality translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames.

[0353] Embodiment 64. The graph neural network of any one of Embodiments 62 or 63, wherein a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation.

[0354] Embodiment 65. A graph neural network for processing a representation of a protein, comprising: a graph transformer comprising: a sparse graph of the protein, comprising: a set of amino acid sequence nodes; a set of atomic nodes comprising a set of backbone atom nodes; a perturbed protein geometry sampled from a learned reverse-time SDE; and a set of edges sparselyAttorney Docket No.59215-726.601 connecting the set of atomic nodes to the set of amino acid sequence nodes, wherein the set of edges are generated based on an inclusion probability function; and one or more stacks of invariant point attention blocks configured to: compute attention scores on the graph; associate each node of the sparsely-connected graph with a plurality of replica coordinate frames, wherein a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation; and output a plurality of translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames; wherein the graph transformer is configured to process an input amino acid sequence, a diffusion time, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0355] Embodiment 66. A method for processing a representation of a protein, comprising: providing a graph transformer comprising: a sparse graph of the protein, comprising: a set of amino acid sequence nodes; a set of atomic nodes; a perturbed protein geometry; and a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes; and one or more stacks of invariant point attention blocks; and processing, using the graph transformer, an input amino acid sequence, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0356] Embodiment 67. The method of Embodiment 66, wherein the set of atomic nodes comprises a set of backbone atom nodes.

[0357] Embodiment 68. The method of Embodiment 67, wherein the set of edges connect the set of backbone atom nodes to the set of amino acid sequence nodes.

[0358] Embodiment 69. The method of Embodiment 67 or 68, wherein the set of edges are generated based on an inclusion probability, the inclusion probability being based on a distance between a backbone atom and a residue in the perturbed protein geometry.

[0359] Embodiment 70. The method of any one of Embodiments 66-69, wherein the graph transformer is configured to process a diffusion time.

[0360] Embodiment 71. The method of Embodiment 70, wherein the diffusion time comprises a random Fourier encoding.

[0361] Embodiment 72. The method of any one of Embodiments 66-71, wherein the set of edges are initialized with a randomly Fourier encoded signed sequence distance between two connected nodes if the two connected nodes are located on the same chain, and zeros if the two connected nodes are located on different chains.

[0362] Embodiment 73. The method of any one of Embodiments 66-72, wherein the perturbed coordinates are sampled from a learned reverse-time SDE.

[0363] Embodiment 74. The method of any one of Embodiments 66-73, wherein the one or more stacks of invariant point attention blocks are configured to compute attention scores on the graph.

[0364] Embodiment 75. The method of any one of Embodiments 66-74, wherein the one or more stacks of invariant point attention blocks are configured to associate each node of the sparsely- connected graph with a plurality of replica coordinate frames.Attorney Docket No.59215-726.601

[0365] Embodiment 76. The method of any one of Embodiments 66-75, wherein the one or more stacks of invariant point attention blocks are configured to output a plurality translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames.

[0366] Embodiment 77. The method of any one of Embodiments 75 or 76, wherein a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation.

[0367] Embodiment 78. A method for processing a representation of a protein, comprising: providing a graph transformer comprising: a sparse graph of the protein, comprising: a set of amino acid sequence nodes; a set of atomic nodes comprising a set of backbone atom nodes; a perturbed protein geometry sampled from a learned reverse-time SDE; and a set of edges sparsely connecting the set of atomic nodes to the set of amino acid sequence nodes, wherein the set of edges are generated based on an inclusion probability function; and one or more stacks of invariant point attention blocks configured to: compute attention scores on the graph; associate each node of the sparsely-connected graph with a plurality of replica coordinate frames, wherein a first replica coordinate frame of the plurality of replica coordinate frames is initialized as a copy of the protein backbone coordinates in frame representation; and output a plurality translation vectors and quaternion variables for updating each subsequent frame in the plurality of replica coordinate frames; processing, using the graph transformer, an input amino acid sequence, a diffusion time, and the perturbed protein geometry, to generate an encoded representation of the protein.

[0368] Embodiment 79. A method for predicting contacts between a protein and a ligand, comprising processing (i) a protein graph, (ii) a ligand graph, and (iii) a set of intermolecular edges connecting the protein graph and the ligand graph, to autoregressively sample the contacts based on a probability distribution output by a neural network that permits a plurality of contact modes.

[0369] Embodiment 80. The method of Embodiment 79, further comprising generating the ligand graph using the graph neural network or the method of any one of Embodiments 1-52.

[0370] Embodiment 81. The method of Embodiment 79, further comprising generating the protein graph using the graph neural network or the method of any one of Embodiments 53-78.

[0371] Embodiment 82. The method of any one of Embodiments 79-81, wherein the neural network outputs the probability distribution based on the (i) the protein graph, (ii) the ligand graph, and (iii) the set of intermolecular edges.

[0372] Embodiment 83. The method of any one of Embodiments 79-82, wherein the neural network comprises one or more neural network blocks.

[0373] Embodiment 84. The method of any one of Embodiments 79-83, wherein the neural network comprises an invariant point attention neural network that processes the protein graph to generate a first set of features.

[0374] Embodiment 85. The method of any one of Embodiments 79-84, wherein the neural network comprises a first multi-head cross-attention neural network that processes the ligand graph toAttorney Docket No.59215-726.601 generate a second set of features, while using edge embeddings of the ligand graph as a relative positional encoding term.

[0375] Embodiment 86. The method of Embodiment 85, wherein the neural network comprises a second multi-head cross-attention neural network that processes the first set of features, the second set of features, and the set of intermolecular edges to generate a third set of features.

[0376] Embodiment 87. The method of Embodiment 86, wherein the neural network comprises a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges, wherein the updated features comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0377] Embodiment 88. A method for predicting contacts between a protein and a ligand, comprising: enumerating a set of intermolecular edges connecting a sparsely-connected protein graph and a ligand graph, wherein the set of intermolecular edges connects residues of the sparsely- connected protein graph and heavy-atoms of the ligand graph; processing, using a neural network, (i) the sparsely-connected protein graph, (ii) the ligand graph, and (iii) the set of intermolecular edges, to predict the contacts, wherein: the sparsely-connected protein graph encodes an amino-acid sequence of the protein, a perturbed geometry of the protein, and a diffusion time; the ligand graph encodes (A) a set of atoms of the ligand, (B) a set of local coordinate frames of the ligand, and (C) a set of stereospecific pairwise embeddings between the set of atoms and the set of local coordinate frames; and the neural network comprises: one or more neural network blocks each comprising: an invariant point attention neural network that processes the sparsely-connected protein graph to generate a first set of features; a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features, while using edge embeddings of the ligand graph as a relative positional encoding term; a second multi-head cross-attention neural network that processes the first set of features, the second set of features, and the pairwise intermolecular edges to generate a third set of features; and a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges, wherein the updated features comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0378] Embodiment 89. The method of any one of Embodiments 79-88, further comprising training the neural network to reduce a difference between (i) the posterior distribution of observed contact maps in training data, and (ii) the predicted contacts.

[0379] Embodiment 90. The method of any one of Embodiments 79-89, further comprising training the neural network to reduce an element-wise difference between (i) the observed contact maps in the training data, and (ii) softmax-transformation of predicted contacts.

[0380] Embodiment 91. The method of any one of Embodiments 79-90, wherein the probability distribution comprises a categorical posterior distribution over a sequence of the protein.Attorney Docket No.59215-726.601

[0381] Embodiment 92. The method of any one of Embodiments 79-91, wherein the processing comprises segmenting and tokenizing the protein graph to generate a first set of patches representing the protein.

[0382] Embodiment 93. The method of any one of Embodiments 79-92, wherein the processing comprises frame sampling the ligand graph to generate a second set of patches representing the ligand.

[0383] Embodiment 94. The method of Embodiment 93, wherein the neural network comprises an intra-patch self-attention mechanism for generating a cross-attention map based on the first set of patches or the second set of patches.

[0384] Embodiment 95. The method of Embodiment 93 or 94, wherein the neural network comprises an inter-patch triangular-gated self-attention mechanism for processing a self-attention map based on the first set of patches and the second set of patches.

[0385] Embodiment 96. The method of any one of Embodiments 79-95, wherein the neural network comprises a graph-attention mechanism for processing a graph attention map based on the set of intermolecular edges.

[0386] Embodiment 97. The method of any one of Embodiments 79-96, wherein the autoregressively sampling comprises generating (i) a plurality of intermediate contact probabilities, (ii) a plurality of intermediate distance histograms, or (iii) both.

[0387] Embodiment 98. The method of Embodiment 97, wherein the autoregressively sampling comprises updating the set of intermolecular edges based on (i) the plurality of intermediate contact probabilities, (ii) the plurality of intermediate distance histograms, or (iii) both.

[0388] Embodiment 99. A method for predicting contacts between a protein and a ligand, comprising: segment tokenizing a protein representation to generate a set of protein representation patches; frame subsampling a ligand representation to generate a set of ligand representation patches; generating an intermolecular representation comprising a plurality of edges that connect a first subset of nodes in the protein representation and a second subset of nodes in the ligand representation; and sampling the contacts from a posterior probability distribution that permits a plurality of modes based on (i) the set of protein representation patches, (ii) the set of ligand representation patches, and (iii) the intermolecular representation, wherein the sampling comprises: using a neural network that comprises (i) an intra-patch attention mechanism between the sparse edges and the dense edges to process the set of protein representation patches and the set of ligand representation patches, (ii) an inter-patch self-attention mechanism to process the set of protein representation patches and the set of ligand representation patches, and (iii) a graph attention mechanism to process the intermolecular representation; autoregressively sampling a plurality of intermediate contact maps for a plurality of iterations, while updating the plurality of edges of the intermolecular representation based on the plurality of intermediate contact maps; and outputting the contacts based on a final contact map of the plurality of intermediate contact maps.Attorney Docket No.59215-726.601

[0389] Embodiment 100. An autoregressive neural network for predicting contacts between a protein and a ligand, comprising: an input comprising (i) protein graph, (ii) a ligand graph, (iii) a set of intermolecular edges connecting the protein graph and the ligand graph; and an output comprising parameters of a probability distribution that permits a plurality of contact modes.

[0390] Embodiment 101. The autoregressive neural network of Embodiment 100, further comprising the graph neural network any one of Embodiments 1-24 to provide the ligand graph.

[0391] Embodiment 102. The autoregressive neural network of Embodiment 100 or 101, further comprising the graph neural network any one of Embodiments 53-65 to provide the protein graph.

[0392] Embodiment 103. The autoregressive neural network of any one of Embodiments 100-102, further comprising one or more neural network blocks.

[0393] Embodiment 104. The autoregressive neural network of any one of Embodiments 100-103, further comprising an invariant point attention neural network that processes the protein graph to generate a first set of features.

[0394] Embodiment 105. The autoregressive neural network of any one of Embodiments 100-104, further comprising a first multi-head cross-attention neural network that processes the ligand graph to generate a second set of features, while using edge embeddings of the ligand graph as a relative positional encoding term.

[0395] Embodiment 106. The autoregressive neural network of Embodiment 105, further comprising a second multi-head cross-attention neural network that processes the first set of features, the second set of features, and the set of intermolecular edges to generate a third set of features.

[0396] Embodiment 107. The autoregressive neural network of Embodiment 106, further comprising a multilayer perceptron that processes the third set of features to generate updated features for the set of intermolecular edges, wherein the updated features comprise the contacts for the last neural network block in the plurality of neural network blocks.

[0397] Embodiment 108. The autoregressive neural network of any one of Embodiments 100-107, wherein the protein graph is segmented and tokenized into a first set of patches representing the protein.

[0398] Embodiment 109. The autoregressive neural network of any one of Embodiments 100-108, wherein the ligand graph is frame-sampled to generate a second set of patches representing the ligand.

[0399] Embodiment 110. The autoregressive neural network of Embodiment 109, further comprising an intra-patch self-attention mechanism for generating a cross-attention map based on the first set of patches or the second set of patches.

[0400] Embodiment 111. The autoregressive neural network of Embodiment 109 or 110, further comprising an inter-patch triangular-gated self-attention mechanism for processing a self-attention map based on the first set of patches and the second set of patches.Attorney Docket No.59215-726.601

[0401] Embodiment 112. The autoregressive neural network of any one of Embodiments 100-111, further comprising a graph-attention mechanism for processing a graph attention map based on the set of intermolecular edges.

[0402] Embodiment 113. The autoregressive neural network of any one of Embodiments 100-112, the output further comprises (i) a plurality of intermediate contact probabilities, (ii) a plurality of intermediate distance histograms, or (iii) both.

[0403] Embodiment 114. The autoregressive neural network of Embodiment 113, configured to autoregressively update the set of intermolecular edges based on (i) the plurality of intermediate contact probabilities, (ii) the plurality of intermediate distance histograms, or (iii) both.

[0404] Embodiment 115. A method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising: sampling an initial geometrical structure of the binding complex from a geometry prior; and denoising, using a machine-learned stochastic differential equation (SDE), the initial geometrical structure to generate the geometrical structure of the binding complex.

[0405] Embodiment 116. A method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising predicting the geometrical structure of the binding complex based on a geometry prior comprising (i) noise structured on a template geometrical structure of the protein and (ii) predicted contacts between the protein and the ligand.

[0406] Embodiment 117. A method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising predicting the geometrical structure of the binding complex based on a sequence representation of the protein and a graph representation of the ligand, and optionally, on a geometry prior comprising predicted contacts between the protein and the ligand.

[0407] Embodiment 118. A method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising: processing (i) a first representation comprising a protein representation and (ii) a second representation comprising a ligand representation to generate a geometry prior comprising (i) noise structured on the geometrical structure of the binding complex and (ii) predicted contacts between the protein and the ligand; and denoising the geometrical structure of the binding complex sampled from the geometry prior.

[0408] Embodiment 119. The method of any one of Embodiments 115-118, wherein the geometry prior is based on a first representation comprising a protein representation.

[0409] Embodiment 120. The method of any one of Embodiments 115-119, wherein the geometry prior is based on a second representation comprising a ligand representation.

[0410] Embodiment 121. The method of any one of Embodiments 118-120, wherein the first representation comprises a protein complex representation of a plurality of proteins.

[0411] Embodiment 122. The method of any one of Embodiments 118-121, wherein the second representation comprises a plurality of ligand representations.Attorney Docket No.59215-726.601

[0412] Embodiment 123. The method of any one of Embodiments 115-122, wherein the geometry prior comprises contacts between the protein and the ligand in the binding complex.

[0413] Embodiment 124. The method of Embodiment 123, further comprising generating the contacts by iteratively sampling binding interface spatial proximity distributions of the binding complex.

[0414] Embodiment 125. The method of Embodiment 123, further comprising generating the contacts using the neural network or the method of any one of Embodiments 79-114.

[0415] Embodiment 126. The method of any one of Embodiments 115-125, wherein the geometry prior comprises an initial geometry of the protein in the binding complex.

[0416] Embodiment 127. The method of Embodiment 126, further comprising generating the initial geometry of the protein using the graph neural network or the method of any one of Embodiments 53-78.

[0417] Embodiment 128. The method of any one of Embodiments 115-127, wherein the geometry prior comprises an initial geometry of the ligand in the binding complex.

[0418] Embodiment 129. The method of Embodiment 128, further comprising generating the initial geometry of the ligand using the graph neural network or the method of any one of Embodiments 1- 52.

[0419] Embodiment 130. The method of Embodiment 116 or 123, wherein the template geometrical structure comprises an experimentally determined geometrical structure.

[0420] Embodiment 131. The method of Embodiment 130, wherein the experimentally determined geometrical structure is determined using X-ray crystallography, nuclear magnetic resonance spectroscopy, or cryo-electron microscopy.

[0421] Embodiment 132. The method of any one of Embodiments 116-131, wherein the template geometrical structure comprises a computationally determined geometrical structure.

[0422] Embodiment 133. The method of Embodiment 132, wherein the computationally determined geometrical structure is determined using an electronic structure calculation, a molecular dynamics simulation, a Monte Carlo simulation, a machine learning model, or any combination thereof.

[0423] Embodiment 134. The method of any one of Embodiments 116-133, wherein the predicting the geometrical structure of the binding complex comprises fixing the geometrical structure of the protein to the template geometrical structure of the protein.

[0424] Embodiment 135. The method of Embodiment 134, wherein the template geometrical structure comprises coordinates of the backbone atoms of the protein.

[0425] Embodiment 136. The method of Embodiment 134, wherein the template geometrical structure comprises coordinates of the atoms of the protein that are distal from the ligand in the binding complex, wherein an atom is distal if it is further than a predetermined distance from the ligand.Attorney Docket No.59215-726.601

[0426] Embodiment 137. The method of any one of Embodiments 116-133, wherein the predicting the geometrical structure of the binding complex comprises allowing the geometrical structure of the protein to depart from the template geometrical structure of the protein.

[0427] Embodiment 138. The method of any one of Embodiments 116-133, wherein the predicting the geometrical structure of the binding complex comprises fixing the geometrical structure of the ligand to a template geometrical structure of the ligand.

[0428] Embodiment 139. The method of any one of Embodiments 116-133, wherein the predicting the geometrical structure of the binding complex comprises allowing the geometrical structure of the ligand to depart from a template geometrical structure of the ligand.

[0429] Embodiment 140. The method of any one of Embodiments 115-139, wherein the binding complex comprises an apoprotein.

[0430] Embodiment 141. The method of any one of Embodiments 115-140, wherein the binding complex comprises a holoprotein.

[0431] Embodiment 142. The method of any one of Embodiments 115-141, further comprising generating a second ligand that is configured to bind to the protein.

[0432] Embodiment 143. The method of Embodiment 142, wherein the generating the second ligand comprises performing a gradient-based design using a differentiable protein sequence and / or a molecular graph generator.

[0433] Embodiment 144. The method of any one of Embodiments 115-143, wherein the geometry prior comprises a finite-time marginal of a SDE configured to inject structured noise into the data distribution.

[0434] Embodiment 145. A method for generating a geometrical structure of a binding complex formed between a protein and a ligand, comprising: performing a E(3)-equivariant forward-time-noising process of a truncated stochastic differential equation (SDE) to ^^ = ^^∗ < ∞ based on acontact map to generate a geometry prior, such that the geometry prior comprises a partially-diffused structured distribution of the geometrical structure of the binding complex, wherein the SDE is configured to: retain information about domain packing of the protein; retain information about ligand binding interfaces of the protein; and erase information about residue-scale local details of the protein; sampling a noisy geometrical structure from the geometry prior; and performing, using a machine-learned reverse-time SDE, a SE(3)-equivariant reverse-time-denoising process on the noisy geometrical structure of the binding complex sampled from the geometry prior, by attracting coordinates of the ligand and the protein based on the contact map, the machine-learned reverse-time SDE comprising: inputs comprising: a protein representation comprising: a set of protein heavy-atom nodes: a set of protein heavy-atom coordinates; a protein residue-wise representation from a protein encoder; a protein atom type; and a first random Fourier encoding of diffusion time step t; a ligand representation comprising: a set of ligand coordinates; a ligand representation from a path-integral graph transformer; and a second random Fourier encoding of the diffusion time step t; a protein-Attorney Docket No.59215-726.601 ligand representation comprising: a protein-ligand graph; outputs comprising a set of displacement vectors for the set of protein heavy-atom coordinates and the set of ligand coordinates.

[0435] Embodiment 146. The method of Embodiment 145, wherein the neural network of the machine-learned reverse-time SDE is trained with a loss function configured to reduce a difference between an observed ground-truth geometrical structure and the denoised geometrical structure of the binding complex.

[0436] Embodiment 147. The method of Embodiment 145 or 146, wherein the SDE retains information about domain packing of the protein in the forward-noising process by diffusing the protein’s backbone atoms around a template protein backbone structure.

[0437] Embodiment 148. The method of any one of Embodiments 145-147, wherein the SDE retains information about ligand binding interfaces of the protein in the forward-noising process by diffusing the ligand atoms around the protein’s residues that are predicted to contact the ligand atoms, the diffusing based on a drift term which is based on a contact map, wherein the contact map comprises predicted contacts between the ligand atoms and the protein’s residues, and wherein the drift term is defined relative to the protein’s residues.

[0438] Embodiment 149. The method of any one of Embodiments 145-148, wherein the SDE erases information about residue-scale local details by diffusing non-backbone atoms of the protein, excluding the protein backbone, around the protein backbone of the binding complex, wherein the drift term of the protein residue coordinates in the SDE is relative to the protein backbone of the binding complex.

[0439] Embodiment 150. The method of any one of Embodiments 145-149, wherein the loss function is time-dependent such that the loss function is normalized by a factor that decreases with reverse-time.

[0440] Embodiment 151. The method of any one of Embodiments 145-150, wherein the protein- ligand representation further comprises: a first set of edges connecting protein heavy-atom nodes and the residue node that the protein atom belongs to; a second set of edges connecting pairs of protein heavy-atom nodes that are within the same residue; a third set of edges connecting pairs of protein heavy-atom nodes that are within a first predetermined distance; and a fourth set of edges connecting protein heavy-atom nodes and ligand atom nodes that are within a second first predetermined distance; wherein the edges in the first, second, third, and fourth set of edges that comprise a protein heavy-atom node are initialized with features that encode (i) whether two nodes of an edge belong to the same residue or the same ligand molecule, and (ii) whether there is a covalent bond between two nodes of the edge, wherein the two nodes are considered be covalently bonded if a distance between the two nodes are less than about an average Van der Waals (VdW) radius of atoms constituting the two nodes.Attorney Docket No.59215-726.601

[0441] Embodiment 152. The method of any one of Embodiments 115-151, wherein the sampling the initial geometrical structure of the binding complex from the geometry prior is based on an inverse temperature parameter.

[0442] Embodiment 153. The method of Embodiment 152, wherein the inverse temperature parameter increases during the sampling process.

[0443] Embodiment 154. The method of any one of Embodiments 115-153, wherein the denoising the initial geometrical structure to generate the geometrical structure of the binding complex is based on an inverse temperature parameter.

[0444] Embodiment 155. The method of Embodiment 154, wherein the denoising comprises annealing the initial geometrical structure.

[0445] Embodiment 156. The method of any one of Embodiments 115-155, further comprising determining a confidence level or an error estimate of the geometrical structure of the binding complex.

[0446] Embodiment 157. The method of Embodiment 156, wherein the determining comprises processing the geometrical structure using autoregressive neural network to generate residue-scale embeddings, wherein the confidence level or the error estimate is based on the residue-scale embeddings.

[0447] Embodiment 158. A method of generating an identification of a ligand predicted to bind with a protein to form a binding complex, comprising: predicting contacts between the ligand and the protein in the binding complex; generating a geometry prior based on the contacts; denoising the geometry prior to generate a structure of the binding complex; and generating a report indicating the identification of the ligand based on the structure of the binding complex.

[0448] Embodiment 159. The method of Embodiment 158, wherein the contacts comprise pairwise proximities among residues of the protein and frames of the graph of the ligand.

[0449] Embodiment 160. The method of Embodiment 158 or 159, wherein the predicting the contacts is based on a sequence of the protein, an embedding of the protein, a graph of the ligand, or any combination thereof.

[0450] Embodiment 161. The method of any one of Embodiments 158-160, wherein the denoising is performed with SE(3) equivariance, reducing temperature, or both.

[0451] Embodiment 162. The method of any one of Embodiments 115-161, further comprising generating a report indicating the geometrical structure of the binding complex.

[0452] Embodiment 163. The method of any one of Embodiments 115-162, further comprising generating a report indicating a representation of the ligand.

[0453] Embodiment 164. A computer-implemented method for identifying a ligand that binds to a protein, comprising: sending instructions to generate an identification of the ligand that binds to the protein to one or more computers, wherein the one or more computers are configured to implementAttorney Docket No.59215-726.601 any one of the methods or the neural networks of Embodiments 1-163 to generate a report indicating the identification; receiving the report from the one or more computers.

[0454] Embodiment 165. A computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted to be executed to implement any one of the methods or the neural networks of Embodiments 1-164.

[0455] Embodiment 166. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to implement any one of the methods or the neural networks of Embodiments 1-164.

[0456] Embodiment 167. A computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to perform any one of the methods or the neural networks of Embodiments 1-164.

[0457] Embodiment 168. A method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: a. processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate a geometry prior; b. sampling an initial geometrical structure of the binding complex based on the geometry prior; and c. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules.

[0458] Embodiment 169. A method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: a. processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate an initial geometrical structure of the binding complex formed by the plurality of macromolecules based on plurality of representations, wherein the plurality of macromolecules comprise at least 32 amino acids and / or nucleotides; and b. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of representations.

[0459] Embodiment 170. A method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules and a plurality of small molecules, comprising: a. processing an input representation comprising a plurality of representations of the plurality of macromolecules and the plurality of small molecules to generate a geometry prior; b. sampling an initial geometrical structure of the binding complex based on the geometry prior; and c. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules and the plurality of small molecules.

[0460] Embodiment 171. The method of any one of Embodiments 168 to 170, wherein the plurality of macromolecules comprise at least 2, 3, 4, 5, or 6 macromolecules.

[0461] Embodiment 172. The method of any one of Embodiments 168 to 171, wherein the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.Attorney Docket No.59215-726.601

[0462] Embodiment 173. The method of any one of Embodiments 168 to 172, wherein the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.

[0463] Embodiment 174. The method of any one of Embodiments 168 to 173, wherein the input representation further comprises a representation of a cofactor of the binding complex.

[0464] Embodiment 175. The method of any one of Embodiments 168 to 174, wherein the plurality of macromolecules comprises a protein, a nucleic acid, or both.

[0465] Embodiment 176. The method of any one of Embodiments 168 to 175, wherein the input representation further comprises a representation of a ligand.

[0466] Embodiment 177. The method of any one of Embodiments 168 to 176, wherein the input representation comprises a plurality of representations of a plurality of ligands.

[0467] Embodiment 178. The method of any one of Embodiments 168 to 177, wherein the input representation comprises at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides.

[0468] Embodiment 179. The method of any one of Embodiments 168 to 178, wherein the input representation comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides.

[0469] Embodiment 180. The method of any one of Embodiments 168 to 179, wherein the processing the input representation comprises segmenting the protein representation into a plurality of segments.

[0470] Embodiment 181. The method of any one of Embodiments 168 to 180, wherein the plurality of segments comprises wherein each segment of the plurality of segments comprises a uniform size

[0471] Embodiment 182. The method of Embodiment 181, wherein each segment in the plurality of segments comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids.

[0472] Embodiment 183. The method of Embodiment 182, wherein each segment in the plurality of segments comprises 8 amino acids.

[0473] Embodiment 184. The method of any one of Embodiments 168 to 183, wherein the processing the input representation comprises predicting a plurality of contacts between a first macromolecule and a second macromolecule of the plurality of macromolecules.

[0474] Embodiment 185. The method Embodiment 184, wherein the predicting the plurality of contacts between the first macromolecule and the second macromolecule is performed by processing the plurality of segments.

[0475] Embodiment 186. The method of any one of Embodiments 168 to 185, wherein the processing the input representation comprises predicting a plurality of contacts between (1) a first macromolecule and a first small molecule, and / or (2) a second macromolecule and the first small molecule or a second small molecule.Attorney Docket No.59215-726.601

[0476] Embodiment 187. The method of any one of Embodiments 184-186, wherein the predicting the plurality of contacts is based on a predetermined set of contacts.

[0477] Embodiment 188. The method of Embodiment 187, wherein the plurality of contacts comprises all or a subset of the predetermined set of contacts.

[0478] Embodiment 189. The method of Embodiment 187 or 188, wherein the predetermined set of contacts is provided by a user.

[0479] Embodiment 190. The method of Embodiment 187 or 188 wherein the predetermined set of contacts is generated from a template structure of the binding complex.

[0480] Embodiment 191. The method of Embodiment 190, wherein the template structure of the binding complex is provided by a user, wherein the predetermined set of contacts specifies a binding site of the macromolecule.

[0481] Embodiment 192. The method of Embodiment 191, wherein the template structure is selected by a user from an array of geometrical structures predicted by a neural network.

[0482] Embodiment 193. The method of any one of Embodiments 190-192, wherein the template structure of the binding complex is generated by molecular mechanics or homology modeling.

[0483] Embodiment 194. The method of any one of Embodiments 168 to 193, wherein the processing the initial geometrical structure to generate the geometrical structure of the binding complex comprises dynamically generating connections between pairs of atoms of the binding complex.

[0484] Embodiment 195. The method of Embodiment 194, wherein the dynamically generating the connections comprises using a mixture of probability distributions to assign the connections between the pairs of atoms.

[0485] Embodiment 196. The method of any one of Embodiments 168 to 195, wherein the total number of connections in the ESDM graph is bounded by a hardware-specific threshold.

[0486] Embodiment 197. The method of any one of Embodiments 168 to 196, wherein the cofactor comprises an organic cofactor or an inorganic cofactor.

[0487] Embodiment 198. The method of Embodiment 197, wherein the organic cofactor comprises flavine, heme, NAD+, thiamin pyrophosphate, pyridoxal phosphate, methylcobalamin, cobalamine, biotin, coenzyme A, tetrahydrofolic acid, menaquinone, ascorbic acid, flavin mononucleotide, flavine adenine dinucleotide, coenzyme F420, Adenosine triphosphate, S-Adenosyl methionine, Coenzyme B, Coenzyme M, Coenzyme Q, Cytidine triphosphate, Glutathione, Lipoamide, Methanofuran, Molybdopterin, a nucleotide sugar, 3'-Phosphoadenosine-5'-phosphosulfate, Tetrahydrobiopterin, Tetrahydromethanopterin, or any combination thereof.

[0488] Embodiment 199. The method of Embodiment 197, wherein the inorganic cofactor comprises a metal ion, metal iron, an iron-sulfur complex, or another transition-metal organometallic complex.Attorney Docket No.59215-726.601

[0489] Embodiment 200. The method of any one of Embodiments 168 to 199, wherein the plurality of macromolecules comprises a post-translation modification or a plurality of amino acids with post- translation modifications.

[0490] Embodiment 201. The method of Embodiment 200, wherein the post-translation modification comprises acylation, alkylation, prenylation, flavination, amination, deamination, carboxylation, decarboxylation, nitrosylation, halogenation, sulfurylation, glutathionylation, oxidation, oxygenation, reduction, ubiquitination, SUMOylation, neddylation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glypiation, glycosylphosphatidylinositol anchor formation, lipoylation, heme functionalization, phosphorylation, phosphopantetheinylation, retinylidene Schiff base formation, diphthamide formation, ethanolamine phosphoglycerol functionalization, hypusine formation, beta-Lysine addition, acetylation, formylation, methylation, amidation, amide bond formation, butyrylation, gamma-carboxylation, glycosylation, polysialylation, malonylation, hydroxylation, iodination, nucleotide addition, phosphate ester formation, phosphoramidate formation, adenylation, uridylylation, propionylation, pyroglutamate formation, gluthathionylation, sulfenylation, sulfinylation, sulfonylation, succinylation, sulfation, glycation, carbonylation, isopeptide bond formation, biotinylation, carbamylation, oxidation, pegylation, citrullination, deamidation, eliminylation, disulfide bond formation, proteolytic cleavage, isoaspartate formation, racemization, protein splicing, chaperone-assisted folding, or any combination thereof.

[0491] Embodiment 202. The method of any one of Embodiments 175 to 201, wherein the nucleic acid comprises a helix, a bulge, a loop, a junction, a stem-loop, a hairpin-loop, a tetraloop, a pseudoknot, a single strand, a double strand, or any combination thereof.

[0492] Embodiment 203. The method of any one of Embodiments 175 to 202, wherein the nucleic acid comprises DNA, RNA, 5mC, 5hmC, 5fC, 5mC, 5hmC, 5fC, 5CaC, 6mA, 4mC, 8-oxoG, Tg, an AP site, DNA ps, or any combination thereof.

[0493] Embodiment 204. The method of any one of Embodiments 168 to 203, wherein the neural network is trained with a training dataset comprising predicted geometrical structures of binding complexes of macromolecules

[0494] Embodiment 205. The method of Embodiment 204, wherein the predicted geometrical structures of binding complexes comprise a plurality of monomers aligned to form the geometric structures.

[0495] Embodiment 206. The method of any one of Embodiments 168 to 205, wherein the neural network is trained using a loss function configured to reduce achiral and steric clashes, wherein the loss function is adjusted with a weighting function that stabilizes the training.

[0496] Embodiment 207. The method of Embodiment 206, wherein the weighting function increases in value as the number of training iterations increases.Attorney Docket No.59215-726.601

[0497] Embodiment 208. The method of any one of Embodiments 168 to 207, wherein the neural network is trained by distributing batches of training data across a plurality of processors based on sizes of the training data to reduce load imbalance between the plurality of processors.

[0498] Embodiment 209. The method of Embodiment 208, wherein the training data comprises at least 100k training samples of geometrical structures of binding complexes.

[0499] Embodiment 210. The method of any one of Embodiments 168 to 209, wherein the neural network comprises at least 100M parameters.

[0500] Embodiment 211. The method of any one of Embodiments 168 to 210, wherein the neural network is further trained with a second training dataset comprising geometrical structures of macromolecules or binding complexes generated from experiments, simulations, or electronic structure calculations.

[0501] Embodiment 212. The method of Embodiment 211, wherein the geometrical structures generated from experiments comprise cryo-EM generated structures, X-ray crystallography generated structures, and / or NMR generated structures.

[0502] Embodiment 213. The method of Embodiment 212, wherein the geometrical structures generated from simulations comprise molecular dynamics simulation or Monte Carlo simulation generated structures.

[0503] Embodiment 214. The method of Embodiment 213, wherein the geometrical structures are generated with homology modeling.

[0504] Embodiment 215. The method of Embodiment 214, wherein the geometrical structures are generated without homology modeling.

[0505] Embodiment 216. The method of Embodiment 215, wherein the geometrical structures generated from electronic structure calculations comprise DFT generated structures.

[0506] Embodiment 217. The method of any one of Embodiments 168 to 216, wherein a computational cost of processing the initial geometrical structure to generate the geometrical structure of the binding complex scales sub-linearly as a function of the number of amino acid residues.

[0507] Embodiment 218. The method of any one of Embodiments 168 to 217, wherein the neural network is configured to provide a mean TM-score accuracy of 0.9 for the geometrical structure of the binding complex.

[0508] Embodiment 219. The method of any one of Embodiments 168 to 218, wherein a ligand root- mean-square deviation of the geometrical structure is below two Angstrom.

[0509] Embodiment 220. The method of any one of Embodiments 168 to 219, wherein the neural network uses site-specific docking.

[0510] Embodiment 221. The method of any one of Embodiments 168 to 220, further comprising generating a confidence score for each amino acid in the geometrical structureAttorney Docket No.59215-726.601

[0511] Embodiment 222. The method of Embodiment 221, wherein the confidence score is above a threshold confidence value.

[0512] Embodiment 223. The method of Embodiment 222, wherein the threshold confidence value is 0.8.

[0513] Embodiment 224. The method of any one of Embodiments 168-223, further comprising processing the geometrical structure of the binding complex using molecular mechanics.

[0514] Embodiment 225. The method of Embodiment 224, wherein the molecular mechanics comprises molecular dynamics.

[0515] Embodiment 226. The method of Embodiment 224 or 225, wherein the molecular mechanics comprises electronic structure calculations.

[0516] Embodiment 227. The method of any one of Embodiments 224-226, wherein the molecular mechanics are performed with restraints or without restraints.

[0517] Embodiment 228. The method of Embodiment 227, wherein the restraints are imposed on (i) covalent bonds formed with a hydrogen atom, (ii) torsional modes of the binding complex, or (iii) both.

[0518] Embodiment 229. The method of Embodiment 227 or 228, wherein the restraints are imposed based on a predicted geometrical structure of the binding complex..

[0519] Embodiment 230. The method of any one of Embodiments 168-229, further comprising predicting experimental binding affinity between at least two molecules, wherein the at least two molecules comprises (a) at least two macromolecules, or (b) at least one macromolecule and at least one small molecule.

[0520] Embodiment 231. The method of Embodiment 230, wherein the input representation comprises a sequence of a macromolecule.

[0521] Embodiment 232. The method of Embodiment 231, wherein the sequence is an amino acid sequence of a protein.

[0522] Embodiment 233. The method of Embodiment 232, wherein the sequence is a nucleic acid sequence of a nucleic acid.

[0523] Embodiment 234. The method of any one of Embodiments 231-234, wherein the input representation comprises a 3-dimensional structure of a molecule.

[0524] Embodiment 235. The method of Embodiment 234, wherein the 3-dimensional structure is of a protein, nucleic acid, or a small molecule.

[0525] Embodiment 236. The method of Embodiment 234 or 235, wherein the input representation comprises a sequence and a 3-dimensional structure of one or more macromolecules.

[0526] Embodiment 237. A method for generating a geometrical structure of a binding complex formed between a macromolecule and one or more of (1) a cofactor, or (2) a ligand of a post translational modification, comprising: a. processing an input representation comprising a representation of the macromolecule and a representation of the cofactor or the ligand of the postAttorney Docket No.59215-726.601 translational modification to generate a geometry prior; b. sampling an initial geometrical structure of the binding complex based on the geometry prior; and c. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the macromolecule and the cofactor or the ligand of the post translational modification.

[0527] Embodiment 238. The method of Embodiment 237, wherein the input representation further comprises a representation of a second macromolecule.

[0528] Embodiment 239. The method of Embodiment 238, wherein a plurality of macromolecules comprises the macromolecule and the second macromolecule.

[0529] Embodiment 240. The method of Embodiment 239, wherein a plurality of representations comprises a representation of each macromolecule of the plurality of macromolecules.

[0530] Embodiment 241. The method of Embodiment 240, wherein the plurality of macromolecules comprise at least 2, 3, 4, 5, or 6 macromolecules.

[0531] Embodiment 242. The method of Embodiment 241, wherein the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.

[0532] Embodiment 243. The method of any one of Embodiments 237 to 242, wherein the input representation comprises the representation of a cofactor of the binding complex.

[0533] Embodiment 244. The method of any one of Embodiments 237 to 243, wherein the macromolecule comprises a protein, a nucleic acid, or both.

[0534] Embodiment 245. The method of any one of Embodiments 237 to 244, wherein the input representation comprises the representation of a ligand of a post translational modification.

[0535] Embodiment 246. The method of any one of Embodiments 237 to 244, wherein the input representation comprises a plurality of representations of a plurality of ligands of post translational modifications.

[0536] Embodiment 247. The method of any one of Embodiments 237 to 246, wherein the input representation comprises at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides.

[0537] Embodiment 248. The method of any one of Embodiments 237 to 247, wherein the input representation comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides.

[0538] Embodiment 249. The method of any one of Embodiments 237 to 248, wherein the processing the input representation comprises segmenting the representation of the macromolecule into a plurality of segments.

[0539] Embodiment 250. The method of Embodiment 249, wherein the plurality of segments comprises wherein each segment of the plurality of segments comprises a uniform size

[0540] Embodiment 251. The method of Embodiment 250, wherein each segment in the plurality of segments comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids.Attorney Docket No.59215-726.601

[0541] Embodiment 252. The method of Embodiment 251, wherein each segment in the plurality of segments comprises 8 amino acids.

[0542] Embodiment 253. The method of any one of Embodiments 239 to 252, wherein the processing the input representation comprises predicting a plurality of contacts between the first macromolecule and the second macromolecule of the plurality of macromolecules.

[0543] Embodiment 254. The method of Embodiment 253, wherein the predicting the plurality of contacts between the first macromolecule and the second macromolecule is performed by processing the plurality of segments.

[0544] Embodiment 255. The method of any one of Embodiments 237 to 254, wherein the processing the input representation comprises predicting a plurality of contacts between (1) the macromolecule and the cofactor or ligand of the post translational modification.

[0545] Embodiment 256. The method of any one of Embodiments 253-255, wherein the predicting the plurality of contacts is based on a predetermined set of contacts.

[0546] Embodiment 257. The method of Embodiment 256, wherein the plurality of contacts comprises all or a subset of the predetermined set of contacts.

[0547] Embodiment 258. The method of Embodiment 256 or 257, wherein the predetermined set of contacts is provided by a user.

[0548] Embodiment 259. The method of Embodiment 256 or 257, wherein the predetermined set of contacts is generated from a template structure of the binding complex.

[0549] Embodiment 260. The method of Embodiment 259, wherein the template structure of the binding complex is provided by a user, wherein the predetermined set of contacts specifies a binding site of the ligand or the cofactor.

[0550] Embodiment 261. The method of Embodiment 260, wherein the template structure is selected by a user from an array of geometrical structures predicted by a neural network.

[0551] Embodiment 262. The method of any one of Embodiments 259 to 261, wherein the template structure of the binding complex is generated by molecular mechanics or homology modeling.

[0552] Embodiment 263. The method of any one of Embodiments 237 to 262, wherein the processing the initial geometrical structure to generate the geometrical structure of the binding complex comprises dynamically generating connections between pairs of atoms of the binding complex.

[0553] Embodiment 264. The method of Embodiment 263, wherein the dynamically generating the connections comprises using a mixture of probability distributions to assign the connections between the pairs of atoms.

[0554] Embodiment 265. The method of any one of Embodiments 237 to 264, wherein the total number of connections in an ESDM graph associated with the binding complex is bounded by a hardware-specific threshold.Attorney Docket No.59215-726.601

[0555] Embodiment 266. The method of any one of Embodiments 237 to 265, wherein the cofactor comprises an organic cofactor or an inorganic cofactor.

[0556] Embodiment 267. The method of Embodiment 266, wherein the organic cofactor comprises flavine, heme, NAD+, thiamin pyrophosphate, pyridoxal phosphate, methylcobalamin, cobalamine, biotin, coenzyme A, tetrahydrofolic acid, menaquinone, ascorbic acid, flavin mononucleotide, flavine adenine dinucleotide, coenzyme F420, Adenosine triphosphate, S-Adenosyl methionine, Coenzyme B, Coenzyme M, Coenzyme Q, Cytidine triphosphate, Glutathione, Lipoamide, Methanofuran, Molybdopterin, a nucleotide sugar, 3'-Phosphoadenosine-5'-phosphosulfate, Tetrahydrobiopterin, Tetrahydromethanopterin, or any combination thereof.

[0557] Embodiment 268. The method of Embodiment 267, wherein the inorganic cofactor comprises a metal ion, metal iron, an iron-sulfur complex, or another transition-metal organometallic complex.

[0558] Embodiment 269. The method of any one of Embodiments 237 to 268, wherein the post- translation modification comprises acylation, alkylation, prenylation, flavination, amination, deamination, carboxylation, decarboxylation, nitrosylation, halogenation, sulfurylation, glutathionylation, oxidation, oxygenation, reduction, ubiquitination, SUMOylation, neddylation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glypiation, glycosylphosphatidylinositol anchor formation, lipoylation, heme functionalization, phosphorylation, phosphopantetheinylation, retinylidene Schiff base formation, diphthamide formation, ethanolamine phosphoglycerol functionalization, hypusine formation, beta-Lysine addition, acetylation, formylation, methylation, amidation, amide bond formation, butyrylation, gamma-carboxylation, glycosylation, polysialylation, malonylation, hydroxylation, iodination, nucleotide addition, phosphate ester formation, phosphoramidate formation, adenylation, uridylylation, propionylation, pyroglutamate formation, gluthathionylation, sulfenylation, sulfinylation, sulfonylation, succinylation, sulfation, glycation, carbonylation, isopeptide bond formation, biotinylation, carbamylation, oxidation, pegylation, citrullination, deamidation, eliminylation, disulfide bond formation, proteolytic cleavage, isoaspartate formation, racemization, protein splicing, chaperone-assisted folding, or any combination thereof.

[0559] Embodiment 270. The method of any one of Embodiments 244 to 269, wherein the nucleic acid comprises a helix, a bulge, a loop, a junction, a stem-loop, a hairpin-loop, a tetraloop, a pseudoknot, a single strand, a double strand, or any combination thereof.

[0560] Embodiment 271. The method of any one of Embodiments 244 to 270, wherein the nucleic acid comprises DNA, RNA, 5mC, 5hmC, 5fC, 5mC, 5hmC, 5fC, 5CaC, 6mA, 4mC, 8-oxoG, Tg, an AP site, DNA ps, or any combination thereof.

[0561] Embodiment 272. The method of any one of Embodiments 237 to 271, wherein the neural network is trained with a training dataset comprising predicted geometrical structures of binding complexes of macromoleculesAttorney Docket No.59215-726.601

[0562] Embodiment 273. The method of Embodiment 272, wherein the predicted geometrical structures of binding complexes comprise a plurality of monomers aligned to form the geometric structures.

[0563] Embodiment 274. The method of any one of Embodiments 237 to 273, wherein the neural network is trained using a loss function configured to reduce achiral and steric clashes, wherein the loss function is adjusted with a weighting function that stabilizes the training.

[0564] Embodiment 275. The method of Embodiment 274, wherein the weighting function increases in value as the number of training iterations increases.

[0565] Embodiment 276. The method of any one of Embodiments 237 to 275, wherein the neural network is trained by distributing batches of training data across a plurality of processors based on sizes of the training data to reduce load imbalance between the plurality of processors.

[0566] Embodiment 277. The method of Embodiment 276, wherein the training data comprises at least 100k training samples of geometrical structures of binding complexes.

[0567] Embodiment 278. The method of any one of Embodiments 237 to 277, wherein the neural network comprises at least 100M parameters.

[0568] Embodiment 279. The method of any one of Embodiments 237 to 278, wherein the neural network is further trained with a second training dataset comprising geometrical structures of macromolecules or binding complexes generated from experiments, simulations, or electronic structure calculations.

[0569] Embodiment 280. The method of Embodiment 279, wherein the geometrical structures generated from experiments comprise cryo-EM generated structures, X-ray crystallography generated structures, and / or NMR generated structures.

[0570] Embodiment 281. The method of Embodiment 280, wherein the geometrical structures generated from simulations comprise molecular dynamics simulation or Monte Carlo simulation generated structures.

[0571] Embodiment 282. The method of Embodiment 281, wherein the geometrical structures are generated with homology modeling.

[0572] Embodiment 283. The method of Embodiment 281, wherein the geometrical structures are generated without homology modeling.

[0573] Embodiment 284. The method of Embodiment 280, wherein the geometrical structures generated from electronic structure calculations comprise DFT generated structures.

[0574] Embodiment 285. The method of any one of Embodiments 237 to 284, wherein a computational cost of processing the initial geometrical structure to generate the geometrical structure of the binding complex scales sub-linearly as a function of the number of amino acid residues.Attorney Docket No.59215-726.601

[0575] Embodiment 286. The method of any one of Embodiments 237 to 285, wherein the neural network is configured to provide a mean TM-score accuracy of 0.9 for the geometrical structure of the binding complex.

[0576] Embodiment 287. The method of any one of Embodiments 237 to 286, wherein a ligand root- mean-square deviation of the geometrical structure is below two Angstrom.

[0577] Embodiment 288. The method of any one of Embodiments 237 to 287, wherein the neural network uses site-specific docking.

[0578] Embodiment 289. The method of any one of Embodiments 237 to 288, further comprising generating a confidence score for each amino acid in the geometrical structure

[0579] Embodiment 290. The method of Embodiment 289, wherein the confidence score is above a threshold confidence value.

[0580] Embodiment 291. The method of Embodiment 290, wherein the threshold confidence value is 0.8.

[0581] Embodiment 292. The method of any one of Embodiments 237-291, further comprising processing the geometrical structure of the binding complex using molecular mechanics.

[0582] Embodiment 293. The method of Embodiment 292, wherein the molecular mechanics comprises molecular dynamics.

[0583] Embodiment 294. The method of Embodiment 292 or 293, wherein the molecular mechanics comprises electronic structure calculations.

[0584] Embodiment 295. The method of any one of Embodiments 292-294, wherein the molecular mechanics performed with restraints or without restraints.

[0585] Embodiment 296. The method of Embodiment 295, restraints are imposed on (i) covalent bonds formed with a hydrogen atom, (ii) torsional modes of the binding complex, or (iii) both.

[0586] Embodiment 297. The method of Embodiment 295 or 296, wherein the restraints are imposed based on a predicted geometrical structure of the binding complex..

[0587] Embodiment 298. The method of any one of Embodiments 237-297, further comprising predicting experimental binding affinity between at least two molecules, wherein the at least two molecules comprises (a) at least two macromolecules, or (b) at least one macromolecule and at least one small molecule.

[0588] Embodiment 299. The method of Embodiment 298, wherein the input representation comprises a sequence of a macromolecule.

[0589] Embodiment 300. The method of Embodiment 299, wherein the sequence is an amino acid sequence of a protein.

[0590] Embodiment 301. The method of Embodiment 299, wherein the sequence is a nucleic acid sequence of a nucleic acid.Attorney Docket No.59215-726.601

[0591] Embodiment 302. The method of any one of Embodiments 298-301, wherein the input representation comprises a 3-dimensional structure of a molecule.

[0592] Embodiment 303. The method of Embodiment 302, wherein the 3-dimensional structure is of a protein, nucleic acid, or a small molecule.

[0593] Embodiment 304. The method of Embodiment 302 or 303, wherein the input representation comprises a sequence and a 3-dimensional structure of one or more macromolecules.

[0594] Embodiment 305. A computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted to be executed to implement any one of the methods of Embodiments 168-305.

[0595] Embodiment 306. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to implement any one of the methods of Embodiments 168-305.

[0596] Embodiment 307. A computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to perform any one of the methods of Embodiments 168-305.

[0597] Embodiment 308. A method comprising: (a) extracting a first slice of a first tensor and a second slice of a second tensor of a plurality of tensors; (b) broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof; (c) performing a tensor operation between the first slice and the second slice; and (d) repeating (a)-(c) with other slices of the first tensor and the second tensor to compute the tensor operation between the first tensor and the second tensor.

[0598] Embodiment 309. The method of Embodiment 308, further comprising obtaining gradients of a result of the tensor operation with respect to the first slice and the second slice.

[0599] Embodiment 310. The method of Embodiment 308 or 309, further comprising obtaining gradients of a result of the tensor operation with respect to the first tensor and the second tensor.

[0600] Embodiment 311. The method of any one of Embodiments 308-310, further comprising storing the first tensor, the second tensor, or both in a first memory.

[0601] Embodiment 312. The method of any one of Embodiments 308-311, further comprising storing the first slice, the second slice, or both in a second memory.

[0602] Embodiment 313. The method of Embodiment 312, wherein the second memory is faster than the first memory.

[0603] Embodiment 314. The method of Embodiment 312 or 313, wherein the second memory has smaller size than the first memory.

[0604] Embodiment 315. The method of any one of Embodiments 312-314, wherein the second memory is more physically localized to transistors of a processing unit than the first memory.Attorney Docket No.59215-726.601

[0605] Embodiment 316. The method of Embodiment 315, wherein the processing unit is a graphical processing unit (GPU).

[0606] Embodiment 317. The method of any one of Embodiments 308-316, wherein the method is performed on a graphical processing unit (GPU).

[0607] Embodiment 318. The method of any one of Embodiments 311-317, wherein the first memory has a speed of at least 1.5 TB / s

[0608] Embodiment 319. The method of any one of Embodiments 311-318, wherein the first memory has a size of at least 40 GB

[0609] Embodiment 320. The method of any one of Embodiments 311-319, wherein the first memory is a high bandwidth memory (HBM).

[0610] Embodiment 321. The method of any one of Embodiments 312-320, wherein the second memory has a speed of at least 19 TB / s.

[0611] Embodiment 322. Th...

Claims

Attorney Docket No.59215-726.601 CLAIMS What is claimed is:

1. A method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: a. processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate a geometry prior; b. sampling an initial geometrical structure of the binding complex based on the geometry prior; and c. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules.

2. A method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules, comprising: a. processing an input representation comprising a plurality of representations of the plurality of macromolecules to generate an initial geometrical structure of the binding complex formed by the plurality of macromolecules based on plurality of representations, wherein the plurality of macromolecules comprise at least 32 amino acids and / or nucleotides; and b. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of representations.

3. A method for generating a geometrical structure of a binding complex formed between a plurality of macromolecules and a plurality of small molecules, comprising: a. processing an input representation comprising a plurality of representations of the plurality of macromolecules and the plurality of small molecules to generate a geometry prior; b. sampling an initial geometrical structure of the binding complex based on the geometry prior; andAttorney Docket No.59215-726.601 c. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the plurality of macromolecules and the plurality of small molecules.

4. The method of any one of claims 1 to 3, wherein the plurality of macromolecules comprise at least 2, 3, 4, 5, or 6 macromolecules.

5. The method of any one of claims 1 to 4, wherein the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.

6. The method of any one of claims 1 to 5, wherein the plurality of representations comprise at least 2, 3, 4, 5, or 6 representations.

7. The method of any one of claims 1 to 6, wherein the input representation further comprises a representation of a cofactor of the binding complex.

8. The method of any one of claims 1 to 7, wherein the plurality of macromolecules comprises a protein, a nucleic acid, or both.

9. The method of any one of claims 1 to 8, wherein the input representation further comprises a representation of a ligand.

10. The method of any one of claims 1 to 9, wherein the input representation comprises a plurality of representations of a plurality of ligands.

11. The method of any one of claims 1 to 10, wherein the input representation comprises at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 amino acids or nucleotides.

12. The method of any one of claims 1 to 11, wherein the input representation comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or 30000 amino acids or nucleotides.

13. The method of any one of claims 1 to 12, wherein the processing the input representation comprises segmenting the protein representation into a plurality of segments.

14. The method of any one of claims 1 to 13, wherein the plurality of segments comprises wherein each segment of the plurality of segments comprises a uniform size 15. The method of claim 14, wherein each segment in the plurality of segments comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids.Attorney Docket No.59215-726.601 16. The method of claim 15, wherein each segment in the plurality of segments comprises 8 amino acids.

17. The method of any one of claims 1 to 16, wherein the processing the input representation comprises predicting a plurality of contacts between a first macromolecule and a second macromolecule of the plurality of macromolecules.

18. The method claim 17, wherein the predicting the plurality of contacts between the first macromolecule and the second macromolecule is performed by processing the plurality of segments.

19. The method of any one of claims 1 to 18, wherein the processing the input representation comprises predicting a plurality of contacts between (1) a first macromolecule and a first small molecule, and / or (2) a second macromolecule and the first small molecule or a second small molecule.

20. The method of any one of claims 17 to 19, wherein the predicting the plurality of contacts is based on a predetermined set of contacts.

21. The method of claim 20, wherein the plurality of contacts comprises all or a subset of the predetermined set of contacts.

22. The method of claim 20 or 21, wherein the predetermined set of contacts is provided by a user.

23. The method of claim 20 or 21 wherein the predetermined set of contacts is generated from a template structure of the binding complex.

24. The method of claim 23, wherein the template structure of the binding complex is provided by a user, wherein the predetermined set of contacts specifies a binding site of the macromolecule.

25. The method of claim 24, wherein the template structure is selected by a user from an array of geometrical structures predicted by a neural network.

26. The method of any one of claims 23 to 25, wherein the template structure of the binding complex is generated by molecular mechanics or homology modeling.

27. The method of any one of claims 1 to 26, wherein the processing the initial geometrical structure to generate the geometrical structure of the binding complex comprises dynamically generating connections between pairs of atoms of the binding complex.Attorney Docket No.59215-726.601 28. The method of claim 27, wherein the dynamically generating the connections comprises using a mixture of probability distributions to assign the connections between the pairs of atoms.

29. The method of any one of claims 1 to 28, wherein the total number of connections in the ESDM graph is bounded by a hardware-specific threshold.

30. The method of any one of claims 1 to 29, wherein the cofactor comprises an organic cofactor or an inorganic cofactor.

31. The method of claim 30, wherein the organic cofactor comprises flavine, heme, NAD+, thiamin pyrophosphate, pyridoxal phosphate, methylcobalamin, cobalamine, biotin, coenzyme A, tetrahydrofolic acid, menaquinone, ascorbic acid, flavin mononucleotide, flavine adenine dinucleotide, coenzyme F420, Adenosine triphosphate, S-Adenosyl methionine, Coenzyme B, Coenzyme M, Coenzyme Q, Cytidine triphosphate, Glutathione, Lipoamide, Methanofuran, Molybdopterin, a nucleotide sugar, 3'-Phosphoadenosine-5'-phosphosulfate, Tetrahydrobiopterin, Tetrahydromethanopterin, or any combination thereof.

32. The method of claim 30, wherein the inorganic cofactor comprises a metal ion, metal iron, an iron-sulfur complex, or another transition-metal organometallic complex.

33. The method of any one of claims 1 to 32, wherein the plurality of macromolecules comprises a post-translation modification or a plurality of amino acids with post-translation modifications.

34. The method of claim 33, wherein the post-translation modification comprises acylation, alkylation, prenylation, flavination, amination, deamination, carboxylation, decarboxylation, nitrosylation, halogenation, sulfurylation, glutathionylation, oxidation, oxygenation, reduction, ubiquitination, SUMOylation, neddylation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glypiation, glycosylphosphatidylinositol anchor formation, lipoylation, heme functionalization, phosphorylation, phosphopantetheinylation, retinylidene Schiff base formation, diphthamide formation, ethanolamine phosphoglycerol functionalization, hypusine formation, beta-Lysine addition, acetylation, formylation, methylation, amidation, amide bond formation, butyrylation, gamma-carboxylation, glycosylation, polysialylation, malonylation, hydroxylation, iodination, nucleotide addition, phosphate ester formation, phosphoramidate formation, adenylation, uridylylation, propionylation, pyroglutamateAttorney Docket No.59215-726.601 formation, gluthathionylation, sulfenylation, sulfinylation, sulfonylation, succinylation, sulfation, glycation, carbonylation, isopeptide bond formation, biotinylation, carbamylation, oxidation, pegylation, citrullination, deamidation, eliminylation, disulfide bond formation, proteolytic cleavage, isoaspartate formation, racemization, protein splicing, chaperone-assisted folding, or any combination thereof.

35. A method for generating a geometrical structure of a binding complex formed between a macromolecule and one or more of (1) a cofactor, or (2) a ligand of a post translational modification, comprising: a. processing an input representation comprising a representation of the macromolecule and a representation of the cofactor or the ligand of the post translational modification to generate a geometry prior; b. sampling an initial geometrical structure of the binding complex based on the geometry prior; and c. processing, using a neural network, the initial geometrical structure to generate the geometrical structure of the binding complex formed by the macromolecule and the cofactor or the ligand of the post translational modification.

36. A method comprising: a. extracting a first slice of a first tensor and a second slice of a second tensor of a plurality of tensors; b. broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof; c. performing a tensor operation between the first slice and the second slice; and d. repeating (a)-(c) with other slices of the first tensor and the second tensor to compute the tensor operation between the first tensor and the second tensor.

37. A method comprising: a. storing a plurality of tensors comprising a first tensor and a second tensor in a first memory of a processing unit; and b. computing a tensor operation between slices of the first tensor and the second tensor by: i. dynamically allocating the slices of the tensors by: 1) extracting a first slice of a first tensor and a second slice of a second tensor; 2) broadcasting the first slice to match a rank and at least one dimension of the second slice or a transpose thereof; andAttorney Docket No.59215-726.601 3) storing the first slice and the second slice in a second memory of the processing unit, wherein the second memory is faster than the first memory, and wherein the second memory is more physically localized to transistors of a processing unit than the first memory; ii. performing the tensor operation between the first slice and the second slice; and iii. repeating (i) and (ii) with other slices of the first tensor and the second tensor to complete the tensor operation between the first tensor and the second tensor.

38. A method of computing a triangular attention tensor, comprising: a. extracting a plurality of slices from a plurality of tensors, wherein the plurality of slices comprises a first slice of a bias tensor of the plurality of tensors, a second slice of a query tensor of the plurality of tensors, a third slice of a key tensor of the plurality of tensors, and a fourth slice of a value tensor of the plurality of tensors; b. broadcasting the first slice to match a rank and a shape of the second slice, the third slice, the fourth slice, or a transpose of the second slice, the third slice, or the fourth slice; c. performing tensor operations using the plurality of slices; and d. repeating (a)-(c) with other slices of the plurality of tensors to compute the triangular attention tensor.

39. An active programming interface comprising a callable function configured to program a processor to perform operations comprising: a. extracting a first slice of a first tensor and a second slice of a second tensor; b. broadcasting the first slice to match the rank and at least one dimension of the second slice or a transpose thereof; c. performing the tensor operation between the first slice and the second slice; and d. repeating (a)-(c) with different slices to compute the tensor operation between the first tensor and the second tensor.

40. A method of computing a tensor operation between tensors of different ranks by dynamically allocating slices of the tensors and computing the tensor operation between the slices.

41. A method for generating a geometrical structure of a molecule, comprising: a. processing a geometry prior of the molecule to obtain a prediction of the geometrical structure; b. processing the prediction and the geometry prior to obtain a vector field; andAttorney Docket No.59215-726.601 c. generating the geometrical structure of the molecule based on the vector field.

42. A method for generating a geometrical structure of a binding complex formed between a plurality of molecules, comprising: a. processing a geometry prior of a plurality of molecules to obtain a prediction of the geometrical structure of the binding complex; b. processing the prediction and the geometry prior to obtain a vector field; and c. generating the geometrical structure of the binding complex based on the vector field.

43. The method of claim 41 or 42, wherein the vector field is time independent or time dependent.

44. The method of any one of claims 41 to 43, wherein the molecule or the plurality of molecules comprises small molecule ligands, proteins, peptides, RNA molecules, DNA molecules, organometallics, glycans, or any combination thereof.

45. The method of any one of claims 41 to 44, wherein the generating the geometrical structure is performed while maintaining optimal transport.

46. The method of any one of claims 41 to 45, wherein the generating the geometrical structure is performed while maintaining SE(3) superposition.

47. The method of claim 46, wherein the SE(3) superposition is maintained by applying a Kabsch transform to the geometrical structure.

48. The method of any one of claims 41 to 47, wherein the generating the geometrical structure is performed while optimizing transport cost.

49. The method of any one of claims 41 to 48, wherein the generating the geometrical structure is performed while maintaining identities of the particles.

50. The method of claim 49, wherein the identities of the particles are maintained by obtaining a graph permutation of a representation of the molecule or the plurality of molecules.

51. The method of any one of claims 41 to 50, further comprising generating the geometry prior.

52. The method of any one of claims 41 to 51, wherein the geometry prior is physics- based.

53. A method of generating a geometrical structure of a molecule, comprising: a. generating a current geometrical structure of the molecule based on a conditioning signal; b. generating a latent variable by encoding the conditioning signal; c. updating the current geometrical structure of the molecule by:Attorney Docket No.59215-726.601 i. obtaining a predicted geometrical structure of the molecule based on the latent variable, a time variable, and the current geometrical structure; ii. applying a global SE(3) superposition to the prediction by aligning the predicted geometrical structure to the current geometrical structure, and wherein the aligning is performed by applying a rotation to the predicted geometrical structure; iii. obtaining an optimal graph permutation of the predicted geometrical structure to find a consistent ordering of nodes between the predicted geometrical structure and the current geometrical structure; and iv. updating the current geometrical structure by interpolating between the current geometrical structure and the predicted geometrical structure; d. repeating (c) until an output geometrical structure is obtained.

54. A non-transitory computer-readable medium comprising executable instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-53.

55. A computer system comprising: a memory comprising executable instructions; and at least one processor configured to execute the instructions, wherein when the at least one processor executes the instructions, the at least one processor causes the system to perform the method according to any one of claims 1-53.

Citation Information

Patent Citations

  • Target molecule-ligand binding mode prediction combining deep learning-based informatics with molecular docking

    US20200342953A1

  • Systems for enteric delivery of therapeutic agents

    US20210196627A1

  • Topical Compositions and Methods to Promote Optimal Dermal White Adipose Tissue Composition in Vivo

    US20210401728A1

  • Adversarial framework for molecular conformation space modeling in internal coordinates

    US20220406404A1

Cited By

  • Double-strategy screening method and system for active ingredients and targets of medicinal and edible food

    CN121641269A

  • Transform-based metal damper anti-seismic performance prediction method and prediction equipment

    CN121835312A

  • A Transformer-based method and device for predicting the seismic performance of metal dampers

    CN121835312B

  • Geometric attention for biological language reasoning

    US12712051B2

  • Geometric attention for biological language reasoning

    US20250378912A1