Machine learning enables prediction of molecular structure and properties

By employing a computational model that modifies structural representations of proteins, the method addresses the computational inefficiencies in protein design, enabling the generation of functional proteins with desired properties efficiently.

JP2025533582APending Publication Date: 2025-10-07GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517891
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-17
Filing Date
2023-09-27
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Current methods for protein design face a computational burden in exploring the vast solution space of amino acid sequences and three-dimensional structures, leading to inefficiencies in generating functional proteins with desired properties.

Method used

A computational model that operates on a structural representation of a molecule, such as a coarse-grained node or backbone torsion angle, applies modifications to reduce the solution space by strategically updating backbone geometry and torsion angles, using machine learning to denoise and refine the molecular structure.

Benefits of technology

This approach efficiently generates protein sequences and structures with desired properties like binding affinity and stability, reducing computational overhead and improving design accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533582000001_ABST
    Figure 2025533582000001_ABST
Patent Text Reader

Abstract

The method may include receiving a molecular structure file specifying an initial three-dimensional structure of the molecule. A representation of the molecule may be determined based on the molecular structure file. For example, the representation of the molecule may include a plurality of coarse-grained nodes, each corresponding to the structure of two or more atoms (e.g., heavy atoms) forming an amino acid residue in the molecule. Alternatively, the representation of the molecule may include a plurality of frames, for each residue of the molecule, specifying the backbone geometry of the residue and one or more torsion angles in the residue's side chain. A design computational model may be applied to determine the three-dimensional structure of the molecule by modifying at least the representation of the molecule. The three-dimensional structure may be associated with desired properties and / or configured for downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 63 / 377,335, entitled "MACHINE LEARNING ENABLED PREDICTION OF MOLECULAR STRUCTURES AND PROPERTIES," filed September 27, 2022; U.S. Provisional Application No. 63 / 387,680, entitled "MACHINE LEARNING ENABLED PREDICTION OF MOLECULAR STRUCTURES AND PROPERTIES," filed December 15, 2022; U.S. Provisional Application No. 63 / 499,333, entitled "MACHINE LEARNING ENABLED PREDICTION OF MOLECULAR STRUCTURES AND PROPERTIES," filed May 1, 2023; and U.S. Provisional Application No. 63 / 499,333, entitled "MACHINE LEARNING ENABLED PREDICTION OF MOLECULAR STRUCTURES AND PROPERTIES," filed May 17, 2023. This application claims priority to U.S. Provisional Application No. 63 / 502,753, entitled "PATENT PROPERTIES," the disclosure of which is incorporated herein by reference in its entirety.

[0002] The subject matter described herein relates generally to molecular design, and more particularly to machine learning-based techniques for predicting molecular structures and properties. [Background technology]

[0003] A molecule is a group of two or more atoms held together by chemical bonds. A molecule forms the smallest identifiable unit that can be divided into a pure substance while still retaining the substance's composition and chemical properties. An example of a molecule is a protein molecule, while examples of non-protein molecules include small molecules, ions, nucleic acids, polysaccharides, glycolipids, and the like. The function and properties of a molecule can depend on its three-dimensional structure. For example, proteins are responsible for many essential cellular functions, including enzymatic reactions, molecular transport, regulation and execution of several biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and the like. A protein's structure can include one or more polypeptides, which are chains of amino acid residues linked together by peptide bonds. The sequence of amino acid residues in the polypeptide chains that form the protein structure determines the protein's three-dimensional structure (e.g., the protein's tertiary structure). Furthermore, the sequence of amino acids in the polypeptide chains that form the protein determines the protein's underlying function. Therefore, one goal of de novo protein design involves constructing one or more sequences of amino acid residues that exhibit desirable properties rather than undesirable ones. For example, in large molecule drug discovery, de novo protein design often seeks to identify sequences of amino acid residues (e.g., antibodies) that can bind to antigens such as viral antigens, tumor antigens, and the like. Summary of the Invention

[0004] Systems, methods, and products, including computer program products, are provided for predicting molecular structure and properties. In one aspect, a system for predicting molecular structure and properties is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, provides operations. The operations may include receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule including a first sequence of amino acid residues; determining a representation of the protein molecule based on at least the molecular structure file, the representation including a plurality of frames for each amino acid residue of the first sequence of amino acid residues, the plurality of frames for each amino acid residue including a first set of frames specifying a backbone geometry of the amino acid residue and a second set of frames specifying one or more torsion angles for a side chain of the amino acid residue; and generating the first three-dimensional structure of the protein molecule by at least applying a design computational model to modify the representation of the protein molecule.

[0005] In another aspect, a method for predicting molecular structures and properties is provided. The method may include receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule including a first sequence of amino acid residues, determining a representation of the protein molecule based on at least the molecular structure file, the representation including a plurality of frames for each amino acid residue of the first sequence of amino acid residues, the plurality of frames for each amino acid residue including a first set of frames specifying a backbone geometry of the amino acid residue and a second set of frames specifying one or more torsion angles for side chains of the amino acid residue, and generating the first three-dimensional structure of the protein molecule by at least applying a design computational model to modify the representation of the protein molecule.

[0006] In another aspect, a computer program product for predicting molecular structures and properties is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations to occur. The operations may include receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule including a first sequence of amino acid residues; determining a representation of the protein molecule based on at least the molecular structure file, the representation including a plurality of frames for each amino acid residue of the first sequence of amino acid residues, the plurality of frames for each amino acid residue including a first set of frames specifying a backbone geometry of the amino acid residue and a second set of frames specifying one or more torsion angles for side chains of the amino acid residue; and generating the first three-dimensional structure of the protein molecule by at least applying a design computational model to modify the representation of the protein molecule.

[0007] In some variations of the methods, systems, and non-transitory computer-readable media, one or more of the following features may optionally be included in any feasible combination.

[0008] In some variations, each frame of the plurality of frames may correspond to a degree of freedom for the design computational model to update the initial three-dimensional structure of the protein molecule.

[0009] In some variations, the first frame set may include a first frame including an affine transformation matrix that specifies the rotation and translation of the backbone of the amino acid residues, and the first frame set may further include a second frame that specifies the torsion angles of the backbone of the amino acid residues.

[0010] In some variations, the first frame set may include a first frame that specifies a first torsion angle for the backbone of the amino acid residues. The first frame set may further include a second frame that specifies a second torsion angle for the backbone of the amino acid residues.

[0011] In some variations, the first torsion angle is the angle between the backbone alpha carbon ( The second torsion angle can be associated with the first rotatable bond between the alpha carbon (C) atom of the backbone of the amino acid residue. TIFF2025533582000003.tif7170) atom and the nitrogen (N) atom.

[0012] In some variations, the first frame set may further include a third frame specifying a third torsion angle present in the backbone of the amino acid residue, the third torsion angle being associated with a third rotatable bond between a carbon (C) atom and a nitrogen (N) atom of the backbone of the amino acid residue.

[0013] In some variations, one or more coordinates of a plurality of backbone atoms of the protein molecule may be determined based at least on a plurality of frames associated with each amino acid residue included in the modified representation of the protein molecule, and one or more coordinates of a plurality of side chain atoms of the protein molecule may be determined based on one or more coordinates of a plurality of backbone atoms of the protein molecule.

[0014] In some variations, the design computational model may include a machine learning model trained to generate a first three-dimensional structure of the protein molecule by denoising at least an initial three-dimensional structure of the protein molecule.

[0015] In some variations, the machine learning model may denoise the initial three-dimensional structure of the protein molecule by performing at least a sequence of updates to the representation of the protein molecule.

[0016] In some variations, the machine learning model may be trained to reduce a loss function and / or energy function associated with each successive update of the initial three-dimensional structure of the protein molecule.

[0017] In some variations, the machine learning model may be a diffusion model that removes some of the noise present in the initial three-dimensional structure of the protein molecule at each time step of a plurality of successive time steps.

[0018] In some variations, the diffusion model may perform a first update to the representation of the protein molecule to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule, and the diffusion model may further perform a second update to the representation of the protein molecule to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule.

[0019] In some variations, the diffusion model may further add a third amount of noise before performing the second update to remove the second amount of noise and a fourth amount of noise after performing the second update to remove the second amount of noise. The third amount of noise and the fourth amount of noise may be determined based on a noise schedule that defines a distribution of noise levels to be added over multiple consecutive time steps.

[0020] In some variations, the distribution of noise levels may correspond to the degrees of freedom present in the representation of the protein molecule in the computational model for modifying the initial three-dimensional structure of the protein molecule.

[0021] In some variations, each update performed by the diffusion model may produce an output equivariant to a special Euclidean group SE(3) transformation.

[0022] In some variations, modifying the representation of the protein molecule may include updating the first frameset to change the backbone geometry of one or more amino acid residues of the protein molecule.

[0023] In some variations, modifying the representation of the protein molecule may include updating the second frameset to change one or more torsion angles of side chains of one or more amino acid residues of the protein molecule.

[0024] In some variations, the first three-dimensional structure of the protein molecule may be associated with one or more desirable properties.

[0025] In some variations, the first three-dimensional structure of the protein molecule may be configured for one or more downstream tasks.

[0026] In some variations, a first sequence of amino acid residues may be determined to exhibit a desired three-dimensional structure and / or desired properties based at least on a first three-dimensional structure of the protein molecule, and in response to determining that the first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties, a second sequence of amino acid residues of a different protein molecule may be generated based at least on the first sequence of amino acid residues.

[0027] In some variations, the representation of the protein molecule may further include a logical vector indicating, for each position in the sequence of amino acid residues forming the protein molecule, the identity of the amino acid residue occupying the position by at least enumerating a probability distribution over the set of possible amino acid residues that may occupy the position.

[0028] In some variations, the design computational model may further generate a first three-dimensional structure of the protein molecule by modifying the identity of at least one amino acid residue in the first residue sequence while modifying a first frameset and / or a second frameset associated with the at least one amino acid residue.

[0029] In some variations, the initial three-dimensional structure of the protein molecule may contain noise in the identity of each amino acid residue and / or the spatial arrangement of the atoms that form each amino acid, which noise can be removed by designing a computational model that modifies the representation of the protein molecule.

[0030] In some variations, the representation of the protein molecule can be further generated to include multiple polymer chains, each of which can include one or more amino acid residues from the first sequence of amino acid residues. The representation of the protein molecule can be modified by a computational protein design model that modifies the positions of one or more amino acids in each of the polymer chains as a group.

[0031] Implementations of the present subject matter may include, but are not limited to, methods according to the descriptions provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described, which may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, etc., one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods according to one or more implementations of the present subject matter may be implemented by one or more data processors in a single computing system or in multiple computing systems. Such multiple computing systems may be connected, for example, via one or more connections, including connections over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via a direct connection between one or more of the multiple computing systems, and may exchange data and / or instructions or other instructions, etc.

[0032] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the subject matter of the present disclosure are described for illustrative purposes in connection with protein design, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure define the scope of the protected subject matter. [Brief explanation of the drawings]

[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed implementations.

[0034] [Figure 1A] 1 depicts a system diagram showing an example of a molecular design system, according to some illustrative embodiments.

[0035] [Figure 1B] 1 depicts a flowchart illustrating an example of a process 160 for predicting protein structure and properties, according to some exemplary embodiments.

[0036] [Figure 2A] 1 depicts an example of a coarse node representation of a protein sequence in which hydrogen (H) atoms are excluded, according to some exemplary embodiments.

[0037] [Figure 2B] 10 depicts another example of a coarse node representation of a protein sequence including hydrogen (H) atoms, according to some exemplary embodiments.

[0038] [Figure 3A]1 depicts a visualization of the relative positions of atoms in a coarse-grained node associated with the amino acid residue tryptophan (Trp), according to some example embodiments.

[0039] [Figure 3B] 1 depicts a visualization of the relative positions of atoms in a coarse-grained node associated with the amino acid residue valine (Val), according to some example embodiments.

[0040] [Figure 4] 1 depicts a visualization of the iterative updates performed by a computational model to determine the three-dimensional structure of a protein molecule, according to some exemplary embodiments.

[0041] [Figure 5A] 1 depicts a flowchart showing an example of a process for predicting protein structure and properties, according to some exemplary embodiments.

[0042] [Figure 5B] 1 depicts a flowchart showing another example of a process for predicting protein structure and properties, according to some exemplary embodiments.

[0043] [Figure 5C] 1 depicts a flowchart showing another example of a process for predicting protein structure and properties, according to some exemplary embodiments.

[0044] [Figure 6A] 1 depicts a schematic diagram showing amino acid structures of example amino acid residues, according to some exemplary embodiments;

[0045] [Figure 6B] 1 depicts a schematic diagram illustrating an example of a diffusion framework, according to some exemplary embodiments;

[0046] [Figure 7A] 1 depicts a screenshot showing an example of a protein molecule undergoing backbone translation, according to some exemplary embodiments.

[0047] [Figure 7B] 1 depicts a screenshot showing an example of a protein molecule undergoing backbone rotation, according to some illustrative embodiments.

[0048] [Figure 7C] 10 depicts a screenshot showing another example of a protein molecule undergoing side chain torsion angle changes, according to some exemplary embodiments.

[0049] [Figure 8A] 10 depicts a graph showing an example of a noise schedule for a diffusion model that modifies torsion angles of a protein structure, according to some exemplary embodiments;

[0050] [Figure 8B] 10 depicts a graph showing another example of a noise schedule for a diffusion model that modifies all degrees of freedom of a protein structure except for the center of mass, according to some exemplary embodiments.

[0051] [Figure 8C] 10 depicts a graph showing another example of a noise schedule for a diffusion model that performs molecular docking between two molecules by modifying the respective centers of mass of the two molecules, according to some exemplary embodiments;

[0052] [Figure 8D] 10 depicts a graph showing another example of a noise schedule for a diffusion model performing molecular docking between two molecules by modifying the center of mass.

[0053] [Figure 9] 1 depicts a block diagram illustrating an example of a computing system, according to some illustrative embodiments.

[0054] Wherever practical, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION OF THE INVENTION

[0055] The properties of a molecule can depend on its composition and structure. For example, the properties of a small molecule, including its safety and effectiveness as a therapeutic agent, can depend on the atomic amount of each constituent element (e.g., molecular formula or empirical formula) and the bond arrangement between these atoms (e.g., structural formula). On the other hand, the properties of a large molecule, such as the binding affinity and developability of a protein molecule, can be determined by the sequence of amino acid residues that form the large molecule and the three-dimensional structure that the amino acid residue sequence adopts. Therefore, developing small and large molecule therapeutic agents involves the composition and structure of each candidate molecule.

[0056] In the case of protein design, where a key objective involves identifying a protein sequence (e.g., a sequence of amino acid residues) that exhibits certain desirable properties, insight into the properties of the protein molecule may be limited without first determining its three-dimensional structure. Thus, the task of protein structure prediction, which involves inferring the three-dimensional structure of a protein molecule based on the arrangement of amino acid sequences that form the protein molecule, may be an important component of protein design. In contrast, structurally agnostic prediction of protein sequences may result in many protein sequences that are unable to adopt a three-dimensional structure capable of binding to a target molecule. In this context, the sequence of amino acid residues that form a protein molecule may also be known as the primary structure of the protein molecule. The three-dimensional structure of a protein molecule includes one or more secondary structures formed when individual amino acid residues are connected by hydrogen bonds, as well as the tertiary structure formed when the secondary structures are arranged along a single polypeptide chain known as the backbone of the protein molecule.

[0057] Although an integrated design approach combining sequence and structure design may result in protein molecules that are more likely to exhibit desired properties such as binding affinity and developability, there are countless variations in the sequence and three-dimensional structure of protein molecules. For example, the solution space occupied by all possible substitutions of amino acid residues that can form a protein molecule is enormous, even though few of those possible substitutions actually correspond to functional protein sequences (e.g., TIFF2025533582000004.tif6170 amino acid residues for the protein sequence TIFF2025533582000005.tif6170). On the other hand, a single sequence of amino acid residues can fold into a multitude of different three-dimensional structures. In some cases, due to its flexible nature, the three-dimensional structure of a protein molecule can even evolve over time, especially when the protein molecule interacts closely with another molecule. Each three-dimensional structure or conformation of a protein molecule can involve a different spatial arrangement of the atoms of each constituent amino acid residue. Thus, the solution space occupied by all possible three-dimensional structures that can be formed by just one sequence of amino acid residues is already incredibly large, even when limited to a few discrete (rather than continuous) structural variability. For example, A protein sequence of 6170 amino acid residues in TIFF2025533582000006.tif can be approximately 100% even if each amino acid residue is constrained to assume one of three distinct geometric states (e.g., rotators). TIFF2025533582000007.tif6170 possible conformations. For protein sequences of any meaningful length, brute-force exploration of the solution space for all possible permutations of amino acid residues combined with the solution space for all possible three-dimensional structures is computationally intensive. The burden on computational resources is further exacerbated by traditional approaches to structural bioinformatics, which are prohibitively slow due to various computational inefficiencies. For at least the aforementioned reasons, current efforts to engineer protein sequences capable of adopting specific three-dimensional structures are constrained by a difficult tradeoff between designing functional output proteins and computational burden.

[0058] The present disclosure eliminates the trade-off between generating functional protein designs and computational burden by strategically reducing the collaborative solution space of integrated sequence and structure design. For example, in some cases, the design of a molecule, such as a protein molecule, may be performed on a representation of the molecule that improves the performance of various design computational models. In this regard, the representation of the molecule may include a structural representation of the molecule's three-dimensional structure, showing the spatial arrangement of atoms forming the molecule. For example, if the molecule is a protein molecule, the structural representation of the molecule may show the spatial arrangement of atoms forming each amino acid residue of the molecule. If the molecule is a protein molecule, the representation of the molecule may also include a sequence representation showing the identity of each amino acid residue forming the molecule.

[0059] As described in more detail below, the sequence (e.g., primary structure) and three-dimensional structure (e.g., secondary structure and tertiary structure) of a protein molecule can be determined by applying a computational model to a representation of the protein molecule. For example, in some cases, the sequence and three-dimensional structure of a protein molecule can be determined by a computational model that applies a series of modifications to the representation of the protein molecule. In some cases, the representation of the protein molecule may impose certain restrictions on the individual modifications that can be made to the sequence and three-dimensional structure of the protein molecule by the computational model. For example, in some cases, the structural representation of the protein molecule may prevent arbitrary and infeasible modifications to the positions of individual atoms in three-dimensional space. Thus, a computational model operating on the representation of the protein molecule may strategically reduce the collaborative solution space explored by the computational model to generate protein molecular sequences and structures with minimal adverse impact on the accuracy of the design results.

[0060] In some exemplary embodiments, a structural representation of a molecule, such as a protein molecule, may be a coarse-grained (CG) node representation (as shown in FIGS. 3A and 3B ). That is, in some cases, the three-dimensional structure of a molecule, including the positions (e.g., three-dimensional coordinates) of all atoms contained in the molecule, may be represented as a set of coarse-grained (CG) nodes. While the molecule may be a protein molecule, it should be understood that the molecule may also be a non-protein molecule (e.g., a small molecule, a nucleic acid, a polysaccharide, a glycolipid, etc.). In the case of a protein molecule, each amino acid residue contained in the protein molecule may include one or more representative structures, including, for example, a rigid, flexible, or variable structure. Furthermore, each amino acid residue in the protein molecule may be represented by a set of coarse-grained nodes, each corresponding to one of the structures that form the amino acid residue. As a rigid structure, a single structure may include two or more groups of atoms whose positions are fixed relative to the structure's coordinates. As a flexible or variable structure, a single structure may include two or more groups of atoms whose positions are somewhat flexible relative to the structure's coordinates. Thus, each coarse-grained node associated with an amino acid residue may include atoms contained in the corresponding structure. In some cases, a single atom may be part of multiple structures of an amino acid residue and therefore may be included in multiple corresponding coarse-grained nodes. Furthermore, the position of the constituent atoms of an amino acid residue may be determined by the rotation of the corresponding coarse-grained node. TIFF2025533582000008.tif6170 and / or translation It can be specified by TIFF2025533582000009.tif6170.

[0061] In some exemplary embodiments, the structure of a molecule, such as a protein molecule, may be represented by a transformation (e.g., a Euclidean transformation) of the coordinates of each coarse-grained node associated with the molecule and a geometric tensor embedding. For example, in some cases, a rotation of each coarse-grained node associated with the molecule TIFF2025533582000010.tif6170 and / or translation TIFF2025533582000011.tif6170 can be represented as a set of geometric tensors (or geometric tensor embeddings), each of which has a maximum configurable degree TIFF2025533582000012.tif6170. As used herein, the term "geometric tensor" may refer to an object (e.g., a scalar, vector, etc.) that transforms when subjected to one or more coordinate transformations (e.g., Euclidean transformations), such as rotation, translation, etc. In some cases, a set of geometric tensors associated with a coarse-grained node may undergo coordinate transformations corresponding to one or more group elements of a three-dimensional rotation group. Such three-dimensional rotation groups may describe the possible rotational symmetries and orientations of a structure in multidimensional space (e.g., three-dimensional space, etc.). In this regard, each group element of a three-dimensional rotation group may be represented as one or more irreducible representations that cannot undergo further decomposition. Thus, a geometric tensor embedding of a coarse-grained node may include a set of geometric tensors, each of which is manipulated according to one or more group elements from the three-dimensional rotation group to describe the current translation and rotation of the corresponding structure. A coarse-grained (CG) nodal representation of a molecule, such as a protein molecule, may reduce the collaborative solution space explored by a computational model by avoiding at least modifying the positions of individual atoms within the molecule. Instead, while the computational model is operating on a coarse-grained (CG) nodal representation of a molecule, the computational model may apply Euclidean transformations (e.g., translations and rotations) to groups of two or more atoms that are more likely to move as a collective.

[0062] In some exemplary embodiments, a structural representation of a molecule, such as a protein molecule, may be a backbone torsion angle (BBT) representation (as shown in FIG. 6A). For example, in the case of a protein molecule, the backbone torsion angle (BBT) representation of the protein molecule may include, for each constituent amino acid residue of the protein molecule, a plurality of frames that define the positions of the atoms forming the amino acid residue by specifying at least the geometrical state of the backbone and side chain of the amino acid residue. In some cases, the plurality of frames may include a first set of frames that specify the geometrical state of the backbone of the corresponding amino acid residue. As used herein, the "backbone" of an amino acid residue refers to the atoms common to all amino acid residues (e.g., nitrogen (N) atoms, alpha carbon ( TIFF2025533582000013.tif7170), and carboxyl carbon (C) atoms. Additionally, the multiple frames may include a second set of frames specifying one or more torsion angles present in the side chains of amino acid residues. As used herein, the term "torsion angle" may be used interchangeably with the term "dihedral angle" and may refer to an angle representing rotation about a central bond of a fragment of a polypeptide chain containing four atoms joined by three consecutive bonds.

[0063] For a single amino acid residue containing multiple atoms, each frame may define a mapping between the position of the atoms in three-dimensional space (e.g., the three-dimensional coordinates of each atom) and one or more internal degrees of freedom (DoF). In this context, internal degrees of freedom (DoF) may refer to constraints on the type and / or extent of modifications that may be made to an amino acid residue, e.g., by a computational model, when determining the sequence and / or three-dimensional structure of a protein molecule containing the amino acid residue. For example, in some cases, certain degrees of freedom (DoF) may limit changes to the identity of the amino acid residue to one of 20 canonical amino acid residues. Alternatively and / or additionally, certain degrees of freedom (DoF), such as backbone translation, backbone rotation, and torsion angle, may impose constraints on the spatial extent to which each atom of the amino acid residue can move as part of the overall three-dimensional structure of the amino acid residue. That is, some frames may restrict the manner and extent to which each atom can be rearranged in three-dimensional space, for example, relative to other atoms in an amino acid residue, preventing the atom from freely moving to any arbitrary position (e.g., coordinate) in three-dimensional space. For example, in the case of backbone translation and rotation, the corresponding degrees of freedom (DoF) may require translating and rotating the backbone atoms of an amino acid residue as a group, thereby preventing individual backbone atoms from moving and changing their relative spatial arrangement. In the case of a torsion angle between two side chain atoms connected by a bond, the corresponding degrees of freedom (DoF) may require one atom to rotate around the other atom without any change in the distance (or bond length) between them. As described in more detail below, each frame may correspond to the degrees of freedom (DoF) of a computational model for updating a protein sequence (e.g., the identities of the constituent amino acid residues) and / or the three-dimensional structure of the protein sequence.

[0064] In some exemplary embodiments, the backbone torsion angle (BBT) representation of a protein molecule may specify the backbone geometry of each amino acid residue in a variety of different ways. For example, in some cases, the backbone geometry of an amino acid residue may be determined by its translation and rotation, as well as the alpha carbon ( The first frame set may be specified based on the second torsion angle of the rotatable bond between the alpha carbon ( TIFF2025533582000014.tif7170) atom and the carbonyl group. Thus, in some cases, the first frame set may include a first frame that specifies the backbone rotation and translation of the amino acid residue. For example, in some cases, the first frame may include an affine transformation matrix that includes a rotation matrix that specifies the backbone rotation of the amino acid residue and a displacement vector that specifies the backbone translation of the amino acid residue. When the first frame specifies the backbone rotation and translation of the amino acid residue, the first frame set may include a first frame that specifies the backbone torsion angle of the amino acid residue (e.g., the alpha carbon ( TIFF2025533582000015.tif7170) atom and the rotatable bond between the carbonyl group.

[0065] In some exemplary embodiments, instead of the backbone geometry of an amino acid residue being specified based on a combination of torsion angles and their translations and rotations, the backbone geometry of an amino acid residue may also be specified based on the torsion angles present in the backbone of the amino acid residue. Thus, in some cases, the first frameset may be a set of alpha carbons ( The first frame may include a first frame that specifies the torsion angle of the rotatable bond between the alpha carbon ( TIFF2025533582000016.tif7170) atom and the carbonyl group. In these cases, the first frame set may further include a third frame and a fourth frame. The third frame specifies the torsion angle of the alpha carbon ( ) of the backbone of the amino acid residue. The third frame may specify a third torsion angle of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom of the amino acid residue, while the fourth frame may specify a fourth torsion angle of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom of the backbone of the amino acid residue.

[0066] In some exemplary embodiments, the three-dimensional structure of a molecule, such as a protein molecule, can be determined by at least applying a computational model to modify an initial three-dimensional structural representation of the molecule. For example, in some cases, the computational model may modify a structural representation of the molecule (e.g., a coarse-grained (CG) node representation, a backbone torsion angle (BBT) representation, etc.) to determine the three-dimensional structure of the molecule. Alternatively and / or additionally, in the case of protein design, the computational model may also modify a sequence representation of the molecule to determine the identities of the amino acid residues that form the molecule, along with modifying the structural representation of the molecule. It should be understood that the initial three-dimensional structure of the molecule may contain varying degrees of entropy, which decreases as the molecule undergoes modification by the computational model. For example, in some cases, the initial three-dimensional structure of the molecule may include at least some noise (e.g., Gaussian noise, etc.) in the positions (e.g., three-dimensional coordinates) of the constituent atoms. If the molecule is a protein molecule, the initial three-dimensional structure of the molecule may include one or more groupings of amino acid residues corresponding to one or more polymer chains present in the molecule.

[0067] In some cases, the three-dimensional structure of a molecule can be determined by at least determining one or more coordinates of each atom in the three-dimensional structure of the molecule based at least on the modified representation of the molecule output by the computational model. For example, if the molecule is a protein molecule whose three-dimensional structure is represented by a backbone torsion angle (BBT) representation, one or more coordinates of each atom in the three-dimensional structure of the molecule can be determined based at least on a plurality of frames associated with each amino acid residue included in the modified representation of the molecule output by the computational model. In some cases, the one or more coordinates of each atom in the three-dimensional structure of the molecule can be determined by at least determining one or more coordinates of a plurality of backbone atoms of the molecule. Furthermore, in some cases, the one or more coordinates of each atom in the three-dimensional structure of the molecule can be further determined by at least determining one or more coordinates of a plurality of side chain atoms of the molecule based on one or more coordinates of a plurality of backbone atoms of the molecule.

[0068] In some exemplary embodiments, when the three-dimensional structure of a protein molecule is rendered in a coarse-grained (CG) node representation, the three-dimensional structure of the protein molecule can be determined by applying a computational model to tensors associated with each coarse-grained node of the molecule. For example, the structural computational model may be applied to, for each coarse-grained (CG) node of the protein molecule, a rotation of the coarse-grained (CG) node in the initial three-dimensional structure of the protein molecule. TIFF2025533582000018.tif6170 and / or translation The computational model may receive input including a set of geometric tensors representing the initial three-dimensional structure of the molecule. TIFF2025533582000020.tif6170 and / or translation TIFF2025533582000021.tif6170. Alternatively, as noted, the computational model may generate a three-dimensional structure of a molecule by modifying at least one or more frames in a backbone torsion angle (BBT) representation of the molecule. For example, in the case of protein design, the computational model may take inputs including, for each amino acid residue in the protein sequence, a representation of the identity of the amino acid residue, a backbone atom translation, a backbone atom rotation, and one or more torsion angles. In some cases, the computational model may also modify a sequence representation of the molecule to determine the sequence of amino acid residues forming the molecule (e.g., the primary structure of the molecule) in addition to modifying a structural representation of the molecule (e.g., a coarse-grained (CG) node or backbone torsion angle (BBT) representation of the molecule) to determine the three-dimensional structure of the molecule. In some cases, the computational model may generate a three-dimensional structure of a molecule that exhibits one or more desirable properties, such as binding affinity for another molecule (e.g., a viral antigen, a tumor antigen, etc.), specificity for another molecule, lack of non-specificity, stability (e.g., conformational stability, thermodynamic stability, robustness to various environmental stresses such as protease resistance, etc.), non-immunogenicity, humanity, lack of self-association (or non-aggregation), lack of chemical hindrance (e.g., aspartic acid isomerization, oxidation, deamidation), developability, etc. Alternatively and / or additionally, the three-dimensional structure of the molecule may be suitable or configured for one or more downstream tasks, such as predictive analysis of various properties exhibited by the molecule.

[0069] In some exemplary embodiments, the computational model may be a machine learning model that can recognize the same three-dimensional structure regardless of the orientation of the three-dimensional structure taken as input. For example, in some cases, the computational model may be implemented as a geometric deep learning model, such as an equivariant neural network (ENN), a many-body, higher-order equivariant message-passing neural network, etc. If the computational model operates on a coarse-grained (CG) node representation of a molecule, the three-dimensional structure of the molecule may be determined by one or more rotations of the coarse-grained nodes associated with the molecule. TIFF2025533582000022.tif6170 and / or translation TIFF2025533582000023.tif6170, thus changing the relative positions of coarse-grained nodes within the molecule. Alternatively, if the computational model operates on a backbone torsion angle (BBT) representation of the molecule, the three-dimensional structure of the molecule can be modified by updating at least the frames of one or more amino acid residues of the molecule. For example, the backbone torsion angle (BBT) representation of the molecule can be modified by modifying at least a first set of frames to change the backbone geometry of one or more amino acid residues and / or a second set of frames to change the torsion angles of the side chains of one or more amino acid residues. However, a change in the orientation of the entire three-dimensional structure of a molecule, whether by rotating or translating that entire three-dimensional structure, without a change in any of the relative positions of the atoms (or groups of atoms) contained therein, does not constitute a change in the three-dimensional structure of the molecule. Thus, a computational model can recognize cases where two three-dimensional structures have different orientations in space but are otherwise identical. Doing so may enable the computational model to generate the correct three-dimensional structure regardless of the orientation of the initial three-dimensional structure taken as input.

[0070] In some exemplary embodiments, the computational model may include a machine learning model trained to determine the three-dimensional structure of a molecule, such as a protein molecule, by denoising at least the initial three-dimensional structure of the molecule. In some cases, denoising may be performed on a structural representation of the molecule's initial three-dimensional structure, including, for example, a coarse-grained (CG) node representation, a backbone torsion angle (BBT) representation, etc. Additionally, in the case of protein design, denoising may be performed on a sequence representation of the initial sequence of amino acid residues forming the molecule. The machine learning model may denoise the initial three-dimensional structure and / or sequence of the molecule by at least performing a sequence of updates to the structural and / or sequence representation of the molecule. For example, in some cases, the machine learning model may be a diffusion model that removes a portion of the noise present in the initial three-dimensional structure and / or sequence of the molecule at each time point over a series of time points. In some cases, the denoising performed at each time point may include incremental updates to the structural and / or sequence representation of the molecule. Additionally, in some cases, removing a portion of the noise from the representation of the molecule at a first time point may add a second amount of noise back to the representation of the molecule before the diffusion model performs the next update at a second time point. The second amount of noise added by the diffusion model may be determined by a noise schedule that defines a distribution of noise levels over successive updates performed by the diffusion model. In some cases, the distribution of noise levels may correspond to the degrees of freedom (DoF) available to the computational model for modifying the initial three-dimensional structure of the molecule. For example, in some cases, more noise may be added to degrees of freedom (DoF) where more entropy may be present than to degrees of freedom (DoF) where less entropy is present. The addition of the second amount of noise may compensate for errors that may be introduced by denoising the representation of the molecule.

[0071] FIG. 1A depicts a system diagram illustrating an example of a molecular design system 100, according to some exemplary embodiments. Referring to FIG. 1A, the molecular design system 100 may include a molecular design engine 110, a molecular analysis engine 120, and a client device 130 having a user interface (UI) 145. As shown in FIG. 1A, the molecular design system 110, the molecular analysis engine 120, and the client device 130 may be communicatively coupled via a network 140. The client device 130 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc. The network 140 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.

[0072] In some exemplary embodiments, the molecular design engine 110 may generate a molecule by determining at least the sequence (or molecular formula) and the corresponding three-dimensional structure of the molecule. As shown in FIG. 1A , the molecular design engine 110 may include a sequence design computational model 113, a representation generator 115, and a molecular design computational model 117. In some cases where the molecular design engine 110 is deployed to perform protein design, the molecular design computational model 117 may determine the corresponding three-dimensional structure of the output protein sequence based on at least the protein sequence generated by the sequence design computational model 113. For example, in some cases, the sequence design computational model 113 may generate a protein sequence based on an input protein sequence (e.g., a seed sequence). Furthermore, in some cases, the molecular design computational model 117 may operate on a structural representation of the protein sequence (e.g., a coarse-grained (CG) node representation, a backbone torsion angle (BBT) representation, etc.) when generating the corresponding three-dimensional structure. Thus, in some cases, the output of the sequence design model 115 ingested by the representation generator 115 may include a molecular structure file (e.g., a protein structure file) describing the initial three-dimensional structure of the protein sequence. The representation generator 115 may generate a corresponding structural representation of the initial three-dimensional structure of the protein sequence for further manipulation by the molecular design computational model 117.

[0073] In some cases, the sequence design computational model 115 may be implemented using one or more machine learning models trained to generate protein sequences based on an input protein sequence (e.g., a seed sequence) by sampling a distribution of data learned by the one or more machine learning models during training. The one or more machine learning models may be trained based on a variety of known (or observed) protein sequences, including protein sequences known to exhibit a particular function as well as protein sequences with no known function. In doing so, the one or more machine learning models may learn a data distribution corresponding to a reduced-dimensional representation of the sequence of amino acid residues that form the known protein sequence.

[0074] Alternatively, the molecular design computational model 117 may determine the sequence and three-dimensional structure of a protein molecule. For example, in some cases, instead of incorporating a structural representation of the initial three-dimensional structure of the protein sequence generated by the sequence design computational model 113, the molecular design computational model 117 may generate the sequence and three-dimensional structure of the protein molecule by modifying at least a hybrid representation of the protein molecule, which includes a sequence representation of the initial sequence of amino acid residues forming the protein molecule and a structural representation of the initial three-dimensional structure of the protein molecule. Furthermore, while the molecular design computational model 117 operates on the sequence representation of the protein molecule to modify the sequence of amino acid residues forming the protein molecule, the molecular design computational model 117 may simultaneously operate on the structural representation of the protein molecule to determine the corresponding three-dimensional structure. In doing so, the molecular design computational model 117 may generate a protein molecule whose sequence and three-dimensional structure are more likely to be associated with certain desirable properties.

[0075] In some cases, the molecular design computational model 117 may include one or more machine learning models that determine the three-dimensional structure of the protein sequence generated by the sequence design computational model 113 by performing successive modifications to the corresponding structural representation of the protein sequence generated by the representation generator 115. For example, in some cases, the one or more machine learning models implementing the molecular design computational model 115 may be an equivariant neural network, a multi-body, higher-order equivariant message-passing neural network, or the like. As described in more detail below, the one or more machine learning models implementing the molecular design computational model 115 may recognize or account for rotational symmetry present in the three-dimensional structure. That is, the one or more machine learning models implementing the molecular design computational model 115 may rotate around an axis of rotation. TIFF2025533582000024.tif6 A three-dimensional structure rotated 170 degrees around the axis of rotation. TIFF2025533582000025.tif6170 degrees rotated three-dimensional structure. In this way, the molecular design computational model 117 may be able to generate a correct three-dimensional structure regardless of the orientation (for example, the rotation angle) of the initial three-dimensional structure taken as input.

[0076] In some exemplary embodiments, the molecular design computational model 117 may be implemented as a machine learning model that recognizes rotational symmetry present in three-dimensional structures. For example, the molecular design computational model 117 may be implemented as a geometric deep learning model such as an equivariant neural network. Recognizing rotational symmetry present in three-dimensional structures may enable the molecular design computational model 117 to recognize when two three-dimensional structures are identical but have different orientations in space. That is, the molecular design computational model 125 may be able to recognize when two three-dimensional structures are structurally identical (or exhibit structural similarity above a threshold), i.e., when the constituent coarse-grained nodes of the two three-dimensional structures have the same relative positions, even when the entire three-dimensional structure has a different orientation in space. In this way, the molecular design computational model 117 may generate the correct final three-dimensional structure regardless of the orientation of the initial three-dimensional structure taken as input.

[0077] In some exemplary embodiments, the molecular design computational model 117 may be trained to reduce or minimize one or more loss functions, including, for example, a frame alignment point error (FAPE) loss function, a structural violation loss function, etc. Alternatively and / or additionally, the molecular design computational model 117 may be trained to reduce or minimize one or more energy functions that quantify the energy of a three-dimensional structure of a molecule (e.g., a protein molecule, a small molecule, an ion, a nucleic acid, a polysaccharide, a glycolipid, etc.). When the molecular computational model 117 is implemented as an equivariant neural network (ENN), the loss and / or energy of the three-dimensional structure that is continuously updated by the molecular design computational model 117 may be calculated based on a transformation of the coordinates of the coarse-grained nodes output from each block of the equivariant neural network (ENN) and an inverse coarse-grained mapping structure.

[0078] In some exemplary embodiments, protein molecules generated by molecular design engine 110 may be subjected to property analysis by molecular analysis engine 120. As shown in FIG. 1A , molecular analysis engine 120 may apply molecular property computational model 125, which may determine one or more properties of the protein molecule based on the sequence and / or three-dimensional structure of the protein molecule determined by molecular design engine 110. For example, in some cases, molecular property computational model 125 may determine, based at least on the sequence and / or three-dimensional structure of the protein molecule, whether the protein molecule exhibits binding affinity for another molecule (e.g., a viral antigen, a tumor antigen, etc.), specificity for another molecule, lack of non-specificity, stability (e.g., conformational stability, thermodynamic stability, robustness to various environmental stresses such as protease resistance, etc.), non-immunogenicity, humanity, lack of self-association (or non-aggregation), lack of chemical disorder (e.g., aspartic acid isomerization, oxidation, deamidation), developability, etc.

[0079] In some exemplary embodiments, the molecular design computational model 125 may be trained to generate three-dimensional structures of protein sequences that exhibit one or more desirable properties, such as a particular energy range, binding affinity and / or binding specificity to another molecule (e.g., a viral antigen, a tumor antigen, etc.) Alternatively and / or additionally, the molecular design computational model 117 may generate three-dimensional structures of protein sequences that are suitable for and / or configured for one or more downstream tasks, such as predictive analysis of various properties exhibited by the protein sequences.

[0080] As noted, in some cases, the molecular design computational model 117 may generate sequences and / or three-dimensional structures of protein molecules such that the protein molecules are more likely to exhibit particular desirable properties. Alternatively and / or additionally, the molecular design computational model 117 may generate sequences and / or three-dimensional structures of protein molecules such that they are more suitable for or configured for one or more downstream tasks, such as property analysis performed by the molecular analysis engine 120.

[0081] In the example shown in FIG. 1A , for example, molecular analysis engine 120 may apply at least molecular property analysis computational model 125 to determine properties of the protein sequence. For example, in some cases, molecular property computational model 125 may determine one or more properties of protein sequence 120 (e.g., expression, affinity, etc.) based on at least the second protein sequence and / or the three-dimensional structure of the second protein sequence determined by molecular design computational model 117. Furthermore, the three-dimensional structure of the second protein sequence determined by molecular design computational model 117 and / or the properties of the second protein sequence determined by molecular property computational model 125 may be used by molecular design engine 110 (e.g., sequence design computational model 113) when generating subsequent additional protein sequences. For example, if the second protein sequence is determined to exhibit a desirable three-dimensional structure and / or desirable properties, molecular design engine 110 may apply sequence design computational model 113 to generate a third one or more additional protein sequences based on at least the second protein sequence (e.g., as a seed sequence). Alternatively, if the second protein sequence does not exhibit the desired three-dimensional structure and / or desired properties, the molecular design engine 110 may instead apply the molecular design computational model 113 to generate a third one or more additional protein sequences based on a fourth, different protein sequence (e.g., as a seed sequence).

[0082] FIG. 1B depicts a flowchart illustrating an example of a process 160 for predicting protein structure and properties, according to some exemplary embodiments. Referring to FIGS. 1A-1B , process 160 may be performed by design system 100, for example, by molecular design engine 110. In some cases, process 160 may implement a generative design process in which molecular design computational model 117 operates on a structural representation, such as a coarse-grained (CG) nodal representation or a backbone torsion angle (BBT) representation, of a protein molecule having a known protein sequence to determine the three-dimensional structure of the protein molecule. Alternatively, in some cases, process 1600 may implement a generative design process in which molecular design computational model 117 operates on a hybrid representation of a molecule, such as a protein molecule, including a sequence representation (e.g., a logit representation, a one-hot coding representation, etc.) and a structural representation of the protein molecule (e.g., a coarse-grained (CG) nodal representation, a backbone torsion angle (BBT) representation, etc.) to determine not only the three-dimensional structure of the protein molecule but also the sequence.

[0083] At 162, the molecular design engine 110 may receive or generate a molecular structure file that specifies an initial three-dimensional structure of a molecule. For example, in some cases, the molecular design engine 120 may apply the sequence design computational model 113 to generate a protein sequence (or a sequence of amino acid residues), in which case the output of the sequence design computational model 113 may be a molecular structure file. Alternatively, in some cases, the molecular design engine 110 may receive the molecular structure file from another source, such as another sequence design platform. In some cases, the molecular structure file may specify the initial three-dimensional structure of a molecule, which may be a protein molecule or a non-protein molecule (e.g., a small molecule, nucleic acid, polysaccharide, glycolipid, etc.). The molecular structure file may specify the initial three-dimensional structure of the molecule by enumerating at least the constituent atoms. If the molecule is a protein molecule, the molecular structure file may specify the initial three-dimensional structure of the protein molecule by enumerating at least the individual atoms (e.g., heavy atoms) that form each amino acid residue of the protein molecule.

[0084] In some exemplary embodiments, the initial three-dimensional structure of the molecule may include various degrees of entropy or randomness in the spatial arrangement of the constituent atoms, which is then removed by the molecular design computational model 117 to determine the actual three-dimensional structure of the molecule. For example, in some cases, the initial three-dimensional structure of the molecule may include at least some noise (e.g., Gaussian noise, etc.), meaning that the positions of the constituent atoms in the initial three-dimensional structure of the molecule may not match the positions of those atoms in the actual three-dimensional structure of the molecule. In some cases, the molecular design computational model 117 may perform successive modifications to correct the positions of the atoms in the initial three-dimensional structure of the molecule. Alternatively and / or additionally, if the molecule is a protein molecule, the initial three-dimensional structure of the molecule may include one or more groups of amino acid residues corresponding to the polymer chains present in the molecule. Such groupings may impose at least some constraints on the modifications made by the molecular design computational model 117. For example, in some cases, the molecular design computational model 117 may avoid modifications that move two or more amino acid residues of a single polymer chain beyond a threshold distance.

[0085] At 164, the molecular design engine 110 may determine a representation of the molecule based at least on the molecular structure file. In some exemplary embodiments, the representation generator 115 of the molecular design engine 120 may determine a representation of the molecule based at least on the molecular structure file. In some cases, the representation of the molecule may include a structural representation of the molecule. As described in more detail below, the structural representation of the molecule may be a coarse-grained (CG) node representation including a collection of coarse-grained (CG) nodes, each of which corresponds to a structure formed by one or more constituent atoms of the molecule. Alternatively, if the molecule is a protein molecule having a sequence of amino acid residues, the representation of the molecule may be a backbone torsion angle (BBT) representation including, for each amino acid residue of the protein molecule, a corresponding plurality of frames specifying the geometric state of the backbone and the side chains of the amino acid residues. If the molecule is a protein molecule, the representation of the molecule may further include a representation of the sequence of the molecule in addition to the structural representation of the molecule. In some cases, the representation of the sequence of a protein molecule may be a logit representation, in which each position in the sequence forming the protein molecule may be a logit vector that represents the identity of the amino acid residue occupying that position by at least enumerating a probability distribution (e.g., a categorical distribution) over a set of possible amino acid residues. Alternatively, the representation of the sequence of a protein molecule may be a one-hot coded representation, in which each position in the sequence forming the protein molecule may be a one-hot coded vector, in which a value of "1" occupies a position in the one-hot coded vector corresponding to the identity of the amino acid residue occupying the position in the sequence, and a value of "0" occupies any other position in the one-hot coded vector.

[0086] As noted, in some exemplary embodiments, the identity of an amino acid residue occupying each position in a protein sequence may constitute one of the degrees of freedom (DoF) associated with that amino acid residue. This particular degree of freedom may limit the change in identity of the amino acid residue occupying each position to, for example, one of 20 canonical amino acid residues. Thus, in some cases, the degrees of freedom (DoF) associated with amino acid residue identity may be expressed as a category probability. In a protein sequence The logit representation of a protein sequence is the probability vector TIFF2025533582000027.tif7170, and each probability vector TIFF2025533582000028.tif7170 is a sequence of sequences spanning a set of possible amino acid residues (e.g., 20 canonical amino acid residues). Enumerate the probability distribution of the identity of the amino acid residue occupying position TIFF2025533582000029.tif6170. For example, the probability vector for the first position in a protein sequence TIFF2025533582000030.tif7170 may include a first probability that the first position is occupied by alanine (Ala), a second probability that the first position is occupied by arginine (Arg), a third probability that the first position is occupied by asparagine (Asn), etc.

[0087] In some cases, each probability vector TIFF2025533582000031.tif7170 means that the configuration probabilities over the set of possible amino acid residues sum to 1. For example, the probabilities that an amino acid residue occupying a particular position in a protein sequence is one of the 20 canonical amino acid residues sum to 1. Probability Simplex TIFF2025533582000032.tif7170 may be a mathematical space in which each point represents a probability distribution among a finite number of mutually exclusive events or categories, in this case corresponding to a set of possible amino acid residues (e.g., the 20 canonical amino acid residues). TIFF2025533582000033.tif7170 Please understand that TIFF2025533582000034.tif7 is a 170-dimensional object. That is, a probability simplex. The points that form TIFF2025533582000035.tif7170 are TIFF2025533582000036.tif7170 dimensional space, TIFF2025533582000037.tif6170 corresponds to the amount of possible amino acid residues (e.g., 20 canonical amino acid residues). The requirement that the probabilities over TIFF2025533582000038.tif6170 sum to 1 is met by the probability The dimensionality of TIFF2025533582000039.tif7170 is reduced by one. As described in more detail below, the molecular design computational model 125 applied to the sequence representation of the protein sequence can be a diffusion model that adds noise during a forward diffusion process and removes noise during a corresponding reverse diffusion process. It is important to define the forward diffusion process for TIFF2025533582000040.tif7170. It is important to define the forward diffusion process for each probability vector Simply adding noise (e.g., Gaussian noise) to TIFF2025533582000041.tif7170 reduces the number of possible amino acid residues (e.g., the 20 canonical amino acid residues). This is because the sum of the constituent probabilities across TIFF2025533582000042.tif6170 may no longer sum to 1. In other words, the probability simplex TIFF2025533582000043.tif7170 TIFF2025533582000044.tif7170 In principle, noise can be added during forward diffusion to maintain dimensionality. For example, the stochastic simplex Logit space from TIFF2025533582000045.tif7170 To maintain a one-to-one mapping to TIFF2025533582000046.tif8170, noise (e.g., Gaussian noise) can be scaled using the logit Adding it to TIFF2025533582000047.tif7170 gives logit TIFF2025533582000048.tif7170 is the same space TIFF2025533582000049.tif7170, and the logit space is dimensioned From TIFF2025533582000050.tif6170 TIFF2025533582000051.tif6170. This constraint means that the average Logit on the zero identity component subspace by subtracting TIFF2025533582000052.tif7170 This is achieved by projecting TIFF2025533582000053.tif7170. By applying the softmax function, we can establish a unique mapping from the logit subspace to probabilities without any degeneracy.

[0088] In some cases, instead of the logit representation described above, the identity of each amino acid residue in a protein sequence may be represented as a one-hot coded representation. Thus, each position in a protein sequence may be associated with a one-hot coded vector, with a value of "1" occupying the position in the one-hot coded vector corresponding to the amino acid residue occupying the position in the protein sequence, and a value of "0" occupying all remaining positions in the one-hot coded vector. For example, if a position in a protein sequence is occupied by alanine (Ala / A), the one-hot coded vector for that position may include a value of "1" at the position corresponding to alanine (Ala / A) and a value of "0" at the positions corresponding to other amino acid residues.

[0089] It should be understood that to improve the learning of the molecular design computational model 117, a logit representation of the amino acid residues forming the protein sequence may be used instead of other representations, such as one-hot encoding. For example, one-hot encoding does not result in a probabilistic representation of the identity of each amino acid residue in the protein sequence. Instead, the identity of an amino acid residue occupying a position in the protein sequence is indicated by a value of "1" at the corresponding position in the one-hot encoding vector, while the remaining positions in the one-hot encoding vector are occupied by values ​​of "0." Adding noise as part of the forward diffusion process does not result in a binary value change to the one-hot encoding vector. The resulting value may no longer be consistent with the one-hot encoding scheme, and a single "1" value in the one-hot encoding vector identifies the identity of an amino acid residue in the protein sequence. Instead, a range of different values ​​may result, but the values ​​still correspond to the same physical state. For example, a one-hot encoding vector TIFF2025533582000054.tif7170 is a one-hot coded vector TIFF2025533582000055.tif7170, which means that the molecular design computational model 117 can be constructed to learn to produce the same output for several one-hot coded vectors with different values. Thus, the performance of the molecular design computational model 117 may degrade when operating with one-hot coded representations of the identities of amino acid residues in protein sequences.

[0090] At 166, the molecular design engine 110 may determine the sequence and / or three-dimensional structure of the molecule by applying at least the molecular design computational model 125 to modify the representation of the molecule. In some exemplary embodiments, the molecular design engine 110 may determine the three-dimensional structure, and possibly the sequence of the molecule, by applying at least the molecular computational model 125 to modify the representation of the molecule. In some cases, the molecular design computational model 125 may perform successive updates to the representation of the molecule to determine the sequence and / or three-dimensional structure of the molecule. For example, if the molecule is represented as a collection of coarse-grained (CG) nodes, the molecular design computational model 125 may perform successive updates, each modifying one or more coarse-grained (CG) nodes that form the structural representation of the molecule. Alternatively, if the molecular design computational model 117 operates on a backbone torsion angle (BBT) representation of the molecule, each successive update may modify one or more frames that form the structural representation of the molecule.

[0091] In some cases, the molecular design computational model 117 may simultaneously determine the sequence and three-dimensional structure of a molecule, for example, by performing continuous updates to the sequence and structural representations of the molecule. As described in more detail below, the sequence and / or three-dimensional structure of a molecule determined by the molecular design computational model 117 may be used for one or more downstream tasks, including, for example, conformer generation, molecular docking, property prediction, etc.

[0092] If the structural representation of the three-dimensional structure of a protein sequence generated by the sequence design computational model 113 is a coarse-grained (CG) node representation, the structural representation of the protein sequence may include a set of coarse-grained nodes, and each amino acid residue in the protein sequence is associated with a corresponding set of coarse-grained (CG) nodes. For example, each amino acid residue in the protein sequence may include one or more structures, each structure being a group of two or more atoms. In the case of a rigid structure, the positions of the two or more atoms may be fixed relative to the coordinates of the structure. In contrast, as a flexible or variable structure, the positions of the two or more atoms may exhibit at least some flexibility relative to the coordinates of the structure.

[0093] Thus, each amino acid residue in a protein sequence may be associated with at least one coarse-grained (CG) node that corresponds to the structure (e.g., rigid, flexible, or flexible) contained in the amino acid residue. If a single atom is part of multiple structures contained in the amino acid residue, the atom may be included in multiple corresponding coarse-grained nodes. In some cases, the coarse-grained (CG) node representation of a protein sequence may include or exclude certain elements, such as hydrogen (H). For example, in some cases, the coarse-grained (CG) node representation of a protein sequence may include hydrogen atoms found in the amino acid residues that form the protein sequence. However, in other instances, the coarse-grained (CG) node representation of a protein sequence may exclude hydrogen atoms found in the amino acid residues that form the protein sequence.

[0094] To explain further, Length where TIFF2025533582000056.tif7170 is an amino acid residue (e.g., one of the 20 canonical amino acids) Protein sequence of TIFF2025533582000057.tif6170 Considering TIFF2025533582000058.tif7170, the three-dimensional structure formed by the protein sequence is TIFF2025533582000059.tif13170, wherein: TIFF2025533582000060.tif7170 is amino acid residue TIFF2025533582000061.tif7170. Thus, in a coarse-grained node representation of a protein sequence, each amino acid TIFF2025533582000062.tif7170 is a graph of each coarse-grained node TIFF2025533582000063.tif8170 is amino acid Subset of atoms forming TIFF2025533582000064.tif7170 Coarse-grained nodes representing TIFF2025533582000065.tif8170 (e.g., heavy atoms) It can be represented by the set TIFF2025533582000066.tif17170.

[0095] In some exemplary embodiments, the representation generator 115 may generate a coarse-grained (CG) node representation of a protein sequence by grouping atoms listed in a molecular structure file into at least one or more coarse-grained (CG) nodes. The coarse-grained node representation of a protein sequence may be generated to satisfy certain characteristics. For example, the representation generator 115 may generate each coarse-grained (CG) node included in the coarse-grained node representation of a protein sequence such that the union of all coarse-grained (CG) nodes representing amino acid residues in the protein sequence includes all of the constituent atoms of that amino acid residue. In this regard, the union of two or more coarse-grained (CG) nodes includes atoms present in all coarse-grained nodes. For example, the union between a first coarse-grained node and a second coarse-grained node includes a first plurality of atoms present in the first coarse-grained node but not in the second coarse-grained node, a second plurality of atoms present in the second coarse-grained node, and a third plurality of atoms present in both the first coarse-grained node and the second coarse-grained node. Thus, all atoms of an amino acid residue of a protein sequence may be included in at least one coarse-grained (CG) node associated with that amino acid residue. Furthermore, when generating the coarse-grained (CG) nodes to be included in the coarse-grained node representation of a protein sequence, the atoms listed in the molecular structure file may be grouped such that each member atom of a coarse-grained node shares at least one covalent bond with another member of the same coarse-grained node. In some cases, the representation generator 115 may generate the coarse-grained node representation of a protein sequence such that each coarse-grained node includes a threshold amount of atoms (e.g., a minimum amount and / or a maximum amount of atoms), whose constituent atoms collectively form at least one structure as described above.

[0096] In some cases, upon grouping the atoms listed in the molecular structure file into one or more coarse-grained (CG) nodes, the representation generator 115 may further generate a coarse-grained (CG) node representation of the protein sequence by mapping the coordinates of the atoms (e.g., heavy atoms) forming each coarse-grained node to a corresponding Euclidean transformation (e.g., translation, rotation, etc.). Each coarse-grained node associated with a protein sequence is generated to include a threshold amount of atoms (e.g., heavy atoms) forming at least one structure, such as an amino acid residue included in the protein sequence. 3D coordinates of TIFF2025533582000067.tif7170 Forward mapping of TIFF2025533582000068.tif8170 to its corresponding coarse-grained node representation TIFF2025533582000069.tif6170 may contain or be defined as:

number

[0097] To further illustrate, Figure 2A shows an example of a coarse-grained node representation 200 of a protein sequence that excludes the hydrogen (H) atoms of each constituent amino acid residue, and Figure 2B depicts another example of a coarse-grained node representation 250 of a protein sequence that includes hydrogen (H) atoms. For example, in the example coarse-grained node representation 200 shown in Figure 2A, the coarse-grained node representation of the amino acid residue tryptophan (Trp) is shown at the first coarse-grained node TIFF2025533582000077.tif6170 ("C", "CA", "CB", "N"), second coarse-grained node TIFF2025533582000078.tif6170 ("C", "CA", "O") and the third coarse-grained node TIFF2025533582000079.tif6170 ("CG", "CD1", "CD2", "CE2", "CE3", "CZ2", "CZ3", "CH2", "NE1"). On the other hand, the coarse-grained node representation of the amino acid residue valine (Val) is the first coarse-grained node TIFF2025533582000080.tif6170("C", "CA", "CB", "N"), second coarse-grained node TIFF2025533582000081.tif6170 ("C", "CA", "O") and the third coarse-grained node TIFF2025533582000082.tif6170 ("CB", "CG1", "CG2"). Coarse-grained nodes are rotated across the coarse-grained nodes. TIFF2025533582000083.tif6170 and / or translation TIFF2025533582000084.tif6170 allows the three-dimensional position of its constituent atoms to be specified. For example, the position of the amino acid residue tryptophan (Trp) TIFF2025533582000085.tif6170 First coarse-grained node Carbon I atom, alpha carbon ( TIFF2025533582000087.tif7170) atom, beta carbon ( TIFF2025533582000088.tif7170) atoms, and the position of nitrogen (N) is determined by the rotation of the first coarse-grained node. TIFF2025533582000089.tif6170 and / or translation It can be specified by TIFF2025533582000090.tif6170.

[0098] As noted, in some exemplary embodiments, the representation generator 115 generates a rotation of each coarse-grained node representing the initial three-dimensional structure of the protein sequence. TIFF2025533582000091.tif6170 and / or translation For example, the representation generator 115 may determine a Euclidean transformation, such as TIFF2025533582000092.tif6170, that is required to transform each coarse-grained node from its current position to the coarse-grained node's template position (e.g., template coordinates) in order to determine the initial three-dimensional structure of the protein sequence. TIFF2025533582000093.tif6170 and / or translation TIFF2025533582000094.tif6170. In some cases, the representation generator 115 may determine the rotation by at least computing a rotation matrix with the minimum root mean square deviation (RMSD) between the template position of the coarse-grained node and the current position of the coarse-grained node. TIFF2025533582000095.tif6170 and / or translation TIFF2025533582000096.tif6170. For example, in some cases, the representation generator 115 may apply the Kabsch algorithm to calculate a rotation matrix and translation with minimum root mean square deviation (RMSD) between the template position of the coarse-grained node and the current position of the coarse-grained node.

[0099] To further illustrate the calculation of each coarse-grained template coordinate in a protein sequence, a protein structure dataset associated with the protein sequence is Consider TIFF2025533582000097.tif6170. Protein structure dataset Each coarse-grained node in TIFF2025533582000098.tif6170 Transformation of coordinates of TIFF2025533582000099.tif6170 (e.g., Euclidean transformation) For example, TIFF2025533582000100.tif7170 is a graph that uses the Gram-Schmidt process shown in Table 2 below to create a coarse-grained node image. It can be calculated by applying it to the three-dimensional coordinates of the first three atoms of the group of atoms in TIFF2025533582000101.tif6170. [Table 2]

[0100] Then, the inverse coordinate transformation (e.g., Euclidean transformation) TIFF2025533582000103.tif8170, then coarse-grained nodes TIFF2025533582000104.tif6170 to transform these three-dimensional coordinates into a corresponding local frame. In this context, the term "frame" may refer to a transformation (e.g., a Euclidean transformation) that defines the coordinates (e.g., three-dimensional coordinates) of at least some of the atoms of a protein molecule. Such a local frame may be a coarse-grained node While atoms contained in a single coarse-grained node such as TIFF2025533582000105.tif6170 may be positioned, a global frame may define the position of the protein molecule as a whole (e.g., by specifying rotations and translations to be applied to the center of mass of the molecule). In some cases, to determine template coordinates, protein structure datasets grouped by coarse-grained node type, as shown in Table 2 below, are used. One can average the three-dimensional coordinates observed in the local frame across all instances of TIFF2025533582000106.tif6170, i.e., the protein structure dataset If TIFF2025533582000107.tif6170 contains multiple coarse-grained nodes representing the same amino acid residue, then calculating the template coordinates may involve determining the average of the three-dimensional coordinates observed over the local frame of these coarse-grained nodes. [Table 2]

[0101] To calculate the transformation (e.g., Euclidean transformation) of the ground truth coordinates of a given coarse-grained node used in the loss function of the molecular design computational model 117 (e.g., a Frame Alignment Point Error (FAPE) loss function), the Kabsch algorithm may be applied to determine a transformation from template coordinates to observation coordinates that reduces or minimizes the root mean square error (RMSE) of the constituent atoms of each coarse-grained node. For example, TIFF2025533582000109.tif6170th amino acid residue TIFF2025533582000110.tif7170 TIFF2025533582000111.tif6170 atoms TIFF2025533582000112.tif6170th coarse-grained node For TIFF2025533582000113.tif8170, use the Kabsch algorithm to find the template coordinates. TIFF2025533582000114.tif18170 and corresponding input coordinates The Kabsch algorithm uses a single vector decomposition to obtain the covariance matrix Decompose TIFF2025533582000116.tif8170 into eigenvectors and values.

number

[0102] The resulting rotation and translation can be expressed as follows:

number

[0103] FIG. 3A shows the first coarse-grained node of the amino acid residue tryptophan (Trp) in the local frame of the template coordinates after fitting the individual atoms to the template coordinates to show their relative positions to each other. TIFF2025533582000125.tif6170, second coarse-grained node TIFF2025533582000126.tif6170, and the third coarse-grained node Figure 3A depicts a visualization of the atom positions in TIFF2025533582000127.tif6170. Figure 3B depicts a visualization of the atom positions at each coarse-grained node of the amino acid valine (Val) in the local frame of template coordinates after fitting individual atoms to the local reference frame template coordinates to show their relative positions to one another. The bars shown in Figures 3A-3B provide a visual indication of the error associated with each atom's position. As shown in Figures 3A-3B, the error is small.

[0104] In some exemplary embodiments, the transformation (e.g., Euclidean transformation) applied to each coarse-grained (CG) node is of a configurable maximum degree. TIFF2025533582000128.tif6170. In some cases, the rotation and / or translation of each geometric tensor associated with a coarse-grained (CG) node may be determined by applying one or more elements from a three-dimensional rotation group. For example, the first coarse-grained node of the amino acid residue tryptophan (Trp) The numerical representation of TIFF2025533582000129.tif6170 is the first coarse-grained node in three-dimensional space. Current translation of TIFF2025533582000130.tif6170 TIFF2025533582000131.tif6170 and / or rotated To describe TIFF2025533582000132.tif6170, it may contain one or more geometric tensors manipulated by one or more elements from a three-dimensional rotation group.

[0105] To further explain, each coarse-grained node associated with a protein sequence has a degree TIFF2025533582000133.tif7170 orders with channels A set of geometric tensor features of TIFF2025533582000134.tif7170 can be assigned. TIFF2025533582000135.tif6170 related to order tensor TIFF2025533582000136.tif6170Assuming that there is a feature, for each coarse-grained node Geometric tensor features of TIFF2025533582000137.tif8170. Initial embedding of coarse-grained nodes. TIFF2025533582000138.tif8170 is TIFF2025533582000139.tif6170 indicates a predetermined set of coarse-grained node types, which may include or be defined as follows:

number

number

[0106] Representing each amino acid residue in the initial three-dimensional structure of a protein sequence as a collection of coarse-grained (CG) nodes, particularly as a geometric tensor embedding, may reduce the computational complexity associated with subsequent manipulation of the initial three-dimensional structure to determine the three-dimensional structure of the protein sequence. The coarse-grained node representation of the protein sequence may omit overly granular details, such as chemical variations and variations in bond angles and bond lengths. Thus, the molecular design computational model 117 may be able to operate more computationally efficiently with coarse-grained nodes as discrete semantic units than with individual atoms. Furthermore, the molecular design computational model 117 may be able to determine the three-dimensional structure of a protein sequence without additional information, such as the co-occurrence frequency of specific amino acid residues at various positions.

[0107] In some exemplary embodiments, the molecular design engine 110 may apply a molecular design computational model 117 to determine a three-dimensional structure of a protein sequence based at least on a geometric tensor embedding of a coarse-grained (CG) node representation of the initial three-dimensional structure of the protein sequence. In some cases, the molecular design computational model 117 may be implemented as a machine learning model (e.g., an equivariant neural network, etc.) having a sequence of blocks, each of which is a subunit of a machine learning model that includes one or more layers of machine learning models. Each block of the machine learning model may be a transformation (e.g., a rotation, etc.) that defines one or more positions of the coarse-grained nodes included in the initial three-dimensional structure of the protein sequence. TIFF2025533582000146.tif6170 and / or translation TIFF2025533582000147.tif6170) to determine updates to the transformations (e.g., rotations, etc.) that are applied to define the positions of one or more of the coarse-grained nodes in the initial three-dimensional structure to derive the actual three-dimensional structure of the protein sequence. TIFF2025533582000148.tif6170 and / or translation Continuous updates can be performed on the image (euclidean transform such as TIFF2025533582000149.tif6170).

[0108] As noted, the coarse-grained (CG) nodal representation of a protein sequence may be instantiated with a transformation of coordinates defined by one or more elements sampled from a three-dimensional rotation group. For example, in some cases, the coarse-grained nodal representation of a protein sequence may be instantiated with a Euclidean transformation, whose translations and rotations are sampled from a normal distribution with zero mean and unit variance and a uniform distribution over the three-dimensional rotation group SO(3). The final three-dimensional structure determined by the molecular design engine 110 may be, for example, TIFF2025533582000150.tif7170 has the same architecture of subblocks The resulting protein sequence may be generated by iterative refinement performed by a molecular design computational model 117 implemented as an equivariant neural network having 170 blocks. Thus, each block of the equivariant neural network may take as input either an initial coarse-grained node representation of the protein sequence or an updated coarse-grained node representation of the protein sequence output by a previous block. Furthermore, each block may assign two coarse-grained nodes to each coarse-grained node associated with the protein sequence. The first geometric tensor is the rotation before the coarse-grained node. Update for TIFF2025533582000153.tif7170 TIFF2025533582000154.tif6170 is used as the vector part of the non-unit quaternion to calculate the second geometric tensor, which is the translational vector before the coarse-grained node. Update for TIFF2025533582000155.tif6170 Used as TIFF2025533582000156.tif6170.

[0109] Thus, the initial coordinate transformation ( TIFF2025533582000157.tif8170 may be subjected to successive updates by an equivariant neural network (ENN) according to the following:

number

[0110] Each block of the equivariant neural network may transform or simply copy the input embeddings of each coarse-grained node. Nevertheless, in either case, the input embeddings of the coarse-grained nodes are updated with the rotation It can be multiplied by the direct sum of the Wigner D matrices corresponding to TIFF2025533582000159.tif6170.

[0111] For example, in some cases, each block of an equivariant neural network shares the same architecture (e.g., a transformer architecture). TIFF2025533582000160.tif7170 number of sub-blocks. Given an input set of coarse-grained nodes and their corresponding coordinate transformations (e.g., to template coordinates), the block is first computed by pairwise distance TIFF2025533582000161.tif7170 and normalized distance vector Calculate TIFF2025533582000162.tif8170 and use the formula TIFF2025533582000163.tif6170 and TIFF2025533582000164.tif6170 indexes coarse-grained nodes. Pairwise distance TIFF2025533582000165.tif7170 is the learnable weights and cutoff distance used in the radius function that parameterizes the tensor product in the equivariant graph attention module. TIFF2025533582000166.tif7170 TIFF2025533582000167.tif7170 can be projected into a radial Bessel basis. Instead of a polynomial envelope function, a soft unit step is used. TIFF2025533582000168.tif11170 can be applied as input. Normalized distance vector Spherical harmonics to tensor product using TIFF2025533582000169.tif8170 When training an equivariant neural network, the gradient is calculated as the pairwise distance TIFF2025533582000171.tif7170 and normalized distance vector It may not propagate through TIFF2025533582000172.tif8170.

[0112] Applying an initial linear layer to the input coarse-grained node embeddings, instead of pairwise additions, a channel-wise fully connected tensor product is generated to produce an output tensor with the same number of channels as the input. All coarse-grained node pairs are then combined before another linear layer is applied to generate TIFF2025533582000173.tif7170 It can then be applied to the embedding of TIFF2025533582000174.tif6170. Then a depth-wise tensor product (DTP) is taken to produce the output tensor TIFF2025533582000175.tif7170, and spherical harmonics with radius functions that take the aforementioned Bessel basis as input. TIFF2025533582000176.tif8170, as well as coarse-grained node pairs clamped at a specific distance (e.g., 32) The edge embedding can be applied to a scalar edge embedding vector corresponding to the amino acid sequence distance of TIFF2025533582000177.tif6170. In some cases, the edge embedding can be implemented as a lookup table with learnable weights and the same dimension as the number of channels in the input tensor. The output of the depthwise tensor stack is uniformly shuffled and then applied to the attention head. The inputs may be grouped by the number of TIFF2025533582000178.tif7170. A linear layer may be applied to generate tensors of various angles with appropriate channel numbers for the rest of the module. The output of each sub-block except the last may be an updated geometric tensor representing the transformation applied to the corresponding coarse-grained node in the molecule's 3D structure. The last sub-block may then generate two tensors for each coarse-grained node. It may output a TIFF2025533582000179.tif6170 tensor. Edge embedding may be shared across sub-blocks of a given block.

[0113] To further illustrate, Figure 4 depicts a visualization of the iterative updates performed by the molecular design computational model 117 to generate a three-dimensional structure of a molecule, according to some exemplary embodiments. In the example shown in Figure 4, the molecular design computational model 117 is a machine learning model having a sequence of four blocks, each of which represents one or more transformations (e.g., rotations) of the coarse-grained nodes of the initial three-dimensional structure of the molecule (shown in block 0). TIFF2025533582000180.tif6170 and / or translation TIFF2025533582000181.tif6170) to derive the three-dimensional structure of the molecule (shown in block 4). As shown in Figure 4, the transformations (e.g., rotations) applied to one or more of the coarse-grained nodes in the initial three-dimensional structure (shown in block 0) are updated. TIFF2025533582000182.tif6170 and / or translation Each update to the Euclidean transform (e.g., TIFF2025533582000183.tif6170) may gradually reduce or minimize a loss, which represents the deviation between the initial three-dimensional structure of the molecule and the ground truth three-dimensional structure of the molecule. In the example shown in Figure 4, the initial three-dimensional structure (shown in block 0) may be associated with a loss of 0.8906, which decreases substantially over subsequent updates such that the final three-dimensional structure of the molecule (shown in block 4) is associated with a loss of 0.1797.

[0114] As noted, protein design engine 110 may implement a generative design process that integrates sequence and structure design in a variety of different ways. Figures 5A-5B illustrate examples in which protein design engine 110 applies sequence design computational model 113 to generate a protein sequence before applying molecular design computational model 117 to determine the three-dimensional structure of the protein sequence based on at least the representation of the protein sequence. In each of these cases, molecular design computational model 117 may operate on different structural representations of the initial three-dimensional structure of the protein sequence (e.g., a coarse-grained (CG) node representation in Figure 5A and a backbone torsion angle (BBT) representation in Figure 5B). Alternatively, Figure 5C illustrates another example in which protein design engine 110 applies molecular design computational model 117 to simultaneously determine the identities of amino acid residues in a protein sequence and the corresponding three-dimensional structure. That is, in some cases, the molecular design computational model 117 may operate on a sequence representation (e.g., a logit representation, a one-hot coded representation, etc.) of the initial sequence of amino acid residues forming the protein sequence as well as a structural representation (e.g., a coarse-grained (CG) node representation or a backbone torsion angle (BBT) representation) of the initial three-dimensional structure of the protein sequence to determine the identities of the amino acid residues in the protein sequence as well as the three-dimensional structure of the protein sequence.

[0115] To further illustrate, the ontology relationship between a representation of a molecule (such as a protein molecule), which may include a structural representation of the three-dimensional structure of the molecule, and, in some cases, a sequence representation of the amino acid residues that form the molecule, is shown in Table 1 below.

[0116] [Table 1]

[0117] FIG. 5A depicts a flowchart illustrating an example of a process 500 for predicting molecular structures and properties, according to some exemplary embodiments. With reference to FIGS. 1A-1B, 2A-2B, 3A-3B, 4, and 5A, process 500 may be performed by design system 100, such as by molecular design engine 110 and molecular analysis engine 120. In some cases, process 500 may implement a generative design process in which molecular design engine 110 operates on coarse-grained (CG) node representations of molecules. Furthermore, process 500 may implement a generative design process that integrates structure and property prediction as part of a pipeline for generating various molecules, including, for example, protein molecules, small molecules, nucleic acids, polysaccharides, glycolipids, and the like.

[0118] At 502, a molecular structure file specifying an initial three-dimensional structure of a molecule may be received. In some exemplary embodiments, the molecular design engine 110 may apply the sequence design computational model 113 to generate a protein sequence (e.g., a sequence of amino acid residues) of a protein molecule, for example, based on another sequence of amino acid residues (e.g., a seed sequence). In some cases, the sequence design computational model 113 generates the protein sequence by at least determining the identity of each amino acid residue in the protein sequence. Further, in some cases, the sequence design computational model 113 may include one or more machine learning models trained to generate a first sequence of amino acid residues based on a second sequence of amino acid residues by sampling data distributions learned by the one or more machine learning models during training. Alternatively, the molecular design engine 110 may receive a protein sequence or a corresponding molecular structure file (e.g., a protein structure file) specifying the initial three-dimensional structure of the protein sequence from a different source, such as a different sequence design platform. In some cases, upon receiving or generating a protein sequence, the molecular design engine 110 may generate a molecular structure file (e.g., a protein structure file) that specifies the initial three-dimensional structure of the protein sequence, including listing the individual atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence.

[0119] At 504, a plurality of coarse-grained nodes may be determined based at least on the molecular structure file. In some exemplary embodiments, the representation generator 115 may generate a coarse-grained (CG) node representation of the initial three-dimensional structure of the protein sequence based at least on the molecular structure file. To generate a coarse-grained node representation of the protein sequence, the representation generator 115 may begin by grouping the atoms (e.g., heavy atoms) that form each of the amino acid residues in the protein sequence into one or more coarse-grained nodes. Thus, each coarse-grained node may correspond to a structure (e.g., rigid, flexible, or variable, etc.) of two or more atoms (e.g., heavy atoms) that form the amino acid residues of the protein sequence. In some cases, the structure may be a rigid structure including a group of two or more atoms whose positions are fixed relative to the structure's coordinates. In other words, each coarse-grained node operates as an isolated semantic unit such that the relative positions of the constituent atoms remain fixed. A change in the position and / or orientation of a coarse-grained node does not change the relative positions of the atoms included in the coarse-grained node.

[0120] At 506, for each coarse-grained node, one or more geometric tensor embeddings may be generated. In some exemplary embodiments, representation generator 115 may further generate a numerical representation for each coarse-grained node associated with the protein sequence. In particular, representation generator 115 may generate, for each coarse-grained node, a numerical representation that describes the rotation of the coarse-grained node in three-dimensional space. For example, in some cases, structural analysis engine 110 may generate, for each coarse-grained node, a numerical representation that describes the rotation of the coarse-grained node in three-dimensional space, up to a configurable maximum degree. A set of one or more geometric tensors in the form of a set of tensors may be determined. In this regard, each geometric tensor may be manipulated by applying one or more elements from a three-dimensional rotation group (e.g., an irreducible representation of the SO(3) group). For example, a rotation of a coarse-grained node in three-dimensional space TIFF2025533582000186.tif6170 may be represented by one or more geometric tensors that have undergone transformations of coordinates defined by one or more elements from a three-dimensional rotation group.

[0121] At 508, a three-dimensional structure of the molecule may be generated by at least updating the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule. In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 117 to determine the three-dimensional structure of the protein sequence based at least on a geometric tensor embedding of the coarse-grained (CG) nodes representing the initial three-dimensional structure of the protein sequence. The molecular design computational model 117 may determine the three-dimensional structure of the protein sequence by performing successive updates to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence. For example, in some cases, the molecular design computational model 117 may translate and / or rotate at least one or more coarse-grained nodes (e.g., rotate one or more of the coarse-grained nodes). TIFF2025533582000187.tif6170 and / or translation TIFF2025533582000188.tif6170). Further, in some cases, the molecular computational model 117 may include a sequence of blocks, each performing a separate incremental update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence. For example, the molecular design computational model 125 may include a first block that performs a first update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence, followed by a second block that performs a second update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence.

[0122] At 510, one or more additional molecules may be generated based on the sequence of amino acid residues forming the molecule or a sequence of different amino acid residues. In some exemplary embodiments, a molecular design computational model 117 may be applied to determine a three-dimensional structure of the protein sequence that is associated with one or more desirable properties, such as a particular energy range, binding affinity and / or binding specificity to another molecule (e.g., a viral antigen, a tumor antigen, etc.). Alternatively and / or additionally, a molecular design computational model 117 may be applied to determine that the three-dimensional structure of the protein sequence is suitable for or configured for one or more downstream tasks. For example, in some cases, the molecular analysis engine 120 may apply a molecular property computational model 125 to determine one or more properties of the protein sequence based at least on the three-dimensional structure of the protein sequence determined by the molecular design computational model 117, including, for example, binding affinity to another molecule (e.g., a viral antigen, a tumor antigen, etc.), specificity to another molecule, non-specificity, stability (e.g., conformational stability, thermodynamic stability, robustness to various environmental stresses such as protease resistance, etc.), non-immunogenicity, humanity, self-association (or non-aggregation), chemical disorder (e.g., aspartic acid isomerization, oxidation, deamidation), developability, etc.

[0123] In some exemplary embodiments, the three-dimensional structure of a protein sequence may be used as part of a generative design process that integrates structure and property prediction as part of a pipeline for generating protein sequences. For example, if a protein sequence is determined to exhibit a desired three-dimensional structure and / or desired properties (e.g., the three-dimensional structure of the protein sequence corresponds to a three-dimensional structure associated with the desired three-dimensional structure and / or desired properties), the molecular design engine 110 may generate one or more additional protein sequences (e.g., sequences of amino acid residues) based on the protein sequence (e.g., as a seed sequence). Alternatively, if a protein sequence is determined to lack the desired three-dimensional structure and / or desired properties, the design engine 110 may generate one or more additional protein sequences (e.g., sequences of amino acid residues) based on a different sequence of amino acid residues (e.g., as a seed sequence) instead of the protein sequence.

[0124] In some cases, the three-dimensional structure of a protein sequence can be generated as part of building a designed library of protein sequences, such as antibodies or nanobodies, that have particular three-dimensional structures and / or desired properties. For example, the designed library can be a combinatorial library that lists, for each position in the protein sequence, a probability distribution of different amino acid residues that may occupy that position, such that the resulting protein sequence has a probability above a threshold that it will exhibit a three-dimensional structure and / or a desired property associated with the three-dimensional structure.

[0125] Alternatively, instead of a coarse-grained (CG) node representation of the three-dimensional structure of a protein sequence, the molecular design engine 110 may generate and operate on a backbone torsion angle (BBT) representation of the three-dimensional structure of the protein sequence. In some exemplary embodiments, the backbone torsion angle (BBT) representation of the protein sequence may include multiple frames for each constituent amino acid residue of the protein sequence. Each frame may correspond to a degree of freedom (DoF) of the molecular design computational model 117 for updating the initial three-dimensional structure of the protein sequence. For example, in some cases, the multiple frames for a single amino acid residue of the protein sequence may include a first set of frames specifying the backbone geometry of the amino acid residue. Additionally, the multiple frames for the amino acid residue may include a second set of frames specifying one or more torsion angles present in the side chain of the amino acid residue.

[0126] 5B depicts a flowchart showing an example of a process 550 for predicting molecular structures and properties, according to some exemplary embodiments. With reference to FIGS. 1A-1B, 2A-2B, 3A-3B, 4, and 5B, process 550 may be performed by design system 100, for example, by molecular design engine 110. The example of process 550 shown in FIG. 5B may involve molecular design engine 110, for example, molecular design computational model 117, performing a generative design process that operates on a backbone torsion angle (BBT) representation of a molecule, such as a protein molecule, to determine the molecule's three-dimensional structure.

[0127] At 552, the molecular design engine 110 may receive or generate a molecular structure file that specifies an initial three-dimensional structure of a protein molecule, including a sequence of amino acid residues. For example, in some cases, the molecular design engine 110 may generate a protein sequence of a protein molecule based on another protein sequence (e.g., a seed sequence), such as by applying the sequence design computational model 113 to determine the identity of each amino acid residue in the protein sequence. In some cases, the molecular design engine 110 may receive a protein sequence or a corresponding molecular structure file that specifies the initial three-dimensional structure of the protein sequence from a different source, such as a different sequence design platform. In some cases, upon receiving or generating the protein sequence, the molecular design engine 110 may generate a molecular structure file (e.g., a protein structure file) that specifies the initial three-dimensional structure of the protein sequence, including enumerating the individual atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence. The molecular structure file may be a protein structure file that specifies the initial three-dimensional structure of the protein sequence by at least enumerating the atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence.

[0128] At 554, the molecular design engine 110 may determine, based at least on the molecular structure file, a representation of the protein molecule including multiple frames for each amino acid residue in the sequence of amino acid residues. In some exemplary embodiments, the representation generator 115 may determine a backbone torsion angle (BBT) representation of the protein molecule to include, for each amino acid residue, multiple frames specifying the geometric state of the backbone and the side chain of the amino acid residue. In some cases, each frame may correspond to a degree of freedom (DoF) for the molecular design computational model 117 to update the three-dimensional structure of the protein molecule. For example, the multiple frames associated with an amino acid residue may include a first set of frames specifying the geometric state of the backbone of the amino acid residue and a second set of frames specifying the torsion angles of the side chain of the amino acid residue. In some cases, the first set of frames may include a first frame specifying the translation and rotation of the backbone and a second frame specifying the torsion angles present therein. Alternatively, in some cases, the first set of frames may include a corresponding frame for each torsion angle present in the backbone of the amino acid residue. As described in more detail below, the second set of frames can include frames for each torsion angle present in the side chains of amino acid residues.

[0129] To further explain, Consider a protein molecule having an amount of amino acid residues equal to TIFF2025533582000189.tif6170. In some cases, the representation generator 115 may use a backbone translation TIFF2025533582000190.tif7170, Skeleton rotation TIFF2025533582000191.tif7170, and five torsion angles (one for oxygen (O) and four for side chain angles): Each residue in TIFF2025533582000192.tif9170 A backbone torsion angle (BBT) representation of the protein molecule may be determined by assigning at least 100 degrees of freedom (DoF) of TIFF2025533582000193.tif6170. The aforementioned degrees of freedom (DoF) may be applicable when the three-dimensional structure of the protein molecule is modified by the molecular design computational model 117. That is, to determine the three-dimensional structure of the protein molecule, the molecular design computational model 117 may be limited to modifying the initial three-dimensional structure of the protein molecule within one or more of the aforementioned degrees of freedom. For example, to determine the three-dimensional structure of the protein molecule, the molecular design computational model 117 may Alternatively and / or additionally, to determine the three-dimensional structure of the protein molecule, the molecular design computational model 117 may perform a translational and / or rotational modification of one or more of the backbones of the amino acid residues in the amount of TIFF2025533582000194.tif6170. The side chain torsion angles at one or more of the amino acid residues in the amount of TIFF2025533582000195.tif6170 may be modified.

[0130] In some cases, the backbone atoms and torsion angles present in each type of amino acid residue are shown in Table 3 below. [Table 3]

[0131] Further illustration is provided in Figure 6A, which depicts exemplary atomic structures of amino acid residues.

[0132] In 556, the molecular design engine 110 may determine a three-dimensional structure of the protein molecule by applying at least the molecular design computational model 117 to modify the representation of the protein molecule. In some exemplary embodiments, the molecular design computational model 117 may modify the backbone torsion angle (BBT) representation of the protein molecule by at least updating a first set of frames to change the geometric state of the backbone of one or more amino acid residues. For example, in some cases, the molecular design computational model 117 may update the first set of frames to change the translation and rotation of the backbone and / or one or more torsion angles present in the backbone. Alternatively and / or additionally, the molecular design computational model 117 may modify the backbone torsion angle (BBT) representation of the protein molecule by at least updating a second set of frames to change one or more of the torsion angles present in the side chains of the amino acid residues. For illustrative purposes, FIG. 7A depicts a screenshot showing an example of a protein molecule in which the backbones of its constituent amino acid residues are translated to determine the three-dimensional structure of the protein molecule. Figure 7B depicts a screenshot showing an example of a protein molecule in which the backbones of its constituent amino acid residues are rotated to determine the three-dimensional structure of the protein molecule, and Figure 7C depicts a screenshot showing an example of a protein molecule in which the side chain torsion angles of its constituent amino acid residues are altered to determine the three-dimensional structure of the protein molecule.

[0133] In some exemplary embodiments, the molecular design computational model 117 may include a machine learning model trained to determine a three-dimensional structure of a protein molecule by at least denoising an initial three-dimensional structure of the protein molecule. In some cases, the machine learning model may denoise the initial three-dimensional structure of the protein molecule by at least performing a sequence of updates to a backbone torsion angle (BBT) representation of the protein molecule. In some cases, the machine learning model may be trained to reduce or minimize a loss function, such as a frame alignment point error (FAPE) loss function, a structure violation loss function, or the like, associated with each successive update to the initial three-dimensional structure of the protein molecule. Alternatively and / or additionally, the machine learning model may be trained to reduce or initialize an energy function associated with each successive update to the initial three-dimensional structure of the protein molecule.

[0134] In some exemplary embodiments, the molecular design computational model 117 may include a diffusion model that performs a sequence of modifications to the initial three-dimensional structure of the protein molecule, each of which removes a portion of the noise present in the initial three-dimensional structure of the protein molecule. An exemplary diffusion model is illustrated in FIG. 6B. Furthermore, in some cases, each successive update performed by the diffusion model may generate an output that is equivariant to a special Euclidean group SE(3) transformation. For example, in some cases, the diffusion model may perform a first update to the backbone torsion angle (BBT) representation of the protein molecule at a first time point to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule, and further, the diffusion model may perform a second update to the backbone torsion angle (BBT) representation of the protein molecule at a second time point to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule. In some cases, the diffusion model may denoise the initial three-dimensional structure of the protein molecule over a large number of successive updates (e.g., 2000 successive denoising operations). Doing so may enable the diffusion model to perform highly nuanced modifications to refine the three-dimensional structure of the protein molecule. In contrast, the above geometric deep learning model, which contains far fewer blocks (e.g., the eight blocks of an equivariant neural network (ENN)), can update the initial three-dimensional structure of a protein molecule with far fewer and far less subtle updates.

[0135] At 558, the molecular design engine 110 may determine one or more coordinates of each atom in the three-dimensional structure of the protein molecule based at least on the modified representation of the protein molecule. In some exemplary embodiments, the molecular design engine 110 may determine one or more coordinates (e.g., three-dimensional coordinates) of each atom in the three-dimensional structure of the protein molecule based at least on multiple frames of each amino acid residue in the modified backbone torsion angle (BBT) representation of the protein molecule. For example, in some cases, the molecular design engine 110 may determine one or more coordinates of backbone atoms of the protein molecule based at least on the modified backbone torsion angle (BBT) representation of the protein molecule. Thereafter, the molecular design engine 110 may determine one or more coordinates of side chain atoms of the protein molecule based at least on the coordinates of the backbone atoms of the protein molecule. An example algorithm for calculating the coordinates of each atom in the three-dimensional structure of the protein molecule is shown in Table 4 below. [Table 4]

[0136] In some exemplary embodiments, instead of simply determining the three-dimensional structure of a fixed sequence of amino acid residues, such as a protein sequence generated by sequence design computational model 113 (or another sequence design platform), molecular design computational model 117 may be applied to determine the sequence and structure of a protein molecule. That is, in the example process 570 shown in FIG. 5C , molecular design computational model 117 may be applied to determine the identities of all amino acid residues forming a protein molecule as well as the spatial arrangement of the constituent atoms. As described in more detail below, molecular design computational model 117 may determine the sequence and three-dimensional structure of a protein molecule by simultaneously modifying corresponding sequence and structural representations (e.g., coarse-grained (CG) node representations, backbone torsion angle (BBT) representations, etc.).

[0137] 5C depicts a flowchart illustrating an example of a process 570 for predicting molecular structures and properties, according to some exemplary embodiments. With reference to FIGS. 1A-1B, 2A-2B, 3A-3B, 4, and 5C, process 570 may be performed by design system 100, e.g., by molecular design engine 120. The example of process 570 shown in FIG. 5C may involve molecular design engine 110, e.g., molecular design computational model 117, implementing a generative design process that operates on sequence and structural representations of protein molecules (e.g., coarse-grained (CG) node representations, backbone torsion angle (BBT) representations, etc.) to determine the sequence as well as the three-dimensional structure of the protein molecule.

[0138] At 572, the molecular design engine 110 may determine a representation of the protein molecule, including a sequence representation of the initial sequence of the protein molecule and a structural representation of the initial three-dimensional structure of the protein molecule. In some exemplary embodiments, the representation generator 115 may generate a representation of the protein molecule, including a sequence representation of the initial sequence of the protein molecule and a structural representation of the initial three-dimensional structure of the protein molecule. In some cases, the structural representation of the initial three-dimensional structure may be a coarse-grained (CG) node representation, in which atoms (e.g., heavy atoms) forming constituent amino acid residues are represented as a collection of coarse-grained (CG) nodes. Alternatively, the structural representation of the initial three-dimensional structure may be a backbone torsion angle (BBT) representation, which includes multiple frames specifying the geometric states of backbone atoms and side chain atoms for each amino acid residue of the protein molecule.

[0139] In some exemplary embodiments, the sequence representation of a protein molecule may be a logit representation having a plurality of logit vectors, each of which corresponds to a position in the initial sequence of the protein molecule and indicates the identity of the amino acid residue occupying that position by at least enumerating a probability distribution over a set of possible amino acid residues that occupy that position. That is, the logit vector for a position in the initial sequence of the protein molecule may include, for each possible amino acid residue, a corresponding probability that the position is occupied by that amino acid residue. Alternatively, the sequence representation of a protein molecule may be a one-hot coded representation having a plurality of one-hot coded vectors, each of which corresponds to a position in the sequence of the initial protein molecule and has a position corresponding to each amino acid residue in the set of possible amino acid residues. The one-hot coded vector for a particular position in the protein sequence indicates the identity of the amino acid residue occupying that position by at least having a value of "1" at the position in the one-hot coded vector that corresponds to the amino acid residue occupying that position in the protein sequence and a value of "0" elsewhere (e.g., TIFF2025533582000198.tif7170). The set of amino acid residues may optionally include "ghost residues" that represent gaps in the sequence of amino acid residues. In this manner, a position in the sequence of amino acid residues includes a gap that is not occupied by any amino acid residues if the probability of the position being occupied by a ghost residue meets one or more thresholds. The inclusion of ghost residues may correspond to a change in length as part of a subsequent generative diffusion process performed, for example, by the molecular design computational model 117.

[0140] As noted, in some exemplary embodiments, the representation generator 115 may generate a representation of a protein molecule to include a structural representation of the initial three-dimensional structure of the protein molecule as well as a sequence representation, if the initial sequence. Thus, in some cases, the representation of a protein molecule may include, for each position within the sequence of amino acid residues forming the protein molecule, a representation of the identity of the amino acid residue occupying that position (e.g., a logit representation, a one-hot coded representation, etc.). Additionally, the representation of a protein molecule may include, for each position within the sequence of amino acid residues forming the protein molecule, a representation of the spatial arrangement of atoms (e.g., heavy atoms) forming the amino acid residue occupying that position (e.g., a coarse-grained (CG) node representation, a backbone torsion angle (BBT) representation, etc.). For example, for a single position within a sequence of amino acid residues forming a protein molecule, the representation of the protein molecule may include a logit vector enumerating a probability distribution (e.g., a categorical distribution) over the set of possible amino acid residues occupying the position (e.g., a first probability that the position is occupied by alanine (Ala / A), a second probability that the position is occupied by arginine (Arg / R), a third probability that the position is occupied by asparagine (Asn / N), etc.). For the same position, the representation of the protein molecule may further include a representation of the spatial arrangement of atoms (e.g., heavy atoms) forming the amino acid residue occupying the position. In the case of a coarse-grained (CG) node representation, the representation may include a collection of coarse-grained (CG) nodes, each containing two or more atoms of the amino acid residue. Alternatively, in the case of a backbone torsion angle (BBT) representation, the representation may include backbone translations, backbone rotations, and torsion angles formed by the atoms of the amino acid residue.

[0141] To further explain, Consider a protein molecule formed by a sequence of 6170 amino acid residues in TIFF2025533582000199.tif. Each residue TIFF2025533582000200.tif6170 may be associated with the following degrees of freedom, which means that the molecular design computational model 117 may perform the following modifications when operating on the representation of the protein molecule: (i) residue identity TIFF2025533582000201.tif7170, in formula TIFF2025533582000202.tif6170 shows the set of possible amino acid residues. (ii) Backbone translation TIFF2025533582000203.tif7170, (iii) Skeleton rotation TIFF2025533582000204.tif7170, and (iv) torsion angles (e.g., one for the backbone oxygen (O) and four for the side chain angles): TIFF2025533582000205.tif9170. Collectively, the aforementioned degrees of freedom are TIFF2025533582000206.tif9170, wherein TIFF2025533582000207.tif9170. Additionally, in some cases, each of the aforementioned degrees of freedom may correspond to a frame that is modified by the molecular design computational model 117 when performing a generative process (e.g., a generative diffusion process) to determine the sequence and three-dimensional structure of the protein molecule.

[0142] At 574, the molecular design engine 110 may apply the molecular design computational model 117 to determine the sequence and three-dimensional structure of the protein molecule by at least modifying a sequence representation and a structural representation of the protein molecule. In some exemplary embodiments, the molecular design computational model 117 may determine the three-dimensional structure of the protein molecule by at least modifying a sequence representation of the initial protein molecule sequence along with a structural representation of the three-dimensional structure of the protein molecule. In some cases, the structural representation may be a coarse-grained (CG) node representation or a backbone torsion angle (BBT) representation of the initial three-dimensional structure of the protein molecule. Further, in some cases, the molecular design computational model 117 may be a diffusion model that determines the sequence and three-dimensional structure of the protein molecule by denoising the initial sequence and initial three-dimensional structure of the protein molecule at least over a series of time points. For example, denoising the initial sequence may include manipulating the sequence representation of the initial sequence to modify the identity of one or more amino acid residues included in the initial sequence. On the other hand, denoising the initial three-dimensional structure may involve a molecular design computational model 117 operating on a structural representation of the initial three-dimensional structure to modify the spatial arrangement of atoms (e.g., heavy atoms) that form the amino acid residues of the protein molecule.

[0143] In some exemplary embodiments, when the molecular design computational model 117 determines the sequence and three-dimensional structure of a protein molecule, the molecular design computational model 117 may perform a diffusion process in which the identity of amino acid residues is modified as another degree of freedom (DoF) in addition to those related to the spatial arrangement of constituent atoms. Thus, in some cases, the molecular design computational model 117 may be a diffusion model including a first diffusion kernel that modifies the identity of individual amino acid residues, a second diffusion kernel that modifies the backbone translation of each amino acid residue, a third diffusion kernel that modifies the backbone rotation of each amino acid residue, and a fourth diffusion kernel that modifies the backbone and side chain torsion angles of each amino acid residue. In some cases, the first diffusion kernel, the second diffusion kernel, the third diffusion kernel, and the fourth diffusion kernel may each be parameterized as a neural network, including, for example, an equivariant neural network (ENN), that recognizes or takes into account the rotational symmetry present in the three-dimensional structure of the protein molecule.

[0144] In some exemplary embodiments, the molecular design computational model 117 may perform a generative diffusion process to determine the sequence and three-dimensional structure of a protein molecule. In some cases, the molecular design computational model 117 may be a diffusion model that determines the initial sequence and initial three-dimensional structure of a protein molecule by at least denoising the initial sequence and initial three-dimensional structure of the protein molecule over successive time steps. For example, in some cases, the molecular design computational model 117 may remove a first amount of noise from the representation of the protein molecule before removing a second amount of noise from the representation of the protein molecule. It should be understood that the generative diffusion process may correspond to a reverse diffusion process, in which the molecular design computational model 117 removes noise from the representation of the protein molecule according to a decreasing noise scale, thereby causing less noise to be present in the sequence and three-dimensional structure of the protein molecule at each successive time step. However, training of the molecular design computational model 117 may also include a forward diffusion process in which noise is added according to an increasing noise scale, so that more noise is present in the three-dimensional structure of the protein molecule at each successive time step. Thus, training a molecular design computational model 117 may involve learning a reverse diffusion process to recover the correct sequence and three-dimensional structure of a protein molecule from a noisy representation of the protein molecule in which the identities of the amino acid residues that form the protein molecule are uncertain and the three-dimensional structure of the protein molecule is random.

[0145] In some exemplary embodiments, training the molecular design computational model 117 may include learning a score function for each diffusion kernel (e.g., parameterized by a neural network) included in the molecular design computational model 117. For example, in some cases, the molecular design computational model 117 may be implemented in a stochastic differential equation (SDE) score matching framework, in which a stochastic differential equation (SDE) is applied to smoothly transform samples from a complex data distribution, which in this case may be occupied by ground truth sequences and three-dimensional structures of various known protein molecules, into corresponding samples in a noise distribution due to noise injection. To restore samples from the original complex data distribution (e.g., the original sequences and three-dimensional structures of protein molecules) by noise removal, a corresponding inverse-time stochastic differential equation (SDE) may be applied. Equations (2) and (3) above are examples of the forward stochastic differential equation (SDE) and inverse stochastic differential equation (SDE) described above.

[0146] Referring to equations (2) and (3), training a diffusion model in a stochastic differential equation (SDE) score matching framework involves determining, for each possible degree of freedom (DoF), a corresponding score function Score-based model approximating TIFF2025533582000208.tif7170 TIFF2025533582000209.tif7170. For example, in some cases, a score function of a first diffusion kernel that modifies the identity of individual amino acid residues may represent the change in logarithmic data density of a complex data distribution associated with a known protein sequence. Learning the score function of the first diffusion kernel may enable drawing samples from regions of the original complex data distribution that are more densely populated by the known protein sequence over successive time steps during the de-diffusion process (e.g., by applying Markov chain Monte Carlo sampling with Langevin dynamics). Similarly, a score function of a second diffusion kernel that modifies the backbone translation of each amino acid residue may represent the change in logarithmic data density of a complex data distribution associated with the backbone translation of a known three-dimensional protein structure. Learning the score function of the second diffusion kernel can allow samples to be drawn from regions of the original complex data distribution that are more densely populated by the backbone torsion angles of known three-dimensional protein structures over successive time steps during the inverse diffusion process (e.g., by applying Markov chain Monte Carlo sampling with Langevin dynamics). Each diffusion kernel's score function can take other degrees of freedom as inputs. For example, the score function of the first diffusion kernel, which modifies the identity of individual amino acid residues, can take as input the backbone translations, backbone rotations, and torsion angles of the amino acid residues determined by the corresponding diffusion kernel, thus allowing multiple diffusion kernel score functions to be simultaneously learned for the same protein molecule.

[0147] In some exemplary embodiments, each diffusion stage may include adding back some noise to the revised representation of the initial sequence and initial three-dimensional structure of the protein molecule following noise removal. For example, in some cases, after removing a first amount of noise from the representation of the protein molecule, the molecular design computational model 117 may add back a third amount of noise to the backbone torsion angle (BBT) representation of the protein molecule before removing a second amount of noise from the representation of the protein molecule. After the second amount of noise is removed from the representation of the protein molecule, a fourth amount of noise may be added back before the representation of the protein molecule is further denoised by the diffusion model. The third and fourth amounts of noise added to the backbone torsion angle (BBT) representation of the protein molecule may be determined based on a noise schedule that defines the distribution of noise levels over the sequence of diffusion operations performed by the diffusion model. The addition of noise may compensate for at least some of the errors that may be present in the noise removal performed each time by the diffusion model.

[0148] FIG. 6A depicts a schematic diagram showing the amino acid structure of an example amino acid residue 600, according to some exemplary embodiments. In some cases, the backbone torsion angle (BBT) representation of the amino acid residue 600 may specify the backbone geometry of the amino acid residue 600 in a variety of different ways. In the example amino acid residue 600 shown in FIG. 6A, the backbone of the amino acid residue 600 is formed by a nitrogen (N), an alpha carbon ( TIFF2025533582000210.tif7170) atom, and a carbonyl group formed by a carbon atom bonded to an oxygen atom. Thus, in some cases, the multiple frames associated with amino acid residue 600 may include a first frame defining a backbone geometry of amino acid residue 600 and may specify a rotation and translation of the backbone of amino acid residue 600. For example, in some cases, the first frame may include an affine transformation matrix including a rotation matrix specifying a rotation of the backbone of amino acid residue 600 and a displacement vector specifying a translation of the backbone of amino acid residue 600. In some cases, the multiple frames associated with amino acid residue 600, together with the first frame specifying a translation and rotation of the backbone of amino acid residue 600, may include a first frame defining a geometry of the alpha carbon ( TIFF2025533582000211.tif7170) Torsion angle of the rotatable bond between an atom and a carbonyl group It may contain a second frame specifying TIFF2025533582000212.tif5170.

[0149] In some cases, instead of translation and rotation of the backbone of amino acid residue 600, the backbone geometry of the amino acid residue may be specified by the torsion angles present therein. Thus, in some cases, multiple frames associated with amino acid residue 600 may be specified by the alpha carbon ( TIFF2025533582000213.tif7170) Torsion angle of the rotatable bond between an atom and a carbonyl group TIFF2025533582000214.tif5170 specifies the first frame, alpha carbon ( TIFF2025533582000215.tif7170) Torsion angle of the rotatable bond between the atom and the nitrogen (N) atom The second frame designating TIFF2025533582000216.tif5170, and the torsion angle of the rotatable bond between the carbon (C) and nitrogen (N) atoms of the backbone of the amino acid residue It may contain a third frame specifying TIFF2025533582000217.tif3170.

[0150] Referring again to FIG. 6A, in addition to the frame specifying the backbone geometry of amino acid residue 600, the plurality of frames associated with amino acid residue 600 may further include one or more additional frames that specify the torsion angles present in the side chain of amino acid residue 600. In the example shown in FIG. 6A, the plurality of frames associated with amino acid residue 600 includes the alpha carbon ( TIFF2025533582000218.tif7170) Atoms and beta carbon ( TIFF2025533582000219.tif7170) Torsion angle of a rotatable bond between atoms TIFF2025533582000220.tif4170, beta carbon ( TIFF2025533582000221.tif7170) Atoms and gamma carbon ( TIFF2025533582000222.tif7170) Torsion angle of a rotatable bond between atoms TIFF2025533582000223.tif4170, gamma carbon ( TIFF2025533582000224.tif7170) Atom and delta carbon ( TIFF2025533582000225.tif7170) Torsion angle of a rotatable bond between atoms TIFF2025533582000226.tif4170, and delta carbon ( TIFF2025533582000227.tif7170) Torsion angle of the rotatable bond between the atom and the nitrogen (N) atom It may contain frames for each of TIFF2025533582000228.tif4170.

[0151] In some cases, the backbone torsion angle (BBT) representation of a protein molecule may be further generated to include multiple polymer chains, each containing one or more amino acid residues in the protein molecule. For example, in some cases, the residues contained in each polymer chain may be treated as a single rigid body. Thus, in some cases, the residues of a polymer chain (including its constituent atoms) may be translated and rotated as a group around a center of mass. That is, the residues of a polymer chain may share a center of mass degree of freedom (DoF) as a group.

[0152] 6B depicts a schematic diagram illustrating an example of a diffusion framework, according to some exemplary embodiments. In some cases, when implemented as the aforementioned diffusion model, the molecular design computational model 117 may perform a reverse diffusion process when denoising the initial three-dimensional structure of the protein molecule. In contrast, the forward diffusion process shown in FIG. 6B may be performed during training of the diffusion model, where the diffusion model performs reverse diffusion to continuously remove noise from the corrupted three-dimensional structure, such that noise is continuously added to the ground truth three-dimensional structure to generate the corrupted three-dimensional structure, before being trained to recover the ground truth three-dimensional structure.

[0153] Referring again to FIG. 6B, in some cases, the forward diffusion process may involve adding noise (to perturb or corrupt the original data) according to an increasing noise scale, so that more noise is present in the three-dimensional structure of the protein molecule at each successive time step. In contrast, the reverse diffusion process may involve removing noise (to restore the original data) according to a decreasing noise scale, so that less noise is present in the three-dimensional structure of the protein molecule at each successive time step. In some cases, the three-dimensional structure of a molecule, such as a protein molecule, may be generated by performing a reverse diffusion process on an initial three-dimensional structure of the molecule in which all constituent atoms occupy random positions in three-dimensional space. As noted, the diffusion model may gradually remove noise from the initial three-dimensional structure of the molecule over a series of time points. If the initial three-dimensional structure of the molecule is represented in a backbone torsion angle (BBT) representation, the initial three-dimensional structure of the molecule may include noise in the spatial arrangement of atoms in the molecule's side chains and backbone. For example, noise may exist in various degrees of freedom (DoF) through which atoms can move within the molecule's three-dimensional structure. Therefore, denoising the initial 3D structure, which involves modifying the positions of these atoms, e.g., the translation of the backbone, TIFF2025533582000229.tif7170, Skeleton rotation TIFF2025533582000230.tif7170, and five torsion angles (one for oxygen (O) and four for side chain angles): TIFF2025533582000231.tif9170 and The diffusion model may be limited to specific degrees of freedom, including TIFF2025533582000232.tif6170. Furthermore, in some cases, the diffusion model may include multiple diffusion kernels, each of which modifies the initial three-dimensional structure of the molecule along a corresponding degree of freedom (DoF). In some cases, each diffusion kernel may be parameterized as a neural network. Furthermore, in some cases, each diffusion kernel may be parameterized as an equivariant neural network that can generate the correct three-dimensional structure regardless of the orientation of the initial three-dimensional structure taken as input.

[0154] Figure 6B shows an example in which the diffusion model is implemented in the framework of stochastic differential equation (SDE) score matching, in which a stochastic differential equation (SDE) is applied to smoothly transform samples from a complex data distribution, which in this case may be occupied by various known ground-truth three-dimensional structures of protein molecules, into corresponding samples in a noise distribution by noise injection. Note that a corresponding inverse-time stochastic differential equation (SDE) can be applied to restore samples from the original complex data distribution (e.g., the original three-dimensional structure of the protein molecule) by noise removal. As a score-based generative model, the inverse stochastic differential equation (SDE) governing the generative process for determining the three-dimensional structure of a protein molecule can be trained by learning a score function of the data distribution (or a function of the gradient of the logarithmic probability density). In this regard, the score of the data distribution at each time point of the diffusion process, determined by the score function, can correspond to the change in the logarithmic data density. Learning the score function not only allows us to approximate the original complex data distribution, but also enables the process of returning samples from the noise distribution to their corresponding samples in the complex data distribution. Unlike the probability density function of the distribution of data, the score function can be calculated without a normalization constant, which requires determining the entire set of possible values, which is often a cumbersome calculation. As described in more detail below, the score function can be estimated by score matching, for example, during training of a stochastic differential equation (SDE)-based diffusion model. Furthermore, it should be understood that there can be a separate score function for each degree of freedom (DoF). During training of a stochastic differential equation (SDE)-based diffusion model, score functions for multiple degrees of freedom (DoF) can be determined simultaneously.

[0155] The following equation (2) converts a complex data distribution into the original distribution. Noise distribution from TIFF2025533582000233.tif6170 This is an example of a forward stochastic differential equation (SDE) that converts to TIFF2025533582000234.tif6170. Equation (3) is a function of the known prior distribution TIFF2025533582000235.tif6170 is a pure noise to original complex data distribution An example of the corresponding inverse stochastic differential equation (SDE) that restores TIFF2025533582000236.tif6170.

number

number

[0156] In some exemplary embodiments, training a diffusion model in a stochastic differential equation (SDE) score matching framework involves calculating, for each possible degree of freedom (DoF), a corresponding score function Score-based model approximating TIFF2025533582000243.tif7170 TIFF2025533582000244.tif7170. In some cases, this may involve training a score function for each degree of freedom (DoF). TIFF2025533582000245.tif7170 is a score-based model The score of the data distribution associated with TIFF2025533582000246.tif7170 and in this case the original distribution TIFF2025533582000247.tif6170) and the ground truth score of the data distribution (e.g., Fisher divergence or squared TIFF2025533582000248.tif7170). TIFF2025533582000249.tif7170 is the original data distribution The density variation of the logarithmic data across TIFF2025533582000250.tif6170 can be determined. Thus, once the score function estimate is calculated, the original data distribution, as guided by the score function, can be calculated. Three-dimensional structures of molecules (e.g., protein molecules) can be generated by sampling (e.g., Markov Chain Monte Carlo sampling with Langevin dynamics) from the data distribution. In doing so, the three-dimensional structure of the molecule is populated by three-dimensional structures that are more consistent with the ground truth three-dimensional molecular structure. It can be generated by sampling from progressively denser regions of TIFF2025533582000252.tif6170.

[0157] 8A-8D depict graphs showing various examples of noise schedules, according to some exemplary embodiments. In some cases, the distribution of noise levels may correspond to the degrees of freedom present in the representation of a protein molecule in the molecular design computational model 117 for modifying the initial three-dimensional structure of the protein molecule. For example, in some cases, the torsion angles and percentage of residue identities of the protein molecule that are uncertain may decrease over time. Thus, the starting point (e.g., At a second time point (e.g., TIFF2025533582000253.tif6170), nearly every torsion angle of the protein molecule can be randomly oriented, but the identity of nearly every residue is ambiguous. TIFF2025533582000254.tif6170), after the representation of the protein molecule has undergone at least some denoising by a molecular design computational model 117, the identity of at least some residues may become more certain, and therefore the randomness of at least some torsion angles may be reduced. TIFF2025533582000255.tif6170), the identity of nearly every residue can be certain, and at that point the corresponding torsion angles are also known.

[0158] At 576, the molecular design engine 110 may determine one or more coordinates of each atom in the three-dimensional structure of the protein molecule based at least on the modified representation of the protein molecule. In some exemplary embodiments, the molecular design engine 110 may determine one or more coordinates (e.g., three-dimensional coordinates) of each atom in the three-dimensional structure of the protein molecule based at least on multiple frames of each amino acid residue in the modified backbone torsion angle (BBT) representation of the protein molecule. For example, in some cases, the molecular design engine 110 may determine one or more coordinates of backbone atoms of the protein molecule based at least on the modified backbone torsion angle (BBT) representation of the protein molecule. Thereafter, the molecular design engine 110 may determine one or more coordinates of side chain atoms of the protein molecule based at least on the coordinates of the backbone atoms of the protein molecule. An example of an algorithm for calculating the coordinates of each atom in the three-dimensional structure of a protein molecule is shown in Table 4 above.

[0159] In view of the above-described embodiments of the subject matter, the present application discloses the following list of examples, which, in combination with one feature of a single example or two or more features of said examples, optionally in combination with one or more features of one or more additional examples, are further examples included in the disclosure of the present application.

[0160] Item 1: A computer-implemented method including: receiving a molecular structure file specifying an initial three-dimensional structure of a molecule; determining a plurality of coarse-grained nodes based at least on the molecular structure file, each coarse-grained node corresponding to a structure of two or more atoms (e.g., heavy atoms) that form an amino acid residue in the molecule; and determining the three-dimensional structure of the molecule using a design computational model, wherein the design computational model determines the three-dimensional structure of the molecule by at least updating the positions of one or more coarse-grained nodes of the initial three-dimensional structure of the molecule.

[0161] Item 2: The method of item 1, wherein the structure of two or more atoms excludes one or more elements that form amino acid residues in the molecule.

[0162] Item 3: The method according to item 1 or 2, wherein the structure of two or more atoms excludes one or more heavy atoms that form amino acid residues in the molecule.

[0163] Item 4: The method of any one of items 1 to 3, further comprising generating, for each coarse-grained node of the plurality of coarse-grained nodes, a geometric tensor embedding corresponding to a numerical representation of the rotation and / or translation of the coarse-grained node.

[0164] Item 5: The method of item 4, wherein the geometric tensor embedding comprises a set of geometric tensors subjected to one or more rotations and / or translations.

[0165] Item 6: The method described in Item 5, wherein the one or more rotations and / or translations correspond to one or more elements from a three-dimensional rotation group that lists all possible significant rotations in three-dimensional space that cannot be further decomposed into combinations of two or more other rotations.

[0166] Item 7: The method of any one of items 1 to 6, wherein the design computational model includes a machine learning model trained to perform continuous updates to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.

[0167] Item 8: The method described in Item 7, wherein the machine learning model includes a sequence of blocks, each block of the sequence of blocks performing an update to the positions of one or more coarse-grained nodes of an initial three-dimensional structure of the molecule.

[0168] Item 9: The method of item 7 or 8, wherein the machine learning model includes a first block that performs a first update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule, and a second block that performs a second update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.

[0169] Item 10: The method of any one of items 7 to 9, wherein the machine learning model is a geometric deep learning model.

[0170] Item 11: The method of any one of items 7 to 10, wherein the machine learning model is an equivariant neural network.

[0171] Item 12: The method of any one of items 7 to 11, wherein the machine learning model is trained to reduce a loss function associated with each successive update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.

[0172] Item 13: The method of item 12, wherein the loss function is a frame alignment point error (FAPE) loss function and / or a structural violation loss function.

[0173] Item 14: The method of any one of items 7 to 13, wherein the machine learning model is trained to reduce an energy function associated with each successive update to the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence.

[0174] Item 15: The method of any one of items 7 to 14, wherein the machine learning model identifies when two or more three-dimensional structures having different orientations in three-dimensional space are identical.

[0175] Item 16: The method of any one of items 1 to 15, wherein the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein are updated by rotating and / or translating at least one or more coarse-grained nodes.

[0176] Item 17: The method of any one of items 1 to 16, wherein the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein are updated without modifying the relative positions of two or more atoms in the structure corresponding to each coarse-grained node.

[0177] Item 18: The method of any one of items 1 to 17, wherein the plurality of coarse-grained nodes is determined such that the bonds of the plurality of coarse-grained nodes include all atoms contained in the molecule.

[0178] Item 19: The method according to any one of items 1 to 18, wherein the plurality of coarse-grained nodes is determined by at least grouping a plurality of atoms included in the molecule such that each atom of a coarse-grained node shares at least one covalent bond with another atom of the same coarse-grained node.

[0179] Item 20: The method of any one of items 1 to 19, wherein the plurality of coarse-grained nodes are determined such that each coarse-grained node includes a threshold amount of atoms (e.g., heavy atoms) that form at least one structure.

[0180] Item 21: The method of any one of items 1 to 20, wherein the three-dimensional structure of the protein molecule is associated with one or more desirable properties.

[0181] Item 22: The method of any one of items 1 to 21, further comprising determining one or more properties of the molecule based on at least the three-dimensional structure of the protein.

[0182] Item 23: The method of any one of items 1 to 22, further comprising generating a second sequence of amino acid residues comprising the molecule based on at least the first sequence of amino acid residues, and generating a fourth sequence of amino acid residues based on at least the second sequence of amino acid residues or the third sequence of amino acid residues.

[0183] Item 24: The method of Item 23, further comprising: determining, based on at least the three-dimensional structure of the molecule, that a second sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties; and, in response to determining that the second sequence of amino acid residues exhibits the desired three-dimensional structure and / or desired properties, generating a fourth sequence of amino acid residues based on at least the second sequence of amino acid residues.

[0184] Item 25: The method of Item 24, further comprising determining, based on at least the three-dimensional structure of the molecule, that the second sequence of amino acid residues lacks a desired three-dimensional structure or a desired property, and generating a fourth sequence of amino acid residues based on at least the sequence of third amino acid residues in response to determining that the second sequence of amino acid residues lacks the desired three-dimensional structure or a desired property.

[0185] Item 26: The method according to any one of items 1 to 25, wherein the molecule is a protein molecule, a small molecule, an ion, a nucleic acid, a polysaccharide or a glycolipid.

[0186] Item 27: The method of any one of items 1 to 26, wherein the structure is a rigid or flexible group of two or more atoms that form an amino acid residue in the molecule.

[0187] Item 28: A system comprising at least one data processor and at least one memory storing instructions, the instructions, when executed by the at least one data processor, causing operations including the methods described in any one of items 1 to 27.

[0188] Item 29: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, produce operations including the methods described in any one of items 1 to 27.

[0189] Item 30: A computer-implemented method, comprising: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule comprising a first sequence of amino acid residues; determining a representation of the protein molecule based on at least the molecular structure file, the representation including a plurality of frames for each amino acid residue of the first sequence of amino acid residues, the plurality of frames for each amino acid residue including a first set of frames specifying a backbone geometry of the amino acid residue, and the plurality of frames for each amino acid residue further including a second set of frames specifying one or more torsion angles at a side chain of the amino acid residue; and generating the first three-dimensional structure of the protein molecule by at least applying a design computational model to modify the representation of the protein molecule.

[0190] Item 31: The method described in Item 30, wherein each frame of the plurality of frames corresponds to a degree of freedom for the design computational model to update the initial three-dimensional structure of the protein molecule.

[0191] Item 32: The method of items 30 or 31, wherein the first frame set includes a first frame that includes an affine transformation matrix that specifies a rotation and translation of a backbone of amino acid residues.

[0192] Item 33: The method described in Item 32, wherein the affine transformation matrix includes a rotation matrix that specifies a rotation of the backbone of the amino acid residues, and the affine transformation matrix further includes a displacement vector that specifies a translation of the backbone of the amino acid residues.

[0193] Item 34: The method of Item 32 or 33, wherein the first frame set further includes a second frame that specifies backbone torsion angles for amino acid residues.

[0194] Item 35: The torsion angle is the alpha carbon ( 35. The method according to item 34, wherein the rotatable bond between the TIFF2025533582000256.tif7170 atom and the carbonyl group of the backbone of the amino acid residue is related to the rotatable bond between the TIFF2025533582000256.tif7170 atom and the carbonyl group of the backbone of the amino acid residue.

[0195] Item 36: The method of any one of items 30 to 35, wherein the first frame set includes a first frame that specifies a first torsion angle for the backbone of the amino acid residues, and the first frame set further includes a second frame that specifies a second torsion angle for the backbone of the amino acid residues.

[0196] Item 37: The first torsion angle is at the alpha carbon ( TIFF2025533582000257.tif7170) atom and a carbon (C) atom, and the second torsion angle is related to the alpha carbon ( 37. The method of claim 36, wherein the second rotatable bond between the 2-amino-2-methyl-1,2-dioxane (TIFF2025533582000258.tif7170) atom and the nitrogen (N) atom is associated with the second rotatable bond between the 2-amino-2-methyl-1,2-dioxane (TIFF2025533582000258.tif7170) atom and the nitrogen (N) atom.

[0197] Item 38: The method of Item 37, wherein the first frame set further includes a third frame that specifies a third torsion angle present in the backbone of the amino acid residues.

[0198] Item 39: The method of Item 38, wherein the third torsion angle is associated with a third rotatable bond between a carbon (C) atom and a nitrogen (N) atom of the backbone of the amino acid residue.

[0199] Item 40: The method of any one of items 30 to 39, further comprising determining one or more coordinates of each atom comprising the first three-dimensional structure of the protein molecule based on at least the corrected representation of the protein molecule.

[0200] Item 41: The method described in Item 40, wherein one or more coordinates of each atom of a first three-dimensional structure of the protein molecule are determined based at least on a plurality of frames associated with each amino acid residue included in the corrected representation of the protein molecule.

[0201] Item 42: The method described in Item 41, wherein one or more coordinates of each atom of the first three-dimensional structure of the protein molecule are determined by at least determining one or more coordinates of a plurality of backbone atoms of the protein molecule.

[0202] Item 43: The method described in Item 42, wherein one or more coordinates of each atom of the first three-dimensional structure of the protein molecule are further determined by at least determining one or more coordinates of a plurality of side chain atoms of the protein molecule based on one or more coordinates of a plurality of backbone atoms of the protein molecule.

[0203] Item 44: The method of any one of items 30 to 43, wherein the design computational model comprises a machine learning model trained to generate a first three-dimensional structure of the protein molecule by denoising at least an initial three-dimensional structure of the protein molecule.

[0204] Item 45: The method described in Item 44, wherein the machine learning model denoises the initial three-dimensional structure of the protein molecule by performing at least a sequence of updates to the representation of the protein molecule.

[0205] Item 46: The method of Item 45, wherein the machine learning model is trained to reduce a loss function associated with each successive update of the initial three-dimensional structure of the protein molecule.

[0206] Item 47: The method of item 46, wherein the loss function is a frame alignment point error (FAPE) loss function and / or a structural violation loss function.

[0207] Item 48: The method of Item 45 or 46, wherein the machine learning model is trained to reduce an energy function associated with each successive update of the initial three-dimensional structure of the protein molecule.

[0208] Item 49: The method of any one of items 30 to 48, wherein the machine learning model is a diffusion model that removes a portion of the noise present in the initial three-dimensional structure of the protein molecule at each time step of a plurality of successive time steps.

[0209] Item 50: The method described in Item 49, wherein the diffusion model performs a first update to the representation of the protein molecule to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule, and the diffusion model further performs a second update to the representation of the protein molecule to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule.

[0210] Item 51: The method described in Item 50, wherein the diffusion model further adds a third amount of noise before performing the second update to remove the second amount of noise and a fourth amount of noise after performing the second update to remove the second amount of noise, and the third amount of noise and the fourth amount of noise are determined based on a noise schedule that defines a distribution of noise levels added over multiple consecutive time steps.

[0211] Item 52: The method according to Item 51, wherein the distribution of noise levels corresponds to the degrees of freedom present in the representation of the protein molecule of the computational model for modifying the initial three-dimensional structure of the protein molecule.

[0212] Item 53: The method of any one of items 49 to 52, wherein the updates performed by the diffusion model produce outputs equivariant to special Euclidean group SE(3) transformations.

[0213] Item 54: The method of any one of items 30 to 53, wherein modifying the representation of the protein molecule includes updating the first frame set to change the backbone geometric state of one or more amino acid residues of the protein molecule.

[0214] Item 55: The method of any one of items 30 to 54, wherein modifying the representation of the protein molecule includes updating the second frameset to change one or more torsion angles of side chains of one or more amino acid residues of the protein molecule.

[0215] Item 56: The method of any one of items 30 to 55, wherein the first three-dimensional structure of the protein molecule is associated with one or more desirable properties.

[0216] Item 57: The method of any one of items 30 to 56, wherein the first three-dimensional structure of the protein molecule is configured for one or more downstream tasks.

[0217] Item 58: The method of Item 57, wherein the one or more downstream tasks include determining one or more properties of the protein molecule based on at least the first three-dimensional structure of the protein molecule.

[0218] Item 59: The method of any one of items 30 to 58, further comprising: determining, based on at least a first three-dimensional structure of the protein molecule, that a first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties; and, in response to determining that the first sequence of amino acid residues exhibits the desired three-dimensional structure and / or desired properties, generating, based on at least the first sequence of amino acid residues, a second sequence of amino acid residues of a different protein molecule.

[0219] Item 60: The method of any one of items 30 to 59, wherein the representation of the protein molecule further comprises a logical vector indicating, for each position in the sequence of amino acid residues forming the protein molecule, the identity of the amino acid residue occupying the position by at least enumerating a probability distribution over the set of possible amino acid residues occupying the position.

[0220] Item 61: The method of Item 60, wherein the design computational model further generates a first three-dimensional structure of the protein molecule by modifying the identity of at least one amino acid residue of the first residue sequence while modifying a first frame set and / or a second frame set associated with the at least one amino acid residue.

[0221] Item 62: The method of any one of items 30 to 61, wherein the initial three-dimensional structure of the protein molecule contains noise in the identity of each amino acid residue and / or the spatial arrangement of multiple atoms forming each amino acid, and the noise is removed by designing a computational model that modifies the representation of the protein molecule.

[0222] Item 63: The method of Item 62, wherein the noise is Gaussian noise.

[0223] Item 64: The method of any one of items 30 to 63, wherein the representation of the protein molecule is further generated to include a plurality of polymer chains, each polymer chain including one or more amino acid residues from the first sequence of amino acid residues.

[0224] Item 65: The method of Item 64, wherein modifying the representation of the protein molecule includes modifying the positions of one or more amino acids in each polymer chain as a group.

[0225] Item 66: A system comprising at least one data processor and at least one memory storing instructions, the instructions, when executed by the at least one data processor, causing operations including the methods described in any one of items 30 to 65.

[0226] Item 67: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, produce operations including the methods described in any one of items 30 to 65.

[0227] 9 is a block diagram depicting an example of a computing system 1100 according to some exemplary embodiments. Referring to FIGS. 1 through 9, the computing system 1100 may be used to implement the molecular design engine 110, the molecular analysis engine 120, the client device 130, and / or any components therein.

[0228] 9, computing system 1100 may include a processor 1110, a memory 1120, a storage device 1130, and an input / output device 1140. The processor 1110, the memory 1120, the storage device 1130, and the input / output device 1140 may be interconnected via a system bus 1150. The processor 1110 is capable of processing instructions for execution within the computing system 1100. Such executed instructions may implement, for example, the molecular design engine 110, the molecular analysis engine 120, one or more components of the client device 130, etc. In some exemplary embodiments, the processor 1110 may be a single-threaded processor. Alternatively, the processor 1110 may be a multi-threaded processor. The processor 1110 is capable of processing instructions stored in the memory 1120 and / or the storage device 1130 to display graphical information for a user interface provided via the input / output device 1140.

[0229] Memory 1120 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 1100. Memory 1120 may store, for example, data structures representing a configuration object database. Storage device 1130 may provide persistent storage for computing system 1100. Storage device 1130 may be a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, or other suitable persistent storage means. Input / output device 1140 provides input / output operations for computing system 1100. In some exemplary embodiments, input / output device 1140 includes a keyboard and / or a pointing device. In various implementations, input / output device 1140 includes a display device for displaying a graphical user interface.

[0230] According to some demonstrative embodiments, the input / output devices 1140 may provide input / output operations for network devices. For example, the input / output devices 1140 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0231] In some exemplary embodiments, computing system 1100 may be used to execute various interactive computer software applications that may be used for organizing, analyzing, and / or storing various types of data. Alternatively, computing system 1100 may be used to execute any type of software application. These applications may be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or other objects), computing functions, communication functions, etc. Applications may include various add-in functions or may be standalone computing products and / or features. When active within an application, functionality may be used to generate a user interface that is provided via input / output devices 1140. The user interface may be generated by computing system 1100 and presented to a user (e.g., on a computer screen monitor, etc.).

[0232] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0233] These computer programs, sometimes referred to as programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.

[0234] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device, such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and a pointing device, such as, for example, a mouse or trackball, by which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, tactile feedback, etc., and input from the user may be received in any form, including acoustic input, voice input, and tactile input. Other possible input devices include touchscreens or other touch-sensitive devices such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, etc.

[0235] In the above specification and claims, phrases such as "at least one of" or "one or more of" may appear before a list of consecutive elements or features. The term "and / or" may also be used in listings of two or more elements or features. Unless otherwise implicitly or explicitly stated by the context in which it is used, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" are intended to mean "A only, B only, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, C," "one or more of A, B, C," and "A, B, and / or C" are intended to mean "A only, B only, C only, A and B together, A and C together, B and C together, or A, B and C together," respectively. Use of the term "based on" above and in the claims means "based at least in part on," and implies that unrecited features or elements are also permitted.

[0236] The subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the above description do not necessarily represent all implementations of the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. While several variations have been detailed above, other modifications and additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features disclosed above. In addition, the logic flow depicted in the accompanying figures and / or described herein does not necessarily require the particular order shown or sequential order to achieve desirable results. Other implementations may be within the scope of the following claims.

Claims

1. 1. A computer-implemented method comprising: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule comprising a first sequence of amino acid residues; determining, based on at least the molecular structure file, a representation of the protein molecule comprising a plurality of frames for each amino acid residue of the first sequence of amino acid residues, the plurality of frames for each amino acid residue comprising a first set of frames specifying a backbone geometry of the amino acid residue, and the plurality of frames for each amino acid residue further comprising a second set of frames specifying one or more torsion angles for a side chain of the amino acid residue; and generating a first three-dimensional structure of the protein molecule by at least applying a design computational model to modify the representation of the protein molecule; 11. A computer-implemented method comprising:

2. The method of claim 1 , wherein each frame of the plurality of frames corresponds to a degree of freedom for the design computational model to update the initial three-dimensional structure of the protein molecule.

3. 3. The method of claim 1, wherein the first frame set includes a first frame including an affine transformation matrix that specifies a rotation and translation of the backbone of the amino acid residue, and the first frame set further includes a second frame that specifies a torsion angle of the backbone of the amino acid residue.

4. 4. The method of claim 1, wherein the first set of frames comprises a first frame that specifies a first torsion angle for the backbone of the amino acid residue, and the first set of frames further comprises a second frame that specifies a second torsion angle for the backbone of the amino acid residue.

5. The first torsion angle is at the alpha carbon of the backbone of the amino acid residue ( ) atom and a carbon (C) atom of the backbone of the amino acid residue, and the second torsion angle is 5. The method of claim 4, wherein the second rotatable bond between the aryl (C) atom and the nitrogen (N) atom is associated with a second rotatable bond between the aryl (C) atom and the nitrogen (N) atom.

6. 6. The method of claim 5, wherein the first frame set further comprises a third frame that specifies a third torsion angle present in the backbone of the amino acid residue, the third torsion angle being associated with a third rotatable bond between the carbon (C) atom and the nitrogen (N) atom of the backbone of the amino acid residue.

7. determining one or more coordinates of a plurality of backbone atoms of the protein molecule based at least on the plurality of frames associated with each amino acid residue included in the modified representation of the protein molecule; and determining one or more coordinates of a plurality of side chain atoms of the protein molecule based on the one or more coordinates of the plurality of backbone atoms of the protein molecule; 7. The method of claim 1, further comprising:

8. 8. The method of claim 1, wherein the design computational model comprises a machine learning model trained to generate the first three-dimensional structure of the protein molecule by at least denoising the initial three-dimensional structure of the protein molecule.

9. 9. The method of claim 8, wherein the machine learning model denoises the initial three-dimensional structure of the protein molecule by at least performing a sequence of updates to the representation of the protein molecule.

10. 10. The method of claim 9, wherein the machine learning model is trained to reduce a loss function and / or an energy function associated with each successive update of the initial three-dimensional structure of the protein molecule.

11. 11. The method of claim 1, wherein the machine learning model is a diffusion model that removes a portion of the noise present in the initial three-dimensional structure of the protein molecule at each time step of a plurality of successive time steps.

12. 12. The method of claim 11 , wherein the diffusion model performs a first update to the representation of the protein molecule to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule, and wherein the diffusion model further performs a second update to the representation of the protein molecule to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule.

13. 13. The method of claim 12, wherein the diffusion model further adds a third amount of noise before performing the second update to remove the second amount of noise and a fourth amount of noise after performing the second update to remove the second amount of noise, the third amount of noise and the fourth amount of noise being determined based on a noise schedule that defines a distribution of noise levels added over the plurality of successive time steps.

14. 14. The method of claim 13, wherein the distribution of noise levels corresponds to degrees of freedom present in the representation of the protein molecule of the computational model for modifying the initial three-dimensional structure of the protein molecule.

15. 15. The method of claim 11, wherein each update performed by the diffusion model produces an output equivariant to a special Euclidean group SE(3) transformation.

16. 16. The method of any one of claims 1 to 15, wherein the modifying the representation of the protein molecule comprises updating the first frameset to change the geometric state of the backbone of one or more amino acid residues of the protein molecule.

17. 17. The method of any one of claims 1 to 16, wherein the modifying the representation of the protein molecule comprises updating the second frameset to change the one or more torsion angles of the side chains of one or more amino acid residues of the protein molecule.

18. 18. The method of any one of claims 1 to 17, wherein the first three-dimensional structure of the protein molecule is associated with one or more desirable properties.

19. 19. The method of any one of claims 1 to 18, wherein the first three-dimensional structure of the protein molecule is configured for one or more downstream tasks.

20. determining, based on at least the first three-dimensional structure of the protein molecule, that the first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties; and generating a second sequence of amino acid residues for a different protein molecule based on at least the first sequence of amino acid residues in response to determining that the first sequence of amino acid residues exhibits the desired three-dimensional structure and / or the desired property.

20. The method of any one of claims 1 to 19, further comprising:

21. 21. The method of any one of claims 1 to 20, wherein the representation of the protein molecule further comprises, for each position in the sequence of amino acid residues forming the protein molecule, a logical vector indicating the identity of the amino acid residue occupying the position by at least enumerating a probability distribution over a set of possible amino acid residues that occupy the position.

22. 22. The method of claim 21, wherein the design computational model further generates the first three-dimensional structure of the protein molecule by modifying the identity of at least one amino acid residue in a first residue sequence while modifying the first frameset and / or the second frameset associated with the at least one amino acid residue.

23. 23. The method of any one of claims 1 to 22, wherein the initial three-dimensional structure of the protein molecule contains noise in the identity of each amino acid residue and / or the spatial arrangement of multiple atoms forming each amino acid, and wherein the noise is removed by the design computational model modifying the representation of the protein molecule.

24. 24. The method of any one of claims 1 to 23, wherein the representation of the protein molecule is further generated to include a plurality of polymer chains, each polymer chain comprising one or more amino acid residues from the first sequence of amino acid residues, and wherein the representation of the protein molecule is modified by a protein design computational model that modifies the positions of the one or more amino acids in each polymer chain as a group.

25. 1. A system comprising: at least one data processor; and At least one memory storing instructions which, when executed by said at least one data processor, cause operations including the method of any one of claims 1 to 24. A system comprising:

26. A non-transitory computer readable medium storing instructions that, when executed by at least one data processor, cause operations including the method of any one of claims 1 to 24.