Machine learning enabled prediction of molecular structures and properties
By receiving the initial three-dimensional structural file of protein molecules, the representation of protein molecules is generated based on the molecular structure file, and the representation of protein molecules is modified through the design calculation model, which solves the problem of difficult to predict the structure and properties of protein molecules in the prior art, and achieves a fast and accurate protein design.
Patent Information
- Application Number
- CN202380068576.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-17
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to effectively predict molecular structures and properties, especially in protein design, and it is difficult to quickly and accurately generate protein sequences and three-dimensional structures with desired properties.
A system and method are provided to receive the initial three-dimensional structural file of a protein molecule through a computer program product, determine the representation of a protein molecule based on the molecular structure file, including the geometric states of the main chain and side chains of multiple frameworks designated by amino acid residues, and modify the representation through design calculation model to generate the three-dimensional structure of the protein molecule.
The three-dimensional structure and properties of protein molecules are achieved quickly and accurately predicted, reducing the pressure on computing resources, and improving the efficiency and effectiveness of protein design.
Smart Images

Figure CN120035862A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 377,335, filed on September 27, 2022, entitled “Machine Learning-Enabled Prediction of Molecular Structure and Properties”; U.S. Provisional Patent Application No. 63 / 387,680, filed on December 15, 2022, entitled “Machine Learning-Enabled Prediction of Molecular Structure and Properties”; U.S. Provisional Patent Application No. 63 / 499,333, filed on May 1, 2023, entitled “Machine Learning-Enabled Prediction of Molecular Structure and Properties”; and U.S. Provisional Patent Application No. 63 / 502,753, filed on May 17, 2023, entitled “Machine Learning-Enabled Prediction of Molecular Structure and Properties”. The disclosures of the above provisional applications are all incorporated herein by reference in their entirety. Technical Field
[0003] The subject matter described herein relates generally to molecular design, and more specifically to machine learning-based techniques for predicting molecular structure and properties. Background Art
[0004] A molecule is an atomic group that is bonded together by two or more atoms through chemical bonds. A molecule forms the smallest recognizable unit into which a pure substance can be divided while still retaining the composition and chemical properties of the substance. An example of a molecule is a protein molecule, while examples of non-protein molecules include small molecules, ions, nucleic acids, polysaccharides, glycolipids, and / or the like. The function and properties of a molecule may depend on its three-dimensional structure. For example, proteins are responsible for many important cellular functions, including, for example, enzymatic reactions, molecular transport, regulation and execution of many biological pathways, cell growth, proliferation, nutrient uptake, morphology, movement, intercellular communication, and / or the like. A protein structure may include one or more polypeptides, which are chains of amino acid residues linked together by peptide bonds. The sequence of amino acid residues in the polypeptide chains that form the protein structure determines the three-dimensional structure of the protein (e.g., the tertiary structure of the protein). In addition, the sequence of amino acids in the polypeptide chains that form the protein determines the basic function of the protein. Therefore, one goal of de novo protein design includes constructing a sequence of one or more amino acid residues that exhibits desired properties, rather than undesirable properties. For example, in the context of macromolecule drug discovery, de novo protein design typically seeks to identify sequences of amino acid residues (e.g., antibodies, and / or the like) that are capable of binding to antigens (such as viral antigens, tumor antigens, and / or the like). Summary of the invention
[0005] Systems, methods and articles, including computer program products, for molecular structure and property prediction are provided. In one aspect, a system for molecular structure and property prediction is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code, which provides operations when executed by at least one processor. These operations may include: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule, the protein molecule comprising a first sequence of amino acid residues; determining a representation of the protein molecule based at least on the molecular structure file, the representation comprising a plurality of frameworks for each amino acid residue in the first sequence of amino acid residues, the plurality of frameworks for each amino acid residue comprising a first set of frameworks specifying a geometric state for the main chain of the amino acid residue, and the plurality of frameworks for each amino acid residue further comprising a second set of frameworks specifying one or more torsion angles in the side chain of the amino acid residue; and generating a first three-dimensional structure of the protein molecule by modifying the representation of the protein molecule by at least applying a design computational model.
[0006] In another aspect, a method for molecular structure and property prediction is provided. The method may include: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule, the protein molecule comprising a first sequence of amino acid residues; determining a representation of the protein molecule based at least on the molecular structure file, the representation comprising a plurality of frameworks for each amino acid residue in the first sequence of amino acid residues, the plurality of frameworks for each amino acid residue comprising a first set of frameworks specifying a geometric state for the main chain of the amino acid residue, and the plurality of frameworks for each amino acid residue further comprising a second set of frameworks specifying one or more torsion angles in the side chain of the amino acid residue; and generating a first three-dimensional structure of the protein molecule by modifying the representation of the protein molecule by at least applying a design computational model.
[0007] In another aspect, a computer program product for molecular structure and property prediction is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that cause operations when executed by at least one data processor. These operations may include: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule, the protein molecule comprising a first sequence of amino acid residues; determining a representation of the protein molecule based at least on the molecular structure file, the representation comprising a plurality of frameworks for each amino acid residue in the first sequence of amino acid residues, the plurality of frameworks for each amino acid residue comprising a first set of frameworks specifying a geometric state for the main chain of the amino acid residue, and the plurality of frameworks for each amino acid residue further comprising a second set of frameworks specifying one or more torsion angles in the side chain of the amino acid residue; and modifying the representation of the protein molecule by at least applying a design computational model to generate a first three-dimensional structure of the protein molecule.
[0008] In some variations of the methods, systems, and non-transitory computer-readable media, one or more of the following features may optionally be included in any feasible combination.
[0009] In some variations, each frame in the plurality of frames corresponds to a degree of freedom used to design a computational model to update an initial three-dimensional structure of a protein molecule.
[0010] In some variations, the first set of frameworks may include a first framework comprising an affine transformation matrix that specifies rotation and translation of the backbone of an amino acid residue. The first set of frameworks may further include a second framework that specifies a torsion angle in the backbone of an amino acid residue.
[0011] In some variations, the first set of frameworks may include a first framework that specifies a first torsion angle in the backbone of an amino acid residue. The first set of frameworks may further include a second framework that specifies a second torsion angle in the backbone of an amino acid residue.
[0012] In some variations, the first torsion angle may be about the alpha carbon (C α ) atom and the carbon (C) atom. The second torsion angle can be associated with the α carbon (C) in the main chain of the amino acid residue. α ) atom and the second rotatable bond between the nitrogen (N) atom.
[0013] In some variations, the first set of frameworks may further include a third framework that specifies a third torsion angle present in the backbone of an amino acid residue. The third torsion angle may be associated with a third rotatable bond between a carbon (C) atom and a nitrogen (N) atom in the backbone of an amino acid residue.
[0014] In some variations, based at least on a plurality of frameworks associated with each amino acid residue included in the modified representation of the protein molecule, one or more coordinates of a plurality of main-chain atoms in the protein molecule may be determined. Based on the one or more coordinates of the plurality of main-chain atoms in the protein molecule, one or more coordinates of a plurality of side-chain atoms in the protein molecule may be determined.
[0015] In some variations, the design computational model may include a machine learning model that is trained to generate a first three-dimensional structure of a protein molecule by denoising at least an initial three-dimensional structure of the protein molecule.
[0016] In some variations, the machine learning model may denoise the initial three-dimensional structure of the protein molecule by performing a series of updates to at least the representation of the protein molecule.
[0017] In some variations, the machine learning model may be trained to reduce a loss function and / or energy function associated with each successive update to the initial three-dimensional structure of the protein molecule.
[0018] In some variations, the machine learning model may be a diffusion model that removes a portion of the noise present in the initial three-dimensional structure of the protein molecule at each time step in a plurality of consecutive time steps.
[0019] In some variations, the diffusion model may perform a first update on the representation of the protein molecule to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule. The diffusion model may further perform a second update on the representation of the protein molecule to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule.
[0020] In some variations, the diffusion model may further add a third amount of noise before performing a second update to remove the second amount of noise, and add a fourth amount of noise after performing a second update to remove the second amount of noise. The third amount of noise and the fourth amount of noise may be determined based on a noise schedule that defines a distribution of noise levels added across a plurality of consecutive time steps.
[0021] In some variations, the distribution of noise levels may correspond to the degrees of freedom present in the representation of the protein molecule that are used by the computational model to modify the initial three-dimensional structure of the protein molecule.
[0022] In some variations, each update performed by the diffusion model generates an output that is equivariant to the special Euclidean group SE(3) transformation.
[0023] In some variations, modifying the representation of the protein molecule may include updating the first set of frames to change the geometry of the backbone of one or more amino acid residues in the protein molecule.
[0024] In some variations, modifying the representation of the protein molecule may include updating the second set of frameworks to change one or more torsion angles in the side chains of one or more amino acid residues in the protein molecule.
[0025] In some variations, the first three-dimensional structure of the protein molecule may be associated with one or more desired properties.
[0026] In some variations, the first three-dimensional structure of the protein molecule may be configured for use in one or more downstream tasks.
[0027] In some variations, based on at least the first three-dimensional structure of the protein molecule, it can be determined that a first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties. In response to determining that the first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties, a second sequence of amino acid residues can be generated for a different protein molecule based on at least the first sequence of amino acid residues.
[0028] In some variations, the representation of the protein molecule may further include, for each position in the sequence of amino acid residues forming the protein molecule, a logic vector indicating the identity of the amino acid residue occupying the position by at least enumerating a probability distribution of a set of possible amino acid residues occupying the position.
[0029] In some variations, the design computational model may further generate a first three-dimensional structure of the protein molecule by modifying the identity of at least one amino acid residue in the first sequence of residues while modifying the first set of frameworks and / or the second set of frameworks associated with the at least one residue.
[0030] In some variations, the initial three-dimensional structure of the protein molecule may include noise in the identification of each amino acid residue and / or the spatial arrangement of the multiple atoms that form each amino acid. The noise is removed by modifying the representation of the protein molecule by designing a computational model.
[0031] In some variations, the representation of the protein molecule can be further generated to include a plurality of polymer chains. Each polymer chain includes one or more amino acid residues from a first sequence of amino acid residues. The representation of the protein molecule is modified by the protein design computational model to modify the position of one or more amino acids in each polymer chain as a group.
[0032] Specific implementations of the current subject matter may include, but are not limited to, methods consistent with the description provided herein and articles including tangibly embodied machine-readable media that are operable to cause one or more machines (e.g., computers, etc.) to cause operations that implement one or more of the features described. Similarly, a computer system that may include one or more processors and one or more memories coupled to the one or more processors is also described. A memory that may include a non-transitory computer-readable or machine-readable storage medium may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method consistent with one or more implementations of the current subject matter may be implemented by one or more data processors present in a single computing system or multiple computing systems. Such multiple computing systems may be connected and may exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.) via a direct connection between one or more computing systems in the multiple computing systems, etc.
[0033] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will become apparent by reference to the description and drawings, and to the claims. Although certain features of the presently disclosed subject matter are described for purposes of illustration related to protein design, it should be readily understood that these features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, help explain some principles associated with the disclosed embodiments.
[0035] According to some exemplary embodiments, Figure 1A depicts a system diagram illustrating an example of a molecular design system;
[0036] According to some exemplary embodiments, Figure 1B depicts a flow chart showing an example of a process 160 for protein structure and property prediction;
[0037] According to some exemplary embodiments, Figure 2A An example of a thick node representation of a protein sequence excluding hydrogen (H) atoms is depicted;
[0038] According to some exemplary embodiments, Figure 2B Another example of a thick node representation of a protein sequence containing hydrogen (H) atoms is depicted;
[0039] According to some exemplary embodiments, Figure 3A Depicted is a visualization of the relative positions of atoms in a coarse-grained node associated with the amino acid residue tryptophan (Trp);
[0040] According to some exemplary embodiments, Figure 3B Depicted is a visualization of the relative positions of atoms in a coarse-grained node associated with the amino acid residue valine (Val);
[0041] According to some exemplary embodiments, Figure 4 A visualization depicting the iterative updates made by the computational model to determine the three-dimensional structure of a protein molecule;
[0042] According to some exemplary embodiments, Figure 5A depicts a flow chart showing an example of a process for protein structure and property prediction;
[0043] According to some exemplary embodiments, Figure 5B depicts a flow chart showing another example of a process for protein structure and property prediction;
[0044] According to some exemplary embodiments, Figure 5C depicts a flow chart showing another example of a process for protein structure and property prediction;
[0045] According to some exemplary embodiments, Fig. 6A depicts schematic diagrams showing atomic structures of examples of amino acid residues;
[0046] According to some exemplary embodiments, Figure 6B depicts a schematic diagram showing an example of a diffusion framework;
[0047] According to some exemplary embodiments, Fig. 7A depicts a screenshot showing an example of a protein molecule undergoing backbone translation;
[0048] According to some exemplary embodiments, Figure 7B depicts a screenshot showing an example of a protein molecule undergoing backbone rotation;
[0049] According to some exemplary embodiments, Figure 7C Screen shot depicting another example of a protein molecule experiencing changes in side chain torsion angles;
[0050] According to some exemplary embodiments, Fig. 8A depicts a diagram showing an example of a noise table for a diffusion model that modifies the torsion angles of a protein structure;
[0051] According to some exemplary embodiments, Figure 8B depicts a diagram showing another example of a noise table for a diffusion model that modifies all degrees of freedom of a protein structure except the center of mass;
[0052] According to some exemplary embodiments, Figure 8C depicts a diagram showing another example of a noise table for performing a diffusion model of molecular docking between two molecules by modifying the respective centers of mass of the two molecules;
[0053] Fig.8D depicts a diagram showing another example of a noise schedule for performing a diffusion model of molecular docking between two molecules by modifying the center of mass; and
[0054] Fig. 9 Depicted is a block diagram illustrating an example of a computing system in accordance with some example embodiments.
[0055] When applicable, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION
[0056] The properties of a molecule may depend on its composition and structure. For example, the properties of a small molecule, including its safety and efficacy as a therapeutic agent, may depend on the number of atoms of each constituent element (e.g., molecular formula or empirical formula) and the bonding arrangement between these atoms (e.g., structural formula). At the same time, the properties of a macromolecule, such as the binding affinity and developability of a protein molecule, may be determined by the sequence of amino acid residues forming the macromolecule and the three-dimensional structure adopted by the sequence of amino acid residues. Therefore, the development of small and macromolecular therapies takes into account the composition and structure of each candidate molecule.
[0057] In the case of protein design, where key goals include identifying a protein sequence (e.g., a sequence of amino acid residues) that exhibits certain desired properties, knowledge of the properties of a protein molecule may be limited if its three-dimensional structure is not first determined. Therefore, the task of protein structure prediction, which involves inferring the three-dimensional structure of a protein molecule based on the sequence of amino acids that form the protein molecule, may be a key component of protein design. In contrast, structure-agnostic predictions of protein sequences may produce many protein sequences that fail to adopt a three-dimensional structure capable of binding to a target molecule. In this context, the sequence of amino acid residues that form a protein molecule may also be referred to as the primary structure of the protein molecule. The three-dimensional structure of a protein molecule includes one or more secondary structures formed when individual amino acid residues are linked by hydrogen bonds, and a tertiary structure formed when the secondary structures are arranged along a single polypeptide chain, called the backbone of the protein molecule.
[0058] Although comprehensive design approaches that combine sequence and structure design can produce protein molecules that are more likely to exhibit desired properties (such as binding affinity and developability), there are countless variations in the sequence and three-dimensional structure of protein molecules. For example, the solution space occupied by each possible arrangement of amino acid residues that can form a protein molecule is huge (for example, for a protein sequence consisting of N amino acid residues, there are approximately 20 N possible arrangements). At the same time, a sequence of a single amino acid residue can fold into many different three-dimensional structures. In some cases, due to its flexible nature, the three-dimensional structure of a protein molecule may evolve over time, especially when the protein molecule is in close proximity and interacting with other molecules. Each three-dimensional structure or conformation of a protein molecule may include a different spatial arrangement of the atoms in each of the constituent amino acid residues. Therefore, even if limited to a few discrete (rather than continuous) structural changes, the solution space occupied by each possible three-dimensional structure that can be formed by a sequence of single amino acid residues is already very large. For example, even if each amino acid residue is limited to assuming one of three discrete geometric states (e.g., rotamers), a protein sequence of N amino acid residues can have approximately 3 N Possible conformations. For protein sequences of meaningful length, brute-force searches of the solution space of every possible arrangement of amino acid residues combined with the solution space of every possible three-dimensional structure are computationally intractable. The strain on computational resources is further exacerbated by conventional structural bioinformatics approaches, which are prohibitively slow due to a variety of computational inefficiencies. For at least the reasons outlined above, current efforts to design protein sequences that are able to adopt a certain three-dimensional structure are limited by the challenging trade-off between designing for a functional output protein and the computational burden.
[0059] The present disclosure eliminates the trade-off between generating functional protein design and computational burden by strategically reducing the joint solution space of integrated sequence and structural design. For example, in some cases, the design of molecules (such as protein molecules) can be performed on a molecular representation that enhances the performance of various design computational models. In this article, the representation of a molecule may include a structural representation of the three-dimensional structure of the molecule, which indicates the spatial arrangement of atoms forming the molecule. For example, when the molecule is a protein molecule, the structural representation of the molecule may indicate the spatial arrangement of atoms forming each amino acid residue in the molecule. In the case where the molecule is a protein molecule, the representation of the molecule may also include a sequence representation indicating the identity of each amino acid residue forming the molecule.
[0060] As described in more detail below, the sequence (e.g., primary structure) and three-dimensional structure (e.g., secondary structure and tertiary structure) of a protein molecule can be determined by applying a computational model to a representation of the protein molecule. For example, in some cases, the sequence and three-dimensional structure of a protein molecule can be determined by applying a series of modified computational models to the representation of the protein molecule. In some cases, the representation of the protein molecule can impose certain restrictions on the individual modifications made to the sequence and three-dimensional structure of the protein molecule by the computational model. For example, in some cases, the structural representation of the protein molecule can prevent arbitrary and unfeasible modifications to the position of a single atom in three-dimensional space. Therefore, the computational model that operates on the representation of the protein molecule can strategically reduce the joint solution space searched by the computational model to generate the sequence and structure of the protein molecule, while minimizing the damage to the accuracy of the design results.
[0061] In some exemplary embodiments, the structural representation of a molecule (such as a protein molecule) may be a coarse-grained (CG) node representation (e.g. Figure 3A and 3B As shown). That is, in some cases, the three-dimensional structure of a molecule, including the position (e.g., three-dimensional coordinates) of each atom included in the molecule, can be represented as a collection of coarse-grained (CG) nodes. Although the molecule may be a protein molecule, it should be understood that the molecule may also be a non-protein molecule (e.g., a small molecule, a nucleic acid, a polysaccharide, a glycolipid, and / or the like). In the case of a protein molecule, each amino acid residue contained in the protein molecule may include one or more representative structures, which include, for example, a rigid body, a flexible or variable body, and / or the like. In addition, each amino acid residue in a protein molecule may be represented by a set of coarse-grained nodes, each of which corresponds to one of the structures forming the amino acid residue. As a rigid body, a single structure may include an atomic group of two or more atoms, the position of which is fixed relative to the coordinates of the structure. As a flexible or variable body, a single structure may include an atomic group of two or more atoms, the position of which is flexible to a certain extent relative to the coordinates of the structure. Therefore, each coarse-grained node associated with an amino acid residue may include the atoms contained in the corresponding structure. In some cases, a single atom may be part of multiple structures of an amino acid residue, and is therefore included in multiple corresponding coarse-grained nodes. In addition, the positions of constituent atoms in an amino acid residue can be specified by a rotation R and / or a translation T of the corresponding coarse-grained node.
[0062] In some exemplary embodiments, the structure of a molecule (such as a protein molecule) may be represented by a coordinate transformation (e.g., Euclidean transformation) and a geometric tensor embedding of each coarse-grained node associated with the molecule. For example, in some cases, the rotation R and / or translation T of each coarse-grained node associated with the molecule may be represented as a set of geometric tensors (or geometric tensor embeddings), each of which has a configurable maximum degree L. As used herein, the term "geometric tensor" may refer to an object (e.g., a scalar, a vector, and / or the like) that undergoes a transformation when undergoing one or more coordinate transformations (e.g., Euclidean transformations) such as (such as rotations, translations, and / or the like). In some cases, a set of geometric tensors associated with a coarse-grained node may undergo a coordinate transformation corresponding to one or more group elements of a three-dimensional rotation group. The above-mentioned three-dimensional rotation group may describe the possible rotational symmetry and orientation of a structure in a multidimensional space (e.g., a three-dimensional space, and / or the like). In this article, each group element of the three-dimensional rotation group may be represented as one or more irreducible representations that cannot be further decomposed. Thus, a geometry tensor embedding of a coarse-grained node may include a set of geometry tensors, each of which is manipulated according to one or more group elements from a three-dimensional rotation group to describe the current translation and rotation of the corresponding structure. A coarse-grained (CG) node representation of a molecule (such as a protein molecule) can reduce the joint solution space searched by a computational model by at least avoiding modifications to the position of each individual atom within the molecule. In contrast, when a computational model is run on a coarse-grained (CG) node representation of a molecule, the computational model may apply Euclidean transformations (e.g., translations and rotations) to groups of two or more atoms that are more likely to move as a collective whole.
[0063] In some exemplary embodiments, the structural representation of a molecule (such as a protein molecule) may be a backbone torsion (BBT) representation (e.g. Fig. 6A For example, in the case of a protein molecule, for each constituent amino acid residue of the protein molecule, the backbone torsion (BBT) representation of the protein molecule may include a plurality of frameworks defining the positions of the atoms forming the amino acid residue by specifying at least the geometric states of the backbone and side chains of the amino acid residue. In some cases, the plurality of frameworks may include a first set of frameworks that specify the geometric states of the backbone of the corresponding amino acid residue. As used herein, the "backbone" of an amino acid residue may include atoms common to each amino acid residue (e.g., nitrogen (N) atoms, alpha carbon (C α ) atom and a carboxyl carbon (C) atom). In addition, the plurality of frameworks may include a second set of frameworks that specify one or more torsion angles present in the side chains of the amino acid residues. As used herein, the term "torsion angle", which may be used interchangeably with the term "dihedral angle", may refer to an angle that describes a rotation around the central bond of a polypeptide chain segment that includes four atoms coupled by three consecutive bonds.
[0064] For a single amino acid residue containing multiple atoms, each framework can define the position of the atom in three-dimensional space (e.g., the three-dimensional coordinates of each atom) to one or more internal degrees of freedom (DoF) mapping. In this article, the internal degrees of freedom (DoF) can refer to the type and / or degree of modification that can be performed on the amino acid residues, such as by a computational model, when determining the sequence and / or three-dimensional structure of a protein molecule containing an amino acid residue. For example, in some cases, certain degrees of freedom (DoF) can limit the change of the amino acid residue identity to one of the 20 standard amino acid residues. Alternatively and / or additionally, some degrees of freedom (DoF), such as main chain translation, main chain rotation and torsion angle, can impose restrictions on the spatial range, within which each atom in the amino acid residue can move as part of the overall three-dimensional structure of the amino acid residue. That is, some frameworks can limit the manner and extent of rearrangement of each atom in three-dimensional space, such as relative to other atoms in the amino acid residue, thereby preventing atoms from being able to freely move to any arbitrary position (e.g., coordinates) in three-dimensional space. For example, in the case of main chain translation and rotation, the corresponding degrees of freedom (DoF) may require the main chain atoms of the amino acid residues to translate and rotate as a group of atoms, thereby preventing the individual main chain atoms from moving to change the relative spatial arrangement between them. For the torsion angle between two side chain atoms connected by a bond, the corresponding degrees of freedom (DoF) may require one atom to rotate around another atom without changing the distance (or bond length) therebetween. As described in more detail below, each framework can correspond to a degree of freedom (DoF) of a computational model to update a protein sequence (e.g., the identities of the constituent amino acid residues) and / or the three-dimensional structure of a protein sequence.
[0065] In some exemplary embodiments, the backbone torsion (BBT) representation of a protein molecule can specify the geometric state of the backbone of each amino acid residue in a variety of different ways. For example, in some cases, the geometric state of the backbone of an amino acid residue can be based on its translation and rotation as well as the alpha carbon (C α ) atom and the carbonyl group. Thus, in some cases, the first set of frameworks may include a first framework that specifies the rotation and translation of the main chain of an amino acid residue. For example, in some cases, the first framework may include an affine transformation matrix that includes a rotation matrix that specifies the rotation of the main chain of an amino acid residue and a displacement vector that specifies the translation of the main chain of an amino acid residue. In the case where the first framework specifies the rotation and translation of the main chain of an amino acid residue, the first set of frameworks may further include specifying the torsion angle in the main chain of the amino acid residue (e.g., the alpha carbon (C α ) the second framework of the torsion angle of the rotatable bond between the ) atom and the carbonyl group.
[0066] In some exemplary embodiments, in addition to specifying the geometric state of the main chain of an amino acid residue based on a torsion angle and a combination of translation and rotation thereof, the geometric state of the main chain of an amino acid residue may also be specified based on the torsion angle present in the main chain of the amino acid residue. Thus, in some cases, the first set of frameworks may include specifying the alpha carbon (C α ) atom and the carbonyl group. In addition, in those cases, the first set of frames may further include a third frame and a fourth spanning frame. The third frame may specify the α carbon (C α ) atom and the nitrogen (N) atom. Meanwhile, the fourth framework may specify the fourth torsion angle of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom in the main chain of the amino acid residue.
[0067] In some exemplary embodiments, the three-dimensional structure of a molecule (such as a protein molecule) can be determined by at least applying a computational model to modify the representation of the initial three-dimensional structure of the molecule. For example, in some cases, the computational model can modify the structural representation of the molecule (e.g., coarse-grained (CG) node representation, main chain torsion (BBT) representation, and / or the like) to determine the three-dimensional structure of the molecule. Alternatively and / or additionally, in the case of protein design, the computational model can modify the sequence representation of the molecule in addition to modifying the structural representation of the molecule to determine the identity of the amino acid residues that form the molecule. It should be understood that the initial three-dimensional structure of the molecule may include varying degrees of entropy, which decreases as the computational model modifies the molecule. For example, in some cases, the initial three-dimensional structure of the molecule may include at least some noise (e.g., Gaussian noise, and / or the like) in the positions (e.g., three-dimensional coordinates) of the constituent atoms. In the case where the molecule is a protein molecule, the initial three-dimensional structure of the molecule may include one or more groups of amino acid residues corresponding to one or more polymer chains present in the molecule.
[0068] In some cases, the three-dimensional structure of the molecule can be determined by at least determining one or more coordinates of each atom in the three-dimensional structure of the molecule based on at least a modified representation of the molecule output by the computational model. For example, in the case where the molecule is a protein molecule whose three-dimensional structure is presented in a main chain torsion (BBT) representation, one or more coordinates of each atom in the three-dimensional structure of the molecule can be determined based on at least a plurality of frameworks, which are associated with each amino acid residue included in the modified representation of the molecule output by the computational model. In some cases, the one or more coordinates of each atom in the three-dimensional structure of the molecule can be determined by at least determining one or more coordinates of a plurality of main chain atoms in the molecule. In addition, in some cases, the one or more coordinates of each atom in the three-dimensional structure of the molecule can be further determined by determining one or more coordinates of a plurality of side chain atoms in the molecule based on at least one or more coordinates of a plurality of main chain atoms in the molecule.
[0069] In some exemplary embodiments, where the three-dimensional structure of a protein molecule is presented in a coarse-grained (CG) node representation, the three-dimensional structure of the protein molecule can be determined by applying a computational model to a tensor associated with each coarse-grained node in the molecule. For example, the structural computational model may receive an input, which, for each coarse-grained (CG) node in the protein molecule, includes a set of geometric tensors representing the rotation R and / or translation T of the coarse-grained (CG) node in the initial three-dimensional structure of the protein molecule. In addition, the computational model may perform continuous updates on the rotation R and / or translation T of one or more coarse-grained nodes in the initial three-dimensional structure in order to derive the three-dimensional structure of the molecule. Alternatively, as noted, the computational model may generate a three-dimensional structure of the molecule by modifying at least one or more frameworks in the main chain torsion (BBT) representation of the molecule. For example, in the case of protein design, the computational model may ingest an input, which, for each amino acid residue in the protein sequence, includes a representation of the amino acid residue identity, a translation of the main chain atoms, a rotation of the main chain atoms, and one or more torsion angles. In some cases, the computational model may modify the sequence representation of the molecule to determine the sequence of amino acid residues that form the molecule (e.g., the primary structure of the molecule), in addition to modifying the structural representation of the molecule (e.g., the coarse-grained (CG) node or backbone torsion (BBT) representation of the molecule) to determine the three-dimensional structure of the molecule. In some cases, the computational model may generate a three-dimensional structure of the molecule to exhibit one or more desired properties, such as binding affinity to another molecule (e.g., a viral antigen, a tumor antigen, and / or the like), specificity to another molecule, lack of nonspecificity, stability (e.g., conformational stability, thermodynamic stability, robustness to different environmental stresses, such as protease resistance, and / or the like), non-immunogenicity, human nature, lack of self-association (or non-aggregation), lack of chemical propensity (e.g., aspartate isomerization, oxidation, deamidation), developability, and / or the like. Alternatively and / or additionally, the three-dimensional structure of the molecule may be a structure suitable for or configured for one or more downstream tasks, such as predictive analysis of various properties exhibited by the molecule.
[0070] In some exemplary embodiments, the computational model may be a machine learning model capable of recognizing the same three-dimensional structure regardless of the orientation of the three-dimensional structure ingested as input. For example, in some cases, the computational model may be implemented as a geometric deep learning model, such as an equivariant neural network (ENN), a multi-body high-order equivariant message passing neural network, and / or the like. In the case where the computational model operates on a coarse-grained (CG) node representation of a molecule, the three-dimensional structure of the molecule may be modified by changing the rotation R and / or translation T of one or more coarse-grained nodes associated with the molecule, thereby changing the relative position of the coarse-grained nodes within the molecule. Alternatively, in the case where the computational model operates on a main-chain torsion (BBT) representation of a molecule, the three-dimensional structure of the molecule may be modified by updating at least the framework of one or more amino acid residues in the molecule. For example, the main-chain torsion (BBT) representation of the molecule may be modified by modifying at least the first set of frameworks to change the geometric state of the main chain of one or more amino acid residues and / or modifying the second set of frameworks to change the torsion angle in the side chain of one or more amino acid residues. However, a change in the orientation of the entire three-dimensional structure of a molecule, whether by rotating or translating the three-dimensional structure as a whole, does not constitute a change in the three-dimensional structure of the molecule if it does not change the relative positions of the atoms (or groups of atoms) contained therein. Therefore, the computational model is able to recognize when two three-dimensional structures have different orientations in space but are otherwise identical. In doing so, the computational model is able to generate the correct three-dimensional structure regardless of the orientation of the initial three-dimensional structure taken as input.
[0071] In some exemplary embodiments, the computational model may include a machine learning model that is trained to determine the three-dimensional structure of a molecule (such as a protein molecule) by denoising at least the initial three-dimensional structure of the molecule. In some cases, the structural representation of the initial three-dimensional structure of the molecule may be denoised, including, for example, a coarse-grained (CG) node representation, a main chain torsion (BBT) representation, and / or the like. In addition, in the case of protein design, the sequence representation of the initial sequence of amino acid residues forming the molecule may be denoised. The machine learning model may denoise the initial three-dimensional structure and / or sequence of the molecule by performing a series of updates on at least the structural representation and / or sequence representation of the molecule. For example, in some cases, the machine learning model may be a diffusion model that removes a portion of the noise present in the initial three-dimensional structure and / or sequence of the molecule at each time point in a series of time points. In some cases, the denoising performed at each time point may include an incremental update to the structural representation and / or sequence representation of the molecule. In addition, in some cases, after removing a portion of the noise from the representation of the molecule at a first time point, before the diffusion model performs its next update at a second time point, a second amount of noise may be added back to the representation of the molecule. The second amount of noise added by the diffusion model can be determined by a noise schedule that defines a distribution of noise levels across successive updates performed by the diffusion model. In some cases, the distribution of noise levels can correspond to degrees of freedom (DoF) that are applicable to the computational model to modify the initial three-dimensional structure of the molecule. For example, in some cases, more noise can be added to degrees of freedom (DoF) where there is more entropy than to degrees of freedom (DoF) where there is less entropy. Adding the second amount of noise can compensate for errors that may be introduced by denoising the molecular representation.
[0072] According to some exemplary embodiments, Figure 1A A system diagram illustrating an example of a molecular design system 100 is depicted. Figure 1A , the molecular design system 100 may include a molecular design engine 110, a molecular analysis engine 120, and a client device 130 with a user interface (UI) 145. Figure 1A As shown, the molecular design engine 110, the molecular analysis engine 120, and the client device 130 may be communicatively coupled via a network 140. The client device 130 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smart phone, a tablet computer, a wearable device, etc. The network 140 may be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.
[0073] In some exemplary embodiments, the molecular design engine 110 may generate a molecule by determining at least a sequence (or molecular formula) and a corresponding three-dimensional structure of the molecule. Figure 1A As shown, the molecular design engine 110 may include a sequence design computational model 113, a representation generator 115, and a molecular design computational model 117. In some cases where the molecular design engine 110 is deployed to perform protein design, the molecular design computational model 117 may determine the corresponding three-dimensional structure of the output protein sequence based at least on the protein sequence generated by the sequence design computational model 113. For example, in some cases, the sequence design computational model 113 may generate a protein sequence based on an input protein sequence (e.g., a seed sequence). In addition, in some cases, when generating the corresponding three-dimensional structure, the molecular design computational model 117 may operate on a structural representation (e.g., a coarse-grained (CG) node representation, a main chain torsion (BBT) representation, and / or the like) of the protein sequence. Therefore, in some cases, the output of the sequence design model 115 ingested by the representation generator 115 may include a molecular structure file (e.g., a protein structure file) describing the initial three-dimensional structure of the protein sequence. The representation generator 115 may generate a corresponding structural representation of the initial three-dimensional structure of the protein sequence for further operation by the molecular design computational model 117.
[0074] In some cases, the sequence design computational model 115 can be implemented using one or more machine learning models that are trained to generate protein sequences based on an input protein sequence (e.g., a seed sequence) by sampling a data distribution that the one or more machine learning models learned during training. The one or more machine learning models can be trained based on a variety of known (or observed) protein sequences, including protein sequences that are known to exhibit certain functions and protein sequences that do not have any known functions. In doing so, the one or more machine learning models can learn a data distribution corresponding to a dimensionality-reduced representation of a sequence of amino acid residues that forms a known protein sequence.
[0075] Alternatively, the molecular design computational model 117 can determine the sequence and three-dimensional structure of a protein molecule. For example, in some cases, instead of ingesting the structural representation of the initial three-dimensional structure of the protein sequence generated by the sequence design computational model 113, the molecular design computational model 117 can generate the sequence and three-dimensional structure of a protein molecule by modifying at least a hybrid representation of the protein molecule that includes a sequence representation of the initial sequence of amino acid residues forming the protein molecule and a structural representation of the initial three-dimensional structure of the protein molecule. Additionally, when the molecular design computational model 117 operates on the sequence representation of a protein molecule to modify the sequence of amino acid residues forming the protein molecule, the molecular design computational model 117 can simultaneously operate on the structural representation of the protein molecule to determine the corresponding three-dimensional structure. In doing so, the molecular design computational model 117 can generate a protein molecule whose sequence and three-dimensional structure are more likely to be associated with certain desired properties.
[0076] In some cases, the molecular design computational model 117 can include one or more machine learning models that determine the three-dimensional structure of the protein sequence generated by the sequence design computational model 113 by performing successive modifications to the corresponding structural representation of the protein sequence generated by the representation generator 115. For example, in some cases, one or more machine learning models implementing the molecular design computational model 115 can be equivariant neural networks, many-body higher-order equivariant message passing neural networks, and / or the like. As described in more detail below, one or more machine learning models implementing the molecular design computational model 115 can recognize or account for the rotational symmetry present in the three-dimensional structure. That is, one or more machine learning models implementing the molecular design computational model 115 can recognize that rotating the three-dimensional structure by x degrees about its axis of rotation is the same as rotating the three-dimensional structure by y degrees about its axis of rotation. In this way, the molecular design computational model 117 is able to generate the correct three-dimensional structure regardless of the orientation (e.g., rotation angle) of the initial three-dimensional structure ingested as input.
[0077] In some exemplary embodiments, the molecular design computational model 117 may be implemented as a machine learning model that recognizes rotational symmetries present in a three-dimensional structure. For example, the molecular design computational model 117 may be implemented as a geometric deep learning model, such as an equivariant neural network, and / or the like. Recognition of rotational symmetries present in a three-dimensional structure may enable the molecular design computational model 117 to recognize when two three-dimensional structures are identical but have different orientations in space. That is, the molecular design computational model 125 is able to recognize when two three-dimensional structures are structurally identical (or exhibit structural similarity above a threshold), which is the case when the constituent coarse-grained nodes of the two three-dimensional structures have the same relative positions, even when the entire three-dimensional structure has a different orientation in space. In this way, the molecular design computational model 117 can generate a correct final three-dimensional structure regardless of the orientation of the initial three-dimensional structure ingested as input.
[0078] In some exemplary embodiments, the molecular design computational model 117 may be trained to reduce or minimize one or more loss functions, including, for example, a framework alignment point error (FAPE) loss function, a structure destruction loss function, and / or the like. Alternatively and / or additionally, the molecular design computational model 117 may be trained to reduce or minimize one or more energy functions that quantify the energy of a three-dimensional structure of a molecule (e.g., a protein molecule, a small molecule, an ion, a nucleic acid, a polysaccharide, a glycolipid, and / or the like). In the case where the molecular computational model 117 is implemented as an equivariant neural network (ENN), the loss and / or energy of the three-dimensional structure that is continuously updated by the molecular design computational model 117 may be calculated based on the coarse-grained node coordinate transformation and the inverse coarse-grained mapping structure output from each individual block of the equivariant neural network (ENN).
[0079] In some exemplary embodiments, the protein molecules generated by the molecular design engine 110 may undergo property analysis by the molecular analysis engine 120. Figure 1A As shown, the molecular analysis engine 120 may apply a molecular property computational model 125, which may determine one or more properties of a protein molecule based on the sequence and / or three-dimensional structure of the protein molecule determined by the molecular design engine 110. For example, in some cases, the molecular property computational model 125 may determine whether the protein molecule exhibits binding affinity for another molecule (e.g., a viral antigen, a tumor antigen, and / or the like), specificity for another molecule, lack of non-specificity, stability (e.g., conformational stability, thermodynamic stability, robustness to different environmental stresses, such as protease resistance, and / or the like), non-immunogenicity, human nature, lack of self-association (or non-aggregation), lack of chemical propensity (e.g., aspartate isomerization, oxidation, deamidation), developability, and / or the like, based at least on the sequence and / or three-dimensional structure of the protein molecule.
[0080] In some exemplary embodiments, the molecular design computational model 125 may be trained to generate three-dimensional structures of protein sequences to exhibit one or more desired properties, such as a specific energy range, binding affinity and / or binding specificity to another molecule (e.g., a viral antigen, a tumor antigen, and / or the like), and / or the like. Alternatively and / or additionally, the molecular design computational model 117 may generate three-dimensional structures of protein sequences to be suitable and / or configured for one or more downstream tasks, such as predictive analysis of various properties exhibited by protein sequences.
[0081] As noted, in some cases, the molecular design computational model 117 may generate a sequence and / or three-dimensional structure of a protein molecule such that the protein molecule is more likely to exhibit certain desired properties. Alternatively and / or additionally, the molecular design computational model 117 may generate a sequence and / or three-dimensional structure of a protein molecule that is more suitable or configured for one or more downstream tasks, such as property analysis performed by the molecular analysis engine 120.
[0082] like Figure 1A In the illustrated example, for example, properties of a protein sequence may be determined by the molecular analysis engine 120 using at least the molecular property analysis computational model 125. For example, in some cases, the molecular property computational model 125 may determine one or more properties (e.g., expression, affinity, and / or the like) of the protein sequence 120 based at least on a second protein sequence determined by the molecular design computational model 117 and / or a three-dimensional structure of the second protein sequence. Additionally, the three-dimensional structure of the second protein sequence determined by the molecular design computational model 117 and / or the properties of the second protein sequence determined by the molecular property computational model 125 may be used by the molecular design engine 110 (e.g., the sequence design computational model 113) in generating subsequent additional protein sequences. For example, in the event that the second protein sequence is determined to exhibit a desired three-dimensional structure and / or desired properties, the molecular design engine 110 may apply the sequence design computational model 113 to generate a third or more additional protein sequences based at least on the second protein sequence (e.g., as a seed sequence). Alternatively, where the second protein sequence fails to exhibit the desired three-dimensional structure and / or desired properties, the molecular design engine 110 may apply the molecular design computational model 113 to instead generate a third or more additional protein sequences based on a fourth different protein sequence (e.g., as a seed sequence).
[0083] According to some exemplary embodiments, Figure 1B A flow chart showing an example of a process 160 for protein structure and property prediction is depicted. Figures 1A to 1B, process 160 may be performed by the design system 100, for example, by the molecular design engine 110. In some cases, process 160 may implement a generative design process in which the molecular design computational model 117 operates on a structural representation (such as a coarse-grained (CG) node representation or a backbone torsion (BBT) representation) of a protein molecule having a known protein sequence in order to determine the three-dimensional structure of the protein molecule. Alternatively, in some cases, process 1600 may implement a generative design process in which the molecular design computational model 117 operates on a hybrid representation of a molecule (such as a protein molecule) that includes a sequence representation (e.g., a logarithmic representation, a one-hot encoding representation, and / or the like) and a structural representation (e.g., a coarse-grained (CG) node representation, a backbone torsion (BBT) representation, and / or the like) of the protein molecule to determine the sequence and three-dimensional structure of the protein molecule.
[0084] At 162, the molecular design engine 110 may receive or generate a molecular structure file specifying the initial three-dimensional structure of the molecule. For example, in some cases, the molecular design engine 120 may apply the sequence design computational model 113 to generate a protein sequence (or a sequence of amino acid residues), in which case the output of the sequence design computational model 113 may be a molecular structure file. Alternatively, in some cases, the molecular design engine 110 may receive a molecular structure file from another source such as another sequence design platform. In some cases, the molecular structure file may specify the initial three-dimensional structure of a molecule, which may be a protein molecule or a non-protein molecule (e.g., a small molecule, a nucleic acid, a polysaccharide, a glycolipid, and / or the like). The molecular structure file may specify the initial three-dimensional structure of the molecule by at least enumerating the constituent atoms. In the case where the molecule is a protein molecule, the molecular structure file may specify the initial three-dimensional structure of the protein molecule by at least enumerating the individual atoms (e.g., heavy atoms) that form each amino acid residue in the protein molecule.
[0085] In some exemplary embodiments, the initial three-dimensional structure of the molecule may include varying degrees of entropy or randomness in the spatial arrangement of the constituent atoms, which is subsequently removed by the molecular design computational model 117 in order to determine the actual three-dimensional structure of the molecule. For example, in some cases, the initial three-dimensional structure of the molecule may include at least some noise (e.g., Gaussian noise and / or the like), which means that the positions of the constituent atoms in the initial three-dimensional structure of the molecule may be inconsistent with the positions of those atoms in the actual three-dimensional structure of the molecule. In some cases, the molecular design computational model 117 may perform continuous modifications to correct the positions of atoms in the initial three-dimensional structure of the molecule. Alternatively and / or additionally, in the case where the molecule is a protein molecule, the initial three-dimensional structure of the molecule may include one or more groups of amino acid residues corresponding to the polymer chains present in the molecule. Such grouping may impose at least some restrictions on the modifications performed by the molecular design computational model 117. For example, in some cases, the molecular design computational model 117 may avoid modifications that move two or more amino acid residues in a single polymer chain apart by more than a threshold distance.
[0086] At 164, the molecular design engine 110 may determine a representation of the molecule based at least on the molecular structure file. In some exemplary embodiments, the representation generator 115 of the molecular design engine 120 may determine a representation of the molecule based at least on the molecular structure file. In some cases, the representation of the molecule may include a structural representation of the molecule. As described in more detail below, the structural representation of the molecule may be a coarse-grained (Cg) node representation including a collection of coarse-grained (CG) nodes, each coarse-grained (CG) node corresponding to a structure formed by one or more constituent atoms in the molecule. Alternatively, when the molecule is a protein molecule having a sequence of amino acid residues, the representation of the molecule may be a main chain torsion (BBT) representation, and for each amino acid residue in the protein molecule, the main chain torsion representation includes a corresponding plurality of frameworks that specify the geometric states of the main chain and side chains of the amino acid residue. In the case where the molecule is a protein molecule, in addition to the structural representation of the molecule, the representation of the molecule may further include a sequence representation of the molecule. In some cases, the sequence representation of the protein molecule can be a logarithmic representation, wherein each position in the sequence forming the protein molecule can be a logarithmic vector that represents the identity of the amino acid residue occupying the position by at least enumerating the probability distribution (e.g., categorical distribution) of the set of possible amino acid residues. Alternatively, the sequence representation of the protein molecule can be a one-hot encoded representation, wherein each position in the sequence forming the protein molecule can be a one-hot encoded vector, wherein the value "1" occupies the position in the one-hot encoded vector corresponding to the identity of the amino acid residue occupying the position in the sequence, and the value "0" occupies the other positions in the one-hot encoded vector.
[0087] As noted, in some exemplary embodiments, the identity of the amino acid residue occupying each position in the protein sequence may constitute one of the degrees of freedom (DoF) associated with the amino acid residue. This particular degree of freedom may restrict the variation of the identity of the amino acid residue occupying each position to, for example, one of the 20 canonical amino acid residues. Thus, in some cases, the degrees of freedom (DoF) associated with the identity of the amino acid residue may be expressed as a classification probability. When there are N number of amino acid residues in the protein sequence, the logarithmic representation of the protein sequence may be a probability vector (p 1 ,…,p N ), where each probability vector p i The probability distribution of the identity of the amino acid residue occupying the i-th position of the set of possible amino acid residues (e.g., 20 canonical amino acid residues) is enumerated. For example, the probability vector p for the first position in the protein sequence is 1 Can include a first probability that the first position is occupied by alanine (Ala), a second probability that the first position is occupied by arginine (Arg), a third probability that the first position is occupied by asparagine (Asn), and / or the like.
[0088] In some cases, each probability vector p i ∈ simplex(D), meaning that the sum of the composition probabilities of the set of possible amino acid residues is 1. For example, the probability that the amino acid residue occupying a particular position in a protein sequence is each of the 20 canonical amino acid residues adds up to 1. A probability simplex simplex(D) can be a mathematical space in which each point represents a probability distribution between a finite number of mutually exclusive events or categories, in this case corresponding to a set of possible amino acid residues (e.g., 20 canonical amino acid residues). It should be understood that a probability simplex simplex(D) is a (D-1) dimensional object. That is, the points forming the probability simplex simplex(D) occupy a (D-1) dimensional space, where D corresponds to the number of possible amino acid residues (e.g., 20 canonical amino acid residues). The requirement that the sum of the probabilities of D number of possible amino acid residues (e.g., 20 canonical amino acid residues) is 1 reduces the dimensionality of the probability simplex simplex(D) by 1. As described in more detail below, the molecular design computational model 125 applied to the sequence representation of a protein sequence can be a diffusion model that adds noise during a forward diffusion process and removes noise during a corresponding reverse diffusion process. Defining the forward diffusion process on the probability simplex simplex (D) is not trivial, at least because in a simple way iAdding noise (e.g., Gaussian noise) may cause the sum of the component probabilities of the D number of possible amino acid residues (e.g., 20 canonical amino acid residues) to no longer be 1. In other words, noise can be added in a principled manner during forward diffusion so as to maintain the (D-1) dimensionality of the probability simplex simplex (D). For example, in order to maintain the (D-1) dimensionality of the probability simplex simplex (D) to the logarithmic space p i ∈Simple A one-to-one mapping of i Add noise (e.g., Gaussian noise) to make the logarithm l i Stay in the same space In the example above, we may want to constrain the logarithmic space from having D dimensions to having D-1 dimensions. This constraint is obtained by applying the logarithm l i Subtract its mean This is achieved by projecting it onto the zero identity component subspace. By applying the soft maximum function, a unique mapping from the logarithmic subspace to probability can be established without any degeneracy.
[0089] In some cases, instead of the aforementioned logarithmic representation, the identity of each amino acid residue in the protein sequence can be presented in a one-hot encoding representation. Thus, each position in the protein sequence can be associated with a one-hot encoding vector, where the value "1" occupies the position in the one-hot encoding vector corresponding to the amino acid residue occupying the position in the protein sequence, and the value "0" occupies each remaining position in the one-hot encoding vector. For example, if a position in the protein sequence is occupied by alanine (Ala / A), the one-hot encoding vector for that position can include a value of "1" in the position corresponding to alanine (Ala / A) and a value of "0" in the positions corresponding to other amino acid residues.
[0090] It should be understood that the logarithmic logical representation of the amino acid residues forming the protein sequence can be used instead of other representations, such as one-hot encoding, in order to enhance the learning of the molecular design computational model 117. For example, one-hot encoding does not provide a probabilistic representation of the identity of each amino acid residue in the protein sequence. Instead, the identity of the amino acid residue occupying a certain position in the protein sequence is represented by the value "1" in the corresponding position in the one-hot encoding vector, while the remaining positions in the one-hot encoding vector are occupied by the value "0". Adding noise as part of the forward diffusion process does not change the binary value of the one-hot encoding vector. The resulting value may no longer be consistent with the one-hot encoding scheme, and a single value "1" in the one-hot encoding vector identifies the identity of the amino acid residue in the protein sequence. Instead, a series of different values may be generated, but these values will still correspond to the same physical state. For example, the one-hot encoding vector [1,0,0.0] may represent the same physical state as the one-hot encoding vector [0,5,0.0], which means that the molecular design computational model 117 can be constructed to learn to generate the same output for some one-hot encoding vectors with different values. Therefore, the performance of the molecular design computational model 117 may degrade when operating on one-hot encoded representations of the identities of amino acid residues in a protein sequence.
[0091] At 166, the molecular design engine 110 may modify the representation of the molecule by applying at least the molecular design computational model 125, determine the sequence and / or three-dimensional structure of the molecule. In some exemplary embodiments, the molecular design engine 110 may modify the representation of the molecule by applying at least the molecular design computational model 125, determine the three-dimensional structure, and in some cases may also determine the sequence of the molecule. In some cases, the molecular design computational model 125 may continuously update the representation of the molecule in order to determine the sequence and / or three-dimensional structure of the molecule. For example, when the molecule is represented as a collection of coarse-grained (CG) nodes, the molecular design computational model 125 may perform continuous updates, wherein each update modifies one or more coarse-grained (CG) nodes that form the structural representation of the molecule. Alternatively, when the molecular design computational model 117 operates on the main chain torsion (BBT) representation of the molecule, each continuous update may modify one or more frameworks that form the structural representation of the molecule.
[0092] In some cases, the molecular design computational model 117 can simultaneously determine the sequence and three-dimensional structure of the molecule, for example by performing continuous updates on the sequence representation and structure representation of the molecule. As described in more detail below, the sequence and / or three-dimensional structure of the molecule determined by the molecular design computational model 117 can be used for one or more downstream tasks, including, for example, conformer generation, molecular docking, property prediction, and / or the like.
[0093] In the case where the structural representation of the three-dimensional structure of the protein sequence generated by the sequence design computational model 113 is a coarse-grained (CG) node representation, the structural representation of the protein sequence may include a set of coarse-grained nodes, wherein each amino acid residue included in the protein sequence is associated with a corresponding coarse-grained (CG) node group. For example, each amino acid residue contained in the protein sequence may include one or more structures, wherein each structure is an atomic group of two or more atoms. In the case of a rigid structure, the positions of two or more atoms may be fixed relative to the coordinates of the structure. In contrast, as a flexible or variable structure, the positions of two or more atoms may exhibit at least some degree of flexibility relative to the coordinates of the structure.
[0094] Thus, each amino acid residue in a protein sequence may be associated with at least one coarse-grained (CG) node corresponding to a structure (e.g., a rigid body, a flexible or variable body, and / or the like) included in the amino acid residue. In the case where a single atom is part of multiple structures included in the amino acid residue, the atom may be contained in multiple corresponding coarse-grained nodes. In some cases, the coarse-grained (CG) node representation of a protein sequence may include or exclude certain elements, such as hydrogen (H), and / or the like. For example, in some cases, the coarse-grained (CG) node representation of a protein sequence may include hydrogen atoms found in the amino acid residues that form the protein sequence. However, in other cases, the coarse-grained (CG) node representation of a protein sequence may exclude hydrogen atoms found in the amino acid residues that form the protein sequence.
[0095] To further illustrate, given a protein sequence a=a of length N 1 a 2 …a n , where a i is an amino acid residue (e.g., one of the 20 canonical amino acids), and the three-dimensional structure formed by the protein sequence can be specified by the three-dimensional coordinates of the constituent atoms, which are grouped by amino acids, i.e. in Indicates amino acid residue a i Therefore, in the coarse-grained node representation of the protein sequence, each amino acid a i A set of coarse-grained nodes Represents that each coarse-grained node Represents the formation of amino acid a i A subset of atoms (e.g., heavy atoms)
[0096] In some exemplary embodiments, a representation generator 115 may generate a coarse-grained (CG) node representation of a protein sequence by at least grouping atoms enumerated in a molecular structure file into one or more coarse-grained (CG) nodes. A coarse-grained node representation of the protein sequence may be generated to satisfy certain properties. For example, the representation generator 115 may generate each coarse-grained node included in the coarse-grained node representation of the protein sequence such that the union of each coarse-grained (CG) node representing an amino acid residue in the protein sequence includes all constituent atoms of the amino acid residue. In such a case, the union of two or more coarse-grained (CG) nodes includes the atoms present in each coarse-grained node. For example, the union of a first coarse-grained node and a second coarse-grained node includes a first plurality of atoms present in the first coarse-grained node but not in the second coarse-grained node, a second plurality of atoms present in the second coarse-grained node, and a third plurality of atoms present in both the first coarse-grained node and the second coarse-grained node. Thus, each atom in an amino acid residue of the protein sequence may be included in at least one coarse-grained (CG) node associated with the amino acid residue. Further, when grouping the atoms enumerated in the molecular structure file into coarse-grained (CG) nodes included in the coarse-grained node representation of the protein sequence, the atoms may be grouped such that each member atom of a coarse-grained node shares at least one covalent bond with another member of the same coarse-grained node. In some cases, the representation generator 115 may generate a coarse-grained node representation of the protein sequence such that each coarse-grained node includes a threshold number of atoms (e.g., a minimum number and / or a maximum number of atoms) and their constituent atoms together form at least one structure, as described above.
[0097] In some cases, when grouping the atoms enumerated in the molecular structure file into one or more coarse-grained (CG) nodes, the representation generator 115 may further generate a coarse-grained (CG) node representation of the protein sequence by mapping the coordinates of the atoms (e.g., heavy atoms) forming each coarse-grained node to a corresponding Euclidean transformation (e.g., translation, rotation, and / or the like). Since each coarse-grained node associated with the protein sequence is generated to include a threshold number of atoms (e.g., heavy atoms) forming at least one structure, the three-dimensional coordinates i of the amino acid residue a included in the protein sequence to the forward mapping F of its corresponding coarse-grained node representation may include or be defined as
[0098]
[0099] where each tuple in the set contains a coarse-grained node identifier and a Euclidean transformation that transforms the atoms in amino acid a i in The template coordinates of a subset of the predefined template coordinates are mapped to the corresponding input atomic coordinates. As used herein, the template coordinates of each coarse-grained (CG) node may correspond to the initial translation and / or rotation applied to the coarse-grained node. That is, the template coordinates of the coarse-grained (CG) node may determine the initial positions of the constituent atoms in the coarse-grained (CG) node representation of the initial three-dimensional structure of the protein sequence. At the same time, the amino acid a i From its coarse-grained (CG) node representation to the corresponding 3D coordinates The reverse map G may include or be defined as the average of the three-dimensional coordinates of the atom specified by any coarse-grained nodes that contain the atom.
[0100] To further illustrate, Figure 2A An example of a coarse-grained node representation 200 of a protein sequence excluding hydrogen (H) atoms in each constituent amino acid residue is depicted. Figure 2B Another example of a coarse-grained node representation 250 of a protein sequence including hydrogen (H) atoms is depicted. For example, in Figure 2A In the example of the coarse-grained node representation 200 shown, the coarse-grained node representation of the amino acid residue tryptophan (Trp) includes a first coarse-grained node CG0 ("C", "CA", "CB", "N"), a second coarse-grained node CG1 ("C", "CA", "O"), and a third coarse-grained node CG2 ("CG", "CD1", "CD2", "CE2", "CE3", "CZ2", "CZ3", "CH2", "NE1"). At the same time, the coarse-grained node representation of the amino acid residue valine (Val) may include a first coarse-grained node CG0 ("C", "CA", "CB", "N"), a second coarse-grained node CG1 ("C", "CA", "O"), and a third coarse-grained node CG2 ("CB", "CG1", "CG2"). A coarse-grained node can specify the three-dimensional position of its constituent atoms by rotating R and / or translating T the coarse-grained node as a whole. For example, the carbon I atom, the α carbon (C α ) atom, β carbon (C β ) atoms and nitrogen (N) can be specified by a rotation R and / or a translation T of the first coarse-grained node CG0.
[0101] As noted, in some exemplary embodiments, the representation generator 115 may determine a Euclidean transformation, such as a rotation R and / or a translation T, of each coarse-grained node representing the initial three-dimensional structure of the protein sequence. For example, the representation generator 115 may determine the rotation R and / or translation T required to transform each coarse-grained node from its current position to the template position (e.g., template coordinates) of the coarse-grained node in order to determine the initial three-dimensional structure of the protein sequence. In some cases, the representation generator 115 may determine the rotation R and / or translation T by at least calculating a rotation matrix having a minimum root mean square deviation (RMSD) between the template position of the coarse-grained node and the current position of the coarse-grained node. For example, in some cases, the representation generator 115 may apply the Karbusch algorithm in order to calculate a rotation matrix and a translation having a minimum root mean square deviation (RMSD) between the template position of the coarse-grained node and the current position of the coarse-grained node.
[0102] To further illustrate the calculation of the template coordinates of each coarse-grained node in a protein sequence, consider a protein structure dataset D associated with the protein sequence. The coordinate transformation (e.g., Euclidean transformation) T of each coarse-grained node q in the protein structure dataset D is q It can be calculated by applying a Gram-Schmidt procedure, such as shown in Table 2 below, to the three-dimensional coordinates of the first three atoms of the atom group of the coarse-grained node q.
[0103] Table 2
[0104]
[0105] Then perform the inverse coordinate transformation (e.g., Euclidean transformation) The three-dimensional coordinates of all atoms in the coarse-grained node q are applied to transform these three-dimensional coordinates into the corresponding local frame. In this case, the term "framework" may refer to a transformation (e.g., a Euclidean transformation) that defines the coordinates (e.g., three-dimensional coordinates) of at least a portion of the atoms in the protein molecule. Although the above-mentioned local frame may define the positions of atoms included in a single coarse-grained node (such as the coarse-grained node q), the global frame may define the position of the protein molecule as a whole (e.g., by specifying the rotation and translation applied to the center of mass of the molecule). In some cases, in order to determine the template coordinates, the three-dimensional coordinates observed in the local frame across all instances in the protein structure dataset D may be averaged, and these instances are grouped according to the coarse-grained node types listed in Table 2 below. That is, when the protein structure dataset D includes multiple coarse-grained nodes representing the same amino acid residue, the calculation of the template coordinates may include determining the average of the three-dimensional coordinates observed in the local frame across these coarse-grained nodes.
[0106] Table 2
[0107]
[0108] In order to calculate the true coordinate transformation (e.g., Euclidean transformation) of a given coarse-grained node to be used in the loss function (e.g., framework alignment point error (FAPE) loss function) of the molecular design computational model 117, the Karbusch algorithm can be applied to determine the transformation from template coordinates to observation coordinates that reduces or minimizes the root mean square error (RMSE) of the constituent atoms of each coarse-grained node. For example, for the ith amino acid residue a i The jth coarse-grained node of the M atom Kabusch algorithm can be applied to template coordinates and the corresponding input coordinates The Kabusch algorithm uses singular vector decomposition to convert the covariance matrix Decomposed into eigenvectors and values, as shown in equation (1) below.
[0109] H=USV T (1)
[0110] Where W c represents the average center template coordinates and X c Represents the mean center input coordinates.
[0111] The resulting rotation and translation can be expressed by the following equation (2).
[0112]
[0113] in and are the mean coordinates of W and X respectively.
[0114] Figure 3A Depicted is a visualization of the positions of atoms in the first, second, and third coarse-grained nodes CG0, CG1, and CG2 of the amino acid residue tryptophan (Trp) in a local frame of template coordinates after fitting the individual atoms to the template coordinates to show their relative positions to each other. Figure 3B Depicted is a visualization of the atomic positions in each coarse-grained node of the amino acid valine (Val) in the local frame of template coordinates after fitting the individual atoms to the local reference frame template coordinates to show their relative positions to each other. Figures 3A to 3B The bar graph shown in provides a visual indication of the error associated with the position of each atom. Figures 3A to 3B As shown, the error is small.
[0115] In some exemplary embodiments, the transformation (e.g., Euclidean transformation) applied to each coarse-grained (CG) node may be further represented digitally as a set of geometric tensors with a configurable maximum degree L. In some cases, the rotation and / or translation of each geometric tensor associated with a coarse-grained (CG) node may be determined by applying one or more elements from a three-dimensional rotation group. For example, the digital representation of the first coarse-grained node CG0 of the amino acid residue tryptophan (Trp) may include one or more geometric tensors controlled by one or more elements from a three-dimensional rotation group to describe the current translation T and / or rotation R of the first coarse-grained node CG0 in three-dimensional space.
[0116] To further illustrate, each coarse-grained node associated with a protein sequence can be assigned a set of orders l = 0, ..., l 最大 The geometric tensor characteristics of each order are n c Assuming there are l features associated with a 2l+1-order tensor, each coarse-grained node has n c ×(l 最大 +1) 2 Geometry tensor features. Coarse-grained nodes The initial embedding of , where C represents a predetermined set of coarse-grained node types, which may include or be defined as
[0117]
[0118] Among them search: represents the embedding function, and The Wigner D matrix D represents the geometric tensor characteristics corresponding to different orders l l The direct sum of is expressed as follows:
[0119]
[0120] Representing each amino acid residue in the initial three-dimensional structure of a protein sequence as a collection of coarse-grained (CG) nodes, particularly as a geometric tensor embedding, can reduce the computational complexity associated with subsequent operations on the initial three-dimensional structure to determine the three-dimensional structure of the protein sequence. The coarse-grained node representation of a protein sequence may ignore overly fine details, such as chemical changes and changes in bond angles and bond lengths. Therefore, the molecular design computational model 117 is able to operate on coarse-grained nodes that are discrete semantic units with higher computational efficiency than on individual atoms. In addition, the molecular design computational model 117 is able to determine the three-dimensional structure of a protein sequence without the need for additional information, such as the co-occurrence frequency of certain amino acid residues at different positions.
[0121] In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 117 to determine the three-dimensional structure of the protein sequence based at least on the geometric tensor embedding of the coarse-grained (CG) node representation of the initial three-dimensional structure of the protein sequence. In some cases, the molecular design computational model 117 may be implemented as a machine learning model (e.g., an equivariant neural network, and / or the like) having a series of modules, each of which is a subunit of a machine learning model including one or more layers of the machine learning model. Each module of the machine learning model may be trained to determine updates to a transformation (e.g., a Euclidean transformation, such as a rotation R and / or a translation T) that defines the position of one or more coarse-grained nodes included in the initial three-dimensional structure of the protein sequence. Therefore, the molecular design computational model 117 may perform continuous updates to the transformation (e.g., a Euclidean transformation, such as a rotation R and / or a translation T) applied to define the position of one or more coarse-grained nodes within the initial three-dimensional structure in order to derive the actual three-dimensional structure of the protein sequence.
[0122] As noted, a coarse-grained (CG) node representation of a protein sequence may be instantiated by a coordinate transformation defined by one or more elements sampled from a three-dimensional rotation group. For example, in some cases, a coarse-grained node representation of a protein sequence may be instantiated by a Euclidean transformation whose translations and rotations are sampled from a normal distribution with mean zero and variance unity and a uniform distribution over the three-dimensional rotation group SO(3). The final three-dimensional structure determined by the molecular design engine 110 may be generated by iterative refinement performed by a molecular design computational model 117, which may be implemented, for example, as a 3D scalar having N scalar structures. 模块 N number of modules, each with the same architecture sub number of submodules. Thus, each module of the equivariant neural network may take in as input an initial coarse-grained node representation of a protein sequence or an updated coarse-grained node representation of a protein sequence output by a previous module. In addition, each module may output two l=1 geometry tensors for each coarse-grained node associated with the protein sequence, wherein the first geometry tensor is used as the vector part of a non-unit quaternion to compute a previous rotation R of the coarse-grained node. 入 The update R' is used, and the second geometry tensor is used as the previous translation of the coarse-grained nodes Update
[0123] Therefore, the initial coordinate transformation applied to the initial template coordinates of the protein sequence ) can be continuously updated by an equivariant neural network (ENN) according to the following:
[0124]
[0125] Each module of the equivariant neural network can transform or simply copy the input embedding of each coarse-grained node. However, in either case, the input embedding of the coarse-grained node can be multiplied by the direct sum of the Wigner D matrix corresponding to the update rotation R'.
[0126] For example, in some cases, each module of an equivariant neural network may include N sub Given an input set of coarse-grained nodes and their corresponding coordinate transformations (e.g., transformed to template coordinates), this module initially computes the pairwise distance r ij and the normalized distance vector where i and j index coarse-grained nodes. The pairwise distance r ij can be projected to a point with learnable weights and a cutoff distance r c D 贝塞尔 The radial Bezier basis is used in the radial function that parameterizes the tensor product in the equivariant graph attention module. Instead of the polynomial envelope function, we can apply A soft unit step function is used as input. The normalized distance vector Can be used to compute spherical harmonics input to the tensor product When training an equivariant neural network, it may happen that the gradient does not pass through the pairwise distance r. ij and the normalized distance vector The situation of transmission.
[0127] When applying an initial linear layer to the embeddings of the input coarse-grained nodes, a channel-wise fully connected tensor product is applied to the embeddings of each coarse-grained node pair ij instead of pairwise summation, followed by another linear layer to produce an output tensor x with the same number of channels as the input ij Afterwards, the output tensor x can be ij and spherical harmonics A deep tensor product (DTP) is applied in conjunction with a radial function that takes as input the above Bezier basis and a scalar edge embedding vector corresponding to the amino acid sequence distance of the coarse-grained node pair ij, where the amino acid sequence distance is constrained to be a certain distance (e.g., 32). In some cases, the edge embedding can be implemented as a lookup table with learnable weights and the same dimension as the number of channels in the input tensor. The output of the depthwise tensor product layer can be obtained by N 头部A number of attention heads are used to evenly shuffle and group. A linear layer can be applied to generate tensors of various degrees with appropriate channel numbers for the rest of the module. The output of each sub-module except the last one is an updated geometric tensor, which represents the transformation applied to the corresponding coarse-grained nodes in the three-dimensional structure of the molecule. The last sub-module can output two l = 1 tensors for each coarse-grained node. Edge embeddings can be shared among the sub-modules of a given module.
[0128] For further illustration, according to some exemplary embodiments Figure 4 depicts a visualization of the iterative updates performed by the molecular design computational model 117 to generate the three-dimensional structure of a molecule. In Figure 4 the example shown, the molecular design computational model 117 is a machine learning model with a sequence of four modules, each module updating the transformation (e.g., Euclidean transformation such as rotation R and / or translation T) of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule (as shown in module 0) to derive the three-dimensional structure of the molecule (as shown in module 4). As Figure 4 shown, each update of the transformation (e.g., Euclidean transformation such as rotation R and / or translation T) applied to one or more coarse-grained nodes in the initial three-dimensional structure (shown in module 0) may gradually reduce or minimize the loss, which represents the deviation between the initial three-dimensional structure of the molecule and the true three-dimensional structure of the molecule. In Figure 4 the example shown, the initial three-dimensional structure (as shown in module 0) may correspond to a loss value of 0.8906, and this loss value drops significantly in subsequent updates, so that the loss value corresponding to the final three-dimensional structure of the molecule (as shown in module 4) is 0.1797.
[0129] As previously mentioned, the protein design engine 110 can implement a generative design process that integrates sequence and structure design in various different ways. Figures 5A to 5B shows an example where the protein design engine 110 first applies the sequence design computational model 113 to generate a protein sequence, and then applies the molecular design computational model 117 to determine the three-dimensional structure of the protein sequence based at least on the representation of the protein sequence. In each case, the molecular design computational model 117 can operate on different structural representations of the initial three-dimensional structure of the protein sequence (e.g., Figure 5A the coarse-grained (CG) node representation in Figure 5B and the backbone torsion (BBT) representation in Figure 5CAnother example is shown in which the protein design engine 110 applies the molecular design computational model 117 to simultaneously determine the identities of the amino acid residues in the protein sequence and the corresponding three-dimensional structure. That is, in some cases, the molecular design computational model 117 may operate on a sequence representation (e.g., a logarithmic representation, a one-hot encoding representation, and / or the like) of an initial sequence of amino acid residues forming a protein sequence and a structural representation (e.g., a coarse-grained (CG) node representation or a backbone torsion (BBT) representation) of the initial three-dimensional structure of the protein sequence to determine the identities of the amino acid residues in the protein sequence and the three-dimensional structure of the protein sequence.
[0130] To further illustrate, ontological relationships between representations of molecules (e.g., protein molecules, and / or the like) are shown below in Table 1, which may include a structural representation of the three-dimensional structure of the molecule and, in some cases, a sequence representation of the amino acid residues that form the molecule.
[0131] Table 1
[0132]
[0133]
[0134] According to some exemplary embodiments, Figure 5A A flow chart showing an example of a process 500 for molecular structure and property prediction is depicted. Figures 1A to 1B , 2A-2B, 3A-3B, 4 and 5A, process 500 can be performed by design system 100, for example, by molecular design engine 110 and molecular analysis engine 120. In some cases, process 500 can implement a generative design process in which molecular design engine 110 operates on a coarse-grained (CG) node representation of a molecule. In addition, process 500 can implement a generative design process that integrates structure and property prediction as part of a process for generating various molecules (including, for example, protein molecules, small molecules, nucleic acids, polysaccharides, glycolipids, and / or the like).
[0135] At 502, a molecular structure file specifying an initial three-dimensional structure of a molecule may be received. In some exemplary embodiments, the molecular design engine 110 may apply the sequence design computational model 113 to generate a protein sequence (e.g., a sequence of amino acid residues) for a protein molecule, for example, based on a sequence of another amino acid residue (e.g., a seed sequence). In some cases, the sequence design computational model 113 generates a protein sequence by at least determining an identity of each amino acid residue in the protein sequence. In addition, in some cases, the sequence design computational model 113 may include one or more machine learning models that are trained to generate a sequence of a first amino acid residue by sampling a data distribution learned by one or more machine learning models during training based on a sequence of a second amino acid residue. Alternatively, the molecular design engine 110 may receive a protein sequence or a corresponding molecular structure file (e.g., a protein structure file) specifying an initial three-dimensional structure of a protein sequence from different sources (such as different sequence design platforms). In some cases, upon receiving or generating a protein sequence, the molecular design engine 110 may generate a molecular structure file (e.g., a protein structure file) that specifies an initial three-dimensional structure of the protein sequence, including by enumerating individual atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence.
[0136] At 504, a plurality of coarse-grained nodes may be determined based at least on the molecular structure file. In some exemplary embodiments, the representation generator 115 may generate a coarse-grained (CG) node representation of the initial three-dimensional structure of the protein sequence based at least on the molecular structure file. In order to generate a coarse-grained node representation of the protein sequence, the representation generator 115 may first group the atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence into one or more coarse-grained nodes. Therefore, each coarse-grained node may correspond to a structure (e.g., a rigid body, a flexible body, or a variable body, and / or the like) of two or more atoms (e.g., heavy atoms) that form the amino acid residues in the protein sequence. In some cases, the structure may be a rigid structure that includes a group of two or more atoms whose positions are fixed relative to the coordinates of the structure. In other words, each coarse-grained node operates as a separate semantic unit so that the relative positions of the constituent atoms remain fixed. Changes in the position and / or direction of the coarse-grained node will not change the relative positions of the atoms contained in the coarse-grained node.
[0137] At 506, one or more geometric tensor embeddings may be generated for each coarse-grained node. In some exemplary embodiments, the representation generator 115 may further generate a numerical representation of each coarse-grained node associated with the protein sequence. Specifically, the representation generator 115 may generate a numerical representation for each coarse-grained node that describes the rotation of the coarse-grained node in three-dimensional space. For example, in some cases, the structure analysis engine 110 may determine a set of numerical representations in the form of one or more geometric tensors with a configurable maximum degree L for each coarse-grained node. In this case, each geometric tensor may be manipulated by applying one or more elements from a three-dimensional rotation group (e.g., an irreducible representation of the SO(3) group). For example, a rotation R of a coarse-grained node in three-dimensional space may be represented by one or more geometric tensors that have undergone a coordinate transformation defined by one or more elements from a three-dimensional rotation group.
[0138] At 508, the three-dimensional structure of the molecule may be generated by updating at least the position of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule. In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 117 to determine the three-dimensional structure of the protein sequence based on at least the geometric tensor embedding of the coarse-grained (CG) nodes representing the initial three-dimensional structure of the protein sequence. The molecular design computational model 117 may determine the three-dimensional structure of the protein sequence by continuously updating the position of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence. For example, in some cases, the molecular design computational model 117 may update the position of one or more coarse-grained nodes by at least translating and / or rotating one or more coarse-grained nodes (e.g., updating the rotation R and / or translation T of one or more coarse-grained nodes). In addition, in some cases, the molecular computational model 117 may include a series of modules, each of which performs a separate and incremental update on the position of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence. For example, the molecular design computational model 125 may include a first block that performs a first update on the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence, followed by a second block that performs a second update on the positions of the one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence.
[0139] At 510, one or more additional molecules may be generated based on the sequence of amino acid residues that form the molecule or a different sequence of amino acid residues. In some exemplary embodiments, the molecular design computational model 117 may be used to determine the three-dimensional structure of a protein sequence associated with one or more desired properties, such as a specific energy range, binding affinity and / or binding specificity to another molecule (e.g., a viral antigen, a tumor antigen, and / or the like), and / or the like. Alternatively and / or additionally, the molecular design computational model 117 may be used to determine the three-dimensional structure of a protein sequence that is suitable or configured for one or more downstream tasks. For example, in some cases, the molecular analysis engine 120 may apply the molecular property computational model 125 to determine one or more properties of the protein sequence based at least on the three-dimensional structure of the protein sequence determined by the molecular design computational model 117, including, for example, binding affinity to another molecule (e.g., a viral antigen, a tumor antigen, and / or the like), specificity to another molecule, non-specificity, stability (e.g., conformational stability, thermodynamic stability, robustness to different environmental stresses, such as protease resistance, and / or the like), non-immunogenicity, humanity, self-association (or non-aggregation), chemical propensity (e.g., aspartate isomerization, oxidation, deamidation), developability, and / or the like.
[0140] In some exemplary embodiments, the three-dimensional structure of a protein sequence may be used as part of a generative design process that integrates structure and property prediction as part of a process for generating a protein sequence. For example, where it is determined that a protein sequence exhibits a desired three-dimensional structure and / or desired properties (e.g., the three-dimensional structure of the protein sequence corresponds to the desired three-dimensional structure and / or a three-dimensional structure associated with the desired properties), the molecular design engine 110 may generate one or more additional protein sequences (e.g., sequences of amino acid residues) based on the protein sequence (e.g., as a seed sequence). Alternatively, where it is determined that a protein sequence lacks a desired three-dimensional structure and / or desired properties, the design engine 110 may generate one or more additional protein sequences (e.g., sequences of amino acid residues) based on a different sequence of amino acid residues (e.g., as a seed sequence) rather than the protein sequence.
[0141] In some cases, the three-dimensional structure of a protein sequence can be generated as part of a design library for constructing protein sequences (such as antibodies or nanobodies) with specific three-dimensional structures and / or desired properties. For example, the design library can be a combinatorial library that enumerates, for each position of a protein sequence, the probability distribution of different amino acid residues that may occupy that position, so that the resulting protein sequence has a probability above a threshold of exhibiting a three-dimensional structure and / or desired properties associated with the three-dimensional structure.
[0142] Alternatively, in addition to the coarse-grained (CG) node representation of the three-dimensional structure of the protein sequence, the molecular design engine 110 may generate a main chain torsion (BBT) representation of the three-dimensional structure of the protein sequence and operate on it. In some exemplary embodiments, for each constituent amino acid residue in the protein sequence, the main chain torsion (BBT) representation of the protein sequence may include multiple frameworks. Each framework may correspond to the degree of freedom (DoF) of the molecular design computational model 117 to update the initial three-dimensional structure of the protein sequence. For example, in some cases, multiple frameworks for a single amino acid residue in a protein sequence may include a first set of frameworks for the geometric state of the main chain of the specified amino acid residue. In addition, multiple frameworks of amino acid residues may include a second set of frameworks for one or more torsion angles present in the side chain of the specified amino acid residue.
[0143] According to some exemplary embodiments, Figure 5B A flow chart showing an example of a process 550 for molecular structure and property prediction is depicted. Figures 1A to 1B , 2A to 2B, 3A to 3B, 4 and 5B, process 550 can be performed by design system 100, such as by molecular design engine 110. Figure 5B The example of process 550 shown in can implement a generative design process in which the molecular design engine 110 (eg, the molecular design computational model 117) operates on a backbone torsion (BBT) representation of a molecule (eg, a protein molecule) to determine a three-dimensional structure of the molecule.
[0144] At 552, the molecular design engine 110 may receive or generate a molecular structure file specifying an initial three-dimensional structure of a protein molecule including a sequence of amino acid residues. For example, in some cases, the molecular design engine 110 may apply the sequence design computational model 113 to generate a protein sequence of a protein molecule based on another protein sequence (e.g., a seed sequence), including by determining the identity of each amino acid residue in the protein sequence. In some cases, the molecular design engine 110 may receive a protein sequence from a different source (such as a differential sequence design platform) or a corresponding molecular structure file specifying an initial three-dimensional structure of a protein sequence. In some cases, when receiving or generating a protein sequence, the molecular design engine 110 may generate a molecular structure file (e.g., a protein structure file) specifying an initial three-dimensional structure of a protein sequence, including by enumerating individual atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence. The molecular structure file may be a protein structure file that specifies an initial three-dimensional structure of a protein sequence by at least enumerating atoms (e.g., heavy atoms) that form each amino acid residue in the protein sequence.
[0145] At 554, the molecular design engine 110 may determine a representation of the protein molecule based at least on the molecular structure file, the representation including multiple frameworks for each amino acid residue in the sequence of amino acid residues. In some exemplary embodiments, the representation generator 115 may determine a main chain torsion (BBT) representation of the protein molecule to include multiple frameworks for each amino acid residue including the geometric state of the main chain and side chain of the amino acid residue. In some cases, each framework may correspond to the degree of freedom (DoF) of the molecular design computational model 117 to update the three-dimensional structure of the protein molecule. For example, a plurality of frameworks associated with an amino acid residue may include a first set of frameworks for the geometric state of the main chain of the amino acid residue and a second set of frameworks for the torsion angle in the side chain of the amino acid residue. In some cases, the first set of frameworks may include a first framework for the translation and rotation of the specified main chain and a second framework for the torsion angle present therein. Alternatively, in some cases, for each torsion angle present in the main chain of the amino acid residue, the first set of frameworks may include a corresponding framework. As described in more detail below, the second set of frameworks may include a framework for each torsion angle present in the side chain of the amino acid residue.
[0146] To further illustrate, consider a protein molecule having N amino acid residues. In some cases, the characterization generator 115 can determine a backbone torsion (BBT) representation of the protein molecule by at least assigning the following degrees of freedom (DoF) to each residue i=1, ..., N: backbone translation x i =R 3 , main chain rotation r i =SO(3) and five torsion angles (one for the oxygen atom (O) and four for the side chains): and θ q ∈SO(2). When the three-dimensional structure of a protein molecule is modified by the molecular design computational model 117, the aforementioned degrees of freedom (DoF) may be applied. That is, in order to determine the three-dimensional structure of a protein molecule, the molecular design computational model 117 may be limited to modifying the initial three-dimensional structure of the protein molecule within one or more of the aforementioned degrees of freedom. For example, in order to determine the three-dimensional structure of a protein molecule, the molecular design computational model 117 may modify the translation and / or rotation of the main chain of one or more of the N number of amino acid residues in the protein molecule. Alternatively and / or additionally, in order to determine the three-dimensional structure of a protein molecule, the molecular design computational model 117 may modify the torsion angle and / or side chain torsion angle of an oxygen (O) atom in the main chain of one or more of the N number of amino acid residues in the protein molecule.
[0147] In some cases, the main chain atoms and torsion angles present in each type of amino acid residue are shown in Table 3 below.
[0148] Table 3
[0149]
[0150] Fig. 6A Further illustrations are provided which depict exemplary atomic structures of amino acid residues.
[0151] At 556, the molecular design engine 110 may modify the representation of the protein molecule by applying at least the molecular design computational model 117 to determine the three-dimensional structure of the protein molecule. In some exemplary embodiments, the molecular design computational model 117 may modify the main chain torsion (BBT) representation of the protein molecule by updating at least a first set of frameworks to change the geometric state of the main chain of one or more amino acid residues. For example, in some cases, the molecular design computational model 117 may update the first set of frameworks to change the translation and rotation of the main chain and / or one or more torsion angles present in the main chain. Alternatively and / or additionally, the molecular design computational model 117 may modify the main chain torsion (BBT) representation of the protein molecule by updating at least a second set of frameworks to change one or more torsion angles present in the side chains of the amino acid residues. For example, Fig. 7A Depicted is a screenshot showing an example of a protein molecule where the backbone of its constituent amino acid residues is translated to determine the three-dimensional structure of the protein molecule. Figure 7B Depicted is a screenshot showing an example of a protein molecule in which the backbone of its constituent amino acid residues is rotated to determine the three-dimensional structure of the protein molecule. Figure 7C Depicted are screen shots showing examples of protein molecules in which the side chain torsion angles of its constituent amino acid residues are varied to determine the three-dimensional structure of the protein molecule.
[0152] In some exemplary embodiments, the molecular design computational model 117 may include a machine learning model that is trained to determine the three-dimensional structure of a protein molecule by denoising at least the initial three-dimensional structure of the protein molecule. In some cases, the machine learning model may denoise the initial three-dimensional structure of the protein molecule by performing a series of updates on at least the backbone torsion (BBT) representation of the protein molecule. In some cases, the machine learning model may be trained to reduce or minimize a loss function (such as a framework alignment point error (FAPE) loss function, a structural violation loss function, and / or the like), which is associated with each successive update of the initial three-dimensional structure of the protein molecule. Alternatively and / or additionally, the machine learning model may be trained to reduce or minimize an energy function that is associated with each successive update of the initial three-dimensional structure of the protein molecule.
[0153] In some exemplary embodiments, the molecular design computational model 117 may include a diffusion model that performs a series of modifications on the initial three-dimensional structure of the protein molecule, each modification removing a portion of the noise present in the initial three-dimensional structure of the protein molecule. Figure 6B An exemplary diffusion model is described. In addition, in some cases, each continuous update performed by the diffusion model may generate an output that is equivariant to the special Euclidean group SE (3) transformation. For example, in some cases, the diffusion model may perform a first update on the main chain torsion (BBT) representation of the protein molecule at a first time point to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule. In addition, the diffusion model may perform a second update on the main chain torsion (BBT) representation of the protein molecule at a second time point to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule. In some cases, the diffusion model may denoise the initial three-dimensional structure of the protein molecule by a large number of continuous updates (e.g., 2000 continuous denoising operations). In doing so, the diffusion model may be able to perform highly detailed modifications to refine the three-dimensional structure of the protein molecule. In contrast, the above-mentioned geometric deep learning model, which contains much fewer modules (e.g., 8 modules in an equivariant neural network (ENN)), can update the initial three-dimensional structure of the protein molecule by fewer and less refined updates.
[0154] At 558, the molecular design engine 110 may determine one or more coordinates of each atom in the three-dimensional structure of the protein molecule based at least on the modified representation of the protein molecule. In some exemplary embodiments, the molecular design engine 110 may determine one or more coordinates (e.g., three-dimensional coordinates) of each atom in the three-dimensional structure of the protein molecule based at least on a plurality of frameworks of each amino acid residue in the modified main chain torsion (BBT) representation of the protein molecule. For example, in some cases, the molecular design engine 110 may determine one or more coordinates of the main chain atoms in the protein molecule based at least on the modified main chain torsion (BBT) representation of the protein molecule. Thereafter, the molecular design engine 110 may determine one or more coordinates of the side chain atoms in the protein molecule based at least on the coordinates of the main chain atoms in the protein molecule. The algorithm for calculating the coordinates of each atom in the three-dimensional structure of the protein molecule is shown in Table 4 below.
[0155] Table 4
[0156]
[0157] In some exemplary embodiments, the molecular design computational model 117 can be used to determine the sequence and structure of a protein molecule, rather than just the three-dimensional structure of a fixed sequence of amino acid residues (such as a protein sequence generated by the sequence design computational model 113 (or another sequence design platform)). Figure 5C In the illustrated example of process 570, the molecular design computational model 117 may be used to determine the identity of each amino acid residue that forms a protein molecule and the spatial arrangement of the constituent atoms. As described in more detail below, the molecular design computational model 117 may determine the sequence and three-dimensional structure of a protein molecule by simultaneously modifying corresponding sequence representations and structural representations (e.g., coarse-grained (CG) node representations, backbone torsion (BBT) representations, and / or the like).
[0158] According to some exemplary embodiments, Figure 5C A flow chart showing an example of a process 570 for molecular structure and property prediction is depicted. Figures 1A to 1B , 2A to 2B, 3A to 3B, 4 and 5C, process 570 can be performed by design system 100, such as by molecular design engine 120. Figure 5C The example of process 570 shown in can implement a generative design process in which the molecular design engine 110 (e.g., the molecular design computational model 117) operates on a sequence representation and a structural representation (e.g., a coarse-grained (CG) node representation, a backbone torsion (BBT) representation, and / or the like) of a protein molecule to determine the sequence and three-dimensional structure of the protein molecule.
[0159] At 572, the molecular design engine 110 may determine a representation of the protein molecule, the representation comprising a sequence representation of the initial sequence of the protein molecule and a structural representation of the initial three-dimensional structure of the protein molecule. In some exemplary embodiments, the representation generator 115 may generate a representation of the protein molecule, the representation comprising a sequence representation of the initial sequence of the protein molecule and a structural representation of the initial three-dimensional structure of the protein molecule. In some cases, the structural representation of the initial three-dimensional structure may be a coarse-grained (CG) node representation, in which the atoms (e.g., heavy atoms) that form the constituent amino acid residues are represented as a collection of coarse-grained (CG) nodes. Alternatively, the structural representation of the initial three-dimensional structure may be a main-chain torsion (BBT) representation, which includes, for each amino acid residue in the protein molecule, a plurality of frameworks that specify the geometric states of the main-chain atoms and the side-chain atoms.
[0160] In some exemplary embodiments, the sequence representation of the protein molecule may be a logarithmic probability representation having a plurality of logarithmic probability vectors, each logarithmic probability vector corresponding to a position in the initial sequence of the protein molecule and indicating the identity of the amino acid residue occupying the position by at least enumerating the probability distribution of the set of amino acid residues occupying the position. That is, for each possible amino acid residue, the logarithmic vector for the position in the initial sequence of the protein molecule may include the corresponding probability that the position is occupied by the amino acid residue. Alternatively, the sequence representation of the protein molecule may be a one-hot encoding representation having a plurality of one-hot encoding vectors, each vector corresponding to a position in the initial sequence of the protein molecule and having a position corresponding to each amino acid residue in the possible set of amino acid residues. The one-hot encoding vector of a particular position in the protein sequence may indicate the identity of the amino acid residue occupying the position by having a value of "1" at least in the one-hot encoding vector corresponding to the amino acid residue occupying the position in the protein sequence and a value of "0" at other positions (e.g., [0,0,0,0,1,0,...,0]). In some cases, the set of amino acid residues may include "dummy residues" representing gaps in the sequence of amino acid residues. Thus, where the probability of a position in the sequence of amino acid residues being occupied by a virtual residue meets one or more thresholds, the position includes a gap that is not occupied by any amino acid residue. The inclusion of virtual residues can accommodate length variations as part of a subsequent generative diffusion process performed, for example, by the molecular design computational model 117.
[0161] As noted, in some exemplary embodiments, the representation generator 115 may generate a representation of the protein molecule to include a sequence representation of the initial sequence and a structural representation of the initial three-dimensional structure of the protein molecule. Thus, in some cases, for each position within the sequence of amino acid residues that form the protein molecule, the representation of the protein molecule may include a representation of the identity of the amino acid residue that occupies the position (e.g., a logarithmic representation, a one-hot encoding representation, and / or the like). In addition, for each position within the sequence of amino acid residues that form the protein molecule, the representation of the protein molecule may include a representation of the spatial arrangement of atoms (e.g., heavy atoms) that form the amino acid residue that occupies the position (e.g., a coarse-grained (CG) node representation, a main-chain torsion (BBT) representation, and / or the like). For example, for a single position within the sequence of amino acid residues forming a protein molecule, the representation of the protein molecule may include a logarithmic vector that enumerates the probability distribution (e.g., categorical distribution) of the set of possible amino acid residues occupying the position (e.g., a first probability that the position is occupied by alanine (Ala / A), a second probability that the position is occupied by arginine (Arg / R), a third probability that the position is occupied by asparagine (Asn / N), and / or the like). For the same position, the representation of the protein molecule may further include a representation of the spatial arrangement of atoms (e.g., heavy atoms) that form the amino acid residue occupying the position. In the case of a coarse-grained (CG) node representation, the representation may include a collection of coarse-grained (CG) nodes, each of which includes two or more atoms in the amino acid residue. Alternatively, in the case of a main-chain torsion (BBT) representation, the representation may include main-chain translations, main-chain rotations, and torsion angles formed by atoms in the amino acid residue.
[0162] To further illustrate, consider a protein molecule formed by a sequence of N amino acid residues. Each residue i=1,...,N can be associated with the following degrees of freedom, which means that the molecular design computational model 117 can perform the following modifications when operating on the representation of the protein molecule: (I) Residue identity c i ∈C, where C represents a set of possible amino acid residues, (ii) main chain translation x i =R 3 , (iii) Main chain rotation r i = SO(3), and (iv) torsion angles (e.g., one for the backbone oxygen (O) and four for the side chain angles): and θ q ∈SO(2). In general, the above degrees of freedom can be expressed as in Furthermore, in some cases, each of the aforementioned degrees of freedom may correspond to a framework that is modified by the molecular design computational model 117 when a generative process (eg, a generative diffusion process) is performed to determine the sequence and three-dimensional structure of a protein molecule.
[0163] At 574, the molecular design engine 110 may apply the molecular design computational model 117 to determine the sequence and three-dimensional structure of the protein molecule by modifying at least the sequence representation and the structural representation of the protein molecule. In some exemplary embodiments, the molecular design computational model 117 may determine the three-dimensional structure of the protein molecule by modifying at least the sequence representation of the initial sequence of the protein molecule and the structural representation of the three-dimensional structure of the protein molecule. In some cases, the structural representation may be a coarse-grained (CG) node representation or a main chain torsion (BBT) representation of the initial three-dimensional structure of the protein molecule. In addition, in some cases, the molecular design computational model 117 may be a diffusion model that determines the sequence and three-dimensional structure of the protein molecule by denoising the initial sequence and the initial three-dimensional structure of the protein molecule at least at a series of time points. For example, denoising the initial sequence may include operating the sequence representation of the initial sequence so as to modify the identity of one or more amino acid residues included in the initial sequence. At the same time, denoising the initial three-dimensional structure may include the molecular design computational model 117 operating the structural representation of the initial three-dimensional structure to modify the spatial arrangement of atoms (e.g., heavy atoms) that form the amino acid residues in the protein molecule.
[0164] In some exemplary embodiments, when the molecular design computational model 117 determines the sequence and three-dimensional structure of the protein molecule, the molecular design computational model 117 may perform a diffusion process in which the identity of the amino acid residues is modified to another degree of freedom (DoF) in addition to the degree of freedom (DoF) associated with the spatial arrangement of the constituent atoms. Therefore, in some cases, the molecular design computational model 117 may be a diffusion model that includes a first diffusion kernel that modifies the identity of each amino acid residue, a second diffusion kernel that modifies the main chain translation of each amino acid residue, a third diffusion kernel that modifies the main chain rotation of each amino acid residue, and a fourth diffusion kernel that modifies the torsion angle in the main chain and side chain of each amino acid residue. In some cases, the first diffusion kernel, the second diffusion kernel, the third diffusion kernel, and the fourth diffusion kernel can each be parameterized as a neural network, including, for example, an equivariant neural network (ENN) that can recognize or take into account the rotational symmetry present in the three-dimensional structure of the protein molecule.
[0165] In some exemplary embodiments, the molecular design computational model 117 may perform a generative diffusion process to determine the sequence and three-dimensional structure of the protein molecule. In some cases, the molecular design computational model 117 may be a diffusion model that determines the sequence and three-dimensional structure of the protein molecule by denoising the initial sequence and initial three-dimensional structure of the protein molecule at least at consecutive time steps. For example, in some cases, the molecular design computational model 117 may remove a first amount of noise from the representation of the protein molecule before removing a second amount of noise from the representation of the protein molecule. It should be understood that the generative diffusion process may correspond to a reverse diffusion process, in which the molecular design computational model 117 removes noise from the representation of the protein molecule according to a decreasing noise level, so that there is less noise in the sequence and three-dimensional structure of the protein molecule at each consecutive time step. However, the training of the molecular design computational model 117 may also include a forward diffusion process, in which noise is added according to an increased noise level, so that there is more noise in the three-dimensional structure of the protein molecule at each consecutive time step. Thus, training of the molecular design computational model 117 may include learning a back diffusion process to recover the correct sequence and three-dimensional structure of a protein molecule from a noisy representation of the protein molecule, where the identities of the amino acid residues forming the protein molecule are uncertain and the three-dimensional structure of the protein molecule is random.
[0166] In some exemplary embodiments, the training of the molecular design computational model 117 may include learning a scoring function for each diffusion kernel (e.g., parameterized by a neural network) included in the molecular design computational model 117. For example, in some cases, the molecular design computational model 117 may be implemented using a stochastic differential equation (SDE) score matching framework, where a stochastic differential equation (SDE) is applied to smoothly transform samples from a complex data distribution (in this case, which may be populated by the real sequences and three-dimensional structures of various known protein molecules) to corresponding samples in a noise distribution by injecting noise. A corresponding inverse stochastic differential equation (SDE) may be applied to restore the samples from the original complex data distribution (e.g., the original sequence and three-dimensional structure of the protein molecule) by removing the noise. The above equations (2) and (3) are examples of the above-mentioned forward stochastic differential equation (SDE) and inverse stochastic differential equation (SDE).
[0167] Referring to equations (2) and (3), the training of the diffusion model in the stochastic differential equation (SDE) score matching framework may include learning a score-based model s θ (x), the model approximates the corresponding score function for each possible degree of freedom (DoF) For example, in some cases, a scoring function for a first diffusion kernel that modifies the identity of each amino acid residue may represent a change in the logarithmic data density of a complex data distribution associated with a known protein sequence. Learning the score function for the first diffusion kernel may enable, at successive time steps during the backdiffusion process (e.g., by applying Markov chain Monte Carlo sampling with Langevin dynamics), to extract samples from regions of the original complex data distribution where known protein sequences are more densely distributed. Similarly, a scoring function for a second diffusion kernel that modifies the backbone translation of each amino acid residue may represent a change in the logarithmic data density of a complex data distribution associated with a backbone translation of a known three-dimensional protein structure. Learning the scoring function for the second diffusion kernel may extract samples from regions of the original complex data distribution where backbone torsions of known three-dimensional protein structures are more densely distributed at successive time steps during the backdiffusion process (e.g., by applying Markov chain Monte Carlo sampling with Langevin dynamics). The scoring function for each diffusion kernel may take in additional degrees of freedom as input. For example, the scoring function of a first diffusion kernel modified for a single amino acid residue identity may take as input the backbone translation, backbone rotation, and torsion angle of the amino acid residue determined by the corresponding diffusion kernel, thereby enabling the simultaneous learning of scoring functions of multiple diffusion kernels for the same protein molecule.
[0168] In some exemplary embodiments, each diffusion stage may include adding some noise back to the modified representation of the initial sequence and initial three-dimensional structure of the protein molecule after noise removal. For example, in some cases, after removing a first amount of noise from the representation of the protein molecule, the molecular design computational model 117 may add a third amount of noise back to the backbone torsion (BBT) representation of the protein molecule before removing a second amount of noise from the representation of the protein molecule. After removing the second amount of noise from the representation of the protein molecule, a fourth amount of noise may be added back before further denoising the representation of the protein molecule by the diffusion model. The third amount and the fourth amount of noise added to the backbone torsion (BBT) representation of the protein molecule may be determined based on a noise schedule that defines the distribution of noise levels across a sequence of diffusion operations performed by the diffusion model. The addition of noise may compensate for at least some of the errors that may exist in the denoising performed by the diffusion model at each time point.
[0169] According to some exemplary embodiments, Fig. 6A A schematic diagram showing an example atomic structure of an amino acid residue 600 is depicted. In some cases, the main chain torsion (BBT) representation of amino acid residue 600 can specify the geometric state of the main chain of amino acid residue 600 in a variety of different ways. Fig. 6A In the example of amino acid residue 600 shown, the backbone of amino acid residue 600 may include nitrogen (N), α carbon (C α) atom and a carbonyl group formed by coupling a carbon atom to an oxygen atom. Therefore, in some cases, the multiple frameworks associated with the amino acid residue 600 may include a first framework that defines the geometric state of the main chain of the amino acid residue 600, and may specify the rotation and translation of the main chain of the amino acid residue 600. For example, in some cases, the first framework may include an affine transformation matrix, which includes a rotation matrix that specifies the rotation of the main chain of the amino acid residue 600 and a displacement vector that specifies the translation of the main chain of the amino acid residue 600. In some cases, in addition to the first framework that specifies the translation and rotation of the main chain of the amino acid residue 600, the multiple frameworks associated with the amino acid residue 600 may include a second framework that is used to specify the α carbon atom (C α ) and the torsion angle ψ of the rotatable bond between the carbonyl group.
[0170] In some cases, instead of translation and rotation of the main chain of amino acid residue 600, the geometric state of the main chain of the amino acid residue can be specified by the torsion angles present therein. Thus, in some cases, the plurality of frameworks associated with amino acid residue 600 may include the alpha carbon (C α ) atom and the carbonyl group, specifying the torsion angle ψ of the first frame of the rotatable bond between the α carbon (C α ) atom and the nitrogen (N) atom, and a third framework of the torsion angle ω of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom in the main chain of the specified amino acid residue.
[0171] Reference again Fig. 6A In addition to the framework that specifies the geometric state of the main chain of the amino acid residue 600, the multiple frameworks associated with the amino acid residue 600 may further include one or more additional frameworks that specify the torsion angles present in the side chains of the amino acid residue 600. Fig. 6A In the example shown, the plurality of frameworks associated with amino acid residue 600 may include frameworks for the following rotatable bond torsion angles: α carbon (C α ) atoms and β carbon (C β ) The torsion angle χ of the rotatable bond between atoms 1 , β carbon (C β ) atoms and γ carbon (C γ ) The torsion angle χ of the rotatable bond between atoms 2 , γ carbon (C γ ) atoms and delta carbon (C δ ) The torsion angle χ of the rotatable bond between atoms 3 and delta carbon (C δ The torsion angle χ of the rotatable bond between the N atom and the N atom 3 .
[0172] In some cases, the backbone torsion (BBT) representation of a protein molecule can be further generated to include multiple polymer chains, each of which includes one or more amino acid residues in the protein molecule. For example, in some cases, the residues included in each polymer chain can be regarded as a single rigid body. Therefore, in some cases, the residues in the polymer chain (including its constituent atoms) can translate and rotate around the center of mass as a group of atoms. That is, the residues in the polymer chain can share the center of mass degrees of freedom (DoF) as a group of atoms.
[0173] According to some exemplary embodiments, Figure 6B A schematic diagram illustrating an example of a diffusion framework is depicted. In some cases, when implemented as the above-described diffusion model, the molecular design computational model 117 may perform an inverse diffusion process when denoising the initial three-dimensional structure of the protein molecule. In contrast, Figure 6B The forward diffusion process shown in may be performed during the training of the diffusion model, in which case noise is sequentially added to the true 3D structure to generate a corrupted 3D structure before the trained diffusion model performs backward diffusion to sequentially remove noise from the corrupted 3D structure and restore the true 3D structure.
[0174] Referring again to 6B, in some cases, the forward diffusion process may include adding noise according to an increased noise scale (to interfere with or destroy the original data), so that at each continuous time step, there is more noise in the three-dimensional structure of the protein molecule. In contrast, the reverse diffusion process may include removing noise according to a reduced noise scale (to restore the original data), so that at each continuous time step, there is less noise in the three-dimensional structure of the protein molecule. In some cases, the three-dimensional structure of a molecule (such as a protein molecule) can be generated by performing a reverse diffusion process on the initial three-dimensional structure of the molecule, wherein each constituent atom occupies a random position in three-dimensional space. As noted, the diffusion model can gradually remove noise from the initial three-dimensional structure of the molecule at a series of time points. In the case where the initial three-dimensional structure of the molecule is presented in the form of a main chain torsion (BBT) representation, the initial three-dimensional structure of the molecule may include noise in the spatial arrangement of atoms in the molecular side chains and the main chain. For example, noise may exist in various degrees of freedom (DoF) in which atoms can move within the three-dimensional structure of the molecule. Therefore, the denoising of the initial three-dimensional structure (including modification of the positions of these atoms) may be limited to certain degrees of freedom, including, for example, main chain translation x i =R 3 , main chain rotation r i = SO(3) and five torsion angles (one for oxygen (O) and four for side chain angles): and θ q∈SO(2). In addition, in some cases, the diffusion model may include multiple diffusion kernels, each of which modifies the initial three-dimensional structure of the molecule along a corresponding degree of freedom (DoF). In some cases, each diffusion kernel may be parameterized as a neural network. In addition, in some cases, each diffusion kernel may be parameterized as an equivariant neural network capable of generating a correct three-dimensional structure regardless of the orientation of the initial three-dimensional structure ingested as input.
[0175] Figure 6B An example of implementing a diffusion model using a stochastic differential equation (SDE) score matching framework is shown, in which a stochastic differential equation (SDE) is applied to smoothly transform samples from a complex data distribution (in this case, it can be filled with the real three-dimensional structures of various known protein molecules) to corresponding samples in a noise distribution by injecting noise. At the same time, the corresponding inverse time stochastic differential equation (SDE) can be applied to recover samples from the original complex data distribution (e.g., the original three-dimensional structure of the protein molecule) by removing noise. As a score-based generative model, the inverse stochastic differential equation (SDE) that controls the generative process of determining the three-dimensional structure of the protein molecule can be learned by learning a score function (or the gradient of the logarithmic probability density function) of the data distribution. In this case, the score of the data distribution at each time point of the diffusion process determined by the score function can correspond to the change in the logarithmic data density. Learning the score function can not only approximate the original complex data distribution, but also realize the process of returning samples in the noise distribution to corresponding samples in the complex data distribution. Unlike the probability density function of the data distribution, the score function can be calculated without a normalization constant, which requires determining a set of all possible values and is usually a tricky calculation. As described in more detail below, for example, during the training of a diffusion model based on a stochastic differential equation (SDE), the score function may be estimated by score matching. In addition, it should be understood that there may be a separate score function for each degree of freedom (DoF). The score functions for multiple degrees of freedom (DoF) may be determined simultaneously during the training of a diffusion model based on a stochastic differential equation (SDE).
[0176] Equation (2) below is an example of a forward stochastic differential equation (SDE) that transforms a complex data distribution from the original distribution x(0) to a noise distribution x(T). Equation (3) is an example of the corresponding inverse stochastic differential equation (SDE) that recovers the original complex data distribution x(0) from pure noise as a known prior distribution x(T).
[0177] dx=f(x,t)dt+g(t)dw (2)
[0178]
[0179] where f(.,t) is a vector-valued function called the drift coefficient of x(t), and g(t) is a scalar function called the diffusion coefficient of x(t).
[0180] In some exemplary embodiments, training of a diffusion model in a stochastic differential equation (SDE) score matching framework may include learning a score-based model s θ (x), the model approximates the corresponding score function log p for each possible degree of freedom (DoF) t In some cases, the score function for each degree of freedom (DoF) is can be approximated by fractional matching, which involves reducing or minimizing the difference (e.g., Fisher divergence or squared distance), the difference is in comparison with the score-based model s θ (x) and the true score of the data distribution, which in this case is the original distribution x(0). As mentioned above, the score function log p t A change in the logarithmic data density across the original data distribution x(0) may be defined. Thus, once an estimate of the score function is calculated, a three-dimensional structure of a molecule (e.g., a protein molecule) may be generated by sampling from the original data distribution x(0) (e.g., Markov chain Monte Carlo sampling with Langevin dynamics) as directed by the score function. In doing so, a three-dimensional structure of a molecule may be generated by sampling from progressively higher density regions of the data distribution x(0), which are occupied by three-dimensional structures that are more consistent with true three-dimensional molecular structures.
[0181] According to some exemplary embodiments, Figures 8A to 8D A diagram showing various examples of noise scheduling is depicted. In some cases, the distribution of noise levels may correspond to the degrees of freedom present in the representation of the protein molecule for the molecular design computational model 117 to modify the initial three-dimensional structure of the protein molecule. For example, in some cases, the proportion of uncertain torsion angles and residue identities in the protein molecule may decrease over time. Therefore, at the starting time point (e.g., t=1), almost every torsion angle in the protein molecule may be randomly oriented, and the identity of almost every residue is unclear. At a second time point (e.g., t=0.5), after the representation of the protein molecule has been subjected to at least some denoising by the molecular design computational model 117, the identities of at least some residues may become more certain, thereby reducing the randomness of at least some torsion angles. At a third time point (e.g., t=0), the identities of almost all residues may be determined, and the corresponding torsion angles are also known at this time.
[0182] At 576, the molecular design engine 110 may determine one or more coordinates of each atom in the three-dimensional structure of the protein molecule based at least on the modified representation of the protein molecule. In some exemplary embodiments, the molecular design engine 110 may determine one or more coordinates (e.g., three-dimensional coordinates) of each atom in the three-dimensional structure of the protein molecule based at least on multiple frameworks of each amino acid residue in the modified main chain torsion (BBT) representation of the protein molecule. For example, in some cases, the molecular design engine 110 may determine one or more coordinates of the main chain atoms in the protein molecule based at least on the modified main chain torsion (BBT) representation of the protein molecule. Thereafter, the molecular design engine 110 may determine one or more coordinates of the side chain atoms in the protein molecule based at least on the coordinates of the main chain atoms in the protein molecule. An example of an algorithm for calculating the coordinates of each atom in the three-dimensional structure of a protein molecule is shown in Table 4 above.
[0183] In view of the above specific implementation of the subject matter, the present application discloses the following list of examples, wherein one feature of a single example or a combination of more than one feature of the example, and optionally a combination with one or more features of one or more other examples, are other examples that also fall within the disclosure scope of the present application:
[0184] Item 1: A computer-implemented method, comprising: receiving a molecular structure file specifying an initial three-dimensional structure of a molecule; determining a plurality of coarse-grained nodes based at least on the molecular structure file, each coarse-grained node corresponding to a structure of two or more atoms (e.g., heavy atoms) that form an amino acid residue in the molecule; and determining the three-dimensional structure of the molecule using a designed computational model, the designed computational model determining the three-dimensional structure of the molecule by at least determining the position of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.
[0185] Item 2: The method according to Item 1, wherein the structure of two or more atoms does not include one or more elements that form amino acid residues in the molecule.
[0186] Item 3: The method according to any one of Items 1 to 2, wherein the structure of two or more atoms excludes one or more heavy atoms forming an amino acid residue in the molecule.
[0187] Item 4: A method according to any one of Items 1 to 3, further comprising: for each coarse-grained node among a plurality of coarse-grained nodes, generating a geometric tensor embedding corresponding to a numerical representation of a rotation and / or translation of the coarse-grained node.
[0188] Item 5: A method according to Item 4, wherein the geometric tensor embedding comprises a set of geometric tensors that undergo one or more rotations and / or translations.
[0189] Item 6: A method according to Item 5, wherein one or more rotations and / or translations correspond to one or more elements from a three-dimensional rotation group, which enumerates every possible non-trivial rotation in three-dimensional space, which non-trivial rotation cannot be further decomposed into a combination of two or more other rotations.
[0190] Item 7: A method according to any one of Items 1 to 6, wherein the design computational model comprises a machine learning model that is trained to perform continuous updates on the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.
[0191] Item 8: A method according to Item 7, wherein the machine learning model comprises a sequence of blocks, and wherein each block in the sequence of blocks performs an update to the position of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.
[0192] Item 9: A method according to any one of Items 7 to 8, wherein the machine learning model includes a first block for performing a first update on the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule, and wherein the machine learning model includes a second block for performing a second update on the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.
[0193] Item 10: A method according to any one of Items 7 to 9, wherein the machine learning model is a geometric deep learning model.
[0194] Item 11: A method according to any one of Items 7 to 10, wherein the machine learning model is an equivariant neural network.
[0195] Item 12: A method according to any one of Items 7 to 11, wherein the machine learning model is trained to reduce a loss function associated with each successive update to the position of one or more coarse-grained nodes in the initial three-dimensional structure of the molecule.
[0196] Item 13: The method of Item 12, wherein the loss function is a framework aligned point error (FAPE) loss function and / or a structure destruction loss function.
[0197] Item 14: A method according to any one of Items 7 to 13, wherein the machine learning model is trained to reduce an energy function associated with each successive update to the position of one or more coarse-grained nodes in the initial three-dimensional structure of the protein sequence.
[0198] Item 15: A method according to any one of Items 7 to 14, wherein the machine learning model identifies when two or more three-dimensional structures having different orientations in three-dimensional space are the same.
[0199] Item 16: A method according to any one of Items 1 to 15, wherein the position of one or more coarse-grained nodes in the initial three-dimensional structure of the protein is updated by at least rotating and / or translating one or more coarse-grained nodes.
[0200] Item 17: A method according to any one of Items 1 to 16, wherein the positions of one or more coarse-grained nodes in the initial three-dimensional structure of the protein are updated without modifying the relative positions of two or more atoms included in the structure corresponding to each coarse-grained node.
[0201] Item 18: A method according to any one of items 1 to 17, wherein a plurality of coarse-grained nodes are determined such that a union of the plurality of coarse-grained nodes includes each atom included in the molecule.
[0202] Item 19: A method according to any one of Items 1 to 18, wherein a plurality of coarse-grained nodes are determined by at least grouping a plurality of atoms included in a molecule such that each atom in the coarse-grained node shares at least one covalent bond with another atom in the same coarse-grained node.
[0203] Item 20: A method according to any one of Items 1 to 19, wherein a plurality of coarse-grained nodes are determined such that each coarse-grained node includes a threshold number of atoms (e.g., heavy atoms) that form at least one structure.
[0204] Item 21: A method according to any one of items 1 to 20, wherein the three-dimensional structure of the protein is associated with one or more desired properties.
[0205] Item 22: A method according to any one of Items 1 to 21, further comprising: determining one or more properties of the molecule based at least on the three-dimensional structure of the protein.
[0206] Item 23: A method according to any one of Items 1 to 22, further comprising: generating a second sequence of amino acid residues comprising a molecule based at least on the first sequence of amino acid residues; and generating a fourth sequence of amino acid residues based at least on the second sequence of amino acid residues or the third sequence of amino acid residues.
[0207] Item 24: The method according to Item 23 further comprises: determining that the second sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties, at least based on the three-dimensional structure of the molecule; and in response to determining that the second sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties, generating a fourth sequence of amino acid residues, at least based on the second sequence of amino acid residues.
[0208] Item 25: The method according to Item 24 further comprises: determining that the second sequence of amino acid residues lacks the desired three-dimensional structure or desired properties based at least on the three-dimensional structure of the molecule; and in response to determining that the second sequence of amino acid residues lacks the desired three-dimensional structure or desired properties, generating a fourth sequence of amino acid residues based at least on the third sequence of amino acid residues.
[0209] Item 26: The method according to any one of Items 1 to 25, wherein the molecule is a protein molecule, a small molecule, an ion, a nucleic acid, a polysaccharide or a glycolipid.
[0210] Item 27: The method according to any one of Items 1 to 26, wherein the structure is a rigid body or a flexible body of two or more atoms forming an amino acid residue in a molecule.
[0211] Item 28: A system comprising: at least one data processor; and at least one memory storing instructions that, when executed by the at least one data processor, cause the operation of the method comprising any one of items 1 to 27.
[0212] Item 29: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, result in operations including those of the method described in any one of Items 1 to 27.
[0213] Item 30: A computer-implemented method, comprising: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule, the protein molecule comprising a first sequence of amino acid residues; determining, based at least on the molecular structure file, a representation of the protein molecule, the representation comprising a plurality of frameworks for each amino acid residue in the first sequence of amino acid residues, the plurality of frameworks for each amino acid residue comprising a first set of frameworks specifying a geometric state for the main chain of the amino acid residue, and the plurality of frameworks for each amino acid residue further comprising a second set of frameworks specifying one or more torsion angles in the side chain of the amino acid residue; and generating a first three-dimensional structure of the protein molecule by modifying the representation of the protein molecule by at least applying a designed computational model.
[0214] Item 31: A method according to Item 30, wherein each of the multiple frameworks corresponds to a degree of freedom, which is used to design a computational model to update the initial three-dimensional structure of the protein molecule.
[0215] Item 32: A method according to any one of items 30 to 31, wherein the first set of frameworks comprises a first framework comprising an affine transformation matrix specifying rotation and translation of the backbone of the amino acid residues.
[0216] Item 33: A method according to Item 32, wherein the affine transformation matrix includes a rotation matrix that specifies the rotation of the main chain of the amino acid residues, and wherein the affine transformation matrix further includes a displacement vector that specifies the translation of the main chain of the amino acid residues.
[0217] Item 34: The method according to any one of items 32 to 33, wherein the first set of frameworks further comprises a second framework specifying the torsion angles in the main chain of the amino acid residues.
[0218] Item 35: A method according to Item 34, wherein the torsion angle is about the α carbon (C α ) atom and the carbonyl group.
[0219] Item 36: A method according to any one of Items 30 to 35, wherein the first group of frameworks includes a first framework that specifies a first torsion angle in the main chain of an amino acid residue, and wherein the first group of frameworks further includes a second framework that specifies a second torsion angle in the main chain of an amino acid residue.
[0220] Item 37: A method according to Item 36, wherein the first torsion angle is with respect to the α carbon (C α ) atom and a carbon (C) atom, and wherein the second torsion angle is associated with a first rotatable bond between the alpha carbon (C) atom in the backbone of the amino acid residue. α ) atom and the second rotatable bond between the nitrogen (N) atom.
[0221] Item 38: The method according to Item 37, wherein the first set of frameworks further comprises a third framework specifying a third torsion angle present in the main chain of the amino acid residue.
[0222] Item 39: A method according to Item 38, wherein the third torsion angle is associated with a third rotatable bond between a carbon (C) atom and a nitrogen (N) atom in the backbone of the amino acid residue.
[0223] Item 40: The method of any one of Items 30 to 39, further comprising: determining one or more coordinates of each atom constituting the first three-dimensional structure of the protein molecule based at least on the modified representation of the protein molecule.
[0224] Item 41: A method according to Item 40, wherein one or more coordinates of each atom in the first three-dimensional structure of the protein molecule are determined based at least on a plurality of frameworks associated with each amino acid residue included in the modified representation of the protein molecule.
[0225] Item 42: A method according to Item 41, wherein one or more coordinates of each atom in the first three-dimensional structure of the protein molecule are determined by at least determining one or more coordinates of multiple backbone atoms in the protein molecule.
[0226] Item 43: A method according to Item 42, wherein one or more coordinates of each atom in the first three-dimensional structure of the protein molecule are further determined by determining one or more coordinates of multiple side chain atoms in the molecule based at least on one or more coordinates of multiple main chain atoms in the molecule.
[0227] Item 44: A method according to any one of Items 30 to 43, wherein the designed computational model includes a machine learning model that is trained to generate a first three-dimensional structure of a protein molecule by at least denoising an initial three-dimensional structure of the protein molecule.
[0228] Item 45: A method according to Item 44, wherein the machine learning model denoises the initial three-dimensional structure of the protein molecule by performing a series of updates on at least the representation of the protein molecule.
[0229] Item 46: A method according to Item 45, wherein the machine learning model is trained to reduce a loss function associated with each successive update to the initial three-dimensional structure of the protein molecule.
[0230] Item 47: A method according to Item 46, wherein the loss function is a framework alignment point error (FAPE) loss function and / or a structural damage loss function.
[0231] Item 48: A method according to any one of Items 45 to 46, wherein the machine learning model is trained to reduce an energy function associated with each successive update to the initial three-dimensional structure of the protein molecule.
[0232] Item 49: A method according to any one of Items 30 to 48, wherein the machine learning model is a diffusion model that removes a portion of the noise present in the initial three-dimensional structure of the protein molecule at each time step in multiple consecutive time steps.
[0233] Item 50: A method according to Item 49, wherein the diffusion model performs a first update on the representation of the protein molecule to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule, and wherein the diffusion model further performs a second update on the representation of the protein molecule to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule.
[0234] Item 51: A method according to Item 50, wherein the diffusion model further adds a third amount of noise before performing a second update to remove the second amount of noise, and adds a fourth amount of noise after performing a second update to remove the second amount of noise, and wherein the third amount of noise and the fourth amount of noise are determined based on a noise schedule, which defines the distribution of noise levels added across multiple consecutive time steps.
[0235] Item 52: A method according to Item 51, wherein the distribution of noise levels corresponds to the degrees of freedom present in the representation of the protein molecule, which degrees of freedom are used by the computational model to modify the initial three-dimensional structure of the protein molecule.
[0236] Item 53: A method according to any one of items 49 to 52, wherein each update performed by the diffusion model generates an output that is equivariant to the special Euclidean group SE(3) transformation.
[0237] Item 54: A method according to any one of Items 30 to 53, wherein the modification of the representation of the protein molecule comprises updating the first set of frameworks to change the geometric state of the main chain of one or more amino acid residues in the protein molecule.
[0238] Item 55: A method according to any one of items 30 to 54, wherein the modification of the representation of the protein molecule comprises updating the second set of frameworks to change one or more torsion angles in the side chains of one or more amino acid residues in the protein molecule.
[0239] Item 56: A method according to any one of Items 30 to 55, wherein the first three-dimensional structure of the protein molecule is associated with one or more desired properties.
[0240] Item 57: The method according to any one of Items 30 to 56, wherein the first three-dimensional structure of the protein molecule is configured for one or more downstream tasks.
[0241] Item 58: The method of Item 57, wherein the one or more downstream tasks include determining one or more properties of the protein molecule based at least on the first three-dimensional structure of the protein molecule.
[0242] Item 59: A method according to any one of Items 30 to 58, further comprising: determining that a first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties based at least on the first three-dimensional structure of the protein molecule; and in response to determining that the first sequence of amino acid residues exhibits the desired three-dimensional structure and / or desired properties, generating a second sequence of amino acid residues for a different protein molecule based at least on the first sequence of amino acid residues.
[0243] Item 60: A method according to any one of Items 30 to 59, wherein the representation of the protein molecule further includes a logical vector for each position in the sequence of amino acid residues forming the protein molecule, which indicates the identity of the amino acid residue occupying the position by at least enumerating the probability distribution of a set of possible amino acid residues occupying the position.
[0244] Item 61: A method according to Item 60, wherein the designed computational model further generates the first three-dimensional structure of the protein molecule by modifying the identification of at least one amino acid residue while modifying the first set of frameworks and / or the second set of frameworks associated with at least one amino acid residue in the first sequence of residues.
[0245] Item 62: A method according to any one of Items 30 to 61, wherein the initial three-dimensional structure of the protein molecule includes noise in the identification of each amino acid residue and / or the spatial arrangement of multiple atoms forming each amino acid, and wherein the noise is removed by modifying the representation of the protein molecule by designing a computational model.
[0246] Item 63: A method according to Item 62, wherein the noise is Gaussian noise.
[0247] Item 64: A method according to any one of items 30 to 63, wherein the representation of the protein molecule is further generated to include a plurality of polymer chains, and wherein each polymer chain includes one or more amino acid residues from the first sequence of amino acid residues.
[0248] Item 65: A method according to Item 64, wherein modifying the representation of the protein molecule includes modifying the position of one or more amino acids as a group in each polymer chain.
[0249] Item 66: A system comprising: at least one data processor; and at least one memory storing instructions that, when executed by the at least one data processor, cause operations of the method comprising any one of items 30 to 65.
[0250] Item 67: A non-transitory computer readable medium storing instructions which, when executed by at least one data processor, result in operations including those of the method described in any one of Items 30 to 65.
[0251] Fig. 9 A block diagram illustrating an example of a computing system 1100 according to some exemplary embodiments is depicted. Referring to Figures 1-9, the computing system 1100 may be used to implement the molecular design engine 110, the molecular analysis engine 120, the client device 130, and / or any components thereof.
[0252] like Fig. 9As shown, the computing system 1100 may include a processor 1110, a memory 1120, a storage device 1130, and an input / output device 1140. The processor 1110, the memory 1120, the storage device 1130, and the input / output device 1140 may be interconnected via a system bus 1150. The processor 1110 is capable of processing instructions for execution within the computing system 1100. Such executed instructions may implement, for example, one or more components of the molecular design engine 110, the molecular analysis engine 120, the client device 130, and / or the like. In some exemplary embodiments, the processor 1110 may be a single-threaded processor. Alternatively, the processor 1110 may be a multi-threaded processor. The processor 1110 is capable of processing instructions stored on the memory 1120 and / or the storage device 1130 to display graphical information for a user interface provided via the input / output device 1140.
[0253] The memory 1120 is a computer-readable medium, such as a volatile or non-volatile computer-readable medium, that stores information within the computing system 1100. For example, the memory 1120 may store a data structure representing a configuration object database. The storage device 1130 is capable of providing persistent storage for the computing system 1100. The storage device 1130 may be a floppy disk device, a hard disk device, an optical disk device, a tape device, or other suitable persistent storage device. The input / output device 1140 provides input / output operations for the computing system 1100. In some exemplary embodiments, the input / output device 1140 includes a keyboard and / or a pointing device. In various specific implementations, the input / output device 1140 includes a display unit for displaying a graphical user interface.
[0254] According to some exemplary embodiments, the input / output device 1140 may provide input / output operations for network devices. For example, the input / output device 1140 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0255] In some exemplary embodiments, the computing system 1100 can be used to execute various interactive computer software applications that can be used to organize, analyze and / or store data in various formats. Alternatively, the computing system 1100 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generating, managing, editing electronic spreadsheet documents, word processing documents and / or any other objects, etc.), computing functions, communication functions, etc. The application may include various additional functions or may be an independent computing product and / or function. After activation within the application, the function can be used to generate a user interface provided via the input / output device 1140. The user interface can be generated by the computing system 1100 and presented to the user (e.g., on a computer screen monitor, etc.).
[0256] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuits, integrated circuits, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include a specific implementation in one or more computer programs, which are executable and / or interpretable on a programmable system, which includes at least one programmable processor (which may be dedicated or general, coupled to receive data and instructions from it and send data and instructions to it), a storage system, at least one input device, and at least one output device. A programmable system or computing system may include a client and a server. Typically, the client and the server are remotely arranged from each other and generally interact through a communication network. The relationship between the client and the server is generated by means of computer programs running on respective computers and the client-server relationship between each other.
[0257] These computer programs may also be referred to as programs, software, applications, applications, components or codes, including machine instructions for programmable processors, and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine languages. As used herein, the term "machine-readable medium" refers to any computer product, device and / or equipment (such as, for example, disks, optical disks, memories and programmable logic devices (PLDs)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor. A machine-readable medium may store such machine instructions non-temporarily (such as, for example, a non-temporary solid-state memory or a magnetic hard drive or any equivalent storage medium). A machine-readable medium may store such machine instructions in a temporary manner (such as, for example, a processor cache or other random access memory associated with one or more physical processor cores) alternatively or additionally.
[0258] To provide interaction with a user, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to a user) and a keyboard and a pointing device (such as, for example, a mouse or trackball, through which a user can provide input to the computer). Other types of devices may also be used to provide interaction with a user. For example, the feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; the input from the user may be received in any form, including sound, voice, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive tracking pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0259] In the above description and claims, phrases such as "at least one" or "one or more" may appear, followed by a list of combinations of elements or features. The term "and / or" may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to represent any element or feature listed alone, or any other described element or feature combined with any other described element or feature. For example, the phrase "at least one of A and B"; "one or more of A and B"; "A and / or B" are each intended to represent "single A, single B, or A and B together". A similar interpretation also applies to lists including three or more items. For example, the phrases "at least one of A, B, and C"; "one or more of A, B, and C" and "A, B, and / or C" are each intended to represent "single A, single B, single C, A and B together, A and C together, B and C together, or A and B and C together". The use of the term "based on" above and in the claims is intended to represent "based at least in part", so that undescribed features or elements are also permissible.
[0260] Depending on the desired configuration, the subject matter described herein may be embodied in systems, devices, methods and / or articles. The embodiments described in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Instead, they are only some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, in addition to those features and / or variations described herein, other features and / or variations may also be provided. For example, the above-mentioned specific implementations may be directed to various combinations and sub-combinations of the disclosed features and / or to combinations and sub-combinations of several further features disclosed above. In addition, the logical flows depicted in the drawings and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other specific implementations may be within the scope of the following claims.
Claims
1. A computer-implemented method, include: receiving a molecular structure file specifying an initial three-dimensional structure of a protein molecule, the protein molecule comprising a first sequence of amino acid residues; determining, based at least on the molecular structure file, a representation of the protein molecule, the representation comprising a plurality of frames for each amino acid residue in the first sequence of amino acid residues, the plurality of frames for each amino acid residue comprising a first set of frames specifying a geometric state for a main chain of the amino acid residue, and the plurality of frames for each amino acid residue further comprising a second set of frames specifying one or more torsion angles in a side chain of the amino acid residue; as well as A first three-dimensional structure of the protein molecule is generated by applying at least a design computational model to modify the representation of the protein molecule.
2. The method according to claim 1, wherein each of the plurality of frameworks corresponds to a degree of freedom, and the degree of freedom is used for the design computational model to update the initial three-dimensional structure of the protein molecule.
3. A method according to any one of claims 1 to 2, wherein the first set of frameworks comprises a first framework, the first framework comprising an affine transformation matrix specifying the rotation and translation of the main chain of the amino acid residue, and wherein the first set of frameworks further comprises a second framework, the second framework specifying the torsion angle in the main chain of the amino acid residue.
4. The method according to any one of claims 1 to 3, wherein the first set of frameworks comprises a first framework specifying a first torsion angle in the main chain of the amino acid residue, and wherein the first set of frameworks further comprises a second framework specifying a second torsion angle in the main chain of the amino acid residue.
5. The method according to claim 4, wherein the first torsion angle is relative to the alpha carbon (C α ) atom and a carbon (C) atom, and wherein the second torsion angle is associated with a first rotatable bond between the α carbon (C) atom in the backbone of the amino acid residue α ) atom and the second rotatable bond between the nitrogen (N) atom.
6. A method according to claim 5, wherein the first set of frameworks further includes a third framework, which specifies a third torsion angle present in the main chain of the amino acid residue, and wherein the third torsion angle is associated with a third rotatable bond between the carbon (C) atom and the nitrogen (N) atom in the main chain of the amino acid residue.
7. The method according to any one of claims 1 to 6, further comprising: include: determining one or more coordinates of a plurality of backbone atoms in the protein molecule based at least on the plurality of frameworks associated with each amino acid residue included in the modified representation of the protein molecule; as well as Based on the one or more coordinates of the plurality of main chain atoms in the protein molecule, one or more coordinates of a plurality of side chain atoms in the protein molecule are determined.
8. The method according to any one of claims 1 to 7, wherein the design computational model comprises a machine learning model, which is trained to generate the first three-dimensional structure of the protein molecule by at least denoising the initial three-dimensional structure of the protein molecule.
9. The method of claim 8, wherein the machine learning model denoises the initial three-dimensional structure of the protein molecule by performing a series of updates on at least the representation of the protein molecule.
10. The method of claim 9, wherein the machine learning model is trained to reduce a loss function and / or an energy function associated with each successive update to the initial three-dimensional structure of the protein molecule.
11. The method according to any one of claims 1 to 10, wherein the machine learning model is a diffusion model, which removes a portion of the noise present in the initial three-dimensional structure of the protein molecule at each time step in multiple consecutive time steps.
12. The method of claim 11, wherein the diffusion model performs a first update on the representation of the protein molecule to remove a first amount of noise present in the initial three-dimensional structure of the protein molecule, and wherein the diffusion model further performs a second update on the representation of the protein molecule to remove a second amount of noise present in the initial three-dimensional structure of the protein molecule.
13. The method of claim 12, wherein the diffusion model further adds a third amount of noise before performing the second update to remove the second amount of noise, and adds a fourth amount of noise after performing the second update to remove the second amount of noise, and wherein the third amount of noise and the fourth amount of noise are determined based on a noise schedule that defines a distribution of noise levels added across the plurality of consecutive time steps.
14. The method of claim 13, wherein the distribution of the noise levels corresponds to the degrees of freedom present in the representation of the protein molecule, the degrees of freedom being used for the computational model to modify the initial three-dimensional structure of the protein molecule.
15. A method according to any one of claims 11 to 14, wherein each update performed by the diffusion model generates an output which is equivariant to a Special Euclidean Group SE(3) transformation.
16. The method according to any one of claims 1 to 15, wherein the modification of the representation of the protein molecule comprises updating the first set of frameworks to change the geometric state of the main chain of one or more amino acid residues in the protein molecule.
17. The method according to any one of claims 1 to 16, wherein the modification of the representation of the protein molecule comprises updating the second set of frameworks to change the one or more torsion angles in the side chains of one or more amino acid residues in the protein molecule.
18. The method of any one of claims 1 to 17, wherein the first three-dimensional structure of the protein molecule is associated with one or more desired properties.
19. The method according to any one of claims 1 to 18, wherein the first three-dimensional structure of the protein molecule is configured for one or more downstream tasks.
20. The method according to any one of claims 1 to 19, further comprising: include: determining, based at least on the first three-dimensional structure of the protein molecule, that the first sequence of amino acid residues exhibits a desired three-dimensional structure and / or desired properties; as well as In response to determining that the first sequence of amino acid residues exhibits the desired three-dimensional structure and / or the desired property, based at least on the first sequence of amino acid residues, A second sequence of amino acid residues is generated for a different protein molecule.
21. The method according to any one of claims 1 to 20, wherein the representation of the protein molecule further comprises a logical vector for each position in the sequence of amino acid residues forming the protein molecule, the logical vector indicating the identity of the amino acid residue occupying the position by at least enumerating a probability distribution of a set of possible amino acid residues occupying the position.
22. The method of claim 21, wherein the designed computational model further generates the first three-dimensional structure of the protein molecule by modifying the identification of at least one amino acid residue while modifying the first set of frameworks and / or the second set of frameworks associated with at least one amino acid residue in the first sequence of residues.
23. The method according to any one of claims 1 to 22, wherein the initial three-dimensional structure of the protein molecule includes noise in the identification of each amino acid residue and / or the spatial arrangement of multiple atoms forming each amino acid, and wherein the noise is removed by modifying the representation of the protein molecule by the designed computational model.
24. The method of any one of claims 1 to 23, wherein the representation of the protein molecule is further generated to include a plurality of polymer chains, wherein each polymer chain includes one or more amino acid residues from the first sequence of amino acid residues, and wherein the representation of the protein molecule is modified by modifying the positions of the one or more amino acids in each polymer chain as a group by a protein design computational model.
25. A system, wherein include: at least one data processor; as well as At least one memory storing instructions which, when executed by said at least one data processor, result in operations including the method according to any one of claims 1 to 24.
26. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, result in operations including the method of any one of claims 1 to 24.