Use of generative artificial intelligence for protein engineering
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- AETHER BIOMACHINES INC
- Filing Date
- 2024-06-20
- Publication Date
- 2026-05-20
AI Technical Summary
Current enzyme design methods are time-consuming, expensive, and limited in predicting novel enzyme structures, often requiring hundreds of human expert hours and relying on trial-and-error experimentation, which hampers the development of new enzymes for biocatalysis, drug discovery, and bioenergy applications.
A computational platform utilizing Transformer neural networks and diffusion models to predict active sites and design enzyme structures, enabling the generation of novel enzymes with precise and efficient enzyme design capabilities, leveraging pre-trained knowledge from scientific literature and experimental datasets to predict active sites and generate three-dimensional configurations of enzymes.
This approach significantly reduces the time and computational costs associated with enzyme design, allowing for rapid prediction and validation of enzyme structures, enhancing the efficiency of enzyme design processes and enabling the exploration of a vast protein design space, thereby accelerating the design-build-test cycles for biocatalysis, drug discovery, and bioenergy applications.
Smart Images

Figure US2024034852_16012025_PF_FP_ABST
Abstract
Description
USE OF GENERATIVE ARTIFICIAL INTELLIGENCE FOR PROTEIN ENGINEERINGFIELD OF INVENTION
[0001] The specification generally relates to the field of biotechnology and artificial intelligence. More particularly, the invention relates to the development of a computational platform for the design of novel enzymes using machine learning models. Enzyme design is a critical aspect of biotechnology with a multitude of applications, including biocatalysis, drag discovery, and bioenergy.
[0002] Machine learning models generally predict an output based on a given input. Some machine learning models employ single layered architectures and rely on statistical analysis to classify or regress a particular dataset. These include models such as linear regression, decision trees, etc. Some other machine learning models have more complex, e.g., multi-layered, architectures. These include large language models (ELM) and diffusion models.BACKGROUND
[0003] Proteins can be useful for catalyzing chemical reactions and as precise binding reagents (e.g., for diagnostics, therapeutics, detection, or separations). Enzymatic catalysis can replace traditional chemical synthesis which often relies on the use of petrochemical-based solvents, hazardous reagents, and metal-based chemical catalysts to bring about chemical transformations. In contrast, enzymes are non-hazardous and nontoxic, biodegradable, operate in aqueous and non-aqueous solvents, and are completely renewable, as they are produced safely and inexpensively through biological processes, hr addition, enzymes can be designed (e.g., as an antibody) to bind to targets with high selectivity, which induces binding of the protein to a desired ligand or epitope while avoiding binding to other closely related ligands or epitopes. This can result in improved precision of detection, treatments with reduced side effects, or selective enrichment of a product. This, in principle, allows the synthesis of products inaccessible by traditional chemical synthesis, as well as the use of alternative raw materials that could dramatically lower infrastructure and operating costs.
[0004] While previous advances in engineering of protein activity have brought improvements in many proteins compared to wild-type (i.e., a variant of the protein found in nature), thesetypically take several months to validate, achieve only modest improvements, and have unpredictable results. Furthermore, previous approaches typically require an initial protein having at least some activity on the desired substrate, and are therefore limited to improving a protein rather than engineering novel activity. There is a need for fundamentally new systems and methods for high-throughput protein design.SUMMARY
[0005] The systems and methods described in this specification satisfy the need for new and improved systems and methods for engineering proteins. This disclosure pertains to an Al-driven enzyme design platform, utilizing the integrative predictive capacities of Transformer neural networks and the conditional generative capabilities of diffusion models to create enzymes that can facilitate desired reactions.
[0006] The system employs Transformer neural networks to generate conditioning inputs that are used as guidance for the generation processes executed using diffusion models. The process involves querying the Transformer neural network to predict an active site configuration that binds to a specific ligand. Subsequently, the diffusion model uses this predicted active site configuration to design the overall structure of the enzyme around the active sites, effectively incorporating the proposed active site into its structure. The symbiotic functioning of these two models allows for precise and efficient enzyme design.
[0007] According to an aspect, there is provided a computer-implemented method comprising: receiving input data that characterizes an enzyme-catalyzed reaction; generating, using a Transformer neural network, predicted locations of active sites of an enzyme based on the input data; generating, using at least a diffusion model and based on the predicted locations of active sites of the enzyme, a predicted structure of the enzyme, the predicted structure of the enzyme defining a three-dimensional configuration of atoms included in the enzyme.
[0008] The input data may comprise data that characterizes a substrate for binding to the enzyme. The Transformer neural network may be configured as a Graph Transformer neural network that operates on graph data defining the enzyme-catalyzed reaction. The Transformer neural netw ork may be configured as a large language model neural network that auto- regressively generates the predicted locations of the active sites of the enzyme. Generating the predicted structure of the enzyme may comprise: generating, using the diffusion model, predictedthree-dimensional locations of the atoms included in the enzyme; generating, using a sequence prediction neur al network, a sequence of amino acids based on the predicted three-dimensional locations of the atoms included in the enzyme; and generating, using a folding neural network or protein structure prediction model, the predicted structure of the enzyme based on the sequence of amino acids. The method may further comprise performing one or more actions with respect to an enzyme having the predicted structure. The one or more actions may comprise: determining whether the enzyme having the predicted structure is valid based on the three-dimensional configuration of atoms defined by the predicted structure. The one or more actions may comprise: performing laboratory synthesis and validation of the enzyme having the predicted structure.
[0009] According to another aspect, there is provided one or more computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the above method aspects.
[0010] According to a further aspect, there is provided a system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of die above method aspects.
[0011] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. The enzyme design platform described in this specification can predict structures of an enzyme based on a target reaction for which the enzyme can aid. The enzyme that has a predicted structure can bind to a substrate to speed up the reaction, allow it to occur at lower energy levels or even help in detection of that substrate (detection of harmful chemicals, for example). More specifically, the enzyme design platform can do this by leveraging (i) the knowledge learned by a Graph Transformer neural network from its large-scale pre-training on scientific literature, protein databases, experimental datasets, and other training data to predict potential active sites and (ii) the generative capability of a diffusion model that is conditioned on the predicted potential active sites. By understanding the intricacies involved in atomic level interactions in the catalytic binding site and its overall impact in the structure of the protein, the enzyme design platform can more efficiently and accurately design novel enzymes compared to, for example, with traditionalenzyme design approaches that attempt to exhaustively search through a very' large enzyme design space.
[0012] Existing systems for enzyme design involve a time-consuming and expensive process of trial-and-error experimentation, and can require significant time and computational costs, sometimes taking hundreds or thousands of human expert hours for the design of enzymes for a single reaction, hi contrast, the enzyme design platform described in this specification can quickly predict structures of novel enzymes with high accuracy. For example, in some implementations, generating a prediction for the structures of an enzyme can take no more than a few minutes or even a few seconds. Additionally, the enzyme design platform is able to incorporate the chemical rules behind the reaction and the required combination of amino acids in the protein to perfomi the intended functions. This shortens the time required to search the vast space of proteins with different combinations of amino acids. Using the enzyme design platform can thus significantly improve the efficiency' of the process of designing new enzymes, allowing human experts to test more enzyme designs for a wider range of applications, including biocatalysis, drug discovery', and bioenergy, and iterate the design-build-test cycles much more quickly.
[0013] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE FIGURES
[0014] FIG. 1 is a diagram of an example enzyme design platform.
[0015] FIG. 2 is a flow diagram of an example process for generating a predicted structure of an enzyme.
[0016] FIG. 3 is an example illustration of perfluoroalkyl and polyfluoroalkyl substances (PFAS).
[0017] FIG. 4 is an example illustration of a part of output data that can be obtained by using a Graph Transformer neural network.
[0018] FIG. 5 is an example illustration of an atomic-level representation of an enzyme.
[0019] FIG. 6 is an example illustration of a sequence of amino acids in a FASTA format.
[0020] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0021] FIG. 1 is a diagram of an example enzyme design platform 100. The enzyme design platform 100 is an example of a platform implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0022] The enzyme design platform 100 obtains input data 102 that characterizes one or more desired properties of an enzyme and uses the input data 102 to generate output data that defines a predicted structure 132 of the enzyme that has the one or more desired properties characterized by the input data 102.
[0023] The input data 102 can be obtained in any of a var iety of ways. For example, the enzyme design platform 100 can receive the input data as a user input from a user of the platform over a data communication network, e.g., using an application programming interface (API) or a graphical user interface (GUI) made available by the system or a web browser through which a user can interact with the platform. As another example, the enzyme design platform 100 can receive a user input from a user specifying which data that is already maintained by the platform or another computer system that is accessible by the platform should be used as the input data that characterizes the one or more desired properties of the enzyme.
[0024] The input data 102 can characterize the desired properties of the enzyme in a variety of ways. In some implementations, the input data 102 can characterize a target (chemical) reaction for which the enzyme can aid in. Such a target reaction may involve the enzyme, a substrate, and one or more products formed after binding the enzyme to the substrate. The target reaction may be part of a pathway or process of interest. As part of the target reaction, the enzyme acts upon the substrate and decreases the activation energy necessary for the chemical reaction to occur by stabilizing the transition state. This stabilization speeds up reaction rates and makes them happen at physiologically significant rates. To act upon the substrate, the enzyme binds the substrate, e.g., using a lock and key mechanism, at key locations in the structure of the enzyme called active sites. Once the substrate is in the active sites of the enzyme, the target reaction may take place.
[0025] The input data 102 can thus characterize aspects of the substrate involved in the target reaction. For example, the input data 102 can define the type of the substrate, e.g., can identify which one of various known types of substrates should be used as the substrate to be involved in the target reaction. As another example, the input data 102 can define the structure of the substrate, e.g., data that defines the three-dimensional (3-D) spatial locations of the atoms in the substrate. As yet another example, the input data 102 can define the concentrations of the substrate, the rate of change of the concentration of the substrate, and the like during the chemical reaction.
[0026] Additionally or alternatively, the input data 102 can characterize aspects of the product(s) that can be formed as a result of the target reaction, i.e., after binding the enzyme to the substrate. For example, the input data 102 can define the type of the product. As another example, the input data 102 can define the structure of the product, e.g., data that defines the three-dimensional (3-D) spatial locations of the atoms in the product. As yet another example, the input data 102 can define the rate of reaction that is needed during the chemical reaction in order to form the product.
[0027] The enzyme design platform 100 includes a Transformer neural network 110, a diffusion model 120, and a post-processing engine 130.
[0028] The Transformer neural network 110 processes input that includes (i) the input data 102, (ii) data derived (generated) from the input data 102, or both (i) and (ii) to generate a neural network output that defines the predicted locations 112 of active sites of the enzyme.
[0029] Being referred to as a “Transformer” neural network means it has a Transformer-based architecture that includes a plurality of Transformer blocks that each apply an attention operation on a Transformer block input to generate a Transformer block output. Each “block” includes a group of one or more neural network layers, e.g., one or more attention layers, e.g., one or more self-attention layers or one or more cross-attention layers, one or more feed-forward layers, or both and possibly other layers, including normalization layers and residual connection layers.
[0030] A protein, e.g., an enzyme or a substrate, consists of a sequence of amino acids (residues). An amino acid is an organic compound which includes an amino functional group and a carboxyl functional group, as well as a side-chain (i.e., group of atoms ) that is specific to the amino acid. These amino acids are linked to one another with a peptide bond to form a protein.When in a sequence linked by peptide bonds, the amino acids may be referred to as amino acid residues.
[0031] In implementations the Transformer neural network 110 can be configured as a Graph Transformer neural network to perform its described functions.
[0032] For example, the Graph Transformer neural network can take as input a substrate graph to generate as output the predicted locations of active sites of the enzyme that would bind to this substrate. The substrate graph is a data structure that includes nodes and edges connecting the nodes. Collectively, the nodes and edges can characterize various properties, such as positions of the atoms and bond interactions within the substrate involved in the target reaction.
[0033] In some implementations, the nodes in the substrate graph can represent amino acid residues, if the substrate is a protein or peptide, and the edges in the substrate graph can represent the spatial relationships between the amino acid residues based on their relative distances between each other.
[0034] In some implementations, the outputs of the Graph Transformer neural network can be in the form of a decoded graph structure, which has nodes and edges representing the amino acids in an active site bound to the substrate graph structure. Initially, each node in the substrate graph is represented by an embedding vector. These embeddings capture the features of the nodes and are used as input to the Graph Transformer neural network.
[0035] Graph Transformer neural network also employs the self-attention mechanism to capture the relationships between nodes in the graph. By applying the self-attention mechanism, each node can attend to other nodes in the graph, and the attention scores represent a measure of attention (or importance) each node assigns to other nodes. Graph Transformer neural network also employs a message passing mechanism, where each node aggregates information from its neighbors based on their importance scores as determined by applying the attention mechanism and updates the embedding of the node accordingly.
[0036] In some implementations, the Graph Transformer neural network typically includes multiple layers, each performing a self-attention mechanism followed by a message passing mechanism. This allow's the model to capture increasingly complex relationships between nodes across multiple hops in the graph.
[0037] In some implementations, the Graph Transformer neural network can also include decoder layers that can process graphs resulting from the message passing updates to generateoutput based on the processed graph representations. In this case, the predicted locations of active sites of the enzyme are generated as a graph structure with locations of each atom in the active site. In some implementations, the decoder layers can receive combined node embeddings for the nodes of the graphs resulting from the message passing updates (e.g., as generated by pooling, averaging, concatenating, etc., the node embeddings for the nodes). For example, the Graph Transformer neural network can generate combined node embeddings based on the graphs resulting from the message passing updates and processing the combined node embeddings using the decoder layers to generate another graph with the predicted locations 112 of the active sites of the enzyme.
[0038] Some example implementations of Graph Transformer neural networks are described in more detail in Yoon, M., Wu, Y., Palowitch, J., Perozzi, B. and Salakhutdinov, R., 2022. Graph generative model for benchmarking graph neural networks. arXiv preprint arXiv;2207.04396 and Muller, L., Galkin, M., Morris, C. and Rampasek, L., 2023. Attending to graph transformers. arXiv preprint arXiv.2302.04181.
[0039] As another example, the large language model can receive an input sequence made up of tokens selected from a vocabulary and auto-regressively generates an output sequence made up of tokens from the vocabulary. The large language model can have any of a variety of Transformer-based neural network architectures, e.g., encoder-only Transformer architectures, encoder-decoder Transformer architectures, decoder-only Transformer architectures, other attention-based architectures, and so on.
[0040] Example implementations of such a large language model are described in more detail in Rohan Anil, et al. “Palm 2 technical report.” arXiv preprint arXiv;2305.10403; and Hugo Touvron, et al. “Llama; Open and efficient foundation language models.” arXiv preprint arXiv;2302.13971, but others may also be used.
[0041] More specifically, the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token.
[0042] The vocabulary of tokens can include text tokens that can represent textual data. For example, the vocabulary of tokens can include any of; characters, sub-words, words, punctuationmarks, sign tokens (e.g., the #, $, and other signs), mathematical symbols, numbers, and so on in a language, e.g., a natural language, a computer programming language, or a markup language.
[0043] The input sequence can include any tokens selected from the vocabulary that can characterize various properties of the substrate involved in the target reaction, e.g., the type of the substrate, the structure of the substrate, and so on. The output sequence can include any tokens selected from the vocabulary that can define the predicted locations 112 of the active sites of the enzyme.
[0044] The diffusion model 120 is a neural network that processes (i) the predicted locations 112 of active sites of the enzyme (that has been generated by the Transformer neural network 110), (ii) data derived from the predicted locations 112 of active sites of the enzyme, or both (i) and (ii) to generate an atomic-level of the enzyme. The atomic-level representation of the enzyme is a representational structure that defines the predicted three-dimensional (3-D) locations 122 of the atoms included in the overall structure of the enzyme.
[0045] In particular, the diffusion model 120 is configured to generate the atomic-level representation of the enzyme across multiple updating iterations by performing a reverse diffusion process, conditioned on the predicted locations 112 of active sites of the enzyme.
[0046] The diffusion model 120 can have any appropriate conditional diffusion neural network architecture that can be used to use information about the active sites of the enzyme to generate a diffusion model output at each of multiple updating iterations. The 3-D atomic locations of the enzyme can be inferred from the diffusion model output generated at the last updating iteration.
[0047] For example, the diffusion model 120 can be configured to, at any given updating iteration, process a diffusion model input for the given updating iteration that includes a current intermediate representation of the atomic-level representation of the enzyme (as of the given updating iteration) to generate a diffusion model output for the given updating iteration from which an updated intermediate representation of the atomic-level representation of the enzyme can be inferred.
[0048] In this way, the enzyme design platform 100 utilizes information about a portion of the structure of the enzyme (active sites) and the diffusion model 120 to predict possible (stable) structures of the enzyme conditioned on these atomic positions on the active site residues.
[0049] Such a diffusion model can have been previously trained on a large corpus of proteins and their 3-D atomic positions. Example implementations of such diffusion models are describedin more detail in Watson, J.L., Juergens, D., Bennett, N.R., Trippe, B.L., Yim, J., Eisenach, H.E., Ahem, W., Borst, A.J., Ragotte, R.J., Milles, L.F. and Wicky, B.L, 2023. De novo design of protein structure and function with RFdiffusion. Nature, 620(7976), pp.1089-1100.
[0050] In some implementations, the diffusion model 120 can be configured through training as a denoising diffusion model, i.e., it is trained to reverse an incremental noising process applied to data specifying protein structures. The diffusion model 120 can have been trained to learn the distribution of this denoising process.
[0051] hi some implementations, self-conditioning is used during training, where the diffusion model 120 conditions its current predictions on its own previous outputs. This can improve coherence and packing of the generated structures.
[0052] In some implementations, the diffusion model input at any given updating iteration includes (i) an intermediate representation of the atomic-level representation of the enzyme (as of the updating iteration), (ii) the predicted locations 112 of active sites of the enzyme, (iii) data derived from the predicted locations 112 of active sites of the enzyme, or some combination of (i)-(iii). For example, the enzyme design platform 100 can provide the predicted atomic locations 112 of active sites of the enzyme as input and conditions the diffusion model on the input to generate the predicted 3-D locations 122 of the atoms included in the enzyme based on the input.
[0053] During the process, if the given updating iteration of the denoising process is the first updating iteration in the reverse diffusion process, the current intermediate representation can be an initial intennediate representation, e.g., an initial intermediate representation that is generated based on sampling the value for each variable included in the atomic-level representation of the enzyme from a predetermined noise distribution, e.g., a Gaussian distribution. For any subsequent updating iteration, the current intermediate representation is the updated intermediate representation that has been generated in the immediately preceding updating iteration.
[0054] The diffusion model output can either define the updated intermediate representation directly, e.g., where the diffusion model output includes a prediction of the updated intermediate representation, or indirectly, e.g., where the diffusion model output includes a prediction of the noise component in the current intermediate representation. For example, the diffusion model output can define a prediction of a noise that needs to be added to the atomic-level representation of the enzyme being generated by the enzyme design platform 100, to generate the currentintermediate representation, and the updated intermediate representation is generated by using the predicted noise to de-noise the current intermediate representation, i.e., by removing the predicted noise from the current intermediate representation.
[0055] For example, the diffusion model 120 can have a Transformer neural network architecture that processes the diffusion model input through a sequence of attention layers to generate the diffusion model output, where some or all of the attention layers that apply an attention operation over a layer input that includes (i) the predicted locations 112 of active sites of the enzyme, (ii) data derived from the predicted locations 112 of active sites of the enzyme, or both (i) and (ii). Example implementations of such a diffusion model 120 are described in more detail in Jonathan Ho, et al. “Denoising diffusion probabilistic models.” Advances in neural information processing systems 33 (2020): 6840-6851, but others may also be used.
[0056] As another example, the diffusion model 120 can have an equivariant neural network architecture that relates 3-D symmetries and captures interatomic interactions effectively. Example implementations of such a diffusion model are described in Igashov, I., Stark, H., Vignac, C., Schneuing, A., Satorras, V.G., Frossard, P., Welling, M., Bronstein, M. and Correia, B., 2024. Equivariant 3D-conditional diffusion model for molecular linker design. Nature Machine Intelligence, pp.1-11, and Hoogeboom, E., Satorras, V.G., Vignac, C. and Welling, M., 2022, June. Equivariant diffusion for molecule generation in 3d. In International conference on machine learning (pp. 8867-8887). PMLR.
[0057] As another example, the diffusion model 120 can have a combination of the Transformer architecture and the equivariant neural network architecture, e.g., in a cascaded architecture, where multiple diffusion models can operate at different atomic resolutions. An example implementation of a cascaded diffusion model is described in Ho, J., Saharia, C., Chan, W., Fleet, D.J., Norouzi, M. and Salimans, T., 2022. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 23(47), pp.1-33.
[0058] The post-processing engine 130 perfomis further processing on the atomic-level representation of the enzyme to generate a predicted structure 132 of the enzyme. In particular, the atomic-level representation of the enzyme needs further processing because it only defines the predicted three-dimensional (3-D) locations 122 of the atoms included in the enzyme, and lacks the information about which amino acid in a plurality of amino acids that each of the atoms may be specific to, i.e., which amino acid each of the atoms may belong to.
[0059] The post-processing engine 130 can include one or more components that can map the atomic-level representation of the enzyme to the predicted structure 132 of the enzyme. In some implementations, the predicted structure 132 of the enzyme can be defined by a set of structure parameters that collectively define a predicted three-dimensional (3-D) configuration of atoms in the amino acids in the structure of the enzyme, e.g., after the enzyme undergoes protein folding. Protein folding refers to a physical process by which one or more sequences of amino acids fold into a three-dimensional (3-D) configuration.
[0060] hr one example, the set of structure parameters can include: (i) location parameters, (ii) rotation parameters, or both (i) and (ii) for each amino acid in the enzyme.
[0061] The location parameters for an amino acid can specify a predicted 3-D spatial location of a specified atom in the amino acid in the structure of the enzyme. The specified atom can be the alpha carbon atom in the amino acid, i.e., the carbon atom in the amino acid to which the amino functional group, the carboxyl functional group, and the side-chain are bonded. The location parameters for an amino acid can be represented in any appropriate coordinate system, e.g., a three-dimensional [x, y, z] Cartesian coordinate system.
[0062] The rotation parameters for an amino acid can specify the predicted “orientation” of the amino acid in the structure of the enzyme. More specifically, the rotation parameters can specify a 3-D spatial rotation operation that, if applied to the coordinate system of the location parameters, causes the three main functional groups in the amino acid to assume fixed positions relative to the rotated coordinate system. The three main chain atoms in the amino acid refer to the linked series of nitrogen, alpha carbon, and carbonyl carbon atoms in the amino acid (or, e.g., the alpha carbon), nitrogen, and oxygen atoms in the amino acid). The rotation parameters for an amino acid may be represented, e.g., as an orthonormal 3x3 matrix with determinant equal to 1.
[0063] Generally, the location and rotation parameters for an amino acid define an egocentric reference frame for the amino acid. In this egocentric reference frame, the side-chain for each amino acid may start at the origin, and the first bond along the side-chain (i.e., the alpha carbonbeta carbon bond) may be along a defined direction.
[0064] In another example, the structure parameters defining the predicted structure 132 of the enzyme can include a “distance map” that characterizes a respective estimated distance (e.g., measured in angstroms) between each pair of amino acids in the enzyme. A distance map cancharacterize the estimated distance between a pair of amino acids, e.g., by a probability distribution over a set of possible distances between the pair of amino acids.
[0065] In yet another example, the structure parameters defining the predicted structure 132 of the enzyme can define a three-dimensional (3-D) spatial location of each atom in each amino acid in the structure of the enzyme.
[0066] In some implementations, the components of the post-processing engine 130 include a sequence prediction neural network 140 and a folding neural network 150. The sequence prediction neural network 140 processes an input that includes (i) the atomic-level representation of the enzyme, (ii) data derived from the atomic-level representation of the enzyme, or both (i) and (ii) to generate a neural network output that defines a sequence of amino acids.
[0067] The sequence of amino acids can be represented in any appropriate format. For example, the sequence of amino acids may be represented in a FASTA format. FASTA format is a textbased format in which amino acids are represented using single-letter codes. When represented in the FASTA format, a sequence of amino acids may begin with a greater-than character (“>”) followed by a description of the sequence. The lines immediately following the description line are the sequence representation, with one letter per amino acid.
[0068] As another example, the sequence of amino acids may be represented as a sequence of one-hot vectors. In this example, each one-hot vector represents a corresponding amino acid in the amino acid sequence. A one-hot vector has a different component for each different amino acid (e.g., of a predetermined number of amino acids). A one-hot vector representing a particular amino acid has value one (or some other predetermined value) in the component corresponding to the particular amino acid and value zero (or some other predetermined value) in the other components.
[0069] In some implementations, the sequence prediction neural network 140 can be configured as a message passing neural network (MPNN). Message passing neural network is a type of neural network that operates on graph data that defines an enzyme graph to generate as output the sequence of amino acids. In these implementations, each amino acid in the protein sequence is an organic compound rvhich includes an amino functional group and a carboxyl functional group, as well as a side-chain, i.e., a subset of the atoms included in the atomic-level representation of the enzyme that has been generated by the diffusion model 120, which is specific to the amino acid.
[0070] The enzyme graph operated by the sequence prediction neural network 140 is a data structure that includes nodes and edges connecting the nodes. For example, the nodes in the enzyme graph can represent the atoms included in the enzyme and the edges in the enzyme graph can represent the relative distances between the atoms included in the enzyme. Such an enzyme graph can be generated by the enzyme design platform 100 based on the atomic-level representation of the enzyme, which defines the predicted 3-D locations 122 of the atoms included in the enzyme, that has been generated by using the diffusion model 120.
[0071] Typically, MPNNs have a two-phase process: a message passing phase and a readout phase. In the message passing phase, the MPNN updates the hidden states for each node in the graph based on messages passed from neighboring nodes. This is done using learned message functions and vertex update functions. The message functions and vertex update functions are differentiable and learned during training, allowing the MPNN to flexibly capture long-range dependencies in the graph. The readout phase then computes a feature vector for the entire graph using a learned readout function that is invariant to pennutations of the node states. This two- phase process makes MPNNs a powerful and versatile class of neural networks that can effectively learn representations of the enzyme graph through iterative message passing and readout operations.
[0072] Example implementations of such a message-passing neural network 140 to predict the protein sequence from the atomic locations of the predicted protein from 122 are described in more detail in Justas Dauparas, et al. “Robust deep learning-based protein sequence design using proteinMPNN.” Science, 378(6615):49-56, but others may also be used.
[0073] The folding neural network 150 is a neural network that processes an input that includes (i) the sequence of amino acids, (ii) data derived from the sequence of amino acids, or both (i) and (ii) to generate an output that defines the predicted structure 132 of the enzyme. The predicted structure can define an estimate of a three-dimensional (3-D) configuration of the atoms in the amino acid sequence of the enzyme (that has been generated by the sequence prediction neural network 140) after the enzyme undergoes protein folding. As mentioned above, in some implementations, the predicted structure 132 of the enzyme can be defined by a set of structure parameters that define a 3-D spatial location and optionally also a 3-D spatial rotation of each amino acid in the enzyme.
[0074] Example implementations of such a folding neural network 150 are described in more detail in John Jumper, et al. “Highly accurate protein structure prediction with AlphaFold.” Nature 596, 583-589 (2021 ) and Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S. and Rives, A., 2022. Language models of protein sequences at the scale of evolution enable accurate structure prediction. BioRxiv, 2022, p.500902, but others may also be used.
[0075] Prior to using the Transformer neural network 110, the diffusion model 120, the sequence prediction neural network 140, and the folding neural network 150 in tandem with each other (as discussed above) to generate the predicted structure 132 of an enzyme from the input data 102, the enzyme design platform 100 or a separate training system trains these neural networks and possibly other trainable components of the platform on the training data so that these neural networks can perform their respective functions well.
[0076] In one example, the Graph Transformer neural network can be trained on known public repositories, e.g., repositories of scientific literature, protein databases, and experimental datasets, that include data defining active site and substrate positions in a 3-D space, based on optimizing a self-supervised objective function. The positions of the active sites or the substrates can be defined, e.g., by respective structure parameters that collectively define a predicted three- dimensional (3-D) configuration of atoms in the amino acids in the structures of the active sites or the substrates, as described above.
[0077] In one similar example, the large language model can be trained on known public repositories that include active site and substrate positions in a 3-D space, based on optimizing a self-supervised objective function. For example, the enzyme design platform 100, or a separate training system, pre-trains the large language model using a large dataset of text, e.g., text that is publicly available from the Internet or another text corpus on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data, and then fme-tunes the pre-trained large language model using the known public repositories that include active site, substrate positions and information about the product formed due to the reaction between the enzyme and substrate, based on optimizing a maximum-likelihood objective function. In this way, the large language model is trained to learn the language and concepts of enzyme biology, biochemistry, structural biology, and so on.
[0078] In another example, the diffusion model 120 can be trained on known public repositories that include known protein structures using a denoising score-matching objective. In one example, such a public repository can be the Protein Data Bank (PDB) repository mentioned above.
[0079] FIG. 2 is a flow chart of an example process 200 for generating a predicted structure of an enzyme. Operations of the process 200 can be performed, for example, by the enzyme design platform 100 of FIG. 1, or another system implemented as computer programs on one or more computers in one or more locations. The operations of the process 200 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 200.
[0080] Input data that characterizes one or more aspects of a target enzyme-catalyzed reaction is obtained (step 202). For example, the enzyme design platform 100 can receive the input data as a user input from a user of the platform over a data communication network, e.g., using an application programming interface (API) or a graphical user interface (GUI) made available by the system or a web browser through which a user can interact with the platform. As another example, the enzyme design platform 100 can receive a user input from a user specifying which data that is already maintained by the platform or another computer system that is accessible by the platform should be used as the data that characterizes the one or more aspects of a target enzyme-catalyzed reaction.
[0081] The target enzyme-catalyzed reaction may involve an enzyme, a substrate, and one or more products formed after binding the enzyme to the substrate. The input data can characterize aspects of the substrate involved in the target enzyme-catalyzed reaction, i.e., can characterize the substrate for binding to the enzyme. For example, the input data can define the type of the substrate. As another example, the input data can define the structure of the substrate. As yet another example, the input data can define the concentrations of the substrate, the rate of change of the concentration of the substrate, and the like during the chemical reaction.
[0082] Additionally or alternatively, the input data can characterize aspects of the product(s) that can be formed as a result of the target enzyme-catalyzed reaction, i.e., after binding the enzyme to the substrate. For example, the input data can define the type of the product. As another example, the input data can define the structure of the product. As yet another example, theinput data can define the rate of reaction that is needed during the chemical reaction in order to form the product.
[0083] As a particular example for illustration, the target enzyme-catalyzed reaction can be a chemical reaction that involves breaking down perfluoroalkyl and polyfluoroalkyl substances (PF AS) with the aid of an enzyme, and the input data can characterize the substance (namely, the PFAS) involved in the chemical reaction. In this particular example, the input data can be generated based on the known 3-D atom locations in PFAS. The input data can have a graph structure, for example.
[0084] FIG. 3 is an example illustration 300 of perfluoroalkyl and polyfluoroalkyl substances (PFAS). As illustrated, PFAS include molecules that are made up of a chain of linked carbon (C) and fluorine (F) atoms. Because of the strength of a carbon-fluorine bond, PFAS can degrade very slowly, or not at all, in the environment. PFAS can be used in nonstick coatings on cookware, stain-resistant clothes and carpets, or firefighting foam, to make it more effective. Further, PFAS can be used in industries such as aerospace, automotive, construction, electronics, and military equipment. Currently, more than 9,000 PFAS have been identified. Due to the fact that PFAS can persist in the environment for an unknown amount of time and gradually accumulate and remain in the human body, there has been an increase in concerns regarding the public health impact of PFAS.
[0085] PFAS include, but are not limited to, perfluorooctanoic acid (PFOA), perfluorooctyl sulfonate (PFOS), hexafluoropropylene oxide (HFPO) dimer acid, and their ammonium, sodium, and potassium salts, or any combination thereof. Further examples of PFAS include N-ethyl perfluorooctanesulfonamidoacetic acid, N-methyl perfluorooctanesulfonamidoacetic acid, perfluorobutanesulfonic acid, perfluorodecanoic acid, perfluorododecanoic acid, perfluoroheptanoic acid, perfluorohexanesulfonic acid, perfluorohexanoic acid, and perfluorononanoic acid. In some examples, PFAS include perfluorooctanoic acid (PFOA), perfluorooctane sulfonate (PFOS), or any combination thereof.
[0086] A Transformer neural network processes a neural network input that includes (i) the input data, (ii) data derived from the input data, or both (i) and (ii) to generate a neural network output that defines predicted locations of active sites of the enzyme (step 204). Active sites refer to specific key locations of an enzyme where a substrate (e.g., PFAS) binds and catalysis takes place, as part of the target enzyme-catalyzed reaction. In some implementations, the predictedlocations of active sites of the enzyme can include predicted 3-D spatial locations of the active sites in the structure of the enzyme.
[0087] For example, when the Transformer neural network is configured as a Graph Transformer neural network, the enzyme design platform 100 can generate, from the input data, graph data that defines a substrate graph, and process die substrate graph using the Graph Transformer neural network to generate as output the predicted locations of active sites of the enzyme.
[0088] In some implementations, the output of the Graph Transformer neural network can be converted to a readable format with information about the 3-D atomic positions.
[0089] FIG. 4 is an example illustration 400 of a part of output data that can be obtained from the further processing of the output of the Graph Transformer neural network. In the example of FIG. 4, the output data defines the predicted amino acids in the active sites and the 3-D atomic positions of the active sites, in the Protein Data Bank (PDB) format. PDB format is a textual format describing the three-dimensional structures of molecules, such as proteins and nucleic acids, maintained in the Worldwide Protein Data Bank (www.pdb.org). The PDB format accordingly provides for description and annotation of protein and nucleic acid structures including atomic coordinates, secondary structure assignments, as well as atomic connectivity.
[0090] The output data in PDB format can define the structure of the predicted active site residues. Specifically, the output data in the PDB format can describe the coordinates of the atoms that are part of a potential enzyme. For example, in FIG. 4, the first ATOM line describes the alpha-N atom is part of the aspartic acid residue in chain L, and its residue (amino acid) ID is “1”; the first three floating point numbers are its x, y and z coordinates and are in units of Angstrom.
[0091] As another example, when the Transformer neural network is configured as a large language model, the enzyme design platform 100 can represent the input data as a sequence of input tokens, and process the sequence of input tokens using the large language model to generate a sequence of output tokens that define the predicted locations of active sites of the enzyme. In this example, in some implementations, the neural network output can include output data that is in the same format as the input data, e.g., the predicted locations of active sites of the enzyme are also in the PDB format.
[0092] A diffusion model generates an atomic-level representation of the enzyme based on the predicted locations of active sites of the enzyme (step 206). The atomic-level representation ofthe enzyme can define predicted 3-D locations of the atoms included in the enzyme. In some implementations, the atomic-level representation of the enzyme is in the same format as the input data to the Transformer neural network, e.g., the atomic-level representation of the enzyme is also in the PDB format.
[0093] To generate the atomic-level representation of the entire enzyme, the diffusion model is configured to perform a reverse diffusion process by updating an intermediate representation of the atomic-level representation of the enzyme at each of multiple updating iterations, conditioned on the predicted locations of active sites of the enzyme. In other words, the atomic-level representation of the enzyme is the updated intermediate representation after the last iteration of the multiple updating iterations.
[0094] FIG. 5 is an example illustration 500 of an atomic-level representation of an enzyme that defines the predicted 3-D locations of the atoms included in the enzyme. In FIG. 5, the atomic- level representation of the enzyme is in the PDB format. For example, in FIG. 5, the first ATOM line describes the alpha-N atom is part of the glycine residue in chain L, and its residue (amino acid) ID is “1”; the first three floating point numbers are tire predicted x, y and z coordinates of the alpha-N atom in units of Angstrom; and the next two numbers are the occupancy and the temperature factor of the alpha-N atom, respectively.
[0095] Note that in the 5th column which is supposed to correspond to the amino acid residue, is indicated as GLY for all rows. This is a default, placeholder residue that is generated after the reverse diffusion process performed by using the diffusion model. Thus, the next step of the process 200 is to predict the sequence, i.e., the amino acids included in the enzyme, based on the 3-D coordinates defined in the atomic-level representation of the enzyme.
[0096] A sequence prediction neural network processes a neural network input that includes (i) the atomic-level representation of the enzyme, (ii) data derived from the atomic-level representation of the enzyme, or both (i) and (ii) to generate a neural network output that defines a sequence of amino acids (step 208). The sequence of amino acids can define a linear order of amino acids that make up the enzyme. The sequence of amino acids can be represented in any appropriate format.
[0097] FIG. 6 is an example illustration 600 of a sequence of amino acids in a FASTA format. In FIG. 6, the sequence of amino acids begins with the greater-than character (“>”) followed by a description of the sequence. The lines immediately following the description line are thesequence representation, with one letter per amino acid. In FIG. 6, for example, “A” represents the amino acid alanine, “C” represents the amino acid cysteine, “T” represents the amino acid threonine, and “G” represents the amino acid glycine, etc.
[0098] A folding neural network processes a neural network input that includes (i) the sequence of amino acids, (ii) data derived from the sequence of amino acids, or both (i) and (ii) to generate a neural network output that defines the predicted structure of the enzyme (step 210). Continuing with the PFAS example mentioned above, the neural network output can define a predicted structure of an enzyme that is capable of catalyzing a chemical reaction that involves breaking down PFAS.
[0099] In some implementations, the predicted structure of the enzyme can define a three- dimensional configuration of atoms included in the enzyme. The predicted structure can define an estimate of a 3-D configuration of the atoms in the amino acid sequence of the enzyme after the enzyme undergoes protein folding. As mentioned above, in some implementations, the predicted structure of the enzyme can be defined by a set of structure parameters that define a 3- D spatial location and optionally also a 3-D spatial rotation of each amino acid in the enzyme.
[0100] In particular, because the reverse diffusion process is conditioned on, or guided by, the predicted locations of active sites of the enzyme, the atomic-level representation of the enzyme and, correspondingly, the predicted structure of the enzyme will incorporate the geometry of these active sites. Thus, for example, the accessibility of the enzyme to the substrate during the target enzyme-catalyzed reaction can be improved.
[0101] The predicted structure of the enzyme generated by the enzyme design platform 100 can be used in any of a variety of ways. For example, the enzyme design platfonn 100 can provide predicted structure of the enzyme for presentation to the user on a user computer or store the predicted structure of the enzyme in an output data repository for later use.
[0102] In some implementations, an enzyme having the predicted structure can be physically synthesized, e.g., using manual or automatic laboratory techniques. Such a physically synthesized protein may be used in any of a variety of possible ways. For example, an enzyme generated by the enzyme design platform 100 can be used to catalyze an industrial process that involves breaking down PFAS, e.g., a chemical reaction where the substrate is PFAS, and the one or more products formed after binding the enzyme to the substrate include useful and / or less harmful substances.
[0103] Physically synthesizing an enzyme having a predicted structure generated by the enzyme design platform 100 can include experimentally validating the enzyme, e.g., by measuring the stability of the enzyme, or by measuring the real-world structure of the enzyme and comparing it to the structure predicted by the enzyme design platform 100, or by evaluating a product formed after binding the enzyme to the substrate against the one or more aspects of the product defined by the input data.
[0104] In some implementations, a relatively large number of predicted structures of an enzyme that can be generated by repeatedly performing multiple iterations of the process 200 can be filtered in accordance with a predetermined set of criteria, and a relatively small number of filtered predicted structures of an enzyme can be synthesized in the laboratory and tested for their performance, for validation. For example, enzymes having these filtered predicted structures can be evaluated based on their predicted performance, structural stability, and other relevant factors. Data defining the predicted structures of the enzyme that are filtered out, i.e., that are not included in the relatively small number of filtered predicted structures, can be discarded.
[0105] For example, the predetermined set of criteria can include distance criteria between the predicted enzyme structures from the diffusion model and the folding neural network 150. from both these models include data a PDB structure that has information on the locations of the alpha-carbon atom for potential amino acid locations. The Root Mean Square Deviation (RMSD) is measured between the alpha-carbon locations of the amino acids in the same order between the diffusion model and the folding neural network. RMSD is a measure used to quantify the similarity or difference between two molecular structures or conformations. It calculates the average distance between the atoms of the two structures after optimal superimposition or alignment, which is an indicator of how deviant the predicted structures from the two models are. In this example, structures with low RMSD can be selected and a list of the selected structures can then be sent to the lab for validation.
[0106] As another example, the predetermined set of criteria can include a Predicted Mean Alignment (PAE) criterion which can be dependent on the outputs of the folding neural network. PAE measures the confidence in the relative positions and orientations of different parts (e.g. domains) of the predicted structure. Low PAE values between residues from different domains indicate that the folding neural network predicts well-defined relative positions andorientations for those domains. In this example, structures with low PAE values can be selected and a list of the selected structures can then be sent to the lab for validation.
[0107] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0108] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0109] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0110] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0111] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.
[0112] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0113] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0114] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memoryor a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks.However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0115] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.
[0116] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory' feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0117] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, e.g., inference, workloads.
[0118] Machine learning models can be implemented and deployed using a machine learning framework, e.g. a TensorFlow framework or a Jax framework.
[0119] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0120] The computing system can include clients and servers. A client and sewer are generally remote from each other and typically interact through a communication network. The relationship of client and sewer arises by virtue of computer programs running on the respective computers and having a client-sewer relationship to each other. In some embodiments, a sewer transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the sewrer from the device.
[0121] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what maybe claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0122] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0123] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: receiving input data that characterizes an enzyme-catalyzed reaction; generating, using a Transformer neural network, predicted locations of active sites of an enzyme based on the input data; generating, using at least a diffusion model and based on the predicted locations of active sites of the enzyme, a predicted structure of the enzyme, the predicted structure of the enzyme defining a three-dimensional configuration of atoms included in the enzyme.
2. The method of claim 1, wherein the input data comprises data that characterizes a substrate for binding to the enzyme.
3. The method of any one of claims 1-2, wherein the Transformer neural network is configured as a Graph Transformer neural network that operates on graph data defining the enzyme- catalyzed reaction.
4. The method of any one of claims 1-2, wherein the Transformer neural network is configured as a large language model neural network that auto-regressively generates the predicted locations of the active sites of the enzyme.
5. The method of any one of claims 1-4, wherein generating the predicted structure of the enzyme comprises: generating, using the diffusion model, predicted three-dimensional locations of the atoms included in the enzyme; generating, using a sequence prediction neural network, a sequence of amino acids based on the predicted three-dimensional locations of the atoms included in the enzyme; and generating, using a folding neural network or protein structure prediction model, the predicted structure of the enzyme based on the sequence of amino acids.
6. The method of any one of claims 1-5, further comprising performing one or more actions with respect to an enzyme having the predicted structure.
7. The method of claim 6, wherein the one or more actions comprise: determining whether the enzyme having the predicted structure is valid based on the three-dimensional configuration of atoms defined by the predicted structure.
8. The method of claim 6, wherein the one or more actions comprise: performing laboratory synthesis and validation of the enzyme having the predicted structure.
9. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-8.
10. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-8.1 1. An Al-driven Enzyme Design Platform, which combines large language models and diffusion models to design novel enzymes.
12. The Al-driven Enzyme Design Platform of claim 11, where the large language model is trained on scientific literature, protein databases, and experimental datasets related to enzyme biology, biochemistry, and structural biology.
13. The Al-driven Enzyme Design Platform of claim 11 or claim 12, where the large language model is used to predict potential active sites for a given enzyme target.
14. The Al-driven Enzyme Design Platform of any one of claims 11-13, where the predicted active sites from the large language model are used to prompt the diffusion model, which designs novel enzymes that incorporate these active sites.