Machine learning-based methods and apparatus for evolutionary data-driven design of proteins and other sequence-defined biomolecules
A data-driven, evolution-based iterative approach using machine learning models effectively optimizes amino acid sequences for desired protein functionality, addressing the challenges of large and complex search spaces in traditional protein design methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- UNIVERSITY OF CHICAGO
- Filing Date
- 2020-09-11
- Publication Date
- 2026-07-24
AI Technical Summary
Existing protein design methods face challenges in efficiently and reliably identifying novel molecules with specific desirable properties due to a large, complex, and unstructured search space, making optimization difficult and costly.
A data-driven, evolution-based iterative approach using machine learning models to combine unsupervised sequence models with supervised functionality models to identify and optimize amino acid sequences for desired protein functionality, leveraging statistical patterns and iterative feedback loops.
This method efficiently and accurately identifies proteins with desired functionality by narrowing the search space and improving the design process through iterative optimization, overcoming the limitations of traditional structure-based approaches.
Smart Images

Figure 0007894644000035 
Figure 0007894644000036 
Figure 0007894644000037
Abstract
Description
[Technical Field]
[0001] Related applications This application claims the interests of U.S. Provisional Application No. 63 / 020,083 (filed May 5, 2020) and U.S. Provisional Application No. 62 / 900,420 (filed September 13, 2019). Each of the above provisional applications in whole is incorporated herein by reference.
[0002] This disclosure relates to data-driven, evolution-based methods for designing sequence-defined molecules such as proteins, and more specifically, to iterative methods for designing proteins to have desired functionality by combining unsupervised sequence models with supervised functionality models. [Background technology]
[0003] The background art provided herein is for the purpose of generally presenting the context of this disclosure. The research of the inventors currently named in this background art section, and any other aspect of the description that may otherwise not be considered prior art at the time of filing, are not considered prior art to this disclosure, either expressly or implicitly.
[0004] Proteins are molecular machines involved in a wide range of biological processes, including those essential for life. For example, they have the ability to catalyze biochemical reactions in the body that would otherwise take years, happening in microseconds. Proteins are fundamental to normal bodily functions, involved in transport (hemoglobin, a blood protein, transports oxygen from the lungs to tissues), motility (flagellas provide sperm motility), information processing (proteins constitute intracellular signaling pathways), and cellular regulatory cues (like the hormone insulin). Antibodies, which provide host immunity, are proteins, and molecular motors such as kinesin and myosin are similarly responsible for muscle contraction and intercellular transport. When light reaches the eye, a membrane protein in the eye called rhodopsin senses the incident photon, then activates a cascade of downstream proteins, ultimately transmitting what is seen to the brain. Thus, proteins perform a wide variety of specialized functions.
[0005] While proteins exhibit a remarkable range of properties, all proteins are polymers constructed from just 20 units called amino acids. Every protein is a linear sequence of amino acids called a protein sequence. In its natural state, the protein molecule itself twists, rotates, and folds to form generally irregular three-dimensional spheres. The precise three-dimensional sequence of amino acids, known as the protein structure, and the interactions between them, give rise to the function of the protein. Using models of protein function, functional properties can be derived by identifying all atomic interactions (i.e., its energy structure) within the protein molecule. Two independent approaches exist for understanding the energy structure of proteins: one based on protein structure, and the other on evolutionary statistics.
[0006] The structural guide view is valuable. For example, the role of binding site residues can be tested by mutagenesis based on the idea that if important amino acids are replaced with residues of low average functional activity (e.g., alanine), the protein is predicted to show a decrease in its ability to function. For example, using this approach, the importance of amino acids including the interface between a protein and a ligand has been demonstrated.
[0007] However, the principle of spatial proximity described above does not fully capture the determinants of biochemical function. For example, amino acids interact in complex cooperative ways within a structure and can affect the function of the binding site even from distant locations, and the structure alone does not provide a general model for understanding how such cooperativity is arranged. Therefore, other approaches to protein design are desired. In particular, the optimization space for protein design is too large and complex to be tractable using only structure-based approaches.
[0008] The goal of protein design is to identify novel molecules with specific desirable properties. This can be regarded as an optimization problem in which a search for proteins that maximize a given set of quantitative desiderata is performed. However, optimization in protein space is very difficult because the search space is large, discrete, and mostly filled with unstructured non-functional sequences. Creating and testing new proteins is costly and time-consuming, and the number of potential candidates is overwhelming.
[0009] Therefore, improved techniques are desired to more efficiently and reliably search for proteins optimized according to a given set of quantitative desiderata.
Brief Description of the Drawings
[0010] A more complete understanding of the present disclosure is provided by reference to the following detailed description when considered in connection with the accompanying drawings. [Figure 1A]A flowchart of method 10 for designing sequence-defined molecules according to one implementation configuration is shown. [Figure 1B] A more detailed flowchart of method 10 implemented on a protein, according to one implementation configuration, is shown. [Figure 1C] A schematic diagram of a computational model for generating the synthetic sequence of a sequence-defined molecule, based on one implementation method, is shown. [Figure 1D] A schematic diagram of Method 10 for data-driven design of sequence-defined molecules, based on one implementation configuration, is shown. [Figure 1E] Figures illustrating examples of designed or candidate proteins resulting from the systems and / or methods described herein are shown, such proteins being offered in one or more alternative forms (e.g., as final products) and / or applicable to various industries. [Figure 2] This shows an example of the paths taken in an iterative search to design a protein using one implementation method. [Figure 3A] A flowchart of another implementation of Method 10, which has an unsupervised learning portion and a supervised learning portion, is shown. [Figure 3B] A flowchart of the iterative loop in the supervised learning portion of Method 10, based on one implementation, is shown. [Figure 3C] A schematic diagram of Method 10, shown from the perspective of the unsupervised learning portion and the supervised learning portion, is provided for one implementation. [Figure 4] A schematic diagram of a variational autoencoder (VAE) in one implementation configuration is shown. [Figure 5] A schematic diagram of a VAE used in combination with a function defined by Gaussian process regression to find Pareto-optimal candidate amino acid sequences, according to one implementation, is shown. [Figure 6A] A schematic diagram of the first implementation method of gene synthesis, based on one implementation method, is shown. [Figure 6B] A schematic diagram of a second implementation of gene synthesis, based on one implementation method, is shown. [Figure 7] This shows the mapping between codons in gene sequences and amino acids. [Figure 8A] A schematic diagram of method 10, which uses a direct join analysis (DCA) model of an array model according to one implementation, is shown. [Figure 8B] This shows a portion of the shikimic acid pathway in bacteria and fungi, which leads to the biosynthesis of the aromatic amino acids tyrosine and phenylalanine. [Figure 8C] The atomic structure of E. coli colismic acid mutase (CM), including dimers with two functionally active sites (e.g., items 800c1, 800c2, and 800c3), is shown. [Figure 9A-C] Figure 9A shows the first-order statistics (i.e., summarizing the empirical first-order statistics MSA) of sequences sampled from a Boltzmann machine direct-coupled analysis (bmDCA) model in one implementation. Figure 9B shows the second-order statistics of sequences sampled from a bmDCA model in one implementation, summarizing the empirical second-order statistics MSA. Figure 9C shows the tertiary statistics of sequences sampled from a bmDCA model in one implementation, summarizing the empirical tertiary statistics MSA. [Figure 9D] The top two principal components of the distance matrix between all natural CM sequences in the MSA (e.g., the shaded circle as item 900d1) and sequences derived from the bmDCA model (e.g., the shaded circle as item 900d2) are shown for one implementation. [Figure 9E] This study presents a quantitative high-throughput functional assay for CM, in which a library of CM mutants is expressed in E. coli strains lacking colismyate mutase, grown as a mixed population under selective conditions, and then subjected to next-generation sequencing to count the frequencies of each CM allele in the input and selected populations. [Figure 9F] The plot shows an approximately linear relationship between the calculated relative enrichment ("re") and the catalytic force ln(kc / Km) over a range of approximately 5-logarithmic order. [Figure 10A-J]Figure 10A shows a histogram plot of the number of natural CM sequences as a function of statistical energy for one implementation. Figure 10B shows a histogram plot of natural CM sequences as a function of re-score for one implementation. Figure 10C shows a histogram plot of bmDCA-generating sequences as a function of statistical energy at a temperature T=0.33 for one implementation. Figure 10D shows a histogram plot of bmDCA-generating sequences as a function of statistical energy at a temperature T=0.66 for one implementation. Figure 10E shows a histogram plot of bmDCA-generating sequences as a function of statistical energy at a temperature T=1.0 for one implementation. Figure 10F shows a histogram plot of bmDCA-generating sequences as a function of re-score at a temperature T=0.33 for one implementation. Figure 10G shows a histogram plot of bmDCA-generating sequences as a function of re-score at a temperature T=0.66 for one implementation. Figure 10H shows a histogram plot of the bmDCA-generated sequences as a function of the re-score at a temperature T=1.0 for one implementation. Figure 10I shows a histogram plot of the sequences generated using only first-order statistics as a function of the statistical energy for one implementation. Figure 10J shows a histogram plot of the sequences generated using only first-order statistics as a function of the re-score for one implementation. [Figure 11A] Scatter plots of all synthetic CM sequences are shown, illustrating the relationship between bmDCA statistical energy and catalytic function, along with functional sequences (e.g., the shaded bar as item 1100a1) and non-functional sequences (e.g., the shaded bar as item 1100a2). [Figure 11B] The scatter plot shows the top two principal components of sequence variations of the native CM sequence within the MSA, along with functional sequences (e.g., the circle or item shaded as item 1100b1) and non-functional sequences (e.g., the circle or item shaded as item 1100b2). [Figure 11C]The scatter plots are shown as a function of the top two principal components of sequence variation for sequences derived from the bmDCA model, along with functional sequences (e.g., the circle or item shaded as item 1100c1) and non-functional sequences (e.g., the circle or item shaded as item 1100c2). [Figure 11D-E] Figure 11D shows a histogram plot of the number of synthetic sequences with EDCA < 40 that do not have any additional statistical conditions (P(x=1|σ)) derived from the functional complementarity patterns of natural CM sequences. Figure 11E shows a histogram plot of the number of synthetic sequences with EDCA < 40 that have all of the additional statistical conditions (P(x=1|σ)) derived from the functional complementarity patterns of natural CM sequences. [Figure 11F] The structure of E. coliCM is shown, along with the position that contributes most to low statistical energy (e.g., the sphere shaded as item 1100f2) and the position that contributes to E. coli-specific function (e.g., the sphere shaded as item 1100f1). [Figure 12] A schematic diagram of a protein optimization system based on one implementation configuration is shown. [Figure 13] This diagram shows a flowchart of a method for training a deep learning (DL) network using one implementation. [Figure 14] An example of an artificial neural network based on one implementation method is shown.
[0011] The figures described herein illustrate various aspects of the systems and methods disclosed herein. Each figure illustrates an embodiment of a particular aspect of the disclosed systems and methods, and it should be understood that each figure is intended to correspond to a possible embodiment thereof. Furthermore, wherever possible, the description herein refers to the reference numbers included in the figures herein, and features shown in multiple figures are designated by corresponding reference numbers. [Overview of the project]
[0012] According to an aspect of one embodiment, a method is provided for designing a protein to have desired functionality. The method includes (i) determining candidate amino acid sequences for a synthetic protein using a machine learning model trained to learn implicit patterns of amino acid sequences in a training dataset of proteins, wherein the machine learning model represents learned implicit patterns in latent space; and (ii) running an iterative loop. Each iteration of the loop includes (i) generating a candidate gene and producing a candidate protein corresponding to each candidate amino acid sequence, wherein each candidate gene encodes the corresponding candidate amino acid sequence; (ii) evaluating the extent to which each candidate protein exhibits desired functionality by measuring values of the candidate proteins using one or more assays; and (iii) when one or more stopping criteria of the iterative loop are not met, calculating a goodness-of-fit function in latent space from the measured values and selecting candidate amino acid sequences for subsequent iterations using a combination of the goodness-of-fit function and the machine learning model.
[0013] According to an aspect of another embodiment, a system is provided for designing proteins to have desired functionality. The system comprises (i) a gene synthesis system, (ii) an assay system, and (iii) a processing circuit. The gene synthesis system is configured to synthesize candidate genes corresponding to each candidate gene sequence encoding candidate amino acid sequences. The assay system is configured to measure values of candidate proteins corresponding to each candidate amino acid sequence, the measured values providing an indicator of desired functionality. The processing circuit is configured to (i) determine candidate amino acid sequences for a synthesized protein using a machine learning model trained to learn implicit patterns in a training dataset of proteins (the machine learning model represents learned implicit patterns in latent space), and (ii) execute an iterative loop. Each iteration of the loop includes (a) submitting a candidate amino acid sequence to be synthesized, (b) receiving measured values generated by measuring the candidate protein using one or more assays, and (iii) when one or more stopping criteria of the iterative loop are not met, calculating a goodness-of-fit function in latent space from the measured values and using a combination of the goodness-of-fit function and the machine learning model to select candidate amino acid sequences for subsequent iterations.
[0014] According to an aspect of the third embodiment, a non-temporary computer-readable storage medium containing executable instructions is provided, and when executed by a circuit, the instructions cause the circuit to perform the following steps: (i) determine candidate amino acid sequences for a synthetic protein using a machine learning model trained to learn implicit patterns of amino acid sequences in a training dataset of proteins, wherein the machine learning model represents learned implicit patterns in latent space; and (ii) execute an iterative loop. Each iteration of the loop includes: (i) synthesize a candidate gene and produce a candidate protein corresponding to each candidate amino acid sequence, wherein each candidate gene codes for the corresponding candidate amino acid sequence; (ii) evaluate the extent to which each candidate protein exhibits desired functionality by measuring values of the candidate proteins using one or more assays; and (iii) when one or more stopping criteria of the iterative loop are not met, calculate a goodness-of-fit function in latent space from the measured values and select a candidate amino acid sequence for subsequent iterations using a combination of the goodness-of-fit function and the machine learning model.
[0015] According to an aspect of the fourth embodiment, a method is provided for designing a protein to have a desired functionality. The method includes (i) determining a candidate gene sequence of a sequence-defined biomolecule, the candidate gene sequence being generated using a machine learning model trained to learn implicit patterns in a training dataset of sequence-defined biomolecules, the machine learning model representing learned implicit patterns in latent space, and (ii) running an iterative loop. Each iteration of the loop includes (i) synthesizing a candidate gene corresponding to each candidate gene sequence, each candidate gene encoding a corresponding candidate biomolecule, (ii) evaluating the extent to which each candidate biomolecule exhibits a desired functionality by measuring values of the candidate biomolecule using one or more assays, and (iii) when one or more stopping criteria of the iterative loop are not met, calculating a goodness-of-fit function in latent space from the measured values, and using a combination of the goodness-of-fit function and the machine learning model to select a candidate gene sequence for subsequent iterations.
[0016] As per the foregoing and the disclosure herein, this disclosure includes improvements to the functionality of the computer or improvements to the art, since the inventions disclosed herein describe an underlying computer or machine that is continuously improved, for example, by updating a machine learning model to more accurately predict or identify a candidate protein or design. In other words, this disclosure describes improvements to the functionality of the computer itself or “any other art or art” so that the predictions or evaluations produced by the underlying computing device improve over time as the machine learning model of the computing device is further trained (across various iterative loops) to better identify or predict a candidate or design protein. This is an improvement over the prior art, since the prior art system could not have been improved without manual coding or development by a human developer.
[0017] Since this disclosure describes at least the use of artificial intelligence (e.g., machine learning models) applied to the design of proteins having desired functionality, this disclosure is relevant to improvements over other technologies or fields.
[0018] This disclosure includes applying the features described herein by using, or by using, certain machines, such as microfluidic devices for measuring the fluorescence of cells corresponding to candidate or designed proteins identified by artificial intelligence based on the systems and methods described herein.
[0019] This disclosure includes converting or reducing specific articles into different states or substances, for example, converting or reducing candidate or designed proteins identified by artificial intelligence-based systems and methods described herein into cells that can be used to generate, produce, or otherwise catalyze final products.
[0020] This disclosure includes, for example, systems and methods for designing proteins having desired functionality for the development, manufacture, or creation of real-world products, by adding certain features beyond routine, conventional activities well understood in the art, or by adding unconventional steps that limit the claims to specific useful applications.
[0021] The advantages will become more apparent to those skilled in the art from the following description of the preferred embodiments shown and described as examples. As will be understood, other different embodiments are possible, and their details are modifiable in various ways. Therefore, the drawings and description should be considered illustrative and not limiting in nature. [Modes for carrying out the invention]
[0022] The methods described herein use data-driven and evolution-based approaches that leverage statistical genomics, machine learning, and artificial intelligence techniques to learn implicit patterns that associate amino acid sequences with protein structure, function, and evolvability, thereby overcoming the shortcomings and challenges of previous protein design methods. In one embodiment, it is desirable that a protein has desired functionality by first training a machine learning model based on protein sequence information in a training dataset (e.g., a database of homologous proteins) and then using the trained machine learning model to generate candidate amino acid sequences.
[0023] Next, an iterative process is performed to select candidate amino acid sequences that are better than the previous ones. This iterative process involves producing proteins of the candidate amino acid sequences and assaying them to measure their functionality. Using the measured values from the assays, a functionality landscape (e.g., a functionality-based model) is generated to predict which candidate sequences are most likely to exhibit the desired functionality. Using a goodness-of-fit function based on the functionality landscape in conjunction with machine learning can determine new candidate sequences that have better functionality than the initial ones. The iterative process is then repeated until a stopping criterion is reached and the final optimized amino acid sequence is output as the designed protein, with each iteration producing a new candidate sequence that is better than the previous iteration. This iterative process is illustrated, for example, in Figure 1A.
[0024] The methods described herein utilize both information provided by the sequence itself (i.e., sequence-based models) and functional information provided by producing and evaluating proteins using assays that measure protein functionality (i.e., functionality-based models). The combination of sequence-based and functionality-based models is called a protein model.
[0025] Sequence-based models are generally unsupervised machine learning models that receive amino acid sequences from protein aggregates as training data. Therefore, sequence-based models are also called machine learning models, unsupervised models, amino acid sequence-based models, or any combination thereof.
[0026] Generally, functionality-based models are supervised models because they are generated / trained using information from additional measurements (i.e., monitoring). Functionality-based models are often expressed as goodness-of-fit functions and / or functionality landscapes. For example, when multi-objective multi-dimensional optimization is used in the design process, the functionality landscape is one of the components of the goodness-of-fit function that identifies which amino acid sequences are suitable candidates when designing the desired functionality. Thus, functionality-based models can be called supervised models, goodness-of-fit functions, functionality landscapes, functionality-based models, functionality models, or any combination thereof. Functionality-based models can be generated via machine learning and may be machine learning models. However, as used herein, the term “machine learning model” generally refers to sequence-based models unless explicitly stated in the context.
[0027] As discussed above, the goal of protein design is to identify novel proteins with specific desirable properties. This can be viewed as an optimization problem in which the search for a protein that maximizes a given quantitative desiderata is performed. However, optimization in the protein space is extremely difficult because the search space is large, discrete, and unstructured (for example, in the context of data mining, amino acid sequences are classified as unstructured data, and all the challenges associated with classification arise). Creating and testing new proteins is costly and time-consuming, and the number of potential candidates is overwhelmingly large.
[0028] For example, the number of possible amino acid sequences of length N is 20 NThis sequence space is largely occupied by unfolded, non-functional molecules. Directed evolution offers one approach to searching this space for functional molecules structured by a round of mutations and functional selection, but it represents only a very localized search of the sequence space around a particular native sequence. Therefore, formal rules are needed to overcome these limitations of directed evolution in order to guide searches that can explore the very large combinatorial complexity of possible functional sequences.
[0029] One set of rules derives from a physical model of the fundamental forces between atoms; this is a physics-based design. However, these methods are limited by (i) the inaccuracies of the force fields, (ii) the imposition of ad-hoc constraints to manage the combinatorial complexity of the search, and (iii) the philosophy that globally optimized sequences are best. Thus, even if sequences generated through physics-based design can form very stable structures, the evidence so far is that without directed evolution, such sequences have poor functional performance.
[0030] In contrast to physics-based design and directed evolution approaches, the methods described herein utilize evolutionary data-driven statistical models. These types of models provide a clear approach to obtaining rules for searching sequence space that do not rely on knowledge of underlying physical forces or the mechanisms of folding, function, or evolution. Furthermore, these models do not require three-dimensional structures as a starting point. Instead, they capture patterns of statistical constraints that define the natural assemblies of protein sequences, and in doing so, indirectly capture the physics of protein folding and function. Thus, rather than globally optimizing proteins for stability, the methods described herein optimize constraints chosen by nature throughout evolutionary history, enabling the retrieval of functional proteins in sequence space with much higher yield and depth than previous methods.
[0031] To address the challenges of previous protein design methods, the systems and methods described herein use a data-driven, evolution-based, iterative approach to identify / select the most promising amino acid sequences for desired protein / enzyme functionality.
[0032] The data-driven aspect of this approach arises in part from using data to train combinations of models, including sequence-based and functionality-based models. Sequence-based models (also called unsupervised models) can be machine learning models trained to represent statistical properties and implicit patterns learned from a training dataset (e.g., multiple sequence alignments (MSAs) of homologous proteins). Functionality-based models (also called supervised models) incorporate feedback measured from assays of synthetic proteins based on amino acid sequences identified as promising candidates in previous iterations. The assays are designed to provide an indicator of the functionality of the desired protein. For example, a desired functionality (e.g., binding, allostery, or catalysis) may quantitatively correlate with changes in the growth rate of an organism, gene expression, or optical properties such as light absorption or fluorescence in various specific assays.
[0033] An evolution-based aspect of this approach arises in part from using iterative feedback loops to identify trends in the highest-performing amino acid sequences and then performing improved amino acid sequence selection based on those trends (i.e., in silico mutagenesis).
[0034] Furthermore, in certain implementations, the training dataset consists of homologs. Thus, the implicit patterns learned by the sequence-based model include sequence patterns of desired functionality learned over time through biological evolution.
[0035] Therefore, the method described herein addresses the challenge of rapidly identifying synthetic proteins that are likely to exhibit desired functionality through an iterative cycle, where a data-driven computational model of a protein is used to generate candidate amino acid sequences, the candidate proteins are then physically produced and evaluated in the laboratory using quantitative assays, and finally, the computational model is improved using information from the assays to generate candidate amino acid sequences that are further optimized for the desired functionality in the next iteration. Thus, a feedback loop is established to iteratively optimize the amino acid sequence for the desired functionality.
[0036] In particular, protein models can be divided into two parts: supervised models and unsupervised models. In the first iteration of the loop, the initial candidate amino acid sequence can be determined using only the unsupervised model. Then, in subsequent iterations, both the supervised and unsupervised models can be used to generate candidate amino acid sequences for the subsequent iterations.
[0037] Here, referring to drawings where similar reference numbers across several figures indicate the same or corresponding parts, Figures 1A and 1B show a flowchart of method 10 for designing synthetic proteins using a combination of sequence models and functional models.
[0038] In process 22 of Method 10, a sequence-based model (also called an unsupervised machine learning model, or more simply, an unsupervised model) is trained using a training dataset of amino acid sequences to generate a new amino acid sequence having a similar statistical structure and / or pattern to that learned from the training dataset. More generally, the methods described herein are applicable not only to proteins but also to any sequence-defined molecule, including biomolecules. For example, nucleic acids such as messenger RNA or microRNA can be designed to have a desired functionality. In this case, the candidate sequence would be a nucleotide sequence rather than an amino acid sequence. Otherwise, Method 10 is the same, with only small, simple changes, as can be understood by those skilled in the art. As another example, polymers defined by a sequence of monomers can be designed to have a desired functionality using the methods described herein. For specificity, the methods and systems described herein are illustrated using a non-limiting example of protein design, but the methods and systems described herein are general and include the design of any sequence-defined molecule.
[0039] In process 32 of method 10, a sequence-defined molecule (e.g., a protein) is physically produced in the laboratory for the candidate sequence determined in process 22.
[0040] In process 42 of method 10, sequence-defined molecules are assayed to measure how well they exhibit the desired functionality. These measured values are then incorporated into a functionality model. This is then used in process 22 in the next iteration to select new, superior candidates based on the measured feedback resulting from process 42.
[0041] Over time, Method 10 converges on candidate sequences that satisfy a predefined stopping criterion (e.g., the degree of desired functionality corresponding to candidate sequences that exceed a predefined threshold, or the rate of increase in desired functionality from iteration to iteration that is slower than another predefined threshold). Method 10 then outputs one or more of the optimized candidate sequences as the designed protein. As used herein, the terms “candidate protein” and “designed protein” may be used interchangeably, and the designed protein represents the output of the candidate protein, including, for example, iterations of iterative loops and / or machine learning models, as described herein, for a given iteration of the system and / or method described herein.
[0042] Returning to process 22, in the non-limiting examples provided herein, the sequence models are exemplified as restricted Boltzmann machines (RBMs), variational autoencoders (VAEs), generative adversarial networks (GANs), statistical combined analysis (SCA), and direct combined analysis (DCA).
[0043] The first three types of sequence models (i.e., RBMs, VAEs, and GANs) are types of artificial neural networks (ANNs) that map inputs at the visible nodes of the ANN to a hidden layer of nodes that is dimensionally reduced compared to the visible layer, compressing information and providing an information bottleneck where implicit patterns are captured, and are therefore often called “generative methods.” This hidden layer of nodes defines a latent space, and the mapping from the visible layer to the hidden space is referred to herein as “encoding,” while the reverse process of mapping points in the latent space to the original space of amino acids is referred to herein as “decoding” or “generation.” The learned turns resulting from the compression into a reduced subspace lead to a random selection of points in the latent space, followed by decoding to the points, resulting in amino acid sequences with patterns similar to those in the training dataset.
[0044] For example, a VAE trained using training images of faces can be used to generate images that are recognizable as faces, even though they are different faces from those in the training images (i.e., a VAE can learn general facial features and then be used to generate new facial images with the learned features). If it is desired to further train the VAE not only to recognize faces but also to recognize faces with desired features (e.g., beautiful faces), the user can examine the faces generated by the VAE and tag the figures according to the desired features. Supervised learning can then be performed by using an index of the desired features to learn patterns of beautiful faces or any other desired features. For example, a VAE trained using only beautiful faces will learn patterns of beautiful faces. Applying a similar principle, patterns of amino acid sequences of a particular type of protein (e.g., homologous proteins) can be learned to generate new candidates, and then a subset of candidate amino acid sequences exhibiting desired functionality can be further learned through supervised learning techniques.
[0045] Returning to the unsupervised sequence models in Process 22, the last two types of models (i.e., SCA and DCA) are often referred to as “statistical methods.” These statistical methods also generate candidate amino acid sequences, but they do so by learning the statistical properties (e.g., primary and secondary statistics) of amino acid sequences in the training dataset and then selecting candidate amino acid sequences that conform to the learned statistical model. In other words, like generative methods, statistical methods generate candidate amino acid sequences, but statistical methods do not use neural networks to map the sequence space to the latent space. Rather, these statistical methods operate within a domain of sequence space and learn patterns in subspaces within sequence space (similar to the latent space in generative methods, but different from the latent space) that are likely to produce proteins similar to those in the training dataset.
[0046] In other words, unsupervised sequence models narrow the search by restricting the search space from the "overwhelmingly large" number of possible candidate amino acid sequences to a much smaller, more manageable subset of sequences that are more likely to exhibit the desired functionality / features. From this subset, several amino acid sequences (e.g., on the order of 1,000) are then selected as candidate amino acid sequences to be produced in the lab and evaluated to provide measurement data / values for supervised learning. It should be understood that the selected candidate amino acid sequences produced can include a variety of numbers and counts, including, as non-limiting examples, sequences with at least 100 amino acids, at least 500 amino acids, at least 1,000 amino acids, at least 1,500 amino acids, and so on.
[0047] Here, considering the supervised functionality model of Process 42, in certain implementation embodiments, the supervised model can be thought of as a functionality landscape in latent space. In this case, peaks correspond to regions in latent space that are more likely to produce candidate amino acid sequences exhibiting the desired functionality, while valleys are less likely to produce superior candidate amino acid sequences. For example, this functionality landscape is generated by performing regression analysis on measured values. The measured values are generated by performing assays on candidate amino acid sequences to measure the desired functionality (e.g., catalysis, binding, or allostery). In the non-limiting examples provided herein, supervised functionality models are exemplified as being generated by fitting measured values to a landscaped functionality using multivariate linear regression, support vector regression (SVR), Gaussian process regression (GPR), random forest (RF), decision tree (DT), or artificial neural network (ANN). For example, Gaussian processes (GP) are discussed in CERasmussen and CKIWilliams, “Gaussian Processes for Machine Learning,” the MIT Press, 2006, ISBN 026218253X (also available at www.GaussianProcess.org / gpml), which is incorporated in its entirety by reference.
[0048] Supervised models are not limited to considering functionality alone; they can also account for other factors that may increase the likelihood of a candidate amino acid sequence being successful, such as similarity and stability. It should be considered that physics-based modeling, such as numerical computation using the Rosseta Commons software suite, can provide indicators of stability. For example, sequences predicted based on numerical computation to maintain the original structure of the natural fold are more likely to be functional.
[0049] Therefore, supervised models can be extended to account for other factors such as stability and similarity. For example, a stability landscape can be defined in the latent space, and a similarity landscape can be defined in the latent space. Then, multidimensional analysis can be used to determine the Pareto frontier (i.e., the convex hull in the multidimensional space of (i) functionality, (ii) similarity, and (iii) stability), and candidate amino acid sequences can be selected from points in the latent space that lie on the Pareto frontier.
[0050] General explanation of the method Figure 1B shows a more detailed flowchart of Method 10 in one non-limiting embodiment.
[0051] In process 22, the amino acid sequence model is trained using the training dataset 15. As a non-restrictive example, a training dataset of N=1388 sequences in multi-segment alignment (MSA) where each sequence has a length L of 245 can be used as the training dataset 15.
[0052] In step 20 of process 22, the sequence model generates a candidate amino acid sequence 25. As schematically shown in Figures 1C and 1D, the sequence model has a compression phase and a generation phase. The compression phase of the model may also be called encoding (e.g., encoding is mapping from amino acid sequence space to latent space), and the generation phase of the model may also be called decoding (e.g., decoding is mapping from latent space to amino acid sequence space).
[0053] Following process 22, in step 30 of process 32, a gene is synthesized to encode a candidate sequence 25, and a candidate protein 35 is produced from the synthesized gene. In some implementations, this search can be accelerated using virtual screening. A virtual library containing thousands to hundreds of millions of candidates can be assayed with first-principles simulations or statistical predictions based on trained proxy models, and only the most promising reads are selected and experimentally tested.
[0054] In step 40 of process 42, an assay is performed to evaluate the extent to which the candidate protein exhibits the desired functionality. The measured values 45 indicating the desired functionality obtained from these assays are considered in step 50 to determine whether the termination criteria are met. If the termination criteria are reached, the optimized amino acid sequence is output as the designed protein 55. Otherwise, method 10 proceeds to step 60.
[0055] In Figure 1D, process 22 is illustrated as computational modeling, step 32 as gene synthesis and protein production, followed by high-throughput functional screening. The measured values of the high-throughput functional screening are fed back into the computational model unless the high-throughput functional screening indicates that the protein design objective has been achieved. In that case, the method is complete, and method 10 outputs a designed protein with the desired features. In the implementation example in Figure 1D, a microfluidic apparatus is used to measure the fluorescence of cells containing each candidate protein, and based on the measured fluorescence, the cells are moved or provided in different bins to illustrate the high-throughput functional screening. However, it should be understood that high-throughput functional screening can also be performed and measured using alternative methods and systems (e.g., alternative to or different from the microfluidic apparatus) as described herein. The measured fluorescence in this implementation is adjusted to be proportional to the protein properties that define the design objective. Thus, the screen generates data for iterative optimization of the computational model.
[0056] In step 60 of process 42, a functional model is generated based on the measured values.
[0057] Design or candidate protein This disclosure illustrates the use of the methods and systems described herein for generating novel enzymes, but the systems and methods are not limited to proteins having a particular type of activity or structure (e.g., the designed or candidate proteins described herein). In fact, in various embodiments, proteins or otherly designed or candidate proteins may be antibodies, enzymes, hormones, cytokines, growth factors, coagulation factors, anticoagulation factors, albumins, antigens, adjuvants, transcription factors, or cell receptors. Indeed, the systems and methods described herein are useful, for example, for generating or identifying novel proteins that exhibit biological activity consistent with antibodies, enzymes, hormones, cytokines, growth factors, coagulation factors, anticoagulation factors, albumins, antigens, adjuvants, or cell receptors. Furthermore, the candidate proteins described herein can be used in a variety of applications or functions. For example, a candidate protein may be used for selective binding to one or more other molecules. Furthermore, or instead, as a further example, a candidate protein may be provided to catalyze one or more chemical reactions. Furthermore, or alternatively, as a further example, candidate proteins may be offered for long-range signaling (e.g., allostery, which involves processes in which biomolecules (primarily proteins) transmit the effects of binding at one site to another, often distal, functional site, enabling the modulation of activity).
[0058] Cytokines include, but are not limited to, chemokines, interferons, interleukins, lymphokines, and tumor necrosis factors. Cellular receptors, such as cytokine receptors, are also considered. Examples of cytokines and cell receptors include, but are not limited to, tumor necrosis factor alpha and beta and their receptors, lipoproteins, colchicine, adrenocorticotropic hormone, vasopressin, somatostatin, ripressin, pancreozymin, leuprolide, alpha-1 antitrypsin, atrial natriuretic factor, thrombin, enkephalinase, RANTES (normally expressed and secreted by T cells and regulated by activation), human macrophage inflammatory protein (MIP-1 alpha), cell-determining factor proteins such as CD-3, CD-4, CD-8, and CD-19, erythropoietin, interferon alpha, interferon beta, interferon gamma, interferon lambda, colony-stimulating factors (CSFs), e.g., M-CSF, GM-CSF, and G-CSF, IL-1 to IL-10, T cell receptors, and prostaglandins.
[0059] Examples of hormones include, but are not limited to, antidiuretic hormone (ADH), oxytocin, growth hormone (GH), prolactin, growth hormone-releasing hormone (GHRH), thyroid-stimulating hormone (TSH), thyrotropin-releasing hormone (TRH), adrenocorticotropic hormone (ACTH), follicle-stimulating hormone (FSH), luteinizing hormone (LH), luteinizing hormone-releasing hormone (LHRH), thyroxine, calcitonin, parathyroid hormone, aldosterone, cortisol, epinephrine, glucagon, insulin, estrogen, progesterone, and testosterone.
[0060] Examples of growth factors include, for instance, vascular endothelial growth factor (VEGF), nerve growth factor (NGF), platelet-derived growth factor (PDGF), fibroblast growth factor (FGF), epidermal growth factor (EGF), transforming growth factor (TGF), bone morphogenetic protein (BMP), and insulin-like growth factors I and II (IGF-I and IGF-II).
[0061] Examples of coagulation factors include factor I, factor II, factor III, factor V, factor VI, factor VII, factor VIII, factor VIIIC, factor IX, factor X, factor XI, factor XII, factor XIII, von Willebrand factor, prekallikrein, heparin cofactor II, antithrombin III, and fibronectin.
[0062] Examples of enzymes include, but are not limited to, angiotensin-converting enzyme, streptokinase, and L-asparaginase. Other examples of enzymes include, for example, nitrate reductase (NADH), catalase, peroxidase, nitrogenase, phosphatase (e.g., acid / alkaline phosphatase), phosphodiesterase I, inorganic diphosphatase (pyrophosphatase), dehydrogenase, sulfatase, arylsulfatase, thiosulfurtransferase, L-asparaginase / L-glutaminase, beta-glucosidase, arylacylaminidase, amidase, invertase, xylanase, cellulose, urease, phytase, carbohydrase, and amylase (alpha-amylase). These include beta-amylase, arabinoxylanase, beta-glucanase, alpha-galactosidase, beta-mannanase, pectinase, non-starch polysaccharide-degrading enzymes, endoproteases, exoproteases, lipases, cellulases, oxidoreductases, ligases, synthetases (e.g., aminoacyltransferase RNA synthetase, glycyl-tRNA synthetase), transferases, hydrolases, lyases (e.g., decarboxylase, dehydratase, deaminase, aldolase), isomerases (e.g., triose phosphate isomerase), and trypsin. Further examples of enzymes include catalase (e.g., alkali-resistant catalase), alkali amylase, pectinase, oxidase, laccase, proxidase, xylanase, mannanase, acylase, alcalase, alkylsulfatase, cellulose-degrading enzymes, cellobiohydrolase, cellobiase, exo-1,4-beta-D-glucosidase, chloroperoxidase, chitinase, cyanidase, cyanidohydratase, l-galactonolactone oxidase, ligninperoxidase, lysozyme, mn-peroxidase, muramidase, parathion hydrolase, pectin esterase, peroxidase, and triosinase. Further examples of enzymes include nucleases (e.g., endonucleases such as zinc finger nucleases, transcription activator-like effector nucleases, Cas nucleases, and modified meganucleases).
[0063] The above disclosures refer to exemplary proteins within various protein classes and subclasses (e.g., progesterone as an example of a hormone). It will be understood that the systems and methods described herein may branch off from the specific proteins described above, as far as their amino acid sequences are concerned, but may produce proteins having similar, equivalent, or improved activity or other desired biological characteristics as shown by the reference proteins.
[0064] Protein production Recombinant proteins are achieved using a variety of biotechnology tools and methods. For example, an expression vector containing nucleic acids encoding a protein of interest is introduced into host cells cultured under conditions that allow for protein production, and the protein is collected, for example, by purifying the protein secreted from the culture medium, or by lysing the cells to release intracellular proteins and collecting the protein of interest from the lysate.
[0065] Methods for introducing nucleic acids into host cells are well known in this art and are described, for example, in Cohen et al (1972) Proc. Natl. Acad. Sci. USA 69, 2110, Sambrook et al (2001) Molecular Cloning, A Laboratory Manual, 3.sup.rd Ed. Cold Spring Harbour Laboratory, Cold Spring Harbour, NY, and Sherman et al (1986) Methods In Yeast Genetics, A Laboratory Manual, Cold Spring Harbour, NY.
[0066] Nucleic acids that encode proteins are typically packaged into expression vectors that contain regulatory sequences that promote protein production and, optionally, facilitate entry into host cells. Plasmids or viral vectors are often used in this context.
[0067] Various host cells are suitable for the production of recombinant proteins. For example, mammalian cells are often used to manufacture therapeutic drugs. Non-exclusive examples of mammalian host cells suitable for recombinant protein production include, but are not limited to, Chinese hamster ovary cells (CHO), monkey kidney CV1 cells transformed with SV40 (COS cells, COS-7, ATCC CRL-1651), human fetal kidney cells (e.g., 293 cells), baby hamster kidney cells (BHK, ATCC CCL-10), monkey kidney cells (CV1, ATCC CCL-70), African green monkey kidney cells (VERO-76, ATCC CRL-1587, VERO, ATCC CCL-81), mouse Sertoli cells, human cervical cancer cells (HELA, ATCC CCL-2), canine kidney cells (MDCK, ATCC CCL-34), human lung cells (W138, ATCC CCL-75), human liver cancer cells (HEP-G2, HB 8065), and mouse mammary cancer cells (MMT060562, ATCC CCL-51). Bacterial cells (Gram-positive or Gram-negative) are prokaryotic host cells suitable for protein production. Yeast cells are also suitable for recombinant protein production.
[0068] Final product Figure 1E illustrates an example 1e10 (e.g., design or candidate proteins 1e20, 1e22, 1e24, and / or 1e26) resulting from the systems and / or methods described herein, such proteins being provided in one or more alternative forms (e.g., as final products) and / or applicable to various industries (e.g., industries 1e30, 1e32, 1e34, 1e36, and / or 1e38).
[0069] Proteins (e.g., candidate proteins) resulting from the systems and / or methods described herein may be applied to any of many industries, including biopharmaceuticals (e.g., therapeutics and diagnostics)1e34, agriculture (e.g., plants and livestock)1e32, veterinary medicine, industrial biotech (e.g., biocatalysts)1e30, environmental protection and remediation1e38, and energy1e36.
[0070] As illustrated in Figure 1E, proteins (e.g., designed or candidate proteins) resulting from the systems and / or methods described herein may be developed or produced for various industries (e.g., 1e30, 1e32, 1e34, 1e36, and / or 1e38) or otherwise used in those industries and may be provided to end users in one or more alternative forms: (1) as purified protein 1e20 supplied in solution, as lyophilized powder, or in any other form commonly used for the transfer of active protein molecules; (2) as a library of synthetic genes 1e22 encoding the designed protein cloned into plasmids, viruses, or cosmid vectors; (3) as genes expressed in modified microorganisms or other host strains 1e24; and / or (4) as genes cloned into gene therapy vectors 1e26. As a non-limiting example, purified protein 1e20 may be used or provided for use in any one or more of the industries of biocatalysis 1e30, agriculture 1e32, and / or biopharmaceuticals 1e34 for the development, manufacture, or creation of final products, as illustrated in Figure 1E. As an additional non-limiting example, a library of synthetic genes 1e22 may be used or provided for use in any one or more of the industries of biocatalysis 1e30, agriculture 1e32, biopharmaceuticals 1e34, and / or environmental protection and restoration 1e38 for the development, manufacture, or creation of final products. As an additional non-limiting example, modified microorganisms or other host strains 1e24 may be used or provided for use in any one or more of the industries of biocatalysis 1e30, agriculture 1e32, biopharmaceuticals 1e34, energy 1e36, and / or environmental protection and restoration 1e38 for the development, manufacture, or creation of final products. As an additional, non-limiting example, gene therapy vector 1e26 may be used or provided for any one or more of the industries of agriculture 1e32 and / or biopharmaceuticals 1e34 for the development, manufacture, or creation of an end product.The following sections of this Specification provide additional details and non-limiting examples of these different approaches for generating, manufacturing, creating, and / or producing modified protein solutions (e.g., designed or candidate proteins) for the development, manufacturing, creation, and / or generation of the final products described herein.
[0071] Purified protein The output of the methods and / or systems described herein may be provided to the end user as a purified protein product (e.g., purified protein 1e20). Proteins can be expressed and purified by any of several methods. Protein expression often involves, but is not limited to, the introduction of a gene encoding the protein of interest into a host microorganism (transformation), growth of the transformed strain to the logarithmic phase, and induction of gene expression using a small molecule that activates transcription from one of many standard promoter sequences. The induced culture is then grown to the saturation phase and harvested for protein purification. Examples of protein purification techniques include, but are not limited to, centrifugation, filtration (e.g., tangential flow microfiltration), precipitation / aggregation, chromatography (e.g., ion exchange chromatography, immobilized metal chelate chromatography, and / or hydrophobic interaction chromatography), thiophilic adsorption, affinity-based purification methods, and crystallization. Purified proteins can be formulated into liquid or lyophilized (or dried) compositions suitable for transport, storage, and final end-use in, for example, therapeutic, agricultural, and / or industrial fields.
[0072] Gene library The output of the methods and / or systems described herein may be provided as a gene or gene library (e.g., gene or gene library 1e22). A gene encoding a designed protein can be created by back-translating an amino acid sequence into a DNA sequence using a standard codon table, and then cloned into one of a variety of host plasmids or viral vectors for replication and amplification. The resulting library can then be used by end-users in a variety of custom manufacturing processes for pharmaceuticals, agricultural products, biocatalysis, and environmental remediation. It is standard practice to prepare gene libraries as pure DNA supplied in a dry (or lyophilized) form for storage and / or transport.
[0073] Modified microorganisms The output of the methods and / or systems described herein may be provided as genes that are incorporated into the chromosomes of a host organism via standard gene transfer methods. In this case, the gene encoding the designed protein can be created by backtranslating the amino acid sequence to a nucleic acid sequence using a standard codon table, ligating it with appropriate 5' and 3' nucleotide sequences to ensure desired expression, regulation, and mRNA stability, and also to be integrated into the genome of the host organism. Such a modified host organism (e.g., modified strain 1e24) may be supplied to end users for use in one of many standard applications. These may include growth in large-scale culture and production facilities, use in multi-step industrial biosynthetic pathways, or use as components of modified microbial communities for industrial, agricultural, pharmaceutical, environmental, or energy extraction processes.
[0074] Gene therapy vectors In various embodiments, the output of the method or system described herein is a nucleic acid having the desired activity (e.g., having the desired activity itself or encoding a peptide or protein having the desired activity). In various embodiments, the nucleic acid is incorporated into an expression vector (e.g., vector 1e26). A “vector” or “expression vector” is any type of gene construct containing nucleic acid (DNA or RNA) for introduction into a host cell. In various embodiments, an expression vector is a viral vector, i.e., a viral particle containing all or part of a viral genome that can function as a nucleic acid delivery vehicle. A viral vector containing an exogenous nucleic acid (maybe plural) encoding a gene product of interest is also referred to as a recombinant viral vector. As will be understood in the art, in some contexts the term “viral vector” (and similar terms) may be used to refer to a vector genome in the absence of a viral capsid. Viral vectors for use in the context of this disclosure include, for example, retroviral vectors, herpes simplex virus (HSV) vectors, parvovirus vectors, for example, adeno-associated virus (AAV) vectors, AAV-adenovirus chimeric vectors, and adenovirus vectors.Any of these viral vectors can be prepared using standard recombinant DNA techniques described in, for example, Sambrook et al., Molecular Cloning, a Laboratory Manual, 2nd edition, Cold Spring Harbor Press, Cold Spring Harbor, NY (1989), Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates and John Wiley & Sons, New York, NY (1994), Coen DM, Molecular Genetics of Animal Viruses in Virology, 2nd Edition, BNFields (editor), Raven Press, NY (1990), and the references cited therein.
[0075] Expression vectors have a variety of industrial applications. For example, expression vectors may be provided to end users for use in gene therapy methods (i.e., administration to patients to treat or prevent a disease or disorder). Expression vectors may also be used by end users to produce proteins of interest encoded by nucleic acids.
[0076] Gaussian process regression (GPR) One non-limiting example for generating functional models is applying Gaussian process regression (GPR) to measured values from assays. For example, GPR can be performed on measured values as discussed in PARomero, et al., “Navigating the protein fitness landscape with GPs,” PNAS, vol.110, pp.E193-E201 (2012), CNBedbrook et al., “Machine learning to design integral membrane channel rhodopsins for efficient eukaryotic expression and plasma membrane localization,” PLoS Comput Biol., vol.13, p.e1005786 (2017), and RGBombarelli, et al., “Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules,” ACS Cent.Sci., vol.4, pp.268-276 (2018) (each of these is incorporated herein in its entirety by reference). GPR is performed on the measured values for the position in the latent space corresponding to each candidate protein, and as a result, a mean and standard deviation (i.e., uncertainty) are assigned to each position in the latent space.
[0077] Using the mean and standard deviation landscape generated by GPR, candidate locations in the latent space can be selected based on different but complementary objectives. On the one hand, the objective in exploitation mode is to select the candidate location most likely to function optimally (e.g., the region in the latent space with the highest mean). On the other hand, the objective in exploration mode is to select the candidate location that best improves the functionality model. These locations may be regions in the latent space with the greatest uncertainty, because the more samples there are in these regions, the narrower the uncertainty becomes and the better the predictive power of the functionality model. Since improving the functionality model can lead to better predictions about which points in the latent space correspond to the optimal functionality, the objectives of exploitation and exploration modes are combined. Furthermore, since numerous candidates are selected during each iteration, both exploitation and exploration objectives can be targeted by selecting a subset of points based on each of these objectives, or a combination thereof.
[0078] The mean value as a function of position can define the functional landscape. As described in more detail below, in utilization mode, the functional landscape can be used to select candidate sequences from regions within the landscape corresponding to peaks of desired functionality. Furthermore, in exploration mode, regions with high uncertainty for exploration can be identified by selecting candidate proteins in these regions and better estimating the degree of desired functionality exhibited in these regions. Both utilization and exploration can be pursued simultaneously by selecting some candidate proteins according to utilization criteria and others according to exploration criteria.
[0079] Furthermore, in step 20, the sequence model can be updated and refined using candidate sequences 25 from previous iterations. For example, candidate sequences 25 from previous iterations can be added to the training data, and the sequence model can be further trained using the expanded / updated training dataset. Updating the sequence model may shift or skew the latent space. Therefore, the sequence model can be updated before computing the functionality model and the goodness-of-fit function 65. In certain implementations, the goodness-of-fit function 65 is the functionality landscape. In other implementations, the goodness-of-fit function 65 includes the functionality landscape combined with other metrics (e.g., stability) that predict which amino acid sequences are promising.
[0080] In certain implementations, the fitness landscape is constructed to map fitness peaks and valleys into the protein sequence space. For example, the descriptor (input) is a protein sequence, and the fitness reading (output) is a measure of a desirable feature (i.e., fitness). In certain implementations, the output may be a multidimensional measure of the desirable feature. These measures of desirable features may include, but are not limited to, activity, similarity to existing native sequences, and protein stability. When the fitness of a protein sequence is multidimensional and includes multiple attributes, a separate landscape can be constructed for each component of the fitness.
[0081] Training data for determining goodness of fitness is obtained from experimental assays on sequences produced through gene synthesis and expression. For example, experimental assays can measure functional activities including, but not limited to, binding, catalytic activity, stability, or substitutes thereof (e.g., growth rate under environmental conditions proportional to one or more of the functional activities). Furthermore, in some implementations, computational modeling can also provide feedback used to determine goodness of fitness; for example, computational modeling of sequences can be used to predict protein stability.
[0082] Using a reduced-dimensional latent space as the domain for input to the goodness-of-fit function may be more effective than using sequence space as the domain. This is because the size of sequence space and the complexity of amino acid interactions make it quite unsuitable to build supervised regression models that take protein sequences directly as input. Therefore, it is preferable to perform regressions by taking a representation of each sequence projected onto a lower-dimensional latent space as input.
[0083] Supervised learning models can be fitted using, but are not limited to, multivariate linear regression, nonlinear regression, support vector regression (SVR), Gaussian process regression (GPR), random forest (RF), and artificial neural networks (ANN).
[0084] In certain implementations, candidate sequences are generated using a two-step approach. In the first step, a sequence-based model (such as SCA, DCA, or VAE) selects a first set of sequences; and in the second step, a goodness-of-fit model selects a subset of the first set of sequences (i.e., sequences corresponding to points close to the peaks of the goodness-of-fit landscape) as candidate sequences based on goodness-of-fit criteria. In other words, the first set of sequences is computationally evaluated according to the goodness-of-fit landscape that defines each component of the goodness-of-fit, and one or more sequences from the first set of sequences generated are identified as top-tier candidates that are passed on to be experimentally synthesized and assayed.
[0085] Multi-purpose / multidimensional optimization in latent space In implementations of multi-objective optimization, the top-performing sequence is identified as one that runs well along all components of the fitness score defined by the various landscapes (such as similarity and stability) discussed above. Since there is no single optimal solution (i.e., a single best sequence) because the fitness score components may serve conflicting objectives (for example, activity may be inversely correlated with stability), the best sequence can be identified using the Pareto frontier, thereby solving the multi-objective optimization problem.
[0086] In other words, multi-objective optimization problems are solved by narrowing down the possible choices for the best sequence to a low-dimensional surface in a multi-dimensional space. In other words, the best sequence lies on the Pareto frontier (also called the optimal frontier or efficient frontier). This frontier and the sequences above it can be identified using, but are not limited to, approaches selected from scalarization, sandwich algorithms, normal boundary crossing (NBI), modified NBI (NBIm), normal constraints (NC), continuous Pareto optimization (SPO), directed search domain (DSD), non-dominant sort genetic algorithm-II (NSGA-II), intensity Pareto evolutionary algorithm 2 (SPEA-2), particle swarm optimization, and simulated annealing.
[0087] For example, one way to perform a multi-objective optimization problem is to maximize all competing components of goodness of fit simultaneously through a Pareto frontier decision. This is essentially an exploitative search, as it focuses on finding the candidate with the highest goodness of fit. In this implementation, the goodness of fit model passes through the first sequence of sequences that the regression model predicts have a high degree of goodness of fit, and blocks through the first sequence of sequences that are not predicted to have a high degree of goodness of fit.
[0088] However, if the selection criterion by the functionality / goodness-of-fit model is only a high degree of goodness-of-fit, undersampled regions of the search space may not be explored because the high uncertainty in these undersampled regions may make it impossible to predict a high degree of goodness-of-fit. Therefore, while the functionality / goodness-of-fit model is solved by multi-objective optimization so as not to maximize each component of goodness-of-fit defined by various regression landscapes, it can also be selected according to a search criterion, which is solved to identify the sequences with the highest uncertainty in the goodness-of-fit prediction model. In other words, the goal of the search criterion is to identify the sequences with the highest uncertainty in the regression model. Based on the search criterion, sequences form an initial set of sequences corresponding to the regions with the most uncertainty and are passed as candidate sequences for experimental synthesis and assay. This search criterion is desirable because collecting experimental data on these sequences is most helpful in improving the model (i.e., reducing uncertainty) when the regression model is highly uncertain about the properties of these sequences. Therefore, exploration is important to provide additional training data to retrain the model and enhance its predictive performance.
[0089] Generally, protocols for carefully selecting sequences based on these utilization / exploration criteria are known as active learning. That is, a machine learning model is instructed to collect new experimental data to (i) guide experiments toward the most promising candidates, and (ii) improve its predictive performance by instructing the collection of new data that is most valuable for retraining the model. In this way, model building and experimental synthesis operate within a positive feedback cycle.
[0090] In a specific implementation of Method 10, the active learning problem is addressed using the multi-objective optimization technique described above, combined with Bayesian optimization techniques, to control the trade-off between exploration and exploitation using an acquisition function that includes, but is not limited to, the probability of improvement, expected improvement, and confidence lower / upper bounds.
[0091] Multi-iteration optimization Figure 2 illustrates the paths that may be traversed in eight iterations of method 10. Region 210(1) represents the locus of points within the set 200 of all possible arrays of length L. In reality, the set 200 can be orders of magnitude larger than the subset 210(i) of candidate arrays in the i-th iteration of method 10. Thus, Figure 2 is not to scale. The subset 210(2) of candidate arrays generated in the second iteration may be shifted with respect to the subset 210(1) of the first iteration, and each subsequent subset 210(i + 1) may be shifted with respect to the previous subset 210(i). Thus, as the iteration number i increases, the subset 210(i) of candidate arrays explores more of the space 200 and evolves towards the peak of the functionality landscape until the desired level of functionality is met.
[0092] Furthermore, method 10 can be used to achieve desired functionality in a new environment. For example, a known enzyme may exist that has a desired catalytic function at temperature X, but a new enzyme with the same catalytic function at temperature Y is desired. To design this new enzyme, a series of intermediate temperatures X < A < B < C < D < Y can be selected. Next, starting from the known enzyme and its homologs, a first set of enzymes can be designed using method 10 to perform an assay at temperature A to measure the desired catalytic function. Next, method 10 can be repeated starting from the first set of enzymes, but now maintaining the assay at temperature B to generate a second set of enzymes that exhibit the desired catalytic function when in an environment at temperature B. This is repeated three more times at temperatures C, D, and Y for the third, fourth, and fifth times, respectively, until the final set of enzymes exhibits the desired catalytic function when at temperature Y. Thus, method 10 can be used to achieve a specific functionality in a new environment (e.g., different temperatures, pressures, light conditions, pH, or different chemical / element concentrations in a solution / environment).
[0093] Recent advances in machine learning have resulted in powerful probabilistic generative models that, after being trained on real-world examples, can generate realistic synthetic samples. Such models typically also generate low-dimensional continuous representations of the data being modeled, enabling interpolation or analogical inference. As discussed above, these generative models are applicable to protein design, for example, by using a pair of deep networks trained as autoencoders to convert proteins, represented as amino acid sequences, into continuous vector representations.
[0094] Supervised learning and unsupervised learning As described in various embodiments of this specification, machine learning models can be trained using supervised or unsupervised machine learning programs or algorithms. Machine learning programs or algorithms may employ neural networks, which may be convolutional neural networks, deep learning neural networks, or composite learning modules or programs that learn on two or more features or feature datasets of a particular area of interest. Machine learning programs or algorithms may also include natural language processing, semantic analysis, automated reasoning, regression analysis, support vector machine (SVM) analysis, decision tree analysis, random forest analysis, K-nearest neighbor analysis, naive Bayes analysis, clustering, reinforcement learning, and / or other machine learning algorithms and / or techniques. Machine learning may involve identifying and recognizing patterns in existing data (such as candidate proteins in a training dataset of protein amino acid sequences) to facilitate prediction, classification, or output of subsequent data (e.g., synthesizing candidate genes and producing candidate proteins corresponding to each candidate amino acid sequence).
[0095] Machine learning models (or models) described herein are created and trained on example inputs or data (sometimes called "features" and "labels") (e.g., "training data") to make valid and reliable predictions for new inputs, such as test-level or production-level data or inputs. In supervised machine learning, a machine learning program running on a server, computing device, or processor (or processor) is provided with example inputs (e.g., "features") and their associated or observed outputs (e.g., "labels") to determine or discover rules, relationships, patterns, or machine learning "models" that map such inputs (e.g., "features") to outputs (e.g., labels), for example, by determining and / or assigning weights or other metrics to the model across various feature categories. Such rules, relationships, or models are then provided with subsequent inputs, and a model running on a server, computing device, or processor (or processor) can predict, classify, or output expected outputs based on the discovered rules, relationships, or models.
[0096] In unsupervised machine learning, a server, computing device, or processor may need to discover its own structure in unlabeled input examples. For example, a server, computing device, or processor may perform multiple training iterations to train multiple generations of models until a satisfactory model is produced, such as a model that provides sufficient predictive accuracy when given test-level or production-level data or input. The disclosures herein may utilize one or both of such supervised or unsupervised machine learning techniques.
[0097] Figures 3A, 3B, and 3C illustrate another non-restrictive implementation of Method 10. For example, Figures 3A, 3B, and 3C illustrate non-restrictive examples of sequence models (i.e., unsupervised models) that are VAEs or RBMs. Method 10 can be divided into two parts: (i) an unsupervised learning process 102, and (ii) a supervised learning process 138. The unsupervised learning process 102 starts with a training dataset 105, from which it generates candidate sequences 135 that are likely to exhibit a given quantitative urgent need (e.g., have the desired functionality). In the unsupervised learning process 102, a machine learning model 115 is trained to map back and forth between visible variables (i.e., the amino acid sequence of a protein) and hidden variables that define the latent space. The machine learning model 115 may be a generative model. The latent space has a reduced dimension compared to the dimension of the amino acid sequence (e.g., an amino acid sequence of length N has a dimension of 20 to occupy 20 possible amino acids in each of the N residues). N (It can have...). To map from a high-dimensional space to a low-dimensional space, the machine learning model 115 learns implicit patterns and correlations between residues of each amino acid sequence in the training dataset to encode the information more compactly. In other words, the mapping from visible variables to hidden variable correlations compresses the information and represents this information more compactly using reduced dimensions. Thus, patterns and correlations between amino acid sequences in the training dataset are learned and implicitly represented by the machine learning model 115.
[0098] Furthermore, a machine learning model 115 can be used to generate new amino acid sequences. A set of proteins with similar functionality can define clusters of points when mapped into latent space. For example, the mean and variance of this cluster can be used to define a multivariate Gaussian distribution to represent the probability density function (pdf) of the cluster. By randomly selecting points from this and then using the machine learning model 115 to map these points back into amino acid sequences, the machine learning model 115 can generate candidate sequences 135 of new synthetic proteins that are likely to have a similar structure to the original proteins used to define the cluster, and therefore similar functionality.
[0099] The supervised learning process 138 starts with a candidate sequence 135, synthesizes a gene sequence, and produces a candidate protein having the candidate sequence 135. Next, the properties of these candidate proteins are evaluated. For example, the candidate proteins can be assayed to measure the extent to which they exhibit desired functionality (e.g., urgent needs). Then, a goodness-of-fit function 145 is determined from the measured values 142, and an iterative cycle is started to search for better proteins by iteratively updating the machine learning model 115, the candidate sequence 135, the measured values 142, and the goodness-of-fit function 145 until some predefined stopping criteria are reached.
[0100] With respect to desired functionality, proteins exhibit the following properties, which, in various combinations, can contribute to the desired design objectives: (i) folding rate and yield (the rate and probability of folding), (ii) thermodynamic stability (the difference in free energy between the unfolded and folded states), (iii) binding affinity (the free energy separating the unbound and bound states), (iv) binding specificity (the difference in binding between desirable and undesirable substrates), (v) catalytic activity (the free energy separating the enzyme-substrate complex from the transition state of the reaction), (vi) allostery (the long-range transfer of amino acids within the protein), and (vii) evolvability (the ability to generate genetic mutations).
[0101] As discussed above, high-throughput assays can be tuned to generate measurements for evaluating desired functionality. This specification provides various examples of assay types corresponding to the respective properties of proteins. Folding rate and yield can be assessed, for example, by gel filtration chromatography or assays aimed at tracking compactness by fluorescence of environmentally sensitive fluorophores. Thermodynamic stability can be measured, for example, using differential scanning calorimetry or 1H-15N HSQC NMR. Binding can be measured by fluorescence or titration calorimetry (e.g., using droplet microfluidics) or by two-hybrid gene expression methods in cells. Catalytic activity can be measured by absorbance / fluorescence methods (e.g., using droplet microfluidics) or by viability or growth rate measurements. Allostery can be measured in various ways, such as by measuring allostery through intracellular regulation. Evolutionary potential can be measured by comparing the sensitivity of amino acids to deep mutation scanning in the context of both current and novel functions.
[0102] In certain implementations, the approach used in Method 10 can be divided into two parts: (i) unsupervised learning of low-dimensional latent space embeddings of protein sequences, and (ii) supervised learning of the functional landscape within the latent space. In other words, Method 10 performs a search to find low-dimensional representations of protein sequences that hold important information about sequence and function, and from these representations predict the potential functionality of new sequences that have even better performance.
[0103] In step 110 of Method 10, the machine learning model 115 is trained using the training dataset 105. In one unrestricted example, the machine learning model 115 is a variational autoencoder (VAE). In another unrestricted example, the machine learning model 115 is a restricted Boltzmann machine (RBM). Both VAEs and RBMs are generative models, and as can be understood by those skilled in the art, other generative models can be used in the machine learning model 115. As discussed above, in addition to VAEs and RBMs, unsupervised learning can be performed using generative adversarial networks (GANs), statistical combined analysis (SCA), or direct combined analysis (DCA). The decision of which unsupervised learning method to use can be based on an empirical assessment of which method provides the best performance in providing a robust latent space representation and accurate generative performance.
[0104] Statistical Joint Analysis (SCA) Model In one non-restrictive implementation, a statistical joint analysis (SCA) model can be used as a sequence-based model, and the SCA model is computed using the training data. This is defined by a conserved weighted correlation matrix (e.g., an SCA matrix of co-evolution between all pairs of amino acids), as described in KARéynolds, et al., “Evolution-Based Design of Proteins,” Methods in Enzymology, vol. 523, pp. 213-235 (2013) and discussed in O. Rivoire, et al., “Evolution-Based Functional Decomposition of Proteins,” PLoS Comput Biol, vol. 12, p. e1004817 (2016) (each of these is incorporated herein by reference in its entirety). In other words, the information in the training dataset is compressed by the SCA model into a single pairwise correlation matrix. Furthermore, singular value decomposition or eigenvalue decomposition of the SCA matrix reveals that while most modes are indistinguishable from sampling noise, some of the higher-level modes (corresponding to latent space) capture statistically significant correlations. These higher-level modes define sectors containing one or more groups of collectively evolving amino acids. SCA-based protein design can be performed by computational simulations that start with a random sequence and evolve a synthetic sequence (in silico) constrained by the observed evolutionary statistics captured in the SCA coevolution matrix.
[0105] For example, SCA-based protein design uses the Metropolis Monte Carlo Simulated Annealing (MCSA) algorithm to search for a sequence space that is consistent with a set of constraints applied between amino acids. The MCSA algorithm is an iterative numerical method for finding the global minimum energy configuration of a system starting from an arbitrary state, and is particularly useful when the number of possible states is very large, the energy landscape is steep, and it is characterized by many local minima. The energy function to be minimized (or "objective function") can generally depend on many parameters of the system and represents constraints that define the size and shape of the final solution space. Essentially, the objective function can be thought of as a set of applied constraints used to test the hypotheses being tested—namely, folding, thermodynamic stability, function, and any other aspects of protein fitness.
[0106] In SCA-based protein design, the system under consideration is MSA (not a single sequence), and the objective function (E) is the sum difference between the correlation of the MSAs of the protein sequences during the iteration of the design process and the target correlation matrix estimated from the native MSAs, provided, for example, by the following equation:
number
[0107] With respect to the above equation, the weighted correlation tensor is given, for example, by the following equation:
number
[0108] In the above formula,
number
number
number
number
number
number
number
number
[0109] Variational Autoencoder (VAE) Model An example of using a VAE as the machine learning model 115 is provided here, and other embodiments in which other types of machine learning are used as the machine learning model 115 will be provided later.
[0110] In one non-restrictive example, machine learning model 115 could be a VAE employing three encoding layers and three decoding layers, which employ batch normalization and dropout per layer, a Tanh activation function, and a Softmax output layer. The VAE neural network can be trained using a training dataset containing order 1,000 protein sequences under one-hot amino acid encoding, employing an Adam optimizer during training, employing a loss function that is the sum of binary cross-entropy and Kullback-Leibler divergence. For example, it should be understood that the protein sequences in the training dataset can include a variety of numbers and counts, such as at least 100 protein sequences, at least 500 protein sequences, at least 1,000 protein sequences, at least 1,500 protein sequences, etc. To prevent overfitting, training is terminated by early stopping when the loss function of the holdout validation partition no longer decreases. The number of VAEs is optimized with various train / test partitions. Furthermore, the dimensionality of the latent space is optimized to select a latent space size where the validation loss no longer decreases as the dimensionality increases. The optimal VAE is validated by the accuracy of reconstructions on test partitions that were not used in training.
[0111] Using a trained model, millions of new sequences can be efficiently generated by sampling the latent space with Gaussian random numbers and converting them into protein sequences through a decoder. Experiments have shown that the latent space encodes natural rules regarding surviving proteins, so it is expected that samples from this latent space will also produce surviving proteins, some of which have not been produced by natural selection.
[0112] Figures 4 and 5 show schematic diagrams of the VAE used as a generative model. In Figure 4, the VAE includes an encoder section, a decoder section, and an autoregressive section, as discussed in Costello, Z. and Garcia Martin, H., “How to hallucinate functional proteins,” Preprint at https: / / arxiv.org / abs / 1903(2019) (which is incorporated in its entirety herein by reference). Within each module, the shaded cubes represent the type of layer, respectively. The shaded layers represent 1D expanding convolutional layers with skip connections in the style of a residual network (resnet). The darkness of the shaded layer indicates the magnitude of the expansion. Gradually darker shading, or cross-hatching, indicates a larger expansion, which is done in the pattern used in Figure 4. The end layers (e.g., end layers 400re1 and 400re2) represent 1D convolutions where the input length is halved with a stride of 2 and the number of channels doubles. The transposed end layers (e.g., end layers 400ge1 and 400ge2) exhibit the inverse operation of the end layers (e.g., end layers 400re1 and 400re2) via transposed one-dimensional (1D) stride convolution.
[0113] Generative models such as VAEs generate data that has the same statistical properties as the data they were trained on. That is, when a VAE is trained on functional protein sequences that fold in their native host, the VAE should produce proteins that are likely to fold and function similarly to the proteins in the training dataset. Therefore, we can verify that the model behaves in a manner consistent with this assumption.
[0114] A VAE is trained to reconstruct its own input. For example, a VAE first encodes a protein sequence into feature vectors (i.e., vectors in the latent space). Feature vectors can be thought of as summaries of the key information in the protein sequence. The vector space in which these feature vectors reside is generally known as the latent space. Next, a variational autoencoder reconstructs the original protein sequence from these feature vectors. The loss function can represent the degree to which the output from the network matches the input. If the VAE is lossless, the output will always match the input. However, generally, VAEs have some degree of loss.
[0115] A well-trained VAE can be used in two modes: as an encoder or as a decoder. Using the encoder, a protein sequence can be taken and its associated feature vectors can be found. These feature vectors can then be used for downstream classification or regression tasks. For example, it can be used to determine where a given protein is likely to localize, given its sequence. Using the decoder, any sequence that is likely to fold and function can be generated by sampling from latent space. Furthermore, latent space samples can be selected so that the sequence is also likely to have the desired phenotype. The remainder of this section details the design and use of the model.
[0116] As shown in Figure 5, various groups of proteins ("A", "B", "C", "D", "E", "F", "G", etc.) are localized and clustered in various regions within the latent space. The latent space also has the advantage of being a continuous space, in contrast to the discretized space of amino acid sequences. As discussed in R. Gomez-Bombarelli et al., "Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules," ACS Cent.Sci. Vol.4, pp.268-276 (2018) and S. Sinai et al., "Variational auto-encoding of protein sequences" (reprinted at https: / / arxiv.org / abs / 1712.03346 (2018)) (both of which are incorporated in their entirety herein and are incorporated in their entirety herein by reference), there are several advantages to continuous, data-driven protein representation methods. First, new compounds can be automatically generated by modifying and then decoding the vector representation, eliminating the need for manually specified mutation rules. Secondly, using differentiable models that map protein representations to desired properties allows for larger steps when searching the protein space using gradient-based optimization. Gradient-based optimization can be combined with Bayesian inference methods to select amino acid sequences that are likely to be beneficial for global optimization. Thirdly, data-driven representations can leverage large sets of proteins (e.g., including proteins that do not exhibit the desired functionality as well as those that do) to automatically build a larger implicit library, and then use a smaller set of proteins that exhibit the desired functionality to build regression models that are incorporated into a goodness-of-fit function (145) from continuous representations to desired properties. Therefore, even if many proteins have unknown properties, large protein databases can be used to train VAEs.
[0117] Limited Boltzmann Machine (RBM) Model A non-restrictive example of a machine learning model for performing unsupervised learning is the RBM (sketched schematically in Figure 3C). In the encoding direction, amino acid sequences are applied as input to the visible layer of neuron nodes, and the hidden layer is computed by taking a biased sigmoid function of the weighted sum of the values input to the visible layer. The RBM is a variation of the Boltzmann machine, with the constraint that the neurons must form a bipartite graph. Pairs of nodes from each of the two unit groups (generally called visible and hidden units, respectively) may have symmetric connections between them, and no connections between nodes within a group. This constraint allows for training algorithms that are more efficient than those available in the general class of Boltzmann machines, particularly gradient-based contrastive divergence algorithms. Restricted Boltzmann machines can also be used in deep learning networks as deep belief networks by stacking RBMs and, optionally, fine-tuning the resulting deep network with gradient descent and backpropagation.
[0118] A restricted Boltzmann machine is trained to maximize the product of probabilities assigned to some training set V (a matrix, where each row is treated as a visible vector v), as provided by the following equation:
number
[0119] Alternatively, and equally, to maximize the expected log probability of the training sample v, we can use, for example, the following formula:
number
[0120] In the above equation(s), P(v) = Σ h P(v,h)=Z -1 Σ hexp(-E(v,h)) is the marginal probability of the boolean visible vector summed over all possible hidden layer configurations, where Z is the partition function that normalizes the probability distribution P(v,h), and the energy function is E(v,h) = -a T v -- b T h - v T Wh, where W is the weight matrix associated with the connections between the values of the hidden units h and the visible units v, a is the bias weight (offset) of the visible units, and b is the bias weight (offset) of the hidden units.
[0121] The RBM can be trained to optimize the weights W between nodes using Contrastive Divergence (CD). That is, Gibbs sampling is used within a gradient descent procedure (similar to how backpropagation is used within such a procedure when training a feedforward neural network), to calculate the weight updates. In a single-step contrastive divergence procedure, the following steps are performed: (i) take a training sample v, calculate the probabilities of the hidden units, and sample a hidden activation vector h from this probability distribution, (ii) calculate the outer product of v and h, which is called the positive gradient, (iii) sample a reconstruction v’ of the visible units from h, and then resample a hidden activation h’ from this. (Gibbs sampling step), (iv) calculate the outer product of v’ and h’, which is called the negative gradient, (v) update the weight matrix W to be the learning rate times the difference between the positive gradient and the negative gradient. For example, it is provided by the following equation: ΔW = ε(vh T - v’h’ T )[[ID=I7]]
[0122] Furthermore, step (vi) may include updating the biases a and b in a similar manner, for example, as provided by the following equations. Δa = ε(v - v’) and Δb = ε(h - h’).
[0123] The hidden layer defines the latent space, and the candidate amino acid sequence is generated by selecting values from the nodes of the hidden layer and applying RBM to decode the values of the hidden layer into an amino acid sequence as the candidate sequence. Variations of this method can also be used, as discussed in Tubiana et al., “Learning protein constitutive motifs from sequence data,” eLife, vol.8, p.e39397 (2019) (which is incorporated herein by reference in its entirety).
[0124] General explanation of the method Returning to Figure 3A, in step 120 of Method 10, candidate points are selected within the latent space. For example, k-means clustering can be used to identify regions / areas within the latent space that are likely to correspond to proteins with the desired functionality / attributes. Alternatively, statistical analysis can be performed to determine the probability density function (pdf) of the desired functionality / attributes. Then, a random number generator can be used to select a statistically representative sample of points within the latent space. Other methods can also be used to determine a sample of points within the latent space that are likely to correspond to proteins with the desired functionality.
[0125] In step 130 of Method 10, candidate amino acid sequences are selected using the machine learning model 115. These candidate sequences are selected based on their similarity to sequences in the training dataset. For example, the training dataset may have a subset that exhibits the desired functionality, and this subset can be clustered to identify a specific region in the latent space. Candidate sequences can then be selected based on the identified region. For example, since the machine learning model 115 is a generative model, points within or adjacent to the identified region can be mapped to amino acid sequences to be used as candidate sequences.
[0126] The search for higher-performing amino acid sequences may have competing / complementary purposes for exploration and utilization. Considering this, the choice of how much to deviate from / differ from amino acid sequences identified as high-performing may depend on whether exploration or utilization is more desired at a given stage of the search.
[0127] In step 140 of Method 10, a gene sequence is synthesized and then used to generate a protein having the candidate amino acid sequence from step 130. The generated protein is then assayed / evaluated to determine its functionality / attributes. The values representing the functionality / attributes of the generated protein are then passed to step 150, where a fitness function is generated to guide the selection of future candidate sequences.
[0128] The synthesis and assay methods are described in more detail herein.
[0129] In step 150 of Method 10, the goodness-of-fit function is determined from the values measured in step 140. For example, the goodness-of-fit function can be determined by performing a regression analysis on the measured values in the latent space. In certain implementations, Gaussian process regression is used to generate the goodness-of-fit function from the measured values. Other non-limiting examples of methods for determining the goodness-of-fit function are discussed herein.
[0130] In process 160 of method 10, the optimization search for candidate sequences continues using an iterative loop, as shown in Figures 2 and 3.
[0131] In step 162 of process 160, the machine learning model is updated using candidate sequences generated in step 130 or in previous iterations of the loop in process 160. By expanding the number of amino acid sequences in the training dataset, the machine learning model 115 can be further trained to improve and enhance its performance.
[0132] In step 164 of process 160, a new candidate sequence is selected based on the machine learning model 115 and the goodness-of-fit function.
[0133] In step 166 of process 160, a gene sequence is synthesized and then used to generate a protein having the new candidate sequence from step 164. The generated protein is then evaluated to measure new values representing the desired functionality. For example, the generated protein can be assayed to measure the critical requirements corresponding to the functionality.
[0134] In step 168 of process 160, various stopping criteria can be evaluated to determine whether the stopping criteria have been met. For example, the stopping criteria may include whether the number of iterations exceeds or is equal to a predetermined maximum number of iterations. Alternatively, the stopping criteria may include whether the functionality of a predetermined number of candidate sequences exceeds a predefined functionality threshold. Alternatively, the stopping criteria may include whether, with each iteration, the rate of improvement in the functionality of the candidate sequences has decreased or converged, and as a result, the rate of improvement has fallen below a predefined improvement threshold.
[0135] If the termination criteria are met, process 160 proceeds to step 172, where the sequence(s) of the most functional protein candidate(s) is saved and / or presented to the user. Otherwise, process 160 proceeds to step 170.
[0136] In step 170 of process 160, the goodness-of-fit function is updated using the new measured values from step 170. For example, the goodness-of-fit function may be updated using a regression analysis of the new measured values. The regression analysis performed can use all measured values of all candidate sequences from the current iteration and all previous iterations.
[0137] The order of steps 162, 164, 166, 168, and 170 can be changed without deviating from the purpose of process 160. For example, the query to determine whether the stopping criteria are met can be performed between different pairs of steps other than steps 166 and 170. Furthermore, step 162 can be omitted in certain iterations. For example, the goodness-of-fit function can be updated for each iteration of the loop in process 160, keeping the machine learning method 115 constant. By improving the goodness-of-fit function with each iteration, the landscape of desired functionality as a function of the latent space can be learned and improved to better select candidates with the desired functionality.
[0138] Furthermore, the parameters for selecting candidates may differ between iterations. For example, exploration may be prioritized over utilization early in the search. Subsequently, candidates can be selected from a broader distribution. This strategy can be useful in avoiding getting stuck at maxima in the functional landscape. That is, the functional landscape has peaks and values, and a global optimization method that encourages exploration can avoid repeating to amino acid sequences of peaks that are maxima but not maximum values. One example of a global optimization method is simulated annealing.
[0139] Other factors besides the functional landscape may be important in selecting the optimal candidate amino acid sequence. For example, these factors may include similarity and stability. Of the millions of possible sequences generated from latent space sampling, the optimal selection method focuses on those most likely to possess the desired properties. In effect, iterative process 160 performs in silico mutagenesis; that is, process 160 performs computational natural selection on sequences generated by the unsupervised model to select sequences predicted to be stable and highly functional. This is achieved by assigning each sequence generated by latent space decoding to a score associated with its desirability.
[0140] In certain implementations, candidate sequence selection is performed using a scoring vector that includes three components: (i) similarity with respect to known native sequences – new sequences that are close to native sequences are expected to have a higher probability of retaining function; (ii) stability predicted by computational modeling using software tools that predict protein structure – sequences predicted to maintain the original structure of the native folding are likely to be more functional; and (iii) functionality measured by experimental assays (e.g., measured values from step 166). The first two scores (i.e., similarity and stability) can be determined computationally with high throughput for each predicted sequence. In certain implementations, the third score is determined approximately by fitting a supervised regression model to all proteins previously synthesized and experimentally assayed (e.g., in step 140 or step 166). The supervised regression model can be thought of as a landscape in latent space. Supervised learning models can be fitted using, but are not limited to, multivariate linear regression, support vector regression (SVR), Gaussian process regression (GPR), random forest (RF), and artificial neural networks (ANN). By scoring each predicted sequence along these three metrics, these sequences can be ranked, and the sequences expected to be the most promising can be selected for experimental synthesis.
[0141] For example, candidate sequence selection can be performed by identifying the Pareto frontier in this 3D multidimensional optimization space and selecting sequences on the frontier as sequences proposed for synthesis. In certain implementations, the subset of sequences on the frontier can be further refined by assigning adjustable weights to three optimization criteria. For example, higher weights associated with stability and homology to known natural sequences may present a more conservative set of candidates, while lower weights allow for more ambitious candidates that are further removed from natural sequences. In other words, these lower weights for similarity and stability may work to the advantage of the search. As the model becomes more accurate through multiple iterations, the reliability of the model tends to increase, allowing for the selection of more ambitious sequences that are further removed from nature.
[0142] For example, the goal of a protein might be to exhibit desired functionality in unnatural environments (e.g., high temperature, high pressure). In that case, it might be reasonable to assume that the protein's structure deviates from what nature has chosen as optimal at room temperature and room pressure. In other words, more ambitious sequences may be advantageous for "computationally evolving" sequences toward functionality in unnatural environments (e.g., high temperature, high pressure) that are valuable for industrial applications but would never be chosen in nature, where most operations operate at room temperature and room pressure.
[0143] Furthermore, regression models can account for uncertainties in predicting functionality. These uncertainties can be leveraged—used in search modalities—to balance the competing interest of synthesizing the most promising sequences identified by the model with the interest of synthesizing sequences to explore a search space where the model has high uncertainty and new experimental data contribute most to improving the model.
[0144] Process 160, using favorable feedback, improves and enhances the predictive power of the goodness-of-fit function in conjunction with the machine learning model 115. For example, sequences produced and assayed in steps 140 and 166 are fed back into supervised and unsupervised learning models (e.g., the goodness-of-fit function and machine learning model 115). The machine learning model 115 can learn better latent spatial representations of a wider variety of proteins. Furthermore, the goodness-of-fit function becomes a better predictor of the function. In this way, the computational model becomes increasingly powerful with each iterative round of sequence production and testing. Moreover, data-driven representations can be used in iterative processes incorporating a new set of candidate proteins to automatically build a larger implicit library, and then use a smaller set of labeled examples to build regression models from continuous representations to desired properties.
[0145] gene synthesis Regarding gene synthesis in steps 140 and 166, the process can be used to synthesize candidate amino acid sequences using tube-synthesized (high-purity) oligonucleotides and automated bacterial cloning techniques. However, this process is generally costly. Therefore, an improved, less expensive gene synthesis process is desired.
[0146] One less expensive approach (e.g., producing genes approximately 500 bp long, i.e., thousands of sequences, on a large scale at an amortization cost of about $2 per gene) involves using oligonucleotide barcode-functionalized beads to sequester copies of all the oligonucleotides required for a given gene into a single droplet of an oil-in-water emulsion, followed by polymerase cycling assembly (PCA) of overlapping oligonucleotides into a full-length gene, as described in W. Stemmer, et al., “Single-Step Assembly of a Gene and Entire Plasmid From Large Numbers of Oligodeoxyribonucleotides,” Gene. 164 1995) 49-53 and C. Plesa, et al., “Multiplexed gene synthesis in emulsions for exploring protein functional landscapes,” Science. 359 (2018) 343-347 (both of which are incorporated herein by reference in their entirety). This approach has a high error rate (e.g., about 5% correct amino acid sequence), but it can produce large-scale genes up to about 500 bp in length (i.e., thousands of sequences) at an amortization cost of, for example, about $2 per gene. This method is referred to herein as the bead-and-barcode method.
[0147] Another, even less expensive approach starts with longer oligonucleotides and uses a separate minipool, thus eliminating the need for bead hybridization in the bead-and-barcode method. This method is referred to herein as the minipool method. In the minipool method, pre-commercial oligonucleotides are synthesized in array form, resulting in two significant improvements over previously available products. First, pre-commercial oligonucleotides are synthesized in lengths up to 300 nt. Second, pre-commercial oligonucleotides are synthesized with an error rate of approximately 1:1300. This increase in length (compared to the previously available 200 nt length) reduces the complexity of the assembly reaction, as 60–80 nt of each oligonucleotide are used for "overhead" sequences that do not contribute to the final gene product. Thus, the available effective sequence per oligonucleotide increases by 75% (approximately 130–230 nt) with 300 nt oligonucleotides, allowing for the assembly of a 1 kb gene with a similar number of oligonucleotides to those previously used in pre-commercial oligonucleotides to make a 500 bp gene.
[0148] Since single-base insertions and deletions of oligonucleotides are the main cause of sequence errors, a lower error rate results in fewer sequence errors in the assembled gene. Furthermore, by providing oligonucleotides as separate "minipools" containing only the oligonucleotides required for each gene, the bead hybridization step in the bead-and-barcode method can be omitted, simultaneously reducing cost and complexity.
[0149] Errors in oligonucleotide synthesis are not randomly distributed but are related to their sequences. For example, purine bases are more susceptible to degradation during synthesis than pyrimidines, and the formation of a compact, folded structure by the growing oligonucleotide chain hinders subsequent nucleotide addition. The flexibility of the genetic code (multiple codons encoding the same amino acid, as shown in Figure 7) allows for the design of oligonucleotides with improved synthesis accuracy. The selection of codons to use to encode specific amino acids can be guided by measuring the accuracy of delivered oligonucleotides using high-throughput (HT) sequencing, identifying sequence patterns associated with performance degradation, and updating oligonucleotide design algorithms to avoid them.
[0150] Regarding the optimization of polymerase cycling assembly (PCA), PCA is similar to the polymerase chain reaction (PCR) technique, but as shown in the schematic diagrams in Figures 6A and 6B, it uses overlapping sequences as primers to anneal the overlapping sequences, followed by polymerase extension, thereby assembling oligonucleotides into a larger gene and generating / synthesizing a gene sequence that expresses a candidate amino acid sequence. Consequently, the design of overlapping sequences is crucial for successful assembly, just as good primers are required for successful PCR amplification. Overlapping sequences are orthogonal to other overlaps within the gene but anneal to each other at similar temperatures. The amino acid sequence of the gene limits the freedom to choose overlaps, but provides limited freedom, as described in previous sections. Furthermore, both of these parameters (i.e., annealing temperature and orthogonality between overlaps) can be optimized to select breakpoints between oligonucleotides in order to achieve highly efficient assembly.
[0151] Variations of gene sequence synthesis methods are within the scope of the methods disclosed herein. For example, gene synthesis can be performed using ligase chain reaction methods, thermodynamically balanced inside-out synthesis, and ligation-based gene synthesis, as well as various error correction methods (e.g., Twebak, annealing, and repair).
[0152] Direct Joint Analysis (DCA) Model As discussed above, Method 10 can be performed using different types of machine learning models 115 (also referred to as unsupervised learning models 115 in certain embodiments). Here, a non-limiting embodiment is provided using DCA for the unsupervised learning model 115.
[0153] Figure 8A shows an evolution-inspired approach to protein design based on direct link analysis (DCA), a method originally devised to predict amino acid contacts in the three-dimensional structure of proteins. Generally, the algorithm begins with multiple sequence alignment of native homologs, from which empirical primary and secondary statistics are calculated. These quantities are used to determine amino acid (h i ) and pairwise interaction (J ij We learn a minimal statistical model consisting of inherent constraints on ). Then, we can use the statistical model to generate more artificial sequences that summarize natural statistics and to screen for desired activity.
[0154] Figure 12 shows a portion of the bacterial and fungal shikimic acid pathway, which leads to the biosynthesis of the aromatic amino acids tyrosine and phenylalanine, with the AroQ family of chorismic acid mutases (CMs) operating at the branching point.
[0155] Figure 8C shows the atomic structure of E. coli CM, dimers with two functional active sites (e.g., items 800c1, 800c2, and 800c3). Each active site consists of amino acids contributed by both protomers. Binding substrate analogs are shown with magenta stick bonds.
[0156] The starting point was the large and diverse multiple sequence alignments (MSAs) of the protein family, from which the observed frequencies of all amino acids were determined.
number
number
[0157] In the above model, P is the amino acid sequence (σ1,···,σ L This is the probability that ) occurs, L is the length of the protein, and H(σ1,···,σ L )=Σ i h i σ i +Σ i<j J ij σ i σ jThis is the statistical energy (or Hamiltonian) that provides a quantitative score of likelihood for each sequence. Lower energies are associated with higher probabilities, and Monte Carlo sampling allows for the generation of a repertoire of non-natural sequences, which can then be screened for desired functional activity. If pairwise correlations are generally sufficient to capture the informational content of protein sequences and the model's inferences are sufficiently accurate, synthetic sequences should outline the functional diversity and properties of native proteins.
[0158] This DCA implementation of Method 10 is illustrated here with a non-restrictive example in which a DCA unsupervised learning model 115 is trained using multiple sequence alignment (MSA) of colismic acid mutase (CM) homologs.
[0159] To illustrate Method 10 using DCA as a sequence model, we performed the method using the AroQ family of chorismic acid mutases (CMs), a classic model for understanding the principles of catalysis and enzyme design. These enzymes are found in bacteria, plants, and fungi and operate at the branching point of the shikimic acid pathway, which leads to the biosynthesis of tyrosine and phenylalanine (as illustrated in Figure 12). CMs catalyze the conversion of the intermediate metabolite chorismic acid to prephenic acid via the Claisen rearrangement, exhibiting a rate acceleration of more than one million times this reaction, which is necessary for growth in bacterial cells. For example, E. coli strains lacking CMs are nutritionally dependent on tyrosine and phenylalanine, and both the degree of supplementation of these amino acids and the expression level of CMs quantitatively determine the growth rate. Structurally, AroQ CMs form domain-swapped dimers of relatively small protomers (approximately 100 amino acids, Figure 1e). This, along with the requirements for bacterial growth and the availability of excellent biochemical assays, makes them excellent design targets for testing the power of statistical models inferred from MSA.
[0160] First, an MSA containing numerous sequences is created. In one implementation, sequences are obtained using residues 1–95 of the E. coli P protein as a starting query for 3 rounds of PSI-Blast (ref) and 1e–4 e-score cutoff. Starting with the structural alignments of PDB entries 1ECM, 2D8E, 3NVT, and 1YBZ, an initial alignment is generated by iteratively generating alignment profiles and aligning the nearest neighbor sequences from the PSI-Blast results to this profile using muscle (ref). The resulting alignment is adjusted and trimmed to match the region found in 1ECM, removing short sequences (less than 82 residues), adding sequences with poorly represented gaps (<30% occupancy), and reducing redundancy (>90% top-hit identity). In this way, an MSA of 1,259 sequences is created. This serves as input for sequence design and for testing experimental procedures for functional sequences. As will be understood by those skilled in the art, by using variations in the procedure, other MSAs of other homologs can be generated in a similar manner.
[0161] Next, a DCA analysis is performed using MSA. For example, a Potts model that assigns probabilities can be inferred using MSA. For example, it would look like this:
number
[0162] Such probabilities are calculated for each aligned sequence of L = 95 amino acids or alignment gaps.
number
number
[0163] The statistical energy (or Hamiltonian) of the Potts model is the direct co-evolutionary bond J between amino acids a and b at positions i and j. ij (a, b), and bias (or field) regarding the use of amino acid a at position i h i (a) is given from this perspective. These parameters are bmDCA(reweighting threshold 0.8, regularization intensity 10 -2 and 10 -3 It is inferred using ). For example,
number
[0164] The objective of the resulting model is to determine the empirical proportion f of the natural sequence having amino acid a at position i. i (a) and the percentage of sequences having amino acids a and b at positions i and j simultaneously f ij The goal is to accurately reproduce (a,b) and therefore combine residue conservation and covariance. For example:
number
[0165] To check the accuracy of the inferred model, the sequence statistics of the natural sequence are calculated using the formula:
number
[0166] Equations [1] and [2] describe, respectively, some of the empirical proportions of 2 and 3 residues, as referenced above, but these cannot be explained by lower-order statistics. Therefore, these are essentially f ij (a,b) and f ijk This is more difficult to reproduce than (a,b,c) and constitutes a more rigorous test of the model's accuracy.
[0167] Regarding the use of regularization, statistical models
number
number
[0168] In this way, a penalty value, such as the one output by the above formula, can be added to the likelihood of the data. This penalty systematically reduces the parameter values of the bmDCA inference, thereby avoiding extremely large parameter values due to undersampled rare events. This improves the model.
number
number
[0169] Statistical energy of the natural sequence (NAT-frequency count f)
number
number
number
number
[0170] In other words, the model shows that the natural sequence has systematically lower energy and therefore higher probability than the sampled sequence, which can be represented, for example, as follows:
number
[0171] To overcome this gap, a lower temperature T<1 is introduced, forcing MCMCs to sample at lower statistical energies compatible with the natural sequence.
[0172] Once the DCA model is trained, it can be used to generate new candidate sequences. Using MCMC sampling,
number
number
number
number
[0173] Using DCA, we created a statistical model for the alignment of 1259 native AroQ CM enzymes, encompassing a broad diversity of bacterial and fungal lineages. Technically, the MSA(f) of any protein was used. i ,f ij From the statistics observed in ), the parameter (h i ,J ij Estimating the parameters (h) of the DCA model is computationally difficult by direct means, but many approximation algorithms are possible. Here, the bmDCA approximation algorithm is used, which is computationally expensive but is a very accurate method based on Boltzmann machine learning. In other implementations, the parameters (h) of the DCA model are estimated. i ,J ij The result can be obtained using the mean-field method, the Monte Carlo gradient descent method, or the pseudo-likelihood maximization method.
[0174] Figures 9A and 9B show sequences sampled from the bmDCA model, summarizing empirical first-order and second-order MSA statistics, respectively. These results demonstrate that the model fit well. Figure 9C shows that the bmDCA model also summarizes third-order correlations of MSAs that were not used in model training, indicating that the model is statistically complete.
[0175] Figure 9D shows the top two principal components of the distance matrix between all native CM sequences in the MSA (e.g., the circular item or point shaded as item 900d1). Here, the structure of the sequence space spanned by the CM family is visualized. Sequences derived from the bmDCA model (e.g., the circular item or point shaded as item 900d2) also fill the sequence space in a manner consistent with the native CM sequences. The location of the E. coli CM sequence is indicated by point 900ds.
[0176] Therefore, it is clear that the sequences generated by Monte Carlo sampling from the model reproduce the empirical primary and secondary statistics of the native sequences used for fitting. The statistics are shown in Figures 9A, 9B, and 9C, representing primary, secondary, and tertiary statistics, respectively. Furthermore, it can be observed that the model outlines higher-order statistical features of the MSA that were not used in the model's inference. These include three-directional residue correlations (see Figure 9C) and the heterogeneously clustered phylogenetic structure of protein families in sequence space (see Figure 9D), suggesting that the statistical model captures essential rules governing the variance of native CM sequences throughout evolution. In contrast, site (h i A simpler model that retains only the intrinsic amino acid tendencies in the sequence and excludes pairwise bonds fails to reproduce even the second-order statistics of the MSA and cannot explain the pattern of sequence variance in natural CM proteins.
[0177] This example illustrates how bmDCA provides a statistical model for generating new candidate sequences. That is, native sequences and sequences sampled from a probability distribution P(σ) are functionally equivalent despite considerable sequence variance. To illustrate this, a high-throughput, quantitative in vivo complementation assay for CMs in E. coli is used to evaluate candidate proteins generated using the bmDCA model for desired functionality. Here, catalytic power is used as the desired functionality to illustrate the method. The high-throughput, quantitative in vivo complementation assay is suitable for studying numerous native and designed CMs in a single, internally controlled experiment.
[0178] Figure 9E shows a quantitative high-throughput functional assay for CM, in which a library of CM mutants is expressed in E. coli strains lacking colismyic acid mutase, grown as a mixed population under selective conditions, and then subjected to next-generation sequencing to count the frequencies of each CM allele in the input and selected populations.
[0179] Figure 9F shows that relative enrichment (re) can be calculated from measurements of a quantitative high-throughput functional assay. Since re is nearly linear with catalytic power (ln(kc / Km)) over a range of approximately 5-logarithmic order, it provides an indicator of catalytic power (e.g., desired functionality). This “standard curve” was constructed using a panel of point mutants of E. coli CM spanning a wide range of specific activities.
[0180] Here, we provide a brief description of the high-throughput assay, with a more detailed explanation provided below. Libraries of CM variants (e.g., cold-start natural and / or synthetic) are constructed using custom de novo gene synthesis protocols that enable rapid and relatively inexpensive assembly of large-scale novel DNA sequences. For example, a library containing all natural CM homologs in MSA (1,259 in total) was constructed, and over 1,900 synthetic variants were created to explore various design parameters for the bmDCA model. These libraries were expressed in a CM-deficient bacterial strain (KA12) and grown together as a single population in a selective medium lacking phenylalanine and tyrosine, and colismyate mutase activity was selected (as shown in Figure 9E). Deep sequencing of the populations before and after selection allowed us to count the log-frequency of each allele relative to the wild type, a quantity called "relative enrichment" (re) that quantitatively and reproducibly reports the catalytic activity of colismyate mutase under specific induction, growth time, and temperature conditions (as shown in Figure 9F). This "select-seq" assay is nearly linear across a wide range of catalytic powers and serves as an effective tool for rigorously comparing numerous natural and synthetic variants for in vivo functional activity in a single, internally controlled experiment.
[0181] The initial study investigated the performance of native CM homologs in a select-seq assay, which serves as a positive control for sequences designed with bmDCA. While native sequences exhibit a unimodal distribution of bmDCA statistical energy centered near the E. coli CM value (defined as zero, see Figure 10A), it is unclear how they perform under specific E. coli strains and experimental conditions used in the assay. For example, members of the CM family may mutate in ways unknown with respect to activity under specific environments, and MSAs contain a proportion of paralogous enzymes that perform related but distinctly different chemical reactions. The select-seq assay revealed that the 1,259 native CM homologs of MSAs exhibited a complementary bimodal distribution in the assay, with one mode containing approximately 31% of the sequences centered at the level of wild-type E. coli CM, and the remainder containing modes centered at the level of the null allele (see Figure 10B). Green fluorescent protein (GFP)-labeled versions of the library suggest that the bimodal complementarity is not clearly related to the difference in expression levels compared to the E. coli variant. Instead, the bimodal likely stems from the amino acid cooperativity in determining the catalytic power of colismic acid, and the possibility of paralogous sequences. For the purposes of this study, the bimodal simply allows us to reduce the full distribution to the probability that the sequence complements the function in the assay. Importantly, the standard curve shows that this amount is a rigorous test of high colismic acid mutase activity (see Figure 9F).
[0182] To evaluate the generative potential of the bmDCA model, Monte Carlo sampling is used to randomly select sequences from the model that span a range of statistical energies relative to the native MSA. For example, Figures 10C-E illustrate how low-energy sequences can become functional colismyate mutases.
[0183] Regarding the sampling process, it should be noted that sequence data is inherently limited (e.g., not all amino acids occur at all positions), and even large sequence families are undersampled for the number of all amino acid combinations for all pairs of sites within the MSA. Given these limitations, regularized inference is used in constructing the bmDCA model to avoid overfitting. Regularization ensures that sequences sampled from the model have, on average, lower probabilities (i.e., higher statistical energies) than native sequences. To sample sequences at low energies, a formal calculation "temperature" T<1 is introduced into the model, provided, for example, by the following equation:
number
[0184] Such models compensate for the effects of regularization. For example, sampling at T{0.33,0.66} produces sequences with statistical energies that more closely reflect the natural distribution and have little dependence on the regularization parameter, as shown, for example, in Figures 10C and 10D. In contrast, sequences sampled at T=1 show a broad distribution of statistical energies that deviates significantly from the natural distribution and is more strongly dependent on the strength of regularization, as shown, for example, in Figure 10E.
[0185] A library of 350 synthetic sequences was created and tested. These libraries were sampled from the bmDCA model, each at T□{0.33,0.66,1.0}. Figures 10F–H show that, overall, these sequences also exhibit a bimodal distribution of complementarity, with many complementarity functions close to the levels of wild-type E. coli sequences. Consistent with the hypothesis, the probability of complementarity is well predicted by the statistical energy of bmDCA, and the low-energy sequences extracted from T□{0.33,0.66} either essentially summarize or even somewhat exceed the performance of the natural sequences (Figures 10f–g). In contrast, sequences extracted from T=1 show degraded performance, consistent with deviations from the bmDCA model (see Figure 10H). In total, 521 synthetic sequences out of 1,050 (approximately 50%) rescued proliferation in the assay and contained a wide range of top-hit identities to any native colismyate mutase, from 44 to 92%. These included 48 sequences with less than 65% identity to the MSA protein, corresponding to at least 33 mutations that were far from the nearest native counterpart. Sequence variance from E. coli CM ranged from 19 to 42%. The representation of protein positions that contribute most to the statistical energy of bmDCA both highlight residues that extend and are distributed through the CM tertiary structure, including the dimer interface, within the active site (see Figure 11F).
[0186] Figures 10A–10I show functional analyses of natural and synthetic CM sequences. Figure 10A shows that the ensemble of natural CM sequences within the MSA contains a unimodal distribution of bmDCA sequences centered near the value of E. coli CM (defined as zero). Figure 10B shows that the relative energy (re) scores of natural CMs exhibit a bimodal distribution, with one model containing approximately 31% of sequences close to the level of E. coli CM (defined as zero re or 1 with respect to the normalized re dashed line 1000b1) and the remaining sequences close to the level of the CM null allele (red dashed line).
[0187] Figures 10C - 10D show the statistical energies of the bmDCA of the arrays sampled at three different computational temperatures (0.33, 0.66, 1). The arrays sampled at temperatures 0.33 and 0.66 reproduce exactly the energy of the native array, while the arrays extracted at T = 1 do not reproduce the energy of the native array.
[0188] In Figures 10F - 10H, the functional analysis shows that the arrays sampled at temperatures T of 0.33 and 0.66 generalize or even exceed the performance of the native array, while the arrays sampled at T = 1 hardly function.
[0189] Figures 10I and 10J show that the arrays created by retaining the first - order statistics but excluding the correlations show large statistical energies and no function at all. Therefore, the bmDCA model encodes the native - like function of the synthetic CM arrays.
[0190] Figure 11A shows a scatter plot of all the synthetic CM arrays, showing the relationship between the statistical energy of bmDCA and the catalytic function. Functional arrays (e.g., bars shaded as item 1100a1) and non - functional arrays (e.g., bars shaded as item 1100a2) are shown. The data show that there are essentially no arrays with E DCA < 40 that complement the CM null phenotype and are predicted by low - function bmDCA energy.
[0191] Figure 11B shows the top two principal components of the array mutations defined by the native CM array as in Figure 9D, shown as in Figure 11A or shaded in another way. The E. coli array is marked by point 1100bs. These data show that the native CM arrays classified as functional in E. coli are localized in specific regions within the overall pattern of the array mutations.
[0192] Figure 11C shows the projection of the synthetic CM into the same space, indicating that the functional arrays are localized in the same clusters. Information on which regions of the PCA-defined domain the functional arrays are localized / clustered in is used to define a score (e.g., a functional landscape), and this can be used to bias the selection of candidate sequences from the regions corresponding to the clusters of functional arrays. For example, a Boltzmann distribution can define a probability density function according to which candidate sequences are randomly drawn. This probability density function can be biased to increase the density of the extracted sequences corresponding to the clusters of functional arrays. In FIGS. 11A - C, the functional versus non-functional determination is binary, but in other implementations, a continuous scale of functionality can be used and GPR can be used to generate, for example, a functional landscape.
[0193] FIGS. 11D and 11E show the complementarity patterns of synthetic sequences with EDCA < 40 that either do not have (FIG. 11D) or have (FIG. 11E) an additional statistical condition (P(x = 1|σ)) derived from the pattern of functional complementarity of natural CM sequences. The data shows that the prediction of synthetic sequences that complement a function in a particular context can be dramatically enhanced using prior knowledge about the function in natural CM.
[0194] FIG. 11F shows the structure of the E. coli CM together with the positions (e.g., the spheres or circles shaded as item 1100f2) that contribute most to the low statistical energy and the positions (e.g., the spheres or circles shaded as item 1100f1) that contribute to the E. coli-specific function. The data shows that the general design constraints focus on both physically continuous patterns of residues at the active site and extending to the dimer interface, and that the additional constraints for the E. coli-specific function are more peripheral.
[0195] As a separate control, 326 sequences were produced that had the same sequence identity distribution as the bmDCA-designed sequences at T=0.66, but only primary statistics were retained and correlations were excluded. These sequences showed high bmDCA energies as expected and showed no complementarity whatsoever (see Figures 10I and 10J), indicating that enzyme function was determined not only by the magnitude of the sequence mutation, but also by J ij This demonstrates that it is fundamentally dependent on the correlation pattern imposed by the combination.
[0196] When all the data is put together, a surprisingly sharp relationship can be observed between bmDCA statistical energy and CM activity. When the statistical energy falls below a threshold (EDCA<50) set by the width of the energy distribution of the native sequence, nearly 50% of the designed sequences rescue the CM null phenotype, and above this value, there are essentially no functional sequences (see Figure 11A). Therefore, bmDCA is an effective generative model, and when the statistical energy is within the range of the native homolog, it is possible to design native-like enzyme activity with considerable sequence diversity.
[0197] The bmDCA model captures the overall statistics of a protein family and does not focus on the specific functional activity of individual members of the family. Therefore, like natural CM homologs, most sequences designed in bmDCA do not complement the function under specific assay conditions (see Figure 11A). Generative models can be improved to estimate additional information that optimizes protein sequences for specific phenotypes. For example, a function-rescuing sequence, the assay occupies the sequence space spanned by natural CM sequences. While natural CMs complementing the function of E. coli are distributed in several diverse clusters (see Figure 11B), interestingly, functional synthetic sequences follow the same pattern (see Figure 11C). This suggests that information about CM function and assay conditions in a specific context of E. coli may be present in and learned from the statistics of natural sequences. In such embodiments, knowledge gained from a single experimental test can be used to formally train a computational model to predict synthetic sequences encoding specific protein phenotypes and biological environments.
[0198] As a test of the above-referenced embodiment, the DCA model was trained / generated from the sequence of the native MSA, but this time annotated with a binary value x indicating its ability to function in the assay (x = 1 for functionality, zero otherwise). From this model, the probability that any synthetic sequence σ complements function in the E. coli select-seq assay can be calculated, i.e., P(x=1|σ). Figures 11D and 11E show an additional condition (83%, see Figure 11e) in which, for a low-energy CM-like synthetic sequence from the naive bmDCA model (see Figure 11D), P(x=1|σ) now efficiently predicts the subset that complements in the context of the assay. Mapping the significantly contributing upper positions to the E. coli-specific CM sequence reveals a clustered arrangement of amino acids around the active site (see Figure 11F). Thus, these positions function allosterically to control catalytic activity, a mechanism that provides context-dependent tuning of reaction parameters. These results support iterative design strategies for specific protein phenotypes. In this strategy, the bmDCA model is updated at each round of selection to optimally target the desired phenotype.
[0199] The results described herein validate and extend the concept that pairwise amino acid correlations in the actually available sequence alignments of protein families are sufficient to determine protein folding and function. The bmDCA model is one approach to capture these correlations.
[0200] Here, the non-restrictive high-throughput gene construction and assays used in Figures 10A-J and 11A-E discussed above are described in more detail. CM genes were constructed using PCR overlap extension of oligos synthesized on a microarray chip. Two oligos (230-mer) were designed for each gene, each having a unique pair of adjacent orthogonal primer annealing sites for "gene-specific primers" (GSPs) and a BtsαI restriction site for removing the adjacent region after amplification. The overlaps were designed to be at least 16 base lengths, have a 3' or C base, and have a melting temperature of at least 59°C, calculated using the nearest neighbor method. PCR was performed in a 384-well plate using Q5 polymerase with 1xQ5 buffer, 0.2 μM dNTPs, and a total of 10 μl of 0.5 μM GSP pairs. Oligonucleotides corresponding to individual genes were thawed, annealed, and extended at 98°C, 61°C, and 72°C, respectively, for 35 cycles of 10 seconds each. To remove the GSP annealing sites and amplify the full-length genes, the amplified products were diluted 500-fold and subjected to a PCR reaction with 0.1 U / μl BtsαI and adjacent primers 5'-AGCGATCTCGGTGACGATGG-3' and 5'-CATTAACGATGCAAGTCTCGTGG-3', incubated at 55°C for 60 minutes before amplification for 10 cycles at an annealing temperature of 61°C and 35 cycles at an annealing temperature of 65°C.
[0201] Cloning: Genes were pooled, digested with NdeI and XhoI, ligated to plasmid pKTCTET, column purified, and transformed into sufficiently electrocompetent NEB10 beta cells to generate more than 1000-fold transformants per gene cloned. After culturing the entire transformation overnight in 500 ml LB containing 100 ug / ml AMP, the plasmid was purified and diluted to 1 ng / ul to minimize the transformation of individual cells with multiple plasmids, and transformed into CM-deficient strain KA12 containing plasmid pKIMP / UAUC to generate more than 1000-fold transformants per gene. The entire transformation was cultured overnight in 500 ml LB containing 100 ug / ml AMP and 30 ug / ml CAM, supplemented with 16% glycerol, and frozen at -80°C.
[0202] Chorismyate mutase selection assay: KA12 glycerol stock was cultured overnight in LB medium at 30°C, diluted to 0.045 OD600 with non-selective M9c FY, grown overnight at 30°C to approximately 0.2 OD600, and washed with M9c (no FY). This pre-selected culture was inoculated into LB containing 100 ug / ml AMP, grown overnight, and collected for plasmid purification to generate the "input" sample. For selection, the culture was diluted in 500 ml of M9c containing 3 ng / ml doxycycline at a calculated starting OD600 of 1e-4 and grown at 30°C for 24 hours. 50 ml of culture was collected by centrifugation, resuspended in 2 ml of LB containing 100 ng / ml AMP, grown overnight, and collected for plasmid purification.
[0203] Sequencing: Plasmids purified from the input and selected cultures were amplified using two rounds of PCR with KOD polymerase to add adapters and indices for Illumina sequencing. In the first round, DNA was amplified using primers that annealed to the plasmid and added 6–9N to aid in initial focusing, and a portion of the i5 or i7 adapter. Examples of primers are 5'-TGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNACGACTCACTATAGGGAGAC-3' and 5'-CACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNTGACTAGTCATTATTAGTGG-3'. In the second round of PCR, the remaining adapters and TruSeq index were added. Examples of primers are 5'-CAAGCAGAAGACGGCATACGAGATCGAGTAATGTGACTGGAGTTCAGACGTG-3' and 5'-AATGATACGGCGACCACCGAGATCTACACTATAGCCTACACTCTTTCCCTACACGAC-3'. Low cycles (16 rounds) and high initial template concentrations were used in both rounds of PCR to minimize amplification bias. The final product was gel-purified, quantified using Qubit, and sequenced in MiSeq with 2x250 cycles.
[0204] Paired-end reads were joined using FLASH(ref), trimmed to NdeI and XhoI cloning sites, and translated. Only those that perfectly matched the designed gene were counted. Finally, the relative enrichment value (re) was calculated.
[0205] Figure 12 also shows the circuitry and hardware for acquiring, storing, processing, and distributing data from the protein assay apparatus 810, the gene synthesis apparatus 820 (also known as the gene synthesis apparatus / system 820), and the gene expression apparatus 830. The circuitry and hardware include a processor 870, a network controller 874, memory 878, and a data acquisition system (DAS) 876. The protein optimization system 800 may include data channels (not shown) that route detection measurement results from each apparatus (e.g., the protein assay apparatus 810, the gene synthesis apparatus 820 (also known as the gene synthesis apparatus / system 820), and the gene expression apparatus 830) to the DAS 876, processor 870, memory 878, and network controller 874. The data acquisition system 876 can control the acquisition, digitization, and routing of detection data from various sensors and detectors. The processor 870 performs functions including training machine learning models 115, fitting functional landscapes, and controlling each apparatus, as discussed herein.
[0206] The processor 870 can be configured to perform various steps of the methods and processes described herein. The processor 870 may include a CPU that can be implemented as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other complex-programmable logic device (CPLD) as a discrete logic gate. The FPGA or CPLD implementation can be coded in VHDL, Verilog, or any other hardware description language, and the code can be stored directly in the electronic memory within the FPGA or CPLD, or as separate electronic memory. Furthermore, the memory may be non-volatile, such as ROM, EPROM, EEPROM, or flash memory. The memory may also be volatile, such as static or dynamic RAM, and a processor, such as a microcontroller or microprocessor, can be provided to manage the electronic memory and the interaction between the FPGA or CPLD and the memory.
[0207] Alternatively, the CPU within the 870 processor can execute a computer program containing a set of computer-readable instructions that perform various steps of Method 10 and / or Method 10'. The program is stored in the non-temporary electronic memory and / or on a hard disk drive, CD, DVD, flash drive, or any other known storage medium. Furthermore, the computer-readable instructions may be provided as utility applications, background daemons, or components of an operating system, or a combination thereof, and run in conjunction with processors such as Intel of America's Xenon processors or AMD of America's Opteron processors, and operating systems such as Microsoft Vista, UNIX®, Solaris, LINUX, Apple, MAC-OS, and other operating systems known to those skilled in the art. Furthermore, the CPU may be implemented as multiple processors working in parallel to execute instructions.
[0208] Memory 878 may be a hard disk drive, CD-ROM drive, DVD drive, flash drive, RAM, ROM, or any other electronic storage known in the art.
[0209] Network controllers 874, such as Intel Ethernet® PRO network interface cards from Intel Corporation of America, can interface between various parts of the protein optimization system 800. Furthermore, network controllers 874 can also interface with external networks. As can be understood, external networks can be public networks such as the Internet, or private networks such as LAN or WAN networks, or any combination thereof, and may also include PSTN or ISDN subnetworks. External networks can also be wired, such as Ethernet® networks, or wireless, such as cellular networks including EDGE, 3G, and 4G wireless cellular systems. Wireless networks can also be WiFi, Bluetooth®, or any other known wireless format.
[0210] Here, a more detailed explanation of training an artificial neural network (e.g., VAE) is provided (e.g., process 310). Here, as mentioned above, the target data is, for example, the output amino acid sequence, and the input data is the same output amino acid sequence.
[0211] Figure 13 shows a flowchart of one implementation of the training process 310. In process 310, the input data and target data are used as training data to train the artificial neural network, and as a result, the trained artificial neural network 370 is output from step 319 of process 310. The offline DL training process 310 trains the artificial neural network using a large number of amino acid sequences from the input data.
[0212] In process 310, a set of training data is obtained and the network is repeatedly updated to reduce an error (e.g., a value generated by a loss function). The artificial neural network infers the mapping implied by the training data, and the loss function generates an error value related to the discrepancy between the target data and the result generated by applying the current embodiment of the artificial neural network to the input data. For example, in certain implementations, the loss function can use the mean squared error to minimize the mean squared error. In the case of a multi-layer perceptron (MLP) neural network, the network can be trained using the backpropagation algorithm by minimizing a mean squared error-based loss function using a (stochastic) gradient descent method.
[0213] In step 316 of process 310, an initial estimate is generated for the coefficients of the artificial neural network. For example, the initial estimate can be based on one of LeCun initialization, Xavier initialization, and Kaiming initialization.
[0214] Steps 316 - 319 of process 310 provide non-limiting examples of optimization methods for training an artificial neural network.
[0215] An error is calculated (e.g., using a loss function or loss functions) to represent a measure of the difference (e.g., a distance measure) between the target data (i.e., the ground truth) and the input data after applying the current version of the network. The error can be calculated using any known loss function or distance measure, including the loss function described above. Further, in certain implementations, the error / loss function can be calculated using one or more of hinge loss and cross-entropy loss. In certain implementations, the loss function can be the l p norm of the difference between the target data and the result of applying the input data to the artificial neural network. l pVarious values of the norm "p" can be used to highlight different aspects of noise. In certain implementations, the difference l between the target data and the result from the input data is used. p Instead of minimizing the norm, the loss function can represent similarity (for example, using the peak signal-to-noise ratio (PSNR) or structural similarity (SSIM) index).
[0216] In certain implementations, the network is trained using backpropagation. Backpropagation can be used to train neural networks and is used in combination with the steepest descent optimization method. During the forward pass, the algorithm computes predictions for the network based on the current parameters (θ). These predictions are then input into a loss function, which is compared to the corresponding ground truth labels (i.e., high-quality target data). During the reverse pass, the model computes the gradient of the loss function with respect to the current parameters. The parameters are then updated by taking steps of a predefined size in the direction of the minimized loss (for example, in accelerated methods, the step size can be chosen to converge more quickly to optimize the loss function, as in the Nesterov momentum method and various adaptive methods).
[0217] Optimization methods that perform backprojection can use one or more of the following: gradient descent, batch gradient descent, stochastic gradient descent, and mini-batch stochastic gradient descent. Forward and backward passes can be performed incrementally through each layer of the network. In the forward pass, execution begins by feeding the input through the first layer, thus creating output activations for subsequent layers. This process is repeated until the loss function of the last layer is reached. During the backward pass, the last layer computes its gradient with respect to its own learnable parameters (if any) and with respect to its own input, which acts as the upstream derivative of the previous layer. This process is repeated until the input layer is reached.
[0218] Returning to Figure 13, step 317 of process 310 determines the error change (e.g., error gradient) as a function of the network change that can be computed, and this error change can be used to select the direction and step size for subsequent changes to the weights / coefficients of the artificial neural network. Calculating the error gradient in this manner is consistent with certain implementations of gradient descent optimization methods. In certain other implementations, as will be understood by those skilled in the art, this step may be omitted or replaced with another step according to a different optimization algorithm (e.g., a non-gradient descent optimization algorithm such as simulated annealing or a genetic algorithm).
[0219] In step 317 of process 310, a new set of coefficients for the artificial neural network is determined. For example, the weights / coefficients can be updated using the changes calculated in step 317, as in a gradient descent optimization method or an over-relaxed acceleration method.
[0220] In step 318 of process 310, a new error value is calculated using the updated weights / coefficients of the artificial neural network.
[0221] In step 319, a predefined stopping criterion is used to determine whether the network training is complete. For example, the predefined stopping criterion could evaluate whether the new error and / or the total number of iterations performed exceeds a predefined value. For example, the stopping criterion may be met if the new error falls below a predefined threshold or if the maximum number of iterations is reached. If the stopping criterion is not met, the training process performed in process 310 returns to step 317 with the new weights and coefficients and continues by repeating (the iteration loop includes steps 317, 318, and 319). Once the stopping criterion is met, the training process performed in process 310 is complete.
[0222] Figure 14 shows an example of the interconnections between layers of an artificial neural network. An artificial neural network may include fully connected layers, convolutional layers, and pooling layers, all of which are described below. In certain preferred implementations of an artificial neural network, the convolutional layers are located near the input layers, while the fully connected layers, which perform high-level inference, are located further down the architecture toward the loss function. Pooling layers can be inserted after convolutions, and a reduction has been proven that it decreases the spatial range of the filter, and therefore the amount of learnable parameters. Activation functions can also be incorporated into various layers to introduce nonlinearity, allowing the network to learn complex predictive relationships. The activation function can be a saturated activation function (e.g., a sigmoid or hyperbolic tangent activation function) or a normalized activation function (e.g., a normalized linear unit (ReLU) applied in the first and second examples discussed above). Batch normalization can also be incorporated into the layers of an artificial neural network, as illustrated in the first and second examples discussed above.
[0223] Figure 14 shows an example of a typical artificial neural network (ANN) with N inputs, K hidden layers, and 3 outputs. Each layer consists of nodes (also called neurons), each node performing a weighted sum of the inputs and comparing the result of the weighted sum to a threshold to produce an output. An ANN constitutes a class of functions from which members of the class are obtained by changing architectural details such as the threshold, the weights of the connections, or the number of nodes and their connectivity. Nodes in an ANN are sometimes called neurons (or neuron nodes), and neurons can have interconnections between different layers of the ANN system. Synapses (i.e., connections between neurons) store values called "weights" (also interchangeably called "coefficients" or "weighting coefficients") that manipulate the data of the computation. The output of an ANN depends on three types of parameters: (i) the interconnection patterns between different layers of neurons, (ii) the learning process for updating the interconnection weights, and (iii) the activation function that translates the weighted inputs of the neurons into output activations.
[0224] Mathematically, the network function m(x) of neurons is a function of another function n i It is defined as a composite of (x), and can be further defined as a composite of other functions. This can be conveniently represented as a network structure using arrows that show the dependencies between variables, as shown in Figure 14. For example, ANN can use a nonlinear weighted sum. Here, m(x) = K(Σ i w i n i (x)) where K (generally called the activation function) is some predefined function, such as the hyperbolic tangent.
[0225] In Figure 14, neurons (i.e., nodes) are represented by circles around a threshold function. In the non-restrictive example shown in Figure 14, the input is depicted as a circle around a linear function, and arrows indicate directed connections between neurons.
[0226] While specific implementations are described, these are presented only as examples and are not intended to limit the teachings of this disclosure. In fact, the novel methods, apparatuses, and systems described herein can be embodied in a variety of other forms. Furthermore, various omissions, substitutions, and modifications in the forms of the methods, apparatuses, and systems described herein can be made without departing from the spirit of this disclosure.
[0227] The nature of this disclosure The following aspects of this disclosure are merely illustrative and are not intended to limit the scope of this disclosure.
[0228] 1. A method for designing a protein having desired functionality, the method comprising: determining candidate amino acid sequences for a synthetic protein using a machine learning model trained to learn implicit patterns of amino acid sequences in a training dataset of proteins, wherein the machine learning model represents the learned implicit patterns in the trained model; executing an iterative loop, wherein each iteration of the loop synthesizes a candidate gene and produces a candidate protein corresponding to each candidate amino acid sequence, wherein each candidate gene encodes the corresponding candidate amino acid sequence; evaluating the extent to which each candidate protein exhibits desired functionality by measuring values that characterize the candidate proteins using one or more assays; and, when one or more stopping criteria in the iterative loop are not met, calculating a goodness-of-fit function assigned to each sequence from the measured values, and using a combination of the goodness-of-fit function and the machine learning model to select a new candidate amino acid sequence for subsequent iterations.
[0229] 2. The method according to aspect 1, wherein the implicit pattern is learned in the latent space and the candidate amino acid sequence is determined, further comprising determining that the latent space has a reduced dimension compared to the feature dimension of the amino acid sequences in the training dataset.
[0230] 3. The training dataset includes multiple sequence alignments of evolutionarily related proteins, and the amino acid sequences of the multiple sequence alignments have a sequence length L, and the feature dimension of the training dataset is 20 amino acids corresponding to the sequence length L. L The method according to any one of embodiments 1 to 2, which is large enough to accommodate the combination of the following.
[0231] 4. The method according to any one of embodiments 1 to 3, wherein the training dataset includes multiple sequence alignments of evolutionarily related proteins, and the feature dimension of the amino acid sequences of the training dataset is the product L × K, where L is the length of one of the amino acid sequences of the training dataset and K is the number of possible types of amino acids.
[0232] 5. The method according to embodiment 4, wherein the amino acid is a natural amino acid and K is 20 or less.
[0233] 6. The method according to embodiment 4, wherein at least one of the possible types of amino acids is a non-natural amino acid.
[0234] 7. The method according to any one of embodiments 1 to 6, wherein the training dataset comprises proteins associated by a common function, which is at least one of the following: (i) common binding function, (ii) common allosteric function, and (iii) common catalytic function.
[0235] 8. The method according to any one of embodiments 1 to 7, wherein the training dataset used to train a machine learning model includes proteins related by at least one of the following: (i) a common ancestor, (ii) a common three-dimensional structure, (iii) a common function, (iv) a common domain structure, and (v) a common evolutionary selection pressure.
[0236] 9. The method according to any one of embodiments 1 to 8, wherein the step of running an iterative loop further includes updating a machine learning model based on an updated training dataset of proteins containing amino acid sequences of candidate proteins when one or more stopping criteria are not met, and selecting a new candidate amino acid sequence for subsequent iterations using a goodness-of-fit function and a combination of the updated machine learning model based on the updated training dataset.
[0237] 10. The method according to any one of embodiments 1 to x, wherein the machine learning model is one of (i) a variational autoencoder (VAE) network, (ii) a restricted Boltzmann machine (RBM) network, (iii) a direct coupled analysis (DCA) model, (iv) a statistical coupled analysis (SCA) model, and (v) a generative adversarial network (GAN).
[0238] 11. The method according to aspect 2, wherein the machine learning model is a network model that performs encoding and decoding / generation, where encoding is performed by mapping an input amino acid sequence to points in a latent space, and decoding / generation is performed by mapping points in the latent space to an output amino acid sequence, the machine learning model is trained to optimize an objective function, the components of the machine learning model represent the degree to which the input amino acid sequence and the output amino acid sequence match, and as a result, when trained with a training dataset, the machine learning model generates an output amino acid sequence that approximately matches the amino acid sequence of the training dataset applied as input to the machine learning model.
[0239] 12. The method according to any one of embodiments 1 to x, wherein the machine learning model is an unsupervised statistics-based model that learns design rules based on primary and secondary statistics of amino acid sequences in a training dataset, and the machine learning model is a generative model trained by a machine learning method to produce output amino acid sequences consistent with the learned design rules.
[0240] 13. The method according to any one of embodiments 1 to 12, further comprising training a machine learning model using a training dataset to learn the external fields and residue-residue bonds of the Potts model to generate a DCA model of the training dataset, wherein the DCA model is used as a machine learning model.
[0241] 14. The method according to aspect 13, wherein the DCA model is trained using one of the following methods: Boltzmann machine learning method, mean-field method, Monte Carlo gradient descent method, and pseudo-likelihood maximization method.
[0242] 15. The method according to aspect 13, wherein the step of determining candidate amino acid sequences further comprises selecting candidate amino acid sequences from a Boltzmann statistical distribution based on Hamiltonians of Potts models trained at one or more predefined temperatures, the candidate amino acid sequences being selected using at least one of the following: Markov chain Monte Carlo (MCMC) methods, simulated annealing methods, simulated heating methods, gene algorithms, basin hopping methods, sampling methods, and optimization methods, and extracting a sample from the Boltzmann statistical distribution.
[0243] 16. The method according to aspect 15, wherein the step of selecting a new candidate amino acid sequence for subsequent iterations further comprises biasing the selection of amino acid sequences from a Boltzmann statistical distribution based on Hamiltonians of Potts models trained at one or more predefined temperatures, wherein the biasing of the selection of amino acid sequences is based on a goodness-of-fit function to increase the number of amino acid sequences selected that more closely match the amino acid sequences of the measured candidate protein, where the measured values indicate that the desired functionality is greater than the mean, median, or mode of the measured values.
[0244] 17. The method according to aspect 15, wherein the step of selecting a new candidate amino acid sequence for subsequent iterations further comprises randomly selecting amino acid sequences from a statistical distribution in which a Boltzmann statistical distribution based on the Hamiltonian of a trained Potts model is weighted by a goodness-of-fit function, thereby increasing the likelihood that a sample is drawn from a region of the latent space that better represents candidate amino acid sequences exhibiting more desired functionality than candidate amino acid sequences corresponding to other regions of the latent space.
[0245] 18. The method according to any one of embodiments 1 to 17, further comprising training a machine learning model using a training dataset to learn a location coevolution matrix and generating an SCA model of the training dataset, wherein the SCA model is used as a machine learning model.
[0246] 19. The method according to aspect 18, further comprising generating a sample set of amino acid sequences by performing simulated annealing or simulated heating using an SCA model, wherein the sample set of amino acid sequences represents learned implicit patterns of a training dataset, and selecting candidate amino acid sequences from the generated sample set of amino acid sequences.
[0247] 20. The method according to any one of embodiments 1 to 19, wherein the step of selecting new candidate amino acid sequences for subsequent iterations further comprises performing linear or nonlinear dimensionality reduction on the candidate amino acid sequences of the candidate protein to rank the components of the low-dimensional model, and biasing the selection of amino acid sequences to increase the number of amino acid sequences selected in one or more regions in the space of leading components of the low-dimensional model that correspond to measured values indicating a high degree of desired functional clusters.
[0248] 21. The method according to aspect 20, wherein the nonlinear dimensionality reduction is principal component analysis, and the leading components of the low-dimensional model are the principal components of the principal component analysis, represented by a set of eigenvectors corresponding to the set of largest eigenvalues of the correlation matrix.
[0249] 22. The method according to aspect 20, wherein nonlinear dimensionality reduction is independent component analysis in which eigenvectors undergo rotation and scaling operations to identify functionally independent modes of sequence variation.
[0250] 23. The method according to aspect 11, further comprising the step of determining candidate amino acid sequences: identifying regions in latent space corresponding to amino acid sequences of proteins selected as likely to exhibit desired functionality; selecting points having the identified regions in latent space; and mapping the selected points to each candidate amino acid sequence using decoding / generation performed by a machine learning model, wherein each candidate amino acid sequence is then used as a candidate amino acid sequence.
[0251] 24. The method according to aspect 11, further comprising the step of selecting new candidate amino acid sequences for subsequent iterations, identifying regions in a latent space that are more likely to exhibit the desired functionality than other regions based on a goodness-of-fit function, or where the sampling is too sparse to make a statistically significant estimate with respect to the desired functionality, selecting points within the identified regions in the latent space, and mapping the selected points to their respective candidate amino acid sequences using decoding / generation performed by a machine learning model, wherein each candidate amino acid sequence is then used as a new candidate amino acid sequence for subsequent iterations, further comprising mapping.
[0252] 25. The method according to embodiment 24, wherein the step of identifying a region in the latent space further comprises generating a density function in the latent space based on a goodness-of-fit function, and the step of selecting points within the identified region in the latent space further comprises selecting points so as to statistically represent the density function.
[0253] 26. The method according to any one of embodiments 1 to 25, wherein the step of calculating the goodness-of-fit function further includes performing supervised learning of a functional landscape that approximates the measured values of candidate proteins as a function of corresponding positions in latent space, and the goodness-of-fit function is at least partially based on the functional landscape.
[0254] 27. The method according to aspect 26, wherein, for a given point in latent space, the functional landscape provides an estimate of the functionality of the amino acid sequence corresponding to the given point, the functionality estimate being at least one of i) the statistical probability of the corresponding amino acid sequence based on a machine learning model, (ii) the statistical or physical energy of folding the corresponding amino acid sequence, computationally predicted based on a statistical scoring function, and (iii) the statistical energy activity when performing a particular structural or functional role, computationally predicted or experimentally measured.
[0255] 28. The method according to embodiment 26, wherein the goodness-of-fit function is a functional landscape.
[0256] 29. The method according to aspect 26, wherein the goodness-of-fit function is based on a functional landscape and at least one other parameter selected from a sequence similarity landscape and a stability landscape, the sequence similarity landscape estimates the degree to which a protein corresponding to a point in latent space is similar to a predefined collection of proteins, and the stability landscape estimates the degree to which a protein corresponding to a point in latent space is stable.
[0257] 30. The method according to aspect 29, wherein the stability landscape is based on numerical simulations of protein folding of proteins corresponding to points in a stable latent space.
[0258] 31. The method according to aspect 29, wherein a functional landscape and at least one other parameter define a multi-objective optimization space, and candidate amino acid sequences for subsequent iterations are selected by determining a convex hull in the multi-objective optimization space as a Pareto frontier, selecting points in a latent space on the Pareto frontier, and using a machine learning model to map the selected points to amino acid sequences, the amino acid sequences then being used as candidate amino acid sequences for subsequent iterations.
[0259] 32. The method according to aspect 29, wherein a functional landscape is generated by performing supervised learning using either supervised classification or regression analysis, the supervised learning being one of the following: (i) multivariate linear, polynomial, stage, lasso, ridge, kernel, or nonlinear regression methods; (ii) support vector regression (SVR) methods; (iii) Gaussian process regression (GPR) methods; (iv) decision tree (DT) methods; (v) random forest (RF) methods; and (vi) artificial neural network (ANN).
[0260] 33. The method according to embodiment 30, wherein the functional landscape further includes uncertainty values as a function of position in latent space, the uncertainty values representing estimated uncertainty about how well the functional landscape approximates measured values.
[0261] 34. The method according to aspect 33, further comprising selecting some of the candidate amino acid sequences for subsequent iterations to correspond to regions of latent space where the uncertainty value is greater than that of other regions, wherein in subsequent iterations, the measured values corresponding to some of the candidate amino acid sequences reduce the large uncertainty value by increasing the sampling in the region with the large uncertainty value.
[0262] 35. The method according to any one of embodiments 1 to 34, wherein the step of measuring the value of a candidate protein comprises measuring the value using at least one of the following: (i) an assay that measures proliferation rate as an indicator of desired functionality; (ii) an assay that measures gene expression as an indicator of desired functionality; and (iii) an assay that measures gene expression or activity as an indicator of desired functionality using microfluidics and fluorescence.
[0263] 36. The method according to any one of embodiments 1 to 34, wherein the step of synthesizing a candidate gene further comprises using a polymerase cycling / linking assembly (PCA) in which oligonucleotides (oligos) having overlapping extensions are provided in solution, and the oligos are cyclically subjected to a series of temperatures, thereby causing the oligos to bind to larger oligos by the steps of (i) denaturation of the oligos, (ii) annealing of the overlapping extensions, and (iii) extension of the non-overlapping extensions.
[0264] 37. The method according to any one of embodiments 1 to 34, wherein the step of performing an iterative loop further comprises evolving one or more assay parameters that have evolved from an initial value to a final value, and as a result, during the first iteration, the candidate gene exhibits the desired functionality when measured at the initial value but not when measured at the final value, and during the final iteration, the candidate gene exhibits the desired functionality when measured at the final value.
[0265] 38. The method according to embodiment 37, wherein the parameter is one of (i) temperature, (ii) pressure, (iii) light conditions, (iv) pH value, and (v) concentration of a substance in a culture medium used in one or more assays.
[0266] 39. The method according to aspect 37, wherein the parameters of one or more assays are selected to evaluate candidate amino acid sequences in relation to a combination of internal phenotype and external environmental conditions.
[0267] 40. The method according to any one of embodiments 1 to 39, wherein the step of executing an iterative loop further comprises stopping the iterative loop when one or more stopping criteria of the iterative loop are met, and outputting information of one or more gene codes corresponding to one or more candidate genes that best exhibit the desired functionality.
[0268] 41. A system for designing a protein having desired functionality, comprising: a gene synthesis system configured to synthesize a gene based on an input gene sequence encoding each amino acid sequence and to produce a protein from the synthesized gene; an assay system configured to measure the value of the protein received from the gene synthesis system, wherein the measured value provides an indicator of desired functionality; and a processing circuit configured to determine candidate amino acid sequences for a synthesized protein and run an iterative loop using a machine learning model trained to learn implicit patterns in a training dataset of amino acid sequences of proteins, wherein the machine learning model represents learned implicit patterns in the trained model, and each iteration of the loop includes: sending a candidate amino acid sequence to the gene synthesis system and producing a candidate protein based on the candidate amino acid sequence; receiving a scale value from the assay system corresponding to the candidate protein based on the candidate amino acid sequence; and, when one or more stopping criteria of the iterative loop are not met, calculating a goodness-of-fit function assigned to each amino acid sequence from the measured value and using a combination of the goodness-of-fit function and the machine learning model to select a new candidate amino acid sequence for subsequent iterations.
[0269] 42. The system according to embodiment 41, wherein the machine learning model represents an implicit pattern learned in the latent space, and the processing circuit is further configured to determine candidate amino acid sequences, and the latent space has a reduced dimension compared to the feature dimension of the amino acid sequences of the training dataset.
[0270] 43. The training dataset includes multiple sequence alignments of homologous proteins, and the amino acid sequences of the multiple sequence alignments have a sequence length L. The feature dimension of the training dataset is 20 amino acids corresponding to the sequence length L. L A system according to any one of embodiments 41 to 42, which is large enough to accommodate the combination of the following.
[0271] 44. The system according to aspect 42, wherein the training dataset includes multiple sequence alignments of evolutionarily related proteins, and the feature dimension of the amino acid sequences of the training dataset is the product L × K, where L is the length of one of the amino acid sequences of the training dataset and K is the number of possible types of amino acids.
[0272] 45. The system according to embodiment 44, wherein the amino acid is a natural amino acid and K is 20 or less.
[0273] 46. The system according to embodiment 44, wherein at least one of the amino acid types is a non-natural amino acid.
[0274] 47. The system according to any one of embodiments 41 to 46, wherein the training dataset comprises proteins associated by a common function, which is at least one of (i) a common binding function, (ii) a common allosteric function, and (iii) a common catalytic function.
[0275] 48. A system according to any one of embodiments 41 to 47, wherein the training dataset used to train a machine learning model comprises proteins related by at least one of (i) a common ancestor, (ii) a common three-dimensional structure, (iii) a common function, (iv) a common domain structure, and (v) a common evolutionary selection pressure.
[0276] 49. The system according to any one of embodiments 41 to 48, wherein the processing circuit is further configured to execute an iterative loop, update a machine learning model based on an updated training dataset of proteins containing amino acid sequences of candidate proteins when one or more stopping criteria are not met, and select a new candidate amino acid sequence for subsequent iterations using a combination of the goodness-of-fit function and the updated machine learning model based on the updated training dataset.
[0277] 50. A system according to any one of embodiments 41 to 49, wherein the machine learning model is one of (i) a variational autoencoder (VAE) network, (ii) a restricted Boltzmann machine (RBM) network, (iii) a direct coupled analysis (DCA) model, (iv) a statistical coupled analysis (SCA) model, and (v) a generative adversarial network (GAN).
[0278] 51. The system according to aspect 42, wherein the machine learning model is a network model that performs encoding and decoding / generation, wherein encoding is performed by mapping an input amino acid sequence to points in a latent space, and decoding / generation is performed by mapping points in the latent space to an output amino acid sequence, the machine learning model is trained to optimize an objective function, the components of the machine learning model represent the degree to which the input amino acid sequence and the output amino acid sequence match, and as a result, when trained with a training dataset, the machine learning model generates an output amino acid sequence that approximately matches the amino acid sequence of the training dataset applied as input to the machine learning model.
[0279] 52. The system according to any one of embodiments 41 to 51, wherein the machine learning model is an unsupervised statistics-based model that learns design rules based on primary and secondary statistics of amino acid sequences of a training dataset, and the machine learning model is a generative model trained by a machine learning method to produce output amino acid sequences consistent with the learned design rules.
[0280] 53. The system according to any one of embodiments 41 to 52, wherein the processing circuit is further configured to train a machine learning model using a training dataset to learn the external fields and residue-residue bonds of the Potts model, and to generate a DCA model of the training dataset, the DCA model being used as a machine learning model.
[0281] 54. The DCA model is trained using one of the following methods: Boltzmann machine learning method, mean-field method, Monte Carlo gradient descent method, and pseudo-likelihood maximization method, according to embodiment 53.
[0282] 55. The system according to embodiment 53, wherein the processing circuit is further configured to determine candidate amino acid sequences by selecting candidate amino acid sequences from a Boltzmann statistical distribution based on Hamiltonians of Potts models trained at one or more predefined temperatures, the candidate amino acid sequences being selected using at least one of the following: Markov chain Monte Carlo (MCMC) methods, simulated annealing methods, simulated heating methods, gene algorithms, basin hopping methods, sampling methods, and optimization methods, and a sample is extracted from the Boltzmann statistical distribution.
[0283] 56. The system according to aspect 55, wherein the processing circuit is further configured to select new candidate amino acid sequences for subsequent iterations by biasing the selection of amino acid sequences from a Boltzmann statistical distribution based on Hamiltonians of Potts models trained at one or more predefined temperatures, the biasing of the selection of amino acid sequences being based on a goodness-of-fit function to increase the number of amino acid sequences selected that more closely match the amino acid sequence of the measured candidate protein, where the measured values indicate that the desired functionality is greater than the mean, median, or mode of the measured values.
[0284] 57. The system according to embodiment 53, wherein the processing circuit is further configured to select new candidate amino acid sequences for subsequent iterations by randomly extracting amino acid sequences from a statistical distribution in which a Boltzmann statistical distribution based on the Hamiltonian of a trained Potts model is weighted by a goodness-of-fit function, increasing the likelihood that a sample is drawn from a region of the latent space that better represents candidate amino acid sequences exhibiting more desired functionality than candidate amino acid sequences corresponding to other regions of the latent space.
[0285] 58. The system according to any one of embodiments 41 to 57, wherein the processing circuit is further configured to train a machine learning model using a training dataset, learn a positional coevolution matrix, and generate an SCA model of the training dataset, the SCA model being used as a machine learning model.
[0286] 59. The system according to embodiment 58, wherein the processing circuit is further configured to generate a sample set of amino acid sequences by performing simulated annealing or simulated heating using an SCA model, the sample set of amino acid sequences representing learned implicit patterns of a training dataset, and the processing circuit is further configured to select candidate amino acid sequences from the sample set of amino acid sequences.
[0287] 60. The system according to any one of embodiments 41 to 59, wherein the processing circuit is further configured to perform linear or nonlinear dimensionality reduction on the candidate amino acid sequences of the measured candidate protein to rank the components of the low-dimensional model, and to bias the selection of amino acid sequences to select new candidate amino acid sequences for subsequent iterations by increasing the number of amino acid sequences selected in one or more regions within the space of leading components of the low-dimensional model that correspond to measured values indicating a high degree of desired functional clusters.
[0288] 61. The system according to embodiment 60, wherein the nonlinear dimensionality reduction is principal component analysis or independent component analysis, and the leading components of the low-dimensional model are the principal components of the principal component analysis or the independent components of the independent component analysis.
[0289] 62. The system according to embodiment 51, wherein the processing circuit is further configured to determine a candidate amino acid sequence by mapping each of the selected points to their respective candidate amino acid sequences, using decoding / generation performed by a machine learning model, wherein each candidate amino acid sequence is then used as a candidate amino acid sequence.
[0290] 63. The system according to embodiment 51, wherein the processing circuit is further configured to select a new candidate amino acid sequence for subsequent iterations by mapping, based on a goodness-of-fit function, a region in the latent space that is more likely to exhibit the desired functionality than other regions, or where the sampling is too sparse to make a statistically significant estimate with respect to the desired functionality; selecting a point within the identified region in the latent space; and mapping the selected point to each candidate amino acid sequence, which is then used as a new candidate amino acid sequence for subsequent iterations.
[0291] 64. The system according to embodiment 63, wherein the processing circuit is further configured to identify a region in the latent space by generating a density function in the latent space based on a goodness-of-fit function, and to select a point in the identified region in the latent space by selecting a point to statistically represent the density function.
[0292] 65. The system according to any one of embodiments 41 to 64, wherein the processing circuit is further configured to perform supervised learning of a functional landscape that approximates the measured values of candidate proteins as a function of corresponding positions in latent space when calculating the goodness-of-fit function, and the goodness-of-fit function is at least partially based on the functional landscape.
[0293] 66. The system according to aspect 65, wherein, for a given point in latent space, the functional landscape provides an estimate of the functionality of the amino acid sequence corresponding to the given point, the functionality estimate being at least one of i) the statistical probability of the corresponding amino acid sequence based on a machine learning model, (ii) the statistical or physical energy of folding the corresponding amino acid sequence, computationally predicted based on a statistical scoring function, and (iii) the statistical energy activity when performing a particular structural or functional role, computationally predicted or experimentally measured.
[0294] 67. The system according to embodiment 65, wherein the goodness-of-fit function is a functional landscape.
[0295] 68. The system according to aspect 65, wherein the goodness-of-fit function is based on a functional landscape and at least one other parameter selected from a sequence similarity landscape and a stability landscape, the sequence similarity landscape estimates the degree to which a protein corresponding to a point in latent space is similar to a predefined collection of proteins, and the stability landscape estimates the degree to which a protein corresponding to a point in latent space is stable.
[0296] 69. The system according to aspect 68, wherein the stability landscape is based on numerical simulations of protein folding of proteins corresponding to points in a stable latent space.
[0297] 70. The system according to aspect 69, wherein a functional landscape and at least one other parameter define a multi-objective optimization space, and new candidate amino acid sequences for subsequent iterations are selected by determining a convex hull in the multi-objective optimization space as a Pareto frontier, selecting points in a latent space on the Pareto frontier, and using a machine learning model to map the selected points to amino acid sequences, the amino acid sequences then being used as new candidate amino acid sequences for subsequent iterations.
[0298] 71. The system according to aspect 65, wherein the functional landscape is generated by performing supervised learning using either supervised classification or regression analysis, the supervised learning being one of the following: (i) multivariate linear, polynomial, stage, lasso, ridge, kernel, or nonlinear regression methods; (ii) support vector regression (SVR) methods; (iii) Gaussian process regression (GPR) methods; (iv) decision tree (DT) methods; (v) random forest (RF) methods; and (vi) artificial neural network (ANN).
[0299] 72. The system according to embodiment 65, wherein the functional landscape further includes uncertainty values as a function of position in latent space, the uncertainty values representing estimated uncertainty about how well the functional landscape approximates measured values.
[0300] 73. The processing circuit is further configured to select some of the new candidate amino acid sequences for subsequent iterations to correspond to regions of latent space where the uncertainty value is larger than in other regions, and as a result, in subsequent iterations, the measured values corresponding to some of the candidate amino acid sequences are reduced by increasing the sampling in regions where the uncertainty value is large, according to any one of embodiments 41 to 72.
[0301] 74. The assay system according to any one of embodiments 41 to 73, further configured to measure the value of a candidate protein using at least one of the following: (i) an assay for measuring proliferation rate as an indicator of desired functionality; (ii) an assay for measuring gene expression as an indicator of desired functionality; and (iii) an assay for measuring gene expression or activity as an indicator of desired functionality using microfluidics and fluorescence.
[0302] 75. The gene synthesis system is further configured to synthesize a candidate gene using polymerase cycling / linking assembly (PCA) in which oligonucleotides (oligos) having overlapping extensions are provided in solution, wherein the oligos are circulated through a series of temperatures, thereby causing the oligos to bind to larger oligos by the steps of (i) denaturation of the oligos, (ii) annealing of the overlapping extensions, and (iii) extension of the non-overlapping extensions, according to any one of embodiments 41 to 74.
[0303] 76. The system according to any one of embodiments 41-x, wherein the processing circuit is further configured to run iterative loops so that the parameters of one or more assays evolve from an initial value to a final value, and as a result, during the first iteration, the candidate gene exhibits the desired functionality when measured at the initial value but not when measured at the final value, and during the final iteration, the candidate gene exhibits the desired functionality when measured at the final value.
[0304] 77. The system according to embodiment 76, wherein the parameter of one or more assays is one of (i) temperature, (ii) pressure, (iii) light conditions, (iv) pH value, and (v) concentration of a substance in a culture medium used in one or more assays.
[0305] 78. The system according to embodiment 76, wherein the parameters of one or more assays evaluate candidate amino acid sequences in relation to a combination of internal phenotype and external environmental conditions.
[0306] 79. A non-temporary computer-readable storage medium containing executable instructions, the instructions causing the circuit to perform a method comprising: determining candidate amino acid sequences for a synthetic protein using a machine learning model trained to learn implicit patterns of a training dataset of proteins, wherein the machine learning model represents learned implicit patterns; and performing an iterative loop, where each iteration of the loop includes determining a candidate gene sequence based on the candidate amino acid sequence; sending the candidate gene sequence to be synthesized into a candidate gene to a gene synthesis system to produce a candidate protein; receiving measured values from an assay system, generated by measuring the candidate protein using one or more assays; and, when one or more stopping criteria of the iterative loop are not met, calculating a goodness-of-fit function from the measured values and using a combination of the goodness-of-fit function and the machine learning model to select additional candidate amino acid sequences for subsequent iterations.
[0307] 80. A method for designing a sequence-defined molecule having desired functionality, the method comprising: determining candidate sequences of a sequence-defined molecule, the candidate sequences being generated using a machine learning model trained to learn implicit patterns from a training dataset of sequence-defined molecules, the machine learning model representing the learned implicit patterns; and executing an iterative loop, each iteration of the loop comprising: synthesizing candidate sequences corresponding to candidate molecules; evaluating the extent to which each candidate molecule exhibits desired functionality by measuring values of the candidate molecules using one or more assays; and, when one or more stopping criteria of the iterative loop are not met, calculating a goodness-of-fit function from the measured values; and using a combination of the goodness-of-fit function and the machine learning model to select additional candidate sequences for subsequent iterations.
[0308] 81. The method according to aspect 80, wherein the step of running an iterative loop further includes updating a machine learning model based on an updated training dataset of molecules containing candidate molecule sequences when one or more stopping criteria are not met, and selecting additional candidate sequences for subsequent iterations using a goodness-of-fit function and a combination of the updated machine learning model based on the updated training dataset.
[0309] 82. The method according to embodiment 80, wherein the molecule is a DNA molecule and the sequence is a sequence of nucleotides.
[0310] 83. The method according to embodiment 80, wherein the molecule is an RNA molecule and the sequence is a sequence of nucleotides.
[0311] 84. The method according to embodiment 80, wherein the molecule is a polymer and the sequence is a sequence of chemical monomers.
[0312] 85. The method according to any one of embodiments 1 to 40, wherein the candidate protein comprises one or more of the following: antibodies, enzymes, hormones, cytokines, growth factors, coagulation factors, anticoagulation factors, albumin, antigens, adjuvants, transcription factors, or cell receptors.
[0313] 86. The method according to any one of embodiments 1 to 40, wherein a candidate protein is provided for selected binding to one or more other molecules.
[0314] 87. The method according to any one of embodiments 1 to 40, wherein a candidate protein is provided to catalyze one or more chemical reactions.
[0315] 88. The method according to any one of embodiments 1 to 40, wherein the candidate protein is provided for long-range signal transduction.
[0316] 89. The method according to any one of embodiments 1 to 40, further comprising generating or manufacturing a final product based on a candidate protein.
[0317] 90. The method according to any one of embodiments 1 to 40, wherein one or more cells are produced from a candidate protein.
[0318] 91. The method according to embodiment 90, wherein cells produced from the candidate protein are directed to or placed in one or more vials.
[0319] 92. The method according to any one of embodiments 1 to 40, wherein the candidate protein is determined by high-throughput functional screening.
[0320] 93. The method according to aspect 92, wherein high-throughput functional screening is implemented by a microfluidic device that measures the fluorescence of cells corresponding to a candidate protein.
[0321] Further consideration While the disclosure herein provides detailed descriptions of numerous different embodiments, it should be understood that the legal scope of the description is defined by the claims and equivalent language set forth at the end of this patent. The detailed descriptions should be interpreted as illustrative only, and not all possible embodiments are described, as it would be impractical to describe all possible embodiments. Many alternative embodiments may be implemented using either the current art or art developed after the filing date of this patent, but would still fall within the scope of the claims.
[0322] The following additional considerations apply to the above discussion. Throughout this specification, multiple examples may implement components, actions, or structures described as a single example. While individual actions of one or more methods are illustrated and described as separate actions, one or more of these actions may be performed simultaneously, and the actions do not need to be performed in the illustrated order. Structures and functions presented as separate components within an illustrative configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as single components may be implemented as separate components. These and other variations, changes, additions, and improvements fall within the scope of the subject matter of this specification.
[0323] Furthermore, specific embodiments described herein include logic or a number of routines, subroutines, applications, or instructions. These can constitute either software (e.g., code embodied on a machine-readable medium or in a transmitted signal) or hardware. In hardware, routines, etc., are tangible units capable of performing specific operations and can be configured or arranged in specific ways. In exemplary embodiments, one or more computer systems (e.g., standalone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., processors or groups of processors), can be configured by software (e.g., an application or part of an application) as hardware modules that operate to perform specific operations described herein.
[0324] In various embodiments, hardware modules can be implemented mechanically or electronically. For example, a hardware module may include a dedicated circuit or logic permanently configured to perform a specific operation (e.g., a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), or other application-specific processor). A hardware module may also include programmable logic or circuit temporarily configured by software to perform a specific operation (e.g., implemented in a general-purpose processor or other programmable processor). It will be understood that the decision of whether to implement a hardware module mechanically, with a dedicated and permanently configured circuit, or with a temporarily configured circuit (e.g., configured by software) can be made considering cost and time.
[0325] Therefore, the term “hardware module” should be understood to encompass tangible entities that are physically constructed, permanently configured (e.g., physically embedded), or temporarily configured (e.g., programmed) for the purpose of operating in a particular manner or performing a particular operation as described herein. When considering embodiments in which a hardware module is temporarily configured (e.g., programmed), each hardware module does not need to be configured or instantiated at any given time. For example, if a hardware module includes a general-purpose processor configured using software, that general-purpose processor can be configured as different hardware modules at different times. Thus, the software may configure the processor, for example, to configure one hardware module at one time and another hardware module at another time.
[0326] Hardware modules can provide and receive information from other hardware modules. Therefore, the described hardware modules can be considered to be communicatively coupled. When such hardware modules exist simultaneously, communication can be achieved through signal transmission (through appropriate circuits and buses) connecting the hardware modules. In embodiments where multiple hardware modules are configured or instantiated at different points in time, communication between such hardware modules can be achieved, for example, through the storage and retrieval of information in a memory structure accessible to the multiple hardware modules. For example, one hardware module may perform an operation and store the output of that operation in a memory device to which that hardware module is communicatively coupled. Further hardware modules can then later access the memory device, retrieve the stored output, and process it. Hardware modules can also initiate communication with input or output devices and operate on resources (e.g., collect information).
[0327] Various operations of the exemplary methods described herein may be performed, at least in part, by one or more processors that are temporarily (e.g., by software) or permanently configured to perform the operations in question. Whether temporarily or permanently configured, such processors may constitute a processor implementation module that operates to perform one or more operations or functions. The modules referred to herein may include processor implementation modules in some exemplary embodiments.
[0328] Similarly, any method or routine described herein can be at least partially processor-implemented. For example, at least part of the operation of a method may be performed by one or more processors or processor-implemented hardware modules. The execution of a particular operation may reside not only within a single machine but also distributed among one or more processors deployed across several machines. In some embodiments, one or more processors may be located in a single location, while in other embodiments, the processors may be distributed across multiple locations.
[0329] Reliable performance can reside not only within a single machine but also distributed across one or more processors deployed across several machines. In some exemplary embodiments, one or more processors or processor implementation modules may be located in a single geographical location (e.g., in a home environment, a work environment, or a server farm). In other embodiments, one or more processors or processor implementation modules may be distributed across multiple geographical locations.
[0330] This detailed description should be interpreted as illustrative only, and not all possible embodiments are described, as it would be impractical, if not impossible, to describe all possible embodiments. Those skilled in the art may implement many alternative embodiments using either the current art or art developed after the filing date of this application.
[0331] Those skilled in the art will recognize that a wide variety of modifications, changes, and combinations can be made with respect to the embodiments described above without departing from the scope of the present invention, and such modifications, changes, and combinations should be considered to fall within the scope of the concept of the present invention.
[0332] Furthermore, the claims at the end of this patent application are not intended to be construed under Section 112(f) of the U.S. Patent Act unless they explicitly enumerate traditional means-plus-function terms such as “means to do” or “steps to do” as explicitly enumerated in the claims. The systems and methods described herein are intended for the improvement of computer functions and the improvement of conventional computer functions.
Claims
1. A method for designing a protein having a desired functionality, wherein the method is A step of determining candidate amino acid sequences for synthetic proteins using a machine learning model trained by supervised learning to learn implicit patterns in a training dataset of protein amino acid sequences, wherein the machine learning model represents the learned implicit patterns in the trained model; The process of executing an iterative loop, where each iteration of the loop is: Synthesizing candidate genes and producing candidate proteins corresponding to each of the candidate amino acid sequences, wherein each of the candidate genes encodes the corresponding candidate amino acid sequence. By measuring values that represent the characteristics of the candidate proteins using one or more assays, the extent to which each candidate protein exhibits the desired functionality is evaluated, and (1) When one or more termination criteria of the iterative loop are not met, calculate a goodness-of-fit function assigned to each sequence from the measured values, and use the combination of the goodness-of-fit function and the machine learning model to select a new candidate amino acid sequence for subsequent iterations, wherein feedback is used to improve and enhance the predictive ability of the goodness-of-fit function together with the machine learning model, wherein the feedback includes the candidate amino acid sequences of the iterative loop. Includes or (2) When one or more termination criteria of the repeating loop are met, output the candidate amino acid sequence as the designed protein. Methods that include...
2. The method according to claim 1, wherein the implicit pattern is learned in a latent space, and determining the candidate amino acid sequence further comprises determining that the latent space has a reduced dimension compared to the feature dimension of the amino acid sequence of the protein.
3. The method according to claim 1 or 2, wherein the training dataset comprises multiple sequence alignments of evolutionarily related proteins, and the feature dimension of the amino acid sequences of the training dataset is a product L × K, where L is the length of one of the amino acid sequences of the training dataset and K is the number of possible types of amino acids.
4. The method according to claim 3, wherein at least one of the possible types of amino acids is a non-natural amino acid.
5. The machine learning model is a network model that performs encoding and decoding / generation, wherein the encoding is performed by mapping the input amino acid sequence to points in the latent space, and the decoding / generation is performed by mapping the points in the latent space to the output amino acid sequence. The method according to claim 2, wherein the machine learning model is trained to optimize an objective function, one component of the objective function representing the degree to which the input amino acid sequence and the output amino acid sequence match, and as a result, when trained using the training dataset, the machine learning model generates an output amino acid sequence that approximately matches the amino acid sequence of the training dataset applied as input to the machine learning model.
6. The method according to any one of claims 1 to 5, further comprising training the machine learning model using the training dataset to learn the external fields and residue-residue bonds of the Potts model to generate a DCA model of the training dataset, wherein the DCA model is used as the machine learning model.
7. The process involves training the machine learning model using the training dataset to learn the location coevolution matrix and generating an SCA model of the training dataset, wherein the SCA model is used as the machine learning model. The process involves generating a sample set of amino acid sequences by performing simulated annealing or simulated heating using the SCA model, wherein the sample set of amino acid sequences represents the learned implicit pattern of the training dataset. Selecting candidate amino acid sequences from the generated sample set of amino acid sequences, The method according to any one of claims 1 to 6, further comprising:
8. The method according to any one of claims 1 to 7, wherein the step of selecting the new candidate amino acid sequences for the subsequent iterations further comprises performing linear or nonlinear dimensionality reduction on the candidate amino acid sequences of the candidate protein to rank the components of a low-dimensional model, and biasing the selection of the amino acid sequences to increase the number of amino acid sequences selected in the neighborhood of one or more leading components of the low-dimensional model, corresponding to measured values that indicate a high degree of desired functional clusters.
9. (a) The nonlinear dimensionality reduction is principal component analysis, and the leading components of the low-dimensional model are the principal components of the principal component analysis, represented by a set of eigenvectors corresponding to the set of largest eigenvalues of the correlation matrix, or (b) The method of claim 8, wherein the nonlinear dimensionality reduction is independent component analysis in which the eigenvectors undergo rotation and scaling operations to identify functionally independent modes of sequence variation.
10. The machine learning model comprises an encoder configured to learn implicit patterns and correlations between residues of each amino acid sequence in a training dataset in order to predict the amino acid sequence in latent space, A decoder that generatively designs a new amino acid sequence from the aforementioned latent space, A supervised regression model that optimizes the latent space by constructing a gradient for identifying the properties of the protein, The method according to any one of claims 1 to 9, comprising a variational autoencoder (VAE)-based artificial neural network (ANN) including
11. The step of selecting the new candidate amino acid sequence for the subsequent repeats is: Based on the goodness-of-fit function, identify regions within the latent space that exhibit or are more likely to exhibit the desired functionality than other regions, or where the sampling is too sparse to make a statistically significant estimate regarding the desired functionality. Selecting a point within the identified region in the latent space, and The method of claim 5, further comprising mapping the selected points to each candidate amino acid sequence using the decoding / generation performed by the machine learning model, wherein each candidate amino acid sequence is then used as the new candidate amino acid sequence for the subsequent iterations.
12. The step of identifying the region in the latent space further includes generating a density function in the latent space based on the goodness-of-fit function, The method according to claim 11, wherein the step of selecting the points within the identified region in the latent space further comprises selecting the points so as to statistically represent the density function.
13. The method according to any one of claims 2, 10, 11, or 12, wherein the step of calculating the goodness-of-fit function further includes performing supervised learning of a functional landscape that approximates the measured values of the candidate protein as a function of corresponding positions in the latent space, and the goodness-of-fit function is at least partially based on the functional landscape.
14. The method according to claim 13, wherein the goodness-of-fit function is based on the functional landscape and at least one other parameter selected from a sequence similarity landscape and a stability landscape, the sequence similarity landscape estimates the degree to which the protein corresponding to a point in the latent space is similar to a predefined aggregate of proteins, and the stability landscape estimates the degree to which the protein corresponding to a point in the latent space is stable, wherein the stability landscape is based on a numerical simulation of the protein folding of the stable protein corresponding to the point in the latent space.
15. A system for designing proteins having desired functionality, wherein the system is A gene synthesis system configured to synthesize genes based on input gene sequences encoding each amino acid sequence, and to produce proteins from the synthesized genes, An assay system configured to measure the value of a protein received from the gene synthesis system, wherein the measured value provides an indicator of desired functionality. A processing circuit, Using a machine learning model trained to supervise learning the implicit patterns of a training dataset of protein amino acid sequences, we determine candidate amino acid sequences for synthetic proteins, and A processing circuit configured to execute an iterative loop, wherein the machine learning model represents the learned implicit patterns in the trained model, and each iteration of the loop is The candidate amino acid sequence is sent to the gene synthesis system, and a candidate protein is generated based on the candidate amino acid sequence. The assay system receives measurement values corresponding to the candidate protein based on the candidate amino acid sequence, and (1) When one or more termination criteria of the iterative loop are not met, calculate a goodness-of-fit function assigned to each amino acid sequence from the measured values, and use the combination of the goodness-of-fit function and the machine learning model to select a new candidate amino acid sequence for subsequent iterations, wherein feedback is used to improve and enhance the predictive ability of the goodness-of-fit function together with the machine learning model, wherein the feedback includes the candidate amino acid sequences of the iterative loop, or (2) When one or more termination criteria of the repeating loop are met, output the candidate amino acid sequence as the designed protein. A system that includes this.