Protein engineering

A neural network-based method using multiple protein sequences for unsupervised learning and a Markov chain Monte Carlo search efficiently identifies improved protein sequences, addressing inefficiencies in existing protein engineering methods by reducing experimental data needs and improving prediction accuracy.

WO2025202661A1PCT designated stage Publication Date: 2025-10-02CAMBRIDGE CONSULTANTS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/050674
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing protein engineering methods are inefficient, unreliable, and require extensive experimental measurements, particularly when few related proteins are available, leading to unpredictable and costly bioengineering processes.

Method used

A neural network-based method that uses multiple protein sequences as input for unsupervised learning to generate protein fitness predictions, reducing the need for fine-tuning and experimental data, and employs a Markov chain Monte Carlo search to identify improved protein sequences.

Benefits of technology

This approach allows for more efficient and accurate protein engineering with fewer experimental measurements, applicable to a wider range of applications, and identifies sequences with higher fitness than the initial protein.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025050674_02102025_PF_FP_ABST
    Figure GB2025050674_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method of generating predictions of protein sequences having an improved fitness for a particular use, the method comprising: determining a first set of protein sequences, based an initial protein sequence, to be used as an input for a neural network; determining a second set of one or more protein sequences for which a corresponding protein fitness prediction is to be generated; using the neural network to generate an output for use by a fitness prediction model, wherein the neural network receives the first set of protein sequences and the second set of protein sequences as an input for a pass through the neural network to generate the output; and determining a predicted fitness for each sequence of the second set of protein sequences using the fitness prediction model and the output of the neural network, to identify one or more protein sequences that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PROTEIN ENGINEERING

[0002] FIELD OF THE INVENTION

[0003] The invention relates to protein engineering, and in particular to methods of generating predictions of one or more protein sequences having an improved fitness for a particular use case. By way of example, the invention relates to generating one or more protein sequences having an improved fitness using a model that receives multiple protein sequences as an input.

[0004] BACKGROUND

[0005] Bioengineering offers significant technical and commercial potential, and can be used to develop improved proteins targeted towards a particular use case. For example, fluorescent proteins are widely used as research tools within life science sensing and imaging applications. Bioengineering can be used to produce proteins having an improved brightness, which beneficially increases the sensitivity of assays. As a further example, cutinases are a class of enzyme able to degrade polyethylene terephthalate (PET) plastic. Naturally occurring wild type sequences based on "Leaf and branch compost cutinase" are insufficiently productive to be industrially relevant. However, bioengineering can be used to develop proteins having improved PET degradation ability that can be used in an industrial environment.

[0006] Realising the potential of bioengineering is challenging. For example, the bioengineering process can be unpredictable, time consuming and expensive, and many of the candidate protein sequences are likely to be entirely non-functional for the intended task. A significant amount of laboratory testing is also required for verification and testing, which is also time consuming and requires significant resources. There is therefore a general need for more efficient and reliable bioengineering methods for developing improved proteins for particular tasks.

[0007] Proteins can be engineered using machine-learning-guided directed evolution. Protein engineering is typically an iterative process, whereby improvements discovered in one round of laboratory experimentation are built upon in subsequent rounds of laboratory experimentation. This is also the case when using a machine learning model to guide the laboratory experimentation. Typically, the model comprises a deep neural network having multiple layers. For example, a 'language model' that is a higher-order model of sequence statistics, and is typically trained on large sets of data, may be used. The model can be trained using unsupervised learning in which the model is trained to predict random amino acids to complete a partial protein sequence. Bidirectional transformers can learn models for the pseudo-likelihood of each amino acid at positions in the protein sequence, based on the statistical structure of the training sequences. However, conventional machine-learning approaches require large amounts of experimentally obtained training data. For example, the number of experimental assays needed may be of the order of ten thousand. Supervised deep learning typically requires of the order of 106data points, which may not be available for many protein engineering tasks.

[0008] Error-prone polymerase chain reactions (PCR) can be used to generate protein variants, starting from an initial 'seed' variant, to measure in the laboratory to generate training data for the machine learning model. However, there is a further problem that many of the protein variants generated using error-prone PCR are likely to be non-functional forthe intended task, and it is disadvantageous for the training data to be based on a large fraction of nonfunctional sequences since this can compromise the reliability of the model predictions.

[0009] A machine-learning guided protein engineering approach requiring fewer experimental measurements has been developed [1], This method uses a foundational protein neural network which has been trained on a wide range of different types of proteins. An important step of this method is the fine tuning of the neural network weights using a smaller set of protein sequences which have been selected based on their relationship to the protein sequence that is being engineered. However, this method requires more than ten thousand related proteins to be provided for the fine tuning, which limits the applicability of the approach to engineering proteins for which such a large number of related proteins exist in nature. Moreover, since in this model mutations can always be introduced into the protein sequence, mutated sequences that the model predicts to have an improved fitness can always be found in each successive iteration. However, in practice, the more mutations that are introduced, the less likely it is that the new protein sequence will have high fitness. Whilst this can be addressed by imposing a limit of the number of mutations that can be introduced, to prevent the final protein sequences from becoming too far removed from the initial sequence, the limit on the number of mutations is somewhat arbitrary, and it is not necessarily the case that a good limit for the number of mutations for one application is also a good limit for other applications.

[0010] The present invention aims to address, or at least partially ameliorate, one or more of the above needs. For example, there is a need for more efficient and reliable methods of protein engineering requiring fewer experimental measurements, and that are computationally efficient.

[0011] SUMMARY OF THE INVENTION

[0012] In a first aspect the invention provides a method of generating predictions of one or more protein sequences having an improved fitness for a particular use, the method comprising: a) determining a first set of protein sequences, based an initial protein sequence, to be used as an input for a neural network; b) determining a second set of one or more protein sequences for which a corresponding protein fitness prediction is to be generated; c) using the neural network to generate an output for use by a fitness prediction model, wherein the neural network receives the first set of protein sequences and the second set of protein sequences as an input for a pass through the neural network to generate the output; and d) determining a predicted fitness for each sequence of the second set of protein sequences using the fitness prediction model and the output of the neural network, to identify one or more protein sequences that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence.

[0013] Advantageously, the amount of experimental data needed to train the fitness prediction model, and therefore the amount of experimental data needed for the overall method, is relatively small, enabling the method to be used for a wider range of protein engineering applications. Moreover, whilst not performing fine tuning of the neural network weights on a set of proteins related to the protein sequence that is being engineered could potentially reduce the accuracy of the fitness predictions, the inventors have realised that the use of a neural network that receives multiple sequences as input greatly reduces the need for such fine tuning, making the method applicable to a significantly larger range of bioengineering applications. The method may comprise performing an iterative search over a search space of protein sequences, by iteratively performing steps b) to d).

[0014] The method may further comprise measuring the fitness of one or more protein sequences that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence, to identify protein sequences having an actual fitness that is higher than the fitness of the initial protein sequence.

[0015] The method may comprise generating a multiple sequence alignment, MSA, using the first set of protein sequences and the second set of protein sequences, for input into the neural network.

[0016] The first set of protein sequences may be determined based on the number of evolutionary couplings between the sequences in the first set.

[0017] The neural network may be trained using unsupervised learning. The fitness prediction model may be trained using supervised learning and experimentally obtained data. Advantageously, the amount of experimental data needed to train the fitness prediction model is relatively small.

[0018] The number of sequences in the first set of sequences may be between 20 and 500 sequences, for example 30 sequences.

[0019] The fitness prediction model may comprise a linear ridge regression model or a gaussian kernel ridge regression model.

[0020] The method may comprise, in each iteration of the search, generating the second set of one or more protein sequences by introducing mutations into each of the sequences of the second set of a previous iteration of the search. The method may comprise selecting the mutations to introduce from a predetermined set of mutations.

[0021] The predetermined set of mutations may comprise mutations present in protein sequences that have been measured to have a fitness that is better than the fitness of the initial protein sequence.

[0022] The output of the neural network for use by the fitness prediction model may be an embedding vector; and the fitness prediction model may be configured for generating a predicted protein fitness based on the embedding vector.

[0023] The iterative search may comprise a Markov chain Monte Carlo search.

[0024] A plurality of the iterative searches may be performed in parallel; wherein the second set of sequences comprises a plurality of sequences; and wherein each sequence in the second set of sequences corresponds to a respective one of the parallel searches.

[0025] The ratio of the number of sequences in the second set of protein sequences to the number of sequences in the first set of protein sequences may be less than or equal to 0.5.

[0026] In a second aspect the invention provides a method of generating the first set of protein sequences of the first aspect, the method comprising: determining a third set of protein sequences based on the initial protein sequence; and reducing the number of sequences in the third set of protein sequences, to generate the first set of protein sequences, based on the evolutionary couplings in the first set. Advantageously, generating the first set of protein sequences based on the evolutionary couplings in the first set results in more accurate predictions of protein fitness.

[0027] Determining the third set of protein sequences may comprise searching a database of sequences based on the initial protein sequence to identify homologous sequences. The method may comprise maximising the number of statistically significant evolutionary couplings in the first set of protein sequences.

[0028] The method may comprise performing filtering of the third set of protein sequences, to reduce the number of sequences in the third set of protein sequences, before the number of sequences in the third set is further reduced.

[0029] The filtering may be based on the similarity of the sequences in the third set of sequences with respect to the initial protein sequence.

[0030] The method may further comprise performing subsampling of the third set of protein sequences to reduce the number of sequences in the third set.

[0031] The subsampling may comprise sequence reweighted sampling. The inventors have found that sequence reweighted sampling is particularly advantageous for generating MSAs that lead to fitness predictions having a good accuracy.

[0032] In a third aspect the invention provides a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to the first aspect or the second aspect.

[0033] In a fourth aspect the invention provides a method of producing a protein sequence having an improved fitness, with respect to an initial protein sequence, for a particular use, the method comprising producing the identified one or more protein sequences of the first aspect that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence. Methods of producing proteins are well-known in the art and include, for example, recombinant expression or chemical synthesis.

[0034] Also disclosed is a computer-implemented neural network for generating an output for use by a protein fitness prediction model to generate a prediction of a protein fitness for a particular use, the neural network comprising: an input layer for receiving a plurality protein sequences for a pass through the network; at least one hidden layer connected to the input layer; and an output layer connected to the at least one hidden layer; wherein the output layer is configured for outputting the output for use by the fitness prediction model for generating a fitness prediction for one or more of the plurality protein sequences. The input layer may be configured for receiving a multiple sequence alignment, MSA. The number of sequences in the plurality of sequences received at the input layer may be between 20 and 500 sequences, for example 30 sequences. The output for use by the fitness prediction model may comprise an embedding vector.

[0035] DESCRIPTION OF THE FIGURES

[0036] Figure 1 shows a simplified schematic diagram of a model for generating an embedding vector based on one or more protein sequences;

[0037] Figure 2 illustrates a table showing exemplary probabilities for amino acids at particular positions along the protein sequence generated using the model;

[0038] Figure 3 shows examples of protein sequences and a protein sequence in which some of the amino acid positions have been masked:

[0039] Figure 4 shows a simplified schematic illustration of a combination of a neural network based model with a regression model for generating a predicted fitness:

[0040] Figure 5 shows a flow diagram of a method of generating a set of improved protein variants;

[0041] Figure 6 shows a simplified schematic illustration of a multiple sequence alignment (MSA):

[0042] Figure 7 shows a flow diagram of a method for generating an MSA to use as input for the neural network:

[0043] Figure 8 shows a flow diagram of a method for generating a predicted fitness based on an MSA;

[0044] Figure 9 shows a flow diagram of a method of searching the protein sequence space for improved variants;

[0045] Figure 10 shows a simplified schematic illustration of searches in the protein sequence space; and

[0046] Figure 11 shows exemplary apparatus for performing the methods illustrated in Figs. 5 to 10. DETAILED DESCRIPTION

[0047] A method of generating protein fitness predictions, for identifying proteins having improved fitness for a particular use case, will now be described. The method uses a neural network architecture that can be provided with multiple protein sequences as input, which leads to more accurate recommendations of beneficial mutations. The neural network is provided with multiple related protein sequences as input each time a new fitness prediction is made. Advantageously, the number of protein sequences for which the fitness need be experimentally measured is much lower than conventional methods (for example, less than thirty sequences, compared to more than ten thousand).

[0048] In the present examples the MSA Transformer model [2] is used, and is trained in an unsupervised manner. The MSATransformer model comprises a neural networkthat receives multiple sequences as input for a single pass through the network. However, any other suitable model could alternatively be used. The MSA transformer model receives MSAs as the input into the model. Unsupervised learning is achieved by training the model to reconstruct partially masked MSAs. The output of the trained model is an embedding vector generated based on the MSA, which can be used as the input into a further model for predicting protein fitness for a particular use case. This further model is referred to as the 'top-model' (or 'fitness prediction model'), and can be trained in a supervised manner using laboratory data. However, advantageously, the amount of experimental data needed to train the top-model, and therefore the amount of experimental data needed for the overall method, is relatively small, enabling the method to be used for a wider range of protein engineering applications.

[0049] Model Architecture and Training

[0050] Fig. 1 shows a simplified schematic diagram of the model for generating the embedding vector based on the protein sequences of the MSA. For simplicity, in the example illustrated in Fig. 1 a single sequence is illustrated as the input into the model. However, as will be described in more detail later, in examples of the present disclosure an MSA is constructed as used as the input into the model. The bottom row of Fig. 1 corresponds to a sequence of amino acids that defines a single protein. Each letter represents a corresponding amino acid. The MSA Transformer model comprises a neural network having a number of layers. Each neural network layer, n, generates an embedding vector, en(xy), for each position, y, in the protein sequence.

[0051] The 'un-embedding' layer of Fig. 1 layer converts the final embedding vector to a set of probabilities for each amino acid at each position. An exemplary table of the generated probabilities for a sequence is illustrated in Fig. 2. As shown in the table, the probability of each of 20 different common amino acids being present at each position in the sequence is illustrated. For example, at position 1 in the sequence a probability of 68% for amino acid 'S' (Serine) and a probability of 32% for amino acid T (Threonine) has been generated using the model illustrated in Fig. 1.

[0052] Unsupervised Learning

[0053] The embedding layers of the model are trained using unsupervised learning, to determine the adjustable parameters that define the embedding layers (for example, using the approach described in reference [1], or using any other suitable unsupervised learning approach). During this unsupervised learning phase, one or more of the amino acids in the sequence is masked, and the partially masked sequence is used as the input for the model. Fig. 3 shows an example of a wild-type sequence xwt, a mutated sequence xmt, and a partially masked sequence x* (the masked positions are indicated by the mask token "?" in the table of Fig. 3). Mask tokens can be introduced at random positions in the sequence, or using any other suitable method (for example, by masking an entire column of the MSA). The percentage of amino acid positions that are masked in the MSA may be, for example, between 15% and 20%, but any other suitable percentage could alternatively be used.

[0054] In the unsupervised learning phase, the model is trained to generate a high probability for the actual amino acid at a particular location in the sequence based on the partially masked sequence, but may also generate high probabilities for other amino acids that could feasibly be located at the same position in the sequence (for example, amino acids having similar characteristics to that of the actual amino acid). The model output for a protein sequence of length L is an L x 20 pseudo-likelihood matrix. The 'zero-shot' fitness is the log-pseudo- likelihood ratio of a mutated sequence relative to the wild-type sequence. The pseudolikelihood extracted from language models is conditional on the input sequence. A masked marginal pseudo-likelihood can be generated by inserting the mask token at each mutated position.

[0055] Not performing fine tuning of the neural network weights on a set of protein sequences related to the protein sequence that is being engineered could potentially reduce the accuracy of the fitness predictions. The inventors have realised that the use of a neural network based models that receive multiple sequences as input greatly reduces the need for such fine tuning, making the method applicable to a significantly larger range of bioengineering applications.

[0056] The number of embedding layers in the model of Fig. 1 may be 12, and the embedding size may be 768. However, it will be appreciated that any other suitable configuration of the layers of the neural network could alternatively be used (for example, an embedding size of 384 could alternatively be used).

[0057] Supervised Learning

[0058] The embedding vectors output from the final layer of the MSA Transformer model are used as input into a 'top-model' for generating a protein fitness prediction. In other words, a first model is trained in an unsupervised manner and used to generate an embedding vector based on a set of sequences, and a second model is used to generate a predicted fitness based on the output of the first model. The top-model is trained in a supervised manner. However, advantageously, the amount of experimental data needed to train the top model is relatively small.

[0059] Fig. 4 shows a simplified schematic illustration of the neural-network based model (trained using unsupervised learning) and the top-model forgenerating a predicted fitness. The output of the final embedding layer of the neural network is converted to a single embedding vector by taking the mean values across all of the positions in the protein sequence. The top-model is then used to map from the mean embedding vector to a predicted protein fitness for each sequence in the MSA.

[0060] The top-model may comprise, for example, any suitable regression model (for example, a linear ridge regression model) for mapping the mean embedding vector to a protein fitness. A particularly advantageous gaussian kernel ridge regression model will be described in more detail later. As will be described in more detail later, the fitness of protein variants can be measured in the laboratory, and the measured fitness can be used to train the top-model.

[0061] Method of Generating Protein Fitness Predictions for Identifying Improved Protein Variants

[0062] Fig. 5 shows a flow diagram illustrating the method of generating a set of improved protein variants using the neural network and the top-model.

[0063] In step S501, an initial 'seed' protein is selected, depending on the particular application. For example, when the aim is to generate a more fluorescent protein, the naturally occurring 'wild' fluorescent protein may be selected as the seed protein.

[0064] In step S502 an MSA is generated. An example of an MSA is illustrated in Fig. 6. As shown in Fig. 6, the MSA comprises a set of evolutionary sequences. The evolutionary sequences are selected based on the initial seed protein. The generation of the MSA will be described in more detail later with reference to Fig. 7.

[0065] In step S503, self-supervised learning (or 'unsupervised learning') for the neural network model is performed, as described above with reference to Fig. 1. The self-supervised learning could be performed as described in reference [2], As described above, during this unsupervised learning phase, one or more of the amino acids in the sequence is masked, and the partially masked sequence is used as the input for the model. The model is trained to generate a high probability for the actual amino acid at a particular location in the sequence based on the partially masked sequence, but may also generate high probabilities for other amino acids that could feasibly be located at the same position in the sequence (for example, amino acids having similar characteristics to that of the actual amino acid).

[0066] In step S504, the fitness of a number of protein sequence variants is measured in the laboratory, for use in training the top model. One or more functional screens can be used to determine the fitness of a protein for a particular task, for example by measuring fluorescence, or using any other suitable method. The protein variants to measure in the laboratory to generate the training data for the top-model may be selected in any suitable manner. For example, random mutations may be introduced into the sequence of the initial seed protein via error-prone PCR. However, this method may produce many variants that are not functional at all, and it is disadvantageous for the majority of the training data to primarily consist of non-functional protein sequences. Alternatively, for example, the Zero-Shot method may be used to generate variants to measure in the laboratory using the neural network trained via unsupervised learning. For example, the method described in reference [3] can be used to calculate a zero-shot fitness score for protein sequences without the need for any training data from lab measurements. The search method of S506 can then be used to identify variants with high zero shot score to measure in the laboratory.

[0067] The number of protein sequences used to generate training data for the supervised learning of the top model, by performing laboratory measurements in step S504, may be between 10 and 500 sequences (e.g., 20 sequences or 360 sequences). However, it will be appreciated that any other suitable number of sequences may be measured in the laboratory to generate the training data. Moreover, after the final set of improved protein variants has been generated in step S507, those improved variants may also be measured in the laboratory, and these additional measurements can be used as additional training data. In other words, after step S507 has been performed, the method of Fig. 5 may return to step S505 in which the top-model is trained, but using a larger set of training data that includes laboratory measurements of the sequences output in step S507.

[0068] In step S505 the top-model is trained to predict the fitness of a protein based on the protein sequence, using the training data from step S504. As described above, the top-model may be any suitable model for mapping the protein sequence to a predicted fitness (for example, a linear ridge regression model).

[0069] In step S506 a search of the protein sequence space is performed, using the neural network model and the top-model, to identify a set of protein sequences having an improved fitness compared to the initial seed protein. In this example, the method comprises a Markov chain Monte Carlo (MCMC) search using the neural network and the top-model. The MCMC search method will be described in more detail later with reference to Figs. 8 and 9. The set of improved proteins can then be tested in a laboratory to verify whether they have an improved fitness. Generating Multiple Sequence Alignments

[0070] As described above, the input into the neural network model is an MSA comprising a plurality of sequences. A method of generating the MSA for input into the neural network will now be described with reference to Fig. 7.

[0071] As illustrated in Fig. 6, the MSA comprises a number of evolutionary sequences and a number of sequence variants to query. The evolutionary sequences are determined based on the starting seed sequence, and are fixed in step S502. In contrast, the sequence variants to query that are included in the MSA are varied each loop through the search method of step S506. When the top-model is trained in step S505 a set of MSAs is generated, where each MSA comprises the evolutionary sequences plus an additional sequence measured in the laboratory.

[0072] Since the size of the MSA and the selection of the evolutionary sequences included in the MSA can affect the accuracy of the fitness predictions, the construction of the MSA must be carefully considered. Reducing the size of the MSA reduces the computational cost of each pass through the neural network, but can also result in less accurate predictions of protein fitness output from the top model. However, as will be described with reference to step S704 and S705, the inventors have found that reducing the size of an initial large MSA based on the evolutionary couplings between the sequences enables a smaller MSA to be generated whilst reducing the risk of adversely affecting the fitness prediction accuracy.

[0073] In step S701 an initial large MSA is generated based on the seed protein. The initial large MSA may be generated, for example, using a tool such as JackHMMER [4, 6], JackHMMER is based on a profile Hidden Markov Model (profile-HMM) search. Profile-HMM scores protein sequences according to their likelihood under a hidden Markov Model that describes the family of proteins being searched for. JackHMMER is an iterative search algorithm in that it performs multiple searches of a sequence database. Once novel homologous sequences are identified, they are used to update the Hidden Markov Model of the query sequence. This new model is then re-searched against the sequence database.

[0074] It will be appreciated that large databases containing billions of protein sequences are available [5], that can be used to generate an initial large MSA based on the seed protein sequence using any suitable method. In step S702, the size of the initial large MSA is reduced using a filtering method. Filtering is used to remove sequences from the MSA that don't achieve certain threshold values for certain metrics (such as minimum coverage with the query, minimum sequence identity with the query, etc.). The filtering may be performed, for example, based on maximising the diversity of the sequences that remain in the MSA after the filtering has been performed (e.g., based on the hamming distance). The filtering may be performed based on sequence similarity to the seed protein and column coverage. For each amino acid in the query sequence, each MSA sequence may or may not have an amino acid in that position (i.e., there may be whole sections of the query sequence that have effectively been deleted from another sequence in the MSA). Column coverage is measure of how many of the query sequence amino acids have a corresponding amino acid in each MSA sequence.

[0075] In step S703 subsampling of the MSA is performed to further reduce the size of the MSA. Subsampling refers to reducing the number of the sequences in the MSA down to a requested number. This may be done either randomly or in a more structured manner such as sequence reweighted sampling. In sequence reweighted sampling, the probability that a given sequence is selected by the subsampling process is set to correct for biases in the original dataset that was used to generate the original set of sequences. Sequences from organisms or classes of organisms that are over-represented in the dataset are given a lower probability. The subsampling may be performed using any suitable method. However, the inventors have found that sequence reweighted sampling is particularly advantageous for generating MSAs that lead to fitness predictions having a good accuracy.

[0076] Step S703 is repeated to generate a plurality of subsampled MSAs from the larger filtered MSA generated in step S702. For example, sequences to maintain in the subsampled MSA could be selected randomly, or could be selected based on the diversity of the sequences in the MSA (for example, to maximise the sequence diversity based on the hamming distance). The method may start from the initial seed protein sequence and then add sequences to the MSA up to a predetermined number of sequences, where each sequence added to the MSA is selected from the larger filtered MSA of step S702 to maximise the sequence diversity.

[0077] The number of subsampled MSAs that are generated may be, for example, ten subsampled MSAs, although any other suitable number of subsampled MSAs could alternatively be used. Steps S701 to S703 can be repeated to generate multiple large MSAs, using a range of different sensitivities, before applying the filtering and a random subsampling. The iterative search algorithms used to generate the large MSAs can be parameterised by a bitscore threshold, which is a sensitivity parameter that controls how similar the profile of evolutionary sequences should be to query profile at each iteration.

[0078] In step S704 the number of statistically significant evolutionary couplings for each of the subsampled MSAs generated in step S703 are calculated. In an MSA of evolutionarily related protein sequences (referred to as 'homologs'), some pairs of positions (MSA columns) are typically more frequently mutated together than would be expected if naturally selected mutations occurred uniformly across the span of the protein sequence. This occurs because some pairs of amino acids in the protein are located close to each other after protein folding occurs, and so a mutation to one amino acid requires a compatible mutation to collocated amino acids for the protein to retain its naturally selected function. Evolutionary couplings are the name assigned to such pairs of positions which mutate together. The plmc tool [7], for example, can be used to assess the statistical significance of the co-occurrence of mutations for all pairs of positions in the MSA. However, it will be appreciated that the cooccurrence of mutations could be identified using any other suitable method. The number of pairs having a significant p-value (e.g. p<0.05) is identified for each MSA. The MSA having the most significant evolutionary couplings can then selected, as that is the MSA with the richest amount of evolutionary context for a given depth of subsampled MSA.

[0079] In step S705, the subsampled MSAs with the most statistically significant evolutionary couplings are selected. The inventors have realised that using the subsampled MSAs having the highest evolutionary couplings as the set of evolutionary sequences used in the MSAs input into the neural network results in more accurate predictions of protein fitness. In other words, the inventors have realised that selecting the evolutionary sequences of Fig. 6 by selecting the sequences related to the seed protein that maximise the evolutionary couplings improves the prediction accuracy for a given size of MSA. Therefore, the size of the MSA can be reduced, reducing the computational cost of each pass of the MSA through the neural network, whilst mitigating against reductions in prediction accuracy that can be caused by reducing the size of the MSA.

[0080] As will be described later with reference to Fig. 9, a plurality of MCMC searches may be performed in parallel (but could also be performed sequentially), in which case step S705 may comprise selecting a plurality of the subsampled MSAs that maximise the evolutionary couplings, where each subsampled MSA is used as the set of evolutionary sequences for the MSA used as input in the neural network for each parallel search of the sequence space. However, this need not necessarily be the case even when multiple MCMC searches are performed in parallel, as each parallel search could nevertheless use the same set of fixed evolutionary sequences in the MSA.

[0081] Fitness Predictions

[0082] The generation of a fitness score using the MSA, neural network and top-model will now be described with reference to Fig. 8.

[0083] Fig. 8 shows a flow diagram of the method of generating a predicted fitness based on an MSA. In step S801, sequences to query are inserted at the top of the MSA. In other words, the sequences to query are added to the MSA that contains a fixed set of subsampled evolutionary sequences generated in step S705, to form an MSA of the type illustrated in Fig. 6.

[0084] When the predicted fitness is being generated as part of the supervised learning of the top- model of step S505, the query protein sequence corresponds to a sequence measured in the laboratory. The selection of the query protein sequences for the case where the predicted fitness is being generated as part of the MCMC search will be described later with reference to Fig. 9. It is noted that the neural network is flexible with respect to the size of the MSA that is input, even after the unsupervised learning of the neural network has been performed. Therefore, the variant sequence to query in step S801 may be a single sequence in the supervised learning stage, but can be a plurality of sequences during the MCMC search.

[0085] In step S802 the MSA including the sequences to query is passed through the neural network to generate an embedding for each query sequence, as illustrated in Fig. 4.

[0086] In step S803 each embedding vector is passed through the top-model to generate a predicted fitness for each of the query sequences in the MSA. As described above, the top-model may comprise any suitable model for mapping the embedding vector to the predicted fitness. For example, the linear regression model used in reference [1] could be used. However, the inventors have found that use of a gaussian kernel ridge regression model is particularly advantageous, since a Gaussian function centred on the training data points can be used to cause the predicted fitness to tend to zero far away from the training data, resulting in more reliable fitness predictions, and allowing the region of searched protein space to adapt naturally to the available training data without imposing an arbitrary limit on the number of mutations that can be introduced during the search. Nevertheless, a 'trust radius' could alternatively, or additionally, be used, to keep the search of the sequence space close to the training data by limiting the number of mutations that are allowed with respect to the seed protein sequence.

[0087] Search Method

[0088] Fig. 9 shows a flow diagram of the method of searching the protein sequence space for improved variants. The flow diagram of Fig. 9 corresponds to step S506 of Fig. 5.

[0089] In step S901, the current sequence variant is selected. When the search method is initiated the selected sequence variant corresponds to the seed protein sequence.

[0090] In step S902 random mutations to the current protein sequence are selected, subject to constraints, to generate a "proposed variant". For example, the method of introducing mutations to protein sequences described in reference [1] may be used. However, the inventors have realised that when the fitness measurements of step S504 are performed, some of those proteins may already have an improved fitness compared to the initial seed protein. In this case, the proposed mutations could be selected exclusively (or primarily) from the set of mutations (compared to the initial seed protein sequence) that are present in improved variants as measured in the lab. The number of mutations made to the sequences may be sampled from a distribution, and need not necessarily the same each time step S902 is performed. The mutated sequence is used as the 'variant to query' in the MSA illustrated in Fig. 6 (that also contains the fixed set of evolutionary sequences selected in the method of Fig. 7). In other words, an MSA is constructed comprising the fixed set of evolutionary sequences selected in the method of Fig. 7 and the mutated sequence (variant to query) determined in step S902. In the case where multiple searches are not being conducted in parallel on the same processor, then there is only a single variant to query in the MSA. Alternatively, a plurality of variants to query can be included in the MSA when multiple searches are being conducted in parallel. In step S903, the MSA constructed in step S902, that includes the variant to query, is input into the neural network, and the corresponding embedding vectors are output from the neural network to the top-model. The top-model outputs a predicted fitness score for each of the sequence variants to query in the MSA.

[0091] In step S904, an acceptance probability, a, for the proposed variant is determined as a function of the predicted fitness of the current and proposed variants, and as a function of the current iteration number, n, based on the parameters of the MCMC search. In general, the acceptance probabilities are used to guide the search of the protein sequence space towards proteins have a higher fitness. In step S905 the proposed variant is accepted as the new current sequence variant based on the acceptance probability a determined in step S904. Otherwise, the current sequence variant remains unchanged. The method then returns to step S902. Optionally, simulated annealing may be incorporated into steps S904 and S905. In this case, for relatively low values of n the acceptance probabilities are calculated such that mutations that result in a reduction in fitness are more likely to be accepted, but are less likely to be accepted for larger values of n. This helps the search of the protein sequence space to escape local minima in the predicted fitness early in the search.

[0092] Steps S902 to S905 are repeated for n iterations, guiding the search of the protein sequence space towards a sequence having a relatively high predicted fitness. The number of iterations may be, for example, between 2000 and 4000. However, whilst the value n could be fixed, the value of n need not necessarily be predetermined. For example, steps S902 to S905 may be repeated until the change in predicted fitness over the last m iterations becomes smaller than a threshold value (many iterations will have no change in predicted fitness due to the proposed sequence being rejected, and so the change in predicted fitness is calculated over multiple iterations).

[0093] An example of searches through the protein sequence space is illustrated schematically in Fig. 10. The method begins at the circular point in the figure (corresponding to the 'current sequence' in step S901), which corresponds to the initial seed protein. Parallel searches are then performed (corresponding to steps S902 to S905), as illustrated by the separate paths extending away from the seed protein. Each parallel search ends at a respective point in the sequence search space. Whilst only four iterations of each search are illustrated in Figure 10, it will be appreciated that the actual number of iterations of each search is much larger (of the order of 103).

[0094] After all of the iterations through steps S902 to S905 have been performed, the protein sequence having the best predicted fitness can be identified. It is noted that for each guided search through the protein sequence space, it is not necessarily the sequences from the final iteration that have the best predicted fitness. Therefore, the final recommended sequence is not necessarily from the final iteration, but the one having the best predicted fitness throughout the whole search. The sequences having the best predicted fitness can then be measured experimentally, to measure the actual fitness.

[0095] The search algorithm used to identify the fittest sequences can require millions of evaluations of the neural network in order to produce the set of recommended protein sequences to be measured in the lab, which is computationally intensive. To mitigate against this, multiple MCMC walks can be performed in parallel (e.g., using the same GPU-enabled compute node). At each iteration of the searches, each walk requires a single protein fitness prediction. Advantageously, in the present example only one evaluation of the neural network for all the walks in the batch is needed, since all of the sequences needed to be evaluated can be inserted into the same MSA as variants to query.

[0096] Each parallelised search loop will result in a different set of explored protein sequences due to the probabilistic nature of step S905, and the random nature of the mutations introduced in step S902 to generate the sequence variants to query. When the same set of evolutionary sequences are used in the MSA for a set of parallel searches, the inventors have realised that predicted fitnesses for the proposed variants from each of the parallel searches can be calculated using just one pass of an MSA through the neural network. This is advantageous since it significantly reduces the computational cost of running a given number of searches, and is achieved by including the sequences from all the proposed variants into a single MSA as illustrated in Fig. 6. While this reduces the accuracy of the fitness predictions compared to evaluating the fitness of each proposed variant separately, the reduction in accuracy can be mitigated against by ensuring that the ratio of the number of query sequences to the number of evolutionary sequences is not too large. The ratio of the number of query sequences to the number of evolutionary sequences may be, for example, 0.5 or smaller. For example, 16 query sequences and, for example, up to 500 (e.g. 32) evolutionary sequences could be used. Exemplary Apparatus

[0097] Figure 11 shows exemplary apparatus 10 for generating one or more protein sequences predicted to have an improved fitness according to any of the methods described above. It will be appreciated that the apparatus may comprise any suitable computer or server. As shown, the apparatus 10 includes a communication interface 407 which is operable to transmit signals to and receive signals from other devices via a network 24. For example, the apparatus 10 may receive the neural network model parameters, top-model parameters or evolutionary sequences for the MSA via the network 24. Alternatively, the neural network model parameters, top-model parameters or evolutionary sequences for the MSA may be loaded from a removable data storage device (RMD), for example.

[0098] A controller 406 controls the overall operation of the apparatus 10 in accordance with software stored in a memory 401, for example to perform any of the methods described above. The software may be pre-installed in the memory 401 and / or may be downloaded via the network 24 or from a removable data storage device (RMD), for example. The software includes, among other things, an operating system 302 and a protein fitness prediction module 403. The protein fitness prediction module 403 is operable to perform any of the methods described above. For example, the protein fitness prediction module 403 may be responsible for performing the unsupervised learning for the neural network, the supervised learning of the top-model, and for performing any of the steps of the methods of Figs. 5 and 7 to 9 to generate a recommendation of a protein sequence having a high predicted fitness.

[0099] The apparatus 10 has been described for ease of understanding as having a number of discrete modules (such as the protein fitness prediction module 403). Whilst these modules may be provided in this way for certain applications, for example where an existing system has been modified to implement the invention, in other applications, for example in systems designed with the inventive features in mind from the outset, these modules may be built into the overall operating system or code and so these modules may not be discernible as discrete entities. These modules may also be implemented in software, hardware, firmware, or a mix of these. As those skilled in the art will appreciate, the software modules may be provided in compiled or un-compiled form and may be supplied to the apparatus 10 as a signal over a computer network, or on a recording medium. Further, the functionality performed by part or all of this software may be performed using one or more dedicated hardware circuits. However, the use of software modules is preferred as it facilitates the updating of the apparatus 10 in order to update the functionalities.

[0100] Each controller may comprise any suitable form of processing circuitry including (but not limited to), for example: one or more hardware implemented computer processors; microprocessors; central processing units (CPUs); graphics processing units (GPUs); arithmetic logic units (ALUs); input / output (IO) circuits; internal memories / caches (program and / or data); processing registers; communication buses (e.g. control, data and / or address buses); direct memory access (DMA) functions; hardware or software implemented counters, pointers and / or timers; and / or the like.

[0101] EXAMPLES

[0102] The invention will be further clarified by the following examples, which are intended to be purely exemplary and are in no way limiting.

[0103] Example 1: Green Fluorescent Proteins

[0104] In this example, the neural network and top-model were used with the method of Fig. 5 to identify a green fluorescent protein having improved brightness. Fluorescent proteins are used widely as research tools within life science sensing and imaging applications, and improved brightness advantageously increases the sensitivity of the assay.

[0105] Using standard wet lab techniques, 23 protein variants were produced, and the brightness of each protein variant was measured. This functional data was then used to train the top-model in step S505 of Fig. 5. Unsupervised learning forthe neural network model was also performed as described above with reference to step S503.

[0106] The neural network and the top-model were then used to identify 20 protein sequences predicted to have a high fitness as described above with reference to Fig. 5 and Figs. 7 to 9. One of these identified fluorescent protein was found to be seven times brighter than the starting initial seed protein.

[0107] Example 2: Cutinase

[0108] Cutinases are a class of enzyme able to degrade polyethylene terephthalate (PET) plastic.

[0109] Naturally occurring wild type sequences based on "Leaf compost cutinase" are insufficiently productive to be industrially relevant. However, bioengineering can be used to develop proteins having improved PET degradation ability that can be used in an industrial environment. In this example, the neural network and top-model were used with the method of Fig. 5 to identify a protein having improved ability to degrade PET and improved thermostability compared to the starting variant.

[0110] The neural network and the top-model were used to identify a protein sequence predicted to have a high fitness as described above with reference to Fig. 5 and Figs. 7 to 9. The performance of the cutinase enzyme generated using the model was found to be improved with respect to two performance metrics. Firstly, the PET degradation performance, measured as the mols of monomer released per 1 mL culture over 72 hours, was found to be improved. Secondly, the thermostability, measured as percent residual activity using a spectrophotometric assay measuring hydrolysis of a proxy substrate, p-nitrophenol butyrate, was found to be improved.

[0111] Modifications and Alternatives

[0112] As those skilled in the art will appreciate, a number of modifications and alternatives can be made to the above embodiments whilst still benefiting from the inventions embodied therein.

[0113] For example, whilst the top-model has been described with respect to a linear regression model or gaussian model, any other suitable type of model or function may be used to map the embedding from the neural network to a predicted fitness. It will also be appreciated that any other suitable configuration of neural network could be used to generate the embedding vectors.

[0114] Whilst the above examples have been described primarily with respect to an iterative search of the protein sequence space, this need not necessarily be the case. For example, in the method illustrated in Fig. 9, rather than introducing mutations into the current sequence, the method could alternatively comprise generating fitness predictions for all sequences in a predetermined set of sequences. In this manner, a set of sequences having mutations selected from a distribution of beneficial mutations identified in a previous round of laboratory experiments can be searched. References

[0115] [1] Biswas S, et al, Low-N protein engineering with data-efficient deep learning. Nat Methods. 2021 Apr;18(4):389-396. doi: 10.1038 / s41592-021-01100-y. Epub 2021 Apr 7. PMID: 33828272.

[0116] [2] Rao, R.M et al. (2021). MSA Transformer. Proceedings of the 38th International Conference on Machine Learning, in Proceedings of Machine Learning Research 139:8844- 8856.

[0117] [3] Meier, J, et al. (2021). Language models enable zero-shot prediction of the effects of mutations on protein function. Advances in Neural Information Processing Systems, 34, 29287-29303.

[0118] [4] Potter, S. C., Luciani, A., Eddy, S. R. & Park, Y. HMMER web server: Nucleic acids (2018).

[0119] [5] UniRef-50 database; Suzek, B. E., et al., "UniRef: Comprehensive and nonredundant UniProt reference clusters". Bioinformatics, 23(10):1282-1288, 5 (2007).

[0120] [6] Johnson, L.S., Eddy, S.R. & Portugaly, E., Hidden Markov model speed heuristic and iterative HMM search procedure. BMC Bioinformatics 11, 431 (2010).

[0121] [7] "plmc" tool, https: / / github.com / debbiemarkslab / plmc

Claims

CLAIMS1. A method of generating predictions of one or more protein sequences having an improved fitness for a particular use, the method comprising: a) determining a first set of protein sequences, based an initial protein sequence, to be used as an input for a neural network; b) determining a second set of one or more protein sequences for which a corresponding protein fitness prediction is to be generated; c) using the neural network to generate an output for use by a fitness prediction model, wherein the neural network receives the first set of protein sequences and the second set of protein sequences as an input for a pass through the neural network to generate the output; and d) determining a predicted fitness for each sequence of the second set of protein sequences using the fitness prediction model and the output of the neural network, to identify one or more protein sequences that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence.

2. The method according to claim 1, wherein the method comprises performing an iterative search over a search space of protein sequences, by iteratively performing steps b) to d).

3. The method according to claim 1 or 2, wherein the method further comprises measuring the fitness of one or more protein sequences that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence, to identify protein sequences having an actual fitness that is higher than the fitness of the initial protein sequence.

4. The method according to any preceding claim, wherein the method comprises generating a multiple sequence alignment, MSA, using the first set of protein sequences and the second set of protein sequences, for input into the neural network.

5. The method according to any preceding claim, wherein the first set of protein sequences is determined based on the number of evolutionary couplings between the sequences in the first set.

6. The method according to any preceding claim, wherein the neural network is trained using unsupervised learning.

7. The method according to any preceding claim, wherein the fitness prediction model is trained using supervised learning and experimentally obtained data.

8. The method according to any preceding claim, wherein the number of sequences in the first set of sequences is between 20 and 500 sequences, for example 30 sequences.

9. The method according to any preceding claim, wherein the fitness prediction model comprises a linear ridge regression model or a gaussian kernel ridge regression model.

10. The method according to any one of claims 2 to 9, wherein the method comprises, in each iteration of the search, generating the second set of one or more protein sequences by introducing mutations into each of the sequences of the second set of a previous iteration of the search.

11. The method according to claim 10, wherein the method comprises selecting the mutations to introduce from a predetermined set of mutations.

12. The method according to claim 11, wherein the predetermined set of mutations comprises mutations present in protein sequences that have been measured to have a fitness that is better than the fitness of the initial protein sequence.

13. The method according to any preceding claim, wherein the output of the neural network for use by the fitness prediction model is an embedding vector; andwherein the fitness prediction model is configured for generating a predicted protein fitness based on the embedding vector.

14. The method according to any preceding claim, wherein the iterative search comprises a Markov chain Monte Carlo search.

15. The method according to any one of claims 2 to 14, wherein a plurality of the iterative searches are performed in parallel; wherein the second set of sequences comprises a plurality of sequences; and wherein each sequence in the second set of sequences corresponds to a respective one of the parallel searches.

16. The method according to any preceding claim, wherein the ratio of the number of sequences in the second set of protein sequences to the number of sequences in the first set of protein sequences is less than or equal to 0.5.

17. A method of generating the first set of protein sequences of claim 1, the method comprising: determining a third set of protein sequences based on the initial protein sequence; and reducing the number of sequences in the third set of protein sequences, to generate the first set of protein sequences, based on the evolutionary couplings in the first set.

18. The method according to claim 17, wherein determining the third set of protein sequences comprises searching a database of sequences based on the initial protein sequence to identify homologous sequences.

19. The method according to claim 17 or 18, wherein the method comprises performing filtering of the third set of protein sequences, to reduce the number of sequences in the third set of protein sequences, before the number of sequences in the third set is further reduced.

20. The method according to claim 19, wherein the filtering is based on the similarity of the sequences in the third set of sequences with respect to the initial protein sequence.

21. The method according to any one of claims 17 to 20, wherein the method further comprises performing subsampling of the third set of protein sequences to reduce the number of sequences in the third set.

22. The method according to claim 21, wherein subsampling comprises sequence reweighted sampling.

23. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of claims 1 to 22.

24. A method of producing a protein sequence having an improved fitness, with respect to an initial protein sequence, for a particular use, the method comprising: producing the identified one or more protein sequences of any one of claims 1 to 16 that are predicted by the fitness prediction model to have higher fitness than the initial protein sequence.