Optimizing viral capsid using machine learning

By combining embedded neural networks and machine learning models, the problems of high laboratory costs and data noise in viral capsid quality prediction were solved, achieving fast, low-cost, and highly accurate capsid screening, thereby reducing the cost of gene therapy.

CN120615216APending Publication Date: 2025-09-09SANOFI SA(FR)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380092507.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-18
Filing Date
2023-11-28
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the process of generating viral capsids using existing technologies, laboratory experiments are costly, time-consuming, and contain high data noise, which results in limited accuracy and robustness of machine learning models in predicting capsid quality, and there is a risk of overfitting.

Method used

An embedded neural network is used to process the amino acid sequence data of capsid monomers to generate embeddings. Combined with the fitness prediction machine learning model and the error prediction machine learning model, unsupervised learning and training data rebalancing are used to generate the fitness score of the capsid and identify high-quality capsid monomers.

Benefits of technology

The accuracy and efficiency of capsid quality prediction are improved, and capsid monomers with high manufacturability and other desired characteristics can be screened quickly and at low cost, thereby reducing the cost of gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120615216A_ABST
    Figure CN120615216A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating fitness scores for garment monomers. In one aspect, a method includes: receiving data defining an amino acid sequence of a capsid monomer; processing data defining an amino acid sequence of the capsid monomer using an embedded neural network according to values of a set of embedded neural network parameters to generate an embedding of the capsid monomer; processing the embedding of the capsid monomers using a fitness prediction machine learning model to generate a fitness score characterizing the predicted quality of the capsid; and outputting a fitness score characterizing the predicted quality of the capsid.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to the optimization of viral capsids using machine learning.

[0002] Virus may refer to a submicroscopic infectious pathogen capable of replicating within the living cells of an organism, and capsid may refer to the protein coat of a virus that encloses the genetic material of the virus.

[0003] The capsid can comprise an assembly of repeating protein structural units, referred to as "capsid monomers" or simply "monomers." The capsid can comprise any suitable number of monomers, e.g., 20 monomers, 50 monomers, 100 monomers, etc. The capsid monomer can comprise an amino acid sequence (chain), wherein each amino acid is selected from a set of possible amino acids, e.g., the standard set of 20 alpha amino acids, e.g., glycine, alanine, valine, etc. The amino acid sequence of a capsid monomer can be of any suitable length (i.e., wherein the "length" of an amino acid sequence refers to the number of positions occupied by amino acids in the amino acid sequence), e.g., a length of 10, 20, 30, etc.

[0004] Adeno-associated virus (AAV) is a small virus that can infect humans but causes a mild immune response, and it is not known whether it causes disease. AAV has properties that make it a suitable candidate for constructing viral vectors for gene therapy, such as its apparent non-pathogenicity, its ability to infect non-dividing cells, and its ability to stably integrate into a specific site in the host cell genome. The AAV capsid can include a mixture of capsid monomers VP1, VP2, and VP3, for a total of 60 monomers, which together assemble to form a capsid with icosahedral symmetry. Summary of the Invention

[0005] This specification describes a fitness prediction system implemented as a computer program on one or more computers at one or more locations that can process data defining the amino acid sequence of a capsid monomer to generate a fitness score that characterizes the predicted quality of the capsid. The quality of the capsid can refer to any suitable property of the capsid, such as the manufacturability of the capsid, the ability of a virus comprising the capsid to avoid neutralization, the immunoreactivity of the capsid, etc., as described in more detail below.

[0006] Throughout this specification, an "embedding" may refer to an ordered collection of values ​​(eg, vectors, matrices, or other numerical tensors).

[0007] Throughout this specification, a "block" in a neural network may refer to a collection of neural network layers in a neural network.

[0008] According to a first aspect, a method is provided, comprising: receiving data defining an amino acid sequence of a capsid monomer; processing the data defining the amino acid sequence of the capsid monomer using an embedded neural network based on values ​​of a set of embedded neural network parameters to generate an embedding of the capsid monomer; processing the embedding of the capsid monomer using a fitness prediction machine learning model to generate a fitness score characterizing a predicted quality of the capsid; and outputting the fitness score characterizing the predicted quality of the capsid.

[0009] In some embodiments, the method further comprises processing the representation of the capsid monomer using an error prediction machine learning model to generate a fitness error, wherein the fitness error is an estimate of the error in the fitness score generated by the fitness prediction machine learning model.

[0010] In some embodiments, the method further includes updating the fitness score by combining: (i) the fitness score generated by the fitness prediction machine learning model, and (ii) the fitness error generated by the error prediction machine learning model.

[0011] In some implementations, updating the fitness score includes summing the fitness score generated by the fitness prediction machine learning model and the fitness error generated by the error prediction machine learning model.

[0012] In some embodiments, the fitness prediction machine learning model has a different model architecture than the error prediction machine learning model.

[0013] In some embodiments, processing a representation of the capsid monomer using the error prediction machine learning model comprises processing a representation of the amino acid sequence of the capsid monomer using the error prediction machine learning model.

[0014] In some embodiments, the error prediction machine learning model processes different model inputs than the fitness prediction machine learning model.

[0015] In some embodiments, processing data defining the amino acid sequence of the capsid using the embedded neural network includes: instantiating a corresponding amino acid embedding corresponding to each position in the amino acid sequence of the capsid monomer; updating each amino acid embedding by processing the amino acid embeddings through one or more neural network layers of the embedded neural network; and generating an embedding for the capsid protein monomer based on one or more of the updated amino acid embeddings.

[0016] In some embodiments, generating an embedding of the capsid protein monomer based on one or more of the updated amino acid embeddings comprises generating an embedding of the capsid protein monomer by combining the updated amino acid embeddings.

[0017] In some embodiments, updating each amino acid embedding by processing the amino acid embeddings through one or more neural network layers of the embedded neural network includes processing the amino acid embeddings using one or more self-attention neural network layers of the embedded neural network.

[0018] In some embodiments, processing data defining the amino acid sequence of the capsid monomer using the embedded neural network includes: instantiating an amino acid embedding sequence including a corresponding amino acid embedding corresponding to each position in the amino acid sequence of the capsid monomer; and sequentially processing each amino acid embedding in the amino acid embedding sequence using a recurrent neural network to generate an updated hidden state of the recurrent neural network; and determining the embedding of the capsid monomer based on the updated hidden state of the recurrent neural network.

[0019] In some embodiments, the embedded neural network has been trained on a set of training proteins to perform a token unmasking task.

[0020] In some embodiments, training the embedded neural network to perform the token unmasking task includes, for each training protein: generating a masked amino acid sequence by masking one or more positions in the amino acid sequence of the training protein; processing data defining the masked amino acid sequence using the embedded neural network to determine a corresponding probability distribution over the set of possible amino acids for each masked position in the amino acid sequence of the training protein; and training the embedded neural network to optimize an objective function that measures, for each masked position in the amino acid sequence of the training protein, the error between: (i) the probability distribution generated for the masked position, and (ii) the identity of the amino acid at the masked position in the training protein.

[0021] In some embodiments, the embedded neural network has been trained on a set of training proteins to perform the next token prediction task.

[0022] In some embodiments, training the embedded neural network to perform the next token prediction task includes, for each training protein: using the embedded neural network, generating a corresponding probability distribution over the set of possible amino acids for each of one or more target positions in the amino acid sequence of the training protein, wherein, for each target position, the probability distribution generated for the target position depends only on the amino acid at the previous position in the amino acid sequence of the training protein; and training the embedded neural network to optimize an objective function that measures, for each target position in the amino acid sequence of the training protein, the error between: (i) the probability distribution generated for the target position, and (ii) the identity of the amino acid at the target position in the training protein.

[0023] In some embodiments, the fitness prediction neural network has been trained on a set of training examples, wherein each training example corresponds to a respective training monomer of a capsid and comprises: (i) an embedding of the training monomer of the capsid, and (ii) a target fitness score for the training monomer of the capsid.

[0024] In some embodiments, training the fitness prediction neural network on the set of training examples includes rebalancing the set of training examples based on target fitness scores specified by the training examples.

[0025] In some embodiments, the fitness prediction machine learning model is a parametric machine learning model.

[0026] In some embodiments, the fitness prediction machine learning model includes one or more of the following: a neural network model, a random forest model, or a support vector regression model.

[0027] In some embodiments, the fitness prediction machine learning model is a non-parametric model.

[0028] In some embodiments, the fitness prediction machine learning model comprises a k-nearest neighbor model.

[0029] In some embodiments, the predicted quality of the capsid is indicative of the predicted manufacturability of the capsid.

[0030] In some embodiments, the manufacturability of the capsid is defined based on the number of capsid instances produced when a plasmid encoding the capsid is transfected into one or more cells.

[0031] In some embodiments, the number of capsid instances generated when a plasmid encoding the capsid is transfected into one or more cells is normalized by the number of wild-type sequences generated.

[0032] In some embodiments, the predicted quality characterization of the capsid comprises the predicted ability of the capsid to prevent viral neutralization.

[0033] In some embodiments, the predicted quality of the capsid is indicative of the predicted immunoreactivity of the capsid.

[0034] In some embodiments, the predicted quality characterization of the capsid comprises the predicted ability of the capsid to penetrate a target tissue.

[0035] In some embodiments, the predicted quality of the capsid is indicative of the packaging ability of the capsid.

[0036] In some embodiments, the predicted quality characterization of the capsid comprises the predicted ability of the capsid to virally integrate into a host genome.

[0037] In some embodiments, the capsid corresponds to the capsid of an adeno-associated virus (AAV).

[0038] In some embodiments, the method further comprises: selecting a capsid monomer from a set of possible capsid monomers using the fitness prediction machine learning model; and physically generating one or more viruses comprising proteins having the selected capsid protein monomer.

[0039] In some embodiments, selecting a capsid monomer from a set of possible capsid monomers using the fitness prediction machine learning model includes: generating a corresponding fitness score for each capsid monomer in the set of possible capsid monomers using the fitness prediction machine learning model; and selecting a capsid monomer based on the fitness scores.

[0040] In some embodiments, selecting a capsid monomer based on the fitness scores comprises selecting the capsid monomer with the highest fitness score from among the capsid monomers included in the set of possible capsid monomers.

[0041] In some embodiments, the method further comprises administering the generated virus to a subject to achieve a therapeutic effect on the subject.

[0042] According to another aspect, a system is provided, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the methods described herein.

[0043] According to another aspect, one or more non-transitory computer storage media are provided that store instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the methods described herein.

[0044] Particular embodiments of the subject matter described in this specification can be implemented to realize one or more of the following advantages.

[0045] The fitness prediction system described in this specification utilizes a machine learning model to predict capsid quality.

[0046] Typically, the accuracy and robustness of a machine learning model may be limited by the amount and quality of training data available to train the machine learning model. In many cases, generating capsid fitness training data for training a fitness prediction system involves conducting laboratory experiments, for example, by measuring the number of capsid instances generated by transfecting cells with plasmids encoding these capsids. These laboratory experiments can be expensive and time-consuming, and may produce "noisy" (e.g., inaccurate or inconsistent) data.

[0047] Furthermore, capsid quality (e.g., manufacturability) may result from a complex interplay of biology, chemistry, and mechanics, and thus, to predict capsid quality, machine learning models must perform complex implicit reasoning. The reasoning power of a machine learning model may depend on the number of parameters of the model, for example, and thus increasing the number of parameters of a machine learning model may improve its reasoning power. Some machine learning models that perform complex prediction tasks may include a large number of parameters, for example, millions or billions of parameters. However, increasing the number of parameters of a machine learning model may increase the likelihood that the machine learning model will "overfit" during training, particularly when some of the residual variation (noise) in the training data is captured as representing variation in the underlying structure in the training data. The risk of overfitting may be exacerbated when the training data is limited in quantity and subject to noise, as was the case with the capsid fitness training data described above.

[0048] The fitness prediction system described in this specification implements various innovations to address these issues.

[0049] To generate a fitness score for the capsid, the fitness prediction system can process data defining the amino acid sequences of the capsid monomers using an embedded neural network to generate embeddings of the capsid monomers. The fitness prediction system can then process the embeddings of the capsid monomers using a fitness prediction machine learning model to generate a fitness score for the capsid.

[0050] The fitness prediction system can use unsupervised learning techniques to train embedded neural networks on large amounts of protein data for which labels (such as capsid fitness scores) are not available. For example, the fitness prediction system can train an embedded neural network to perform a token unmasking task or a next token prediction task, as will be described in more detail below. Training the embedded neural network to perform the unsupervised task enables the embedded neural network to learn to generate capsid monomer embeddings that compactly encode the biological / chemical / mechanical properties of the capsid monomers, which can be used by the fitness prediction machine learning model to generate accurate fitness scores. In addition, training the embedded neural network on the unsupervised task does not rely on capsid fitness training data and is therefore not affected by the quantity and quality limitations of the available capsid fitness training data.

[0051] After generating a fitness score for the capsid using the fitness prediction machine learning model, the fitness prediction system can use a separate error prediction machine learning model to generate a fitness error that predicts errors in the fitness score. The fitness prediction system can then use the fitness error to update the fitness score, for example, to correct errors in the fitness score, which can improve the overall accuracy of the fitness prediction system. More specifically, the fitness prediction system can train the error prediction machine learning model to learn to correct systematic errors in the fitness scores generated by the fitness prediction machine learning model. The error prediction machine learning model can detect and avoid certain systematic errors of the fitness prediction machine learning model, for example, by adopting a different model architecture than the fitness prediction machine learning model, or by processing inputs in a different form than the fitness prediction machine learning model, as will be described in more detail below.

[0052] The fitness prediction system can rebalance the capsid fitness training data to improve the accuracy and learning rate of the machine learning model included in the fitness prediction system. For example, if the training data is not rebalanced, the number of training examples with low capsid fitness scores may be significantly higher than the number of training examples with high capsid fitness scores. Training on unbalanced training data may limit the ability of training examples with high capsid fitness scores to influence the parameter values ​​of the machine learning model, thereby reducing the ability of the machine learning model to learn the characteristics of capsid monomers that produce high capsid fitness scores. The fitness prediction system can rebalance the training data to enhance the influence of training examples with high capsid fitness scores. For example, the fitness prediction system can resample the training data to increase the proportion of training examples with high capsid fitness scores, or can modify the objective function used during training to increase the penalty for prediction errors of training examples corresponding to high capsid fitness scores.

[0053] After training, the fitness prediction system can be used to screen large libraries of capsid monomers (e.g., comprising millions or billions of possible capsid monomers) to identify capsid monomers that can produce capsids with high fitness scores (e.g., capsids with high manufacturability). That is, the fitness prediction system enables automated searches in the space of possible capsid monomers to identify capsid monomers with desired properties. The fitness prediction system can be able to explore capsids that are orders of magnitude larger than those that can be physically generated and tested in laboratory experiments. In addition, the fitness prediction system operates quickly (e.g., it takes only 1 second or less to generate a predicted fitness score), is less expensive, and consumes relatively fewer operating resources than laboratory experiments.

[0054] The fitness prediction system can be used to identify the capsids of viral vectors for gene therapy.Gene therapy has the potential to revolutionize healthcare, for example, by enabling single-dose treatment with therapeutic viral vectors (e.g., AAV vectors) to address diseases (such as hemophilia, leukemia, melanoma, etc.). However, gene therapy treatments are usually expensive, and in some cases, it costs $1 million or more to treat one patient. High costs are a major obstacle to the full realization of the potential transformative effects of gene therapy. A major factor in the cost of gene therapy is the manufacturing cost of viral vector capsids. The fitness prediction system described in this specification enables the discovery of viral capsids with a high level of manufacturability, thereby significantly reducing the cost of gene therapy.

[0055] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 An example environment for screening a library of capsid monomers using a fitness prediction system is shown.

[0057] Figure 2 An example fitness prediction system is shown.

[0058] Figure 3 is a flow chart of an example process for generating a fitness score that characterizes the predicted quality of a capsid.

[0059] Figure 4 is a flow chart of an example process for generating embeddings of capsid monomers using an embedding neural network with a recurrent neural network architecture.

[0060] Figure 5 is a flow chart of an example process for generating embeddings of capsid monomers using an embedding neural network having a neural network architecture that includes a sequence of update blocks.

[0061] Figure 6 is a flowchart of an example process for training an embedded neural network to perform the next token prediction task.

[0062] Figure 7 is a flowchart of an example process for training an embedded neural network to perform a token unmasking task.

[0063] Figure 8 A graphical representation of the accuracy achieved by the fitness prediction system described in this specification on the task of predicting a fitness score characterizing capsid manufacturability is provided.

[0064] Figure 9A comparison of the accuracy achieved by the fitness prediction system with and without the error prediction machine learning model is provided.

[0065] Figure 10 A comparison of the prediction accuracy achieved by the fitness prediction system under different model architecture choices of the fitness prediction machine learning model is provided.

[0066] Figure 11 A comparison of the prediction accuracy achieved by the fitness prediction system when no training data rebalancing is performed and when training data rebalancing is performed is provided.

[0067] Figure 12 A comparison of the prediction accuracy achieved by the fitness prediction machine learning model when the embedded neural network generates 1900-dimensional capsid monomer embeddings and when the embedded neural network generates 256-dimensional capsid monomer embeddings is provided.

[0068] Like reference numbers and designations throughout the various drawings indicate like elements. DETAILED DESCRIPTION

[0069] Figure 1 An example environment 100 for screening a capsid monomer library 102 using a fitness prediction system 200 is shown.

[0070] The capsid monomer library 102 defines a set of capsid monomers, wherein each capsid monomer is represented by a corresponding amino acid sequence. The capsid monomer library 102 can include any appropriate number of capsid monomers, for example, one thousand, one million, or one billion capsid monomers.

[0071] The capsid monomer library 102 can be generated in any of a variety of possible ways. For example, some or all of the capsid monomers in the capsid monomer library 102 can be mutations of the "original" capsid monomers (e.g., VP1, VP2, or VP3 capsid monomers of AAV). More specifically, each capsid monomer in the capsid monomer library 102 can be generated by modifying the identity of the corresponding amino acids at one or more positions in the amino acid sequence of the original capsid monomer. The positions in the amino acid sequence of the original capsid monomer can be selected in any appropriate manner for mutation, for example, by random selection or by selection according to predefined rules. The identity of the new amino acids at the positions in the amino acid sequence of the original capsid monomer can be selected in any appropriate manner, for example, by random selection from a probability distribution of possible amino acid sets. The corresponding amino acid sequence of each capsid monomer in the capsid monomer library can differ from the amino acid sequence of the original capsid monomer at any appropriate number of positions, for example, 1 position, 3 positions, or 10 positions.

[0072] The fitness prediction system 200 is configured to process capsid monomers (e.g., from the capsid monomer library 102) to generate a fitness score 104 for the capsid monomer that characterizes the predicted quality of the corresponding capsid. The fitness score 104 for the capsid monomer can characterize any suitable quality of the corresponding capsid. Several examples of possible fitness scores 104 are described below.

[0073] In some embodiments, the fitness prediction system 200 generates a fitness score 104 for a capsid monomer that characterizes the manufacturability of the corresponding capsid. The manufacturability of a capsid characterizes the rate at which a capsid can be generated when a plasmid encoding the capsid is transfected into one or more cells. (More specifically, the manufacturability of a capsid can be based, at least in part, on the number of capsid instances generated when a plasmid encoding the capsid is transfected into one or more cells). In other words, the manufacturability of a capsid characterizes the efficiency with which a capsid can be generated. The manufacturability of a capsid indirectly characterizes various structural and biochemical properties of the capsid that affect its manufacture, for example, the ease with which the capsid monomers of the capsid assemble into the three-dimensional structure of the capsid, the stability of the capsid, and the like.

[0074] In some embodiments, the fitness prediction system 200 can generate a fitness score 104 for a capsid monomer, which represents the predicted ability of a virus comprising the corresponding capsid to avoid neutralization in an organism. Viral neutralization refers to the process of reducing the ability of a virus to infect cells, such as by antibodies binding to epitopes on the surface of the virus.

[0075] In some embodiments, the fitness prediction system 200 can generate a fitness score 104 for a capsid monomer, the fitness score representing the predicted immunoreactivity of the corresponding capsid. The immunoreactivity of a capsid is a measure of the immune response elicited by the capsid.

[0076] In some embodiments, the fitness prediction system 200 can generate a fitness score 104 for a capsid monomer, which represents the predicted ability of a virus comprising the capsid to penetrate a target tissue in a living organism. The target tissue can be, for example, a tissue corresponding to a specific organ, such as liver tissue, brain tissue, eye tissue, etc.

[0077] In some embodiments, the fitness prediction system 200 can generate a fitness score 104 for a capsid monomer, which represents the predicted packaging capacity of the corresponding capsid. The packaging capacity of a capsid is a measure of the amount of genetic material (e.g., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA)) that the capsid can accommodate.

[0078] In some embodiments, the fitness prediction system 200 can generate a fitness score 104 for a capsid monomer, the fitness score representing the predicted ability of a virus comprising the corresponding capsid to integrate into a host genome.

[0079] The fitness prediction system 200 can screen the capsid monomer library 102 to identify capsid monomers corresponding to capsids having desired properties. More specifically, the fitness prediction system 200 can generate a corresponding fitness score 104 for each capsid monomer in the capsid monomer library 102. The fitness prediction system 200 can designate an appropriate subset of capsid monomers in the capsid monomer library 102 as "target" capsid monomers 106 based at least in part on the fitness scores 104. Then, for each target capsid monomer 106, a virus including a capsid corresponding to the target capsid monomer 106 can be manufactured 108 (i.e., physically generated) using an appropriate manufacturing technique.

[0080] The generated viruses can be used in any of a variety of applications. For example, the generated viruses can be applied to a subject 112 as a therapeutic agent 110 to achieve a therapeutic effect on the subject. For example, the generated viruses can be used as viral vectors for gene therapy, as described above. The generated viruses can be applied to any suitable subject, such as a mouse, cat, dog, pig, or human, to achieve any suitable therapeutic effect, such as, for example, treating a disease such as hemophilia, leukemia, melanoma, etc.

[0081] In certain examples, a viral vector (e.g., an AAV vector, such as a recombinant AAV vector) comprising a capsid of interest selected using a fitness prediction system can be generated and administered to a subject for gene therapy, as described above.

[0082] The fitness prediction system 200 can select an appropriate subset of capsid monomers from the capsid monomer library in any of various possible ways to designate them as target capsid monomers. For example, the fitness prediction system 200 can designate any capsid monomer whose fitness score 104 meets a predefined threshold as a target capsid monomer. As another example, the fitness prediction system 200 can designate a predefined number of capsid monomers with the highest fitness scores 104 as target capsid monomers.

[0083] The fitness prediction system 200 can designate any suitable number of capsid monomers (e.g., 10 capsid monomers, 100 capsid monomers, or 1000 capsid monomers) in the capsid monomer library 102 as target capsid monomers 106. In some cases, the fitness prediction system 200 designates only a small fraction of the total number of capsid monomers in the capsid monomer library (e.g., <1%, <0.1%, or less than <0.01% of the total number of capsid monomers in the capsid monomer library) as target capsid monomers.

[0084] Figure 2An example fitness prediction system 200 is shown. Fitness prediction system 200 is an example of a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described below are implemented.

[0085] The fitness prediction system 200 is configured to process data defining the amino acid sequence of a capsid monomer 202 to generate a fitness score 104 for the capsid monomer 202. The fitness score 104 characterizes the predicted quality of the capsid corresponding to the capsid monomer 202. For example, as described above with reference to Figure 1 As described above, the fitness score 104 can characterize the predicted manufacturability of the capsid, the predicted ability of a virus comprising the capsid to avoid neutralization, the predicted immunoreactivity of the capsid, the predicted ability of a virus comprising the capsid to penetrate a target tissue, the predicted packaging capacity of the capsid, or the predicted ability of a virus comprising the capsid to integrate into a host genome. Figure 1 As described, the fitness scores 104 generated by the fitness prediction system 200 can be used to screen a library of capsid monomers to identify target capsid monomers with desired properties.

[0086] The fitness prediction system 200 includes an embedded neural network 204, a fitness prediction machine learning model 208, and an error prediction machine learning model 212, each of which will be described in more detail below.

[0087] The embedded neural network 204 is configured to process data defining the amino acid sequence of the capsid monomer 202 to generate an embedding 206 of the capsid monomer 202 .

[0088] The embedded neural network 204 can be configured to process any suitable representation of the amino acid sequence of the capsid monomer 202. For example, the embedded neural network 204 can process the representation of the amino acid sequence of the capsid monomer 202 as an amino acid embedding sequence. The amino acid embedding at each position in the amino acid embedding sequence can represent the identity of the amino group at the corresponding position in the amino acid sequence of the capsid monomer. The amino acid embedding representing the amino acid can be, for example, a one-hot embedding, or any other suitable predefined embedding.

[0089] Embedded neural network 204 can have any suitable neural network architecture that enables embedded neural network 204 to perform its described functions. In particular, embedded neural network 204 can include any suitable number (e.g., 5 layers, 10 layers, 50 layers, etc.) of any suitable type of neural network layers (e.g., fully connected layers, convolutional layers, recurrent layers, attention layers, etc.) connected in any suitable configuration (e.g., as a linear sequence of layers).

[0090] In some embodiments, the embedded neural network 204 can be a recurrent neural network configured to sequentially process each amino acid embedding in the amino acid embedding sequence representing the amino acid sequence of the capsid monomer 202. Figure 4 An example process for generating an embedding 206 of a capsid monomer 202 using an embedding neural network 204 having a recurrent neural network architecture is described in more detail.

[0091] In some embodiments, the embedded neural network 204 may include a sequence of update blocks, wherein each update block is configured to process a current set of embeddings representing a capsid monomer to generate an updated set of embeddings representing the capsid monomer. Each update block may include, for example, one or more self-attention neural network layers. Figure 5 An example process for generating an embedding 206 of a capsid monomer 202 using an embedded neural network 204 including a sequence of update blocks is described in greater detail.

[0092] Fitness prediction system 200 can train embedded neural network 204 on a set of training examples, where each training example corresponds to a corresponding protein and defines the amino acid sequence of the protein. Fitness prediction system 200 can train embedded neural network 204 on the set of training examples using unsupervised learning techniques (i.e., training techniques that do not rely on training examples labeled with fitness scores). Several example techniques for training embedded neural network 204 on the set of training examples using unsupervised learning techniques are described below.

[0093] In some embodiments, the fitness prediction system 200 trains the embedded neural network 204 to perform a token unmasking task. More specifically, for each training example, the fitness prediction system 200 can "mask" the identity of the amino acid at one or more positions in the protein amino acid sequence. (Masking the identity of the amino acid can refer to replacing the amino acid embedding representing the amino acid with a default embedding, as will be referenced below. Figure 5 ). The fitness prediction system 200 can then train the embedded neural network 204 to process the masked amino acid sequence to predict the identity of the amino acid at the masked position in the amino acid sequence. The token unmasking task requires the embedded neural network 204 to learn to infer the identity of the masked amino acid from the contextual information provided by the remaining unmasked portion of the amino acid sequence. Figure 7 An example process for training an embedded neural network to perform a token unmasking task is described in more detail.

[0094] In some embodiments, the fitness prediction system 200 trains an embedded neural network to perform a next token prediction task. More specifically, for each training example, the fitness prediction system 200 may designate one or more positions in a protein amino acid sequence as "target" positions. For each target position, the fitness prediction system 200 may train an embedded neural network 204 to learn to predict the identity of the amino acid at the target position based solely on the identity of the amino acid at the previous position in the amino acid sequence. The next token prediction task requires the embedded neural network 204 to learn to infer the identity of the amino acid at a position in the amino acid sequence based on the identity of the amino acid at the previous position in the amino acid sequence. Figure 6 Describes an example process for training an embedded neural network to perform the next token prediction task in more detail.

[0095] Training the embedded neural network 204 using unsupervised learning techniques, for example, to perform tasks such as token unmasking, next token prediction, or both, can enable the embedded neural network 204 to learn to generate embeddings that densely encode rich information characterizing a protein (e.g., a capsid monomer).

[0096] In some cases, in conjunction with or as an alternative to training the embedded neural network 204 using unsupervised learning techniques, the fitness prediction system 200 can use supervised learning techniques to train the embedded neural network 204 on an auxiliary set of training examples. Each training example can correspond to a corresponding protein and can define: (i) the amino acid sequence of the protein, and (ii) a label associated with the protein. The label can represent any suitable characteristic of the protein, such as the structure of the protein, the stability of the protein, etc.

[0097] The fitness prediction system 200 can jointly train an embedded neural network and a "projection" neural network on a set of supervised training examples. The projection neural network can be configured to process the protein embeddings generated by the embedded neural network to generate predictions for labels associated with the proteins.

[0098] To train the embedded neural network on the supervised training examples, the fitness prediction system 200 can process the amino acid sequence of the protein corresponding to the training example using the embedded neural network 204 to generate a protein embedding. The fitness prediction system 200 can then process the protein embedding using a projection neural network to generate a predicted label associated with the protein. The fitness prediction system 200 can evaluate a supervised objective function (which measures the error of the predicted label generated for the protein) and backpropagate the gradient of the supervised objective function into the embedded neural network 204 through the projection neural network. After training, the projection neural network can be discarded, and the fitness prediction system 200 can use the trained embedded neural network 204 to generate an embedding for generating the fitness score 104.

[0099] The supervised objective function can measure the error of the predicted labels generated by the embedded neural network, for example, in the form of cross-entropy error, squared error, or in any other suitable manner. Backpropagating the gradients of the objective function through the neural network can refer to: determining the gradients of the objective function with respect to the parameters of the neural network, and then using these gradients (for example, using an appropriate gradient descent optimization technique such as Adam or RMSprop) to adjust the parameter values ​​of the neural network.

[0100] The fitness prediction machine learning model 208 is configured to process the embeddings 206 of the capsid monomers 202 to generate a fitness score 104 that characterizes the predicted quality of the capsid corresponding to the capsid monomers 202 .

[0101] The fitness prediction machine learning model 208 can have any suitable machine learning model architecture that enables the fitness prediction machine learning model 208 to perform its described functions. For example, the fitness prediction machine learning model can be a parametric machine learning model, such as a neural network model, a random forest model, or a support vector regression model. As another example, the fitness prediction machine learning model can be a non-parametric model, such as a k-nearest neighbor model.

[0102] The fitness prediction system 200 can train a fitness prediction machine learning model on a set of training examples. Each training example corresponds to a corresponding capsid monomer and defines: (i) an embedding of the capsid monomer amino acid sequence generated using an embedded neural network; and (ii) a target fitness score associated with the capsid monomer. The target fitness score for the capsid monomer represents the prediction target of the fitness prediction system, i.e., the "true" fitness that the fitness prediction system should generate for the capsid monomer.

[0103] The fitness prediction system 200 can train the fitness prediction machine learning model 208 on a set of training examples and their associated target fitness scores using any appropriate supervised machine learning training technique. More specifically, for each training example, the fitness prediction system 200 can train the fitness prediction machine learning model to process the capsid monomer embeddings included in the training example to generate a predicted fitness score that matches the target fitness score specified by the training example. The fitness prediction system 200 can train the fitness prediction machine learning model 208 to optimize an objective function that, for each training example, measures the error between: (i) the target fitness score, and (ii) the predicted fitness score generated by the fitness prediction machine learning model 208. The objective function can measure the error in any appropriate manner, for example, in the form of squared error or absolute error.

[0104] During training of the fitness prediction machine learning model 208, the fitness prediction system 200 can rebalance the set of training examples based on the target fitness scores associated with the training examples. In particular, the set of training examples can include significantly more training examples associated with low fitness scores than it includes training examples associated with high fitness scores. Training the fitness prediction machine learning model 208 without rebalancing the set of training examples may reduce the accuracy of the fitness prediction machine learning model 208. The fitness prediction system 200 can rebalance the set of training examples in any of a variety of possible ways. Several example techniques for rebalancing the set of training examples are described below.

[0105] In some implementations, the fitness prediction system 200 can rebalance the set of training examples by replicating training examples associated with high fitness scores (e.g., fitness scores exceeding a threshold) until at least a threshold portion of the set of training examples are associated with high fitness scores.

[0106] In some embodiments, the fitness prediction system 200 can rebalance the training data set by determining a corresponding weight factor associated with each training example based on the fitness score associated with the training example. For example, the fitness prediction system 200 can generate a weight factor for the training example by processing the fitness score of the training example using a sigmoid function. The weight factor of the training example can control the influence of the training example on the parameter value of the fitness prediction machine learning model during training, for example, so that the training examples with higher weight factors have a greater influence on the fitness prediction machine learning model. For example, if the fitness prediction machine learning model is a neural network model, then as part of training the fitness prediction machine learning model on the training example, the fitness prediction system 200 can scale the gradient of the objective function by the weight factor. Therefore, the weight factor can define the personalized learning rate of the corresponding training example during training of the neural network that implements the fitness prediction machine learning model.

[0107] In some embodiments, the fitness prediction system 200 can rebalance the training data set by determining a corresponding penalty factor associated with each training example based on the fitness score associated with the training example. For example, the fitness prediction system 200 can generate a penalty factor for the training example by processing the fitness score of the training example using a sigmoid function. The penalty factor for the training example can control the penalty imposed on the fitness prediction machine learning model for incorrectly predicting the fitness score of the training example during training. For example, the penalty factor for the training example can multiply and scale the term in the objective function that measures the error in the predicted fitness score generated by the fitness prediction machine learning model 208 for the training example.

[0108] The target fitness scores for capsid monomers in the training examples can be generated in any suitable manner (e.g., through physical experiments). Several example techniques for generating target fitness scores for capsid monomers are described below.

[0109] In some embodiments, the fitness score of a capsid monomer characterizes the manufacturability of the corresponding capsid. To generate a target fitness score for a capsid monomer, a plasmid encoding the capsid can be transfected into one or more cells. Experimental measurements can then be performed to determine: (i) the number of capsid instances generated by the one or more cells, and (ii) the number of wild-type capsid instances generated by the one or more cells. The fitness score of a capsid can be defined, for example, as the ratio of: (i) the number of capsid instances generated, and (ii) the number of wild-type capsid instances generated.

[0110] In some embodiments, the fitness score of a capsid monomer characterizes the ability of a virus comprising the corresponding capsid to avoid neutralization. To generate a target fitness score for a capsid monomer, a virus comprising the corresponding capsid can be introduced into an organism, and experimental measurements can be performed to determine the number of viruses that are subsequently neutralized. The fitness score of a capsid can be defined, for example, based on the ratio of: (i) the number of virus instances (having the capsid) that are neutralized, and (ii) the total number of virus instances (having the capsid).

[0111] In some embodiments, the fitness score of a capsid monomer characterizes the immunoreactivity of the capsid. To generate a target fitness score for a capsid monomer, the corresponding capsid can be introduced into an organism, and experimental measurements can be performed to determine the immune response generated by the organism. A target fitness score for the capsid can then be defined based on the measured immune response.

[0112] In some embodiments, the fitness score of a capsid monomer characterizes the ability of a virus comprising the corresponding capsid to penetrate a target tissue. To generate a target fitness score for a capsid monomer, a virus comprising the corresponding capsid can be introduced into an organism, and experimental measurements can be performed to determine the number of viruses that penetrate the target tissue. The target fitness score for a capsid monomer can then be defined based on the ratio of: (i) the number of virus instances (having the capsid) that penetrate the target tissue, and (ii) the total number of virus instances (having the capsid) that have been introduced into the organism.

[0113] In some embodiments, the fitness score of a capsid monomer characterizes the packaging capacity of the corresponding capsid. To generate a target fitness score for a capsid monomer, one or more instances of the corresponding capsid can be generated, and the amount of genetic material contained in the generated capsids can be experimentally measured. The target fitness score for the capsid monomer can then be defined based on the experimentally measured amount of genetic material contained in the corresponding capsid.

[0114] In some embodiments, the fitness score of a capsid monomer characterizes the ability of a virus comprising the corresponding capsid to integrate into the host genome. To generate a target fitness score for a capsid monomer, one or more instances of a virus comprising the corresponding capsid can be introduced into an organism. Experimental measurements can be performed to determine the number of viral instances that successfully integrate into the host genome. The target fitness score for a capsid monomer can then be defined based on the ratio of: (i) the number of viral instances (having the capsid) that are successfully integrated into the host genome of the organism, and (ii) the total number of viral instances (having the capsid) that are introduced into the organism.

[0115] The error prediction machine learning model 212 is configured to process the representation of the amino acid sequence of the capsid monomer 202 to generate a fitness error 214. The fitness error 214 is an estimate of the error in the fitness score 104 generated by the fitness prediction machine learning model 208.

[0116] The fitness prediction system 200 updates the fitness score 104 generated by the fitness prediction machine learning model 208 by combining: (i) the fitness score 104, and (ii) the fitness error 214. The fitness prediction system 200 can combine the fitness score 104 and the fitness error 214, for example, by summing the fitness score 104 and the fitness error 214. The fitness prediction system 200 can then output the updated fitness score 104 as the predicted fitness score for the input capsid monomer 202.

[0117] In some implementations, the error prediction machine learning model 212 can have a different model architecture than the fitness prediction machine learning model 208. For example, the error prediction machine learning model 212 can be a linear regression model, while the fitness prediction machine learning model 208 can be a random forest model.

[0118] In some embodiments, the error prediction machine learning model 212 can be configured to process different inputs (e.g., in a different format) than the fitness prediction machine learning model 208. For example, the error prediction machine learning model 212 can process representations of capsid monomers as amino acid embedding sequences, while the fitness prediction machine learning model 208 can process embeddings 206 of capsid monomers generated using the embedding neural network 204.

[0119] Configuring the error prediction machine learning model 212 to have a different architecture than the fitness prediction machine learning model 208 or to process different inputs than the fitness prediction machine learning model can cause the two machine learning models to be "orthogonal" to each other. More specifically, configuring these models differently can cause each model to perform implicit reasoning independent of and different from the other model, thereby reducing the mutual information between their outputs and maximizing the information gain achieved by combining their outputs to generate the updated fitness score 104.

[0120] The error prediction machine learning model 212 can have any suitable machine learning model architecture that enables the error prediction machine learning model 212 to perform its described functions. For example, the error prediction machine learning model can be a parametric machine learning model, such as a neural network model, a random forest model, or a support vector regression model. As another example, the fitness prediction machine learning model can be a non-parametric model, such as a k-nearest neighbor model.

[0121] The fitness prediction system 200 can train an error prediction machine learning model on a set of training examples. Each training example corresponds to a corresponding capsid monomer and defines: (i) the amino acid sequence of the capsid monomer; and (ii) a target fitness error associated with the capsid monomer. The target fitness error of the capsid monomer represents the prediction target of the error prediction machine learning model, that is, the "true" fitness error that the error prediction system should generate for the capsid monomer. The fitness prediction system 200 can generate a target fitness error for a training example by determining the error between: (i) the target fitness score for the training example, and (ii) the predicted fitness score generated by the (trained) fitness prediction machine learning model 208 for the prediction example.

[0122] The fitness prediction system 200 can train the error prediction machine learning model 212 based on a set of training examples and their associated target fitness errors using any appropriate supervised machine learning training technique. More specifically, for each training example, the fitness prediction system 200 can train the error prediction machine learning model to process data defining the amino acid sequence of the capsid monomer contained in the training example to generate a predicted fitness error that matches the target fitness error specified for the training example. The fitness prediction system 200 can train the error prediction machine learning model 212 to optimize an objective function that, for each training example, measures the error between: (i) the target fitness error, and (ii) the predicted fitness error generated by the error prediction machine learning model 212. The objective function can measure the error in any appropriate manner, for example, in the form of a squared error or an absolute error.

[0123] Figure 3 is a flow chart of an example process 300 for generating a fitness score that characterizes the predicted quality of a capsid. For convenience, process 300 will be described as being performed by a system consisting of one or more computers located in one or more locations. For example, a fitness prediction system (e.g., Figure 2 The fitness prediction system 200) can perform process 300.

[0124] The system receives data defining the amino acid sequence of a capsid monomer (302).

[0125] The system processes data defining the amino acid sequence of the capsid monomer using the embedded neural network according to values ​​of a set of embedded neural network parameters to generate an embedding of the capsid monomer (304).

[0126] The system processes the embeddings of the capsid monomers using a fitness prediction machine learning model to generate a fitness score (306) that characterizes the predicted quality of the capsid.

[0127] The system processes the representation of the capsid monomers using the error prediction machine learning model to generate a fitness error (308). The fitness error is an estimate of the error in the fitness score generated by the fitness prediction machine learning model.

[0128] The system updates the fitness score by combining: (i) the fitness score generated by the fitness prediction machine learning model, and (ii) the fitness error generated by the error prediction machine learning model (310).

[0129] The system outputs an updated fitness score (312).

[0130] Figure 4 is a flow chart of an example process 400 for generating an embedding of capsid monomers using an embedded neural network having a recurrent neural network architecture. For convenience, process 400 will be described as being performed by a system consisting of one or more computers located in one or more locations. For example, a fitness prediction system (e.g., Figure 2 The fitness prediction system 200) can perform process 400.

[0131] The system instantiates an amino acid embedding sequence representing the amino acid sequence of the capsid monomer (402). The amino acid embedding sequence includes a corresponding amino acid embedding corresponding to each position in the amino acid sequence of the capsid monomer. The amino acid embedding corresponding to a position in the amino acid sequence of the capsid monomer can represent the identity of the amino acid at that position, for example, by one-hot embedding.

[0132] The system sequentially processes each amino acid embedding in the amino acid embedding sequence using an embedded neural network having a recurrent neural network architecture to generate an updated hidden state of the embedded neural network (404). More specifically, before processing the first amino acid embedding in the amino acid embedding sequence, the system initializes the hidden state of the embedded neural network to, for example, a default hidden state. The embedded neural network then sequentially processes each amino acid embedding in the amino acid embedding sequence, starting with the first amino acid embedding, according to the order of the amino acid embeddings in the amino acid embedding sequence. Each time the embedded neural network processes an amino acid embedding, the embedded neural network uses the amino acid embedding to update the hidden state of the embedded neural network.

[0133] The embedded neural network includes one or more recurrent neural network layers, and the hidden states of the recurrent neural network layers collectively define the hidden states of the embedded neural network. Each recurrent neural network layer can have any suitable architecture, for example, a long short-term memory (LSTM) architecture or a gated recurrent unit (GRU) architecture.

[0134] The system determines a capsid monomer embedding for the capsid based on an updated hidden state of the embedded neural network after processing the last amino acid embedding in the amino acid embedding sequence (406). For example, the system can generate the capsid monomer embedding by processing the updated hidden state of the embedded neural network using one or more neural network layers (e.g., fully connected layers).

[0135] Figure 5 is a flow chart of an example process 500 for generating an embedding of capsid monomers using an embedded neural network having a neural network architecture that includes an update block sequence. For convenience, process 500 will be described as being performed by a system consisting of one or more computers located in one or more locations. For example, a fitness prediction system (e.g., Figure 2 The fitness prediction system 200) can perform process 500.

[0136] The system instantiates an amino acid embedding sequence representing the amino acid sequence of the capsid monomer (502). The amino acid embedding sequence includes a corresponding amino acid embedding corresponding to each position in the amino acid sequence of the capsid monomer. The amino acid embedding corresponding to a position in the amino acid sequence of the capsid monomer can, for example, represent the identity of the amino acid at that position by a one-hot embedding. Optionally, the system can combine each amino acid embedding with a position embedding representing the corresponding position in the amino acid sequence of the capsid monomer.

[0137] The system processes the amino acid embedding sequence using a sequence of one or more blocks (referred to as update blocks) included in the embedded neural network (504). Each update block is configured to process the amino acid embedding sequence using one or more neural network layers included in the update block to generate an updated amino acid embedding sequence. The first update block can be configured to receive an initial amino acid embedding sequence, and each subsequent update block can receive an updated amino acid embedding sequence generated by the previous update block. Each update block can have any suitable neural network architecture. In some embodiments, some or all update blocks include one or more self-attention neural network layers.

[0138] The system generates an embedding for the capsid monomer based on one or more of the updated amino acid embeddings generated by the final update block in the sequence of update blocks (506). For example, the system can designate the updated amino acid embedding corresponding to the last amino acid in the amino acid sequence of the capsid monomer as representing the embedding for the capsid monomer. As another example, the system can combine (e.g., average or sum) the updated amino acid embeddings generated by the final update block to generate the embedding for the capsid monomer.

[0139] Figure 6is a flow chart of an example process 600 for training an embedded neural network to perform a next token prediction task. For convenience, process 600 will be described as being performed by a system consisting of one or more computers located in one or more locations. For example, a fitness prediction system (e.g., Figure 2 The fitness prediction system 200) can perform process 600.

[0140] The system receives a training protein set (602).

[0141] The system performs steps 604 to 608 for each training protein. For convenience, steps 604 to 608 will be described with reference to a specific training protein in the set of training proteins.

[0142] The system designates one or more positions in the amino acid sequence of the training protein as target positions (604). The system can select the target positions, for example, randomly or according to a predefined selection rule.

[0143] The system uses an embedded neural network to generate a corresponding probability distribution over a set of possible amino acids for each target position in the amino acid sequence of the training protein (606). For each target position, the probability distribution generated for the target position depends only on the amino acid at the previous position in the amino acid sequence of the training protein. To generate the probability distribution for the target position, the system can use the embedded neural network to generate an embedding for the target position and then process the embedding using one or more neural network layers to generate the probability distribution. In some embodiments, the embedded neural network is a recurrent neural network, and to generate the embedding for the target position, the embedded neural network sequentially processes the amino acid embedding corresponding to each position before the target position to generate an updated hidden state. The system can then assign the updated hidden state of the embedded neural network as the embedding for the target position.

[0144] The system trains the embedded neural network to optimize an objective function that measures, for each target position in the amino acid sequence of the training protein, the error between: (i) a probability distribution generated for the target position, and (ii) the identity of the amino acid at the target position in the training protein (608). The objective function can measure the error, for example, in the form of a cross-entropy error.

[0145] Figure 7 is a flow chart of an example process 700 for training an embedded neural network to perform a token unmasking task. For convenience, process 700 will be described as being performed by a system consisting of one or more computers located in one or more locations. For example, a fitness prediction system (e.g., Figure 2The fitness prediction system 200) can perform process 700.

[0146] The system receives a training protein set (702).

[0147] The system performs steps 704 to 708 for each training protein. For convenience, steps 704 to 708 will be described with reference to a specific training protein in the set of training proteins.

[0148] The system generates a masked amino acid sequence by masking one or more positions in the amino acid sequence of the training protein (704). Masking a position in the amino acid sequence of the training protein can refer to replacing the amino acid embedding representing the identity of the amino acid at that position with a default embedding. The default embedding can be, for example, a predefined embedding, for example, an embedding in which each entry has a value of zero. The system can determine the positions to be masked in the amino acid sequence, for example, by randomly selecting the positions to be masked or by selecting the positions to be masked according to a predefined rule.

[0149] The system processes data defining a masked amino acid sequence using an embedded neural network to determine a corresponding probability distribution over a set of possible amino acids for each masked position in the amino acid sequence of the training protein (706). To generate a probability distribution for the masked positions, the system can use the embedded neural network to generate an embedding for the masked position and then process the embedding using one or more neural network layers to generate a probability distribution. In some embodiments, the embedded neural network includes one or more update blocks, each of which is configured to process an input embedding sequence to generate an updated embedding sequence. In these embodiments, to generate an embedding for the masked position, the embedded neural network can use a sequence of update blocks of the embedded neural network to process the masked amino acid sequence to generate an updated amino acid embedding sequence. Then, for each masked position, the system can specify the updated amino acid embedding for the masked position as the embedding that is processed to generate the probability distribution for the masked position.

[0150] The system trains the embedded neural network to optimize an objective function that measures, for each masked position in the amino acid sequence of the training protein, the error between: (i) a probability distribution generated for the masked position, and (ii) the identity of the amino acid at the masked position in the training protein (708). The objective function can measure the error, for example, in the form of a cross-entropy error.

[0151] Figure 8A graphical representation of the accuracy achieved by the fitness prediction system described herein in predicting a fitness score that characterizes capsid manufacturability is provided. The horizontal axis of the graph represents the experimental (true) fitness score, while the vertical axis of the graph represents the predicted fitness score. It will be appreciated that the experimental fitness score is strongly correlated with the predicted fitness score.

[0152] Figure 9 A comparison of the accuracy achieved by the fitness prediction system with and without an error prediction machine learning model is provided. More specifically, the bar chart on the left represents the prediction accuracy when the fitness prediction system does not use an error prediction machine learning model, while the bar chart on the right represents the prediction accuracy when the fitness prediction system uses an error prediction machine learning model. This accuracy is measured by the Pearson correlation coefficient between the experimental fitness score and the predicted fitness score.

[0153] Figure 10 A comparison of the prediction accuracy achieved by the fitness prediction system under different model architecture choices of the fitness prediction machine learning model is provided. The accuracy is measured by the Pearson correlation coefficient between the experimental fitness score and the predicted fitness score.

[0154] Figure 11 A comparison of the prediction accuracy achieved by the fitness prediction system when no training data rebalancing is performed (left bar graph) and when training data rebalancing is performed (right bar graph) is provided. The accuracy is measured by the Pearson correlation coefficient between the experimental fitness score and the predicted fitness score.

[0155] Figure 12 A comparison of the prediction accuracy achieved by the fitness prediction machine learning model when the embedded neural network generates 1900-dimensional capsid monomer embeddings (left bar graph) and when the embedded neural network generates 256-dimensional capsid monomer embeddings (right bar graph) is provided. The accuracy is measured by the Pearson correlation coefficient between the experimental fitness score and the predicted fitness score.

[0156] This specification uses the term "configuration" when referring to systems and computer program components. For a system consisting of one or more computers to be configured to perform a particular operation or action, it means that the system has installed thereon software, firmware, hardware, or a combination thereof that, when run, causes the system to perform those operations or actions. For one or more computer programs to be configured to perform a particular operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform those operations or actions.

[0157] The embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their structural equivalents), or a combination of one or more thereof. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs (i.e., one or more computer program instruction modules encoded on a tangible, non-transitory storage medium) for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more thereof. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical signal, optical signal, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by a data processing apparatus.

[0158] The term "data processing apparatus" refers to data processing hardware and encompasses various devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus may also be or further include dedicated logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus may optionally include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof.

[0159] A computer program (which may also be referred to or described as a program, software, software application, application, module, software module, script, or code) may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files that store one or more modules, subroutines, or code portions). A computer program may be deployed to execute on one computer, or on multiple computers located in one location or distributed across multiple locations and interconnected by a data communications network.

[0160] Throughout this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines may be installed on or run on the same computer or computers.

[0161] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry (e.g., an FPGA or ASIC), or by a combination of special purpose logic circuitry and one or more programmed computers.

[0162] A computer suitable for executing a computer program can be based on a general-purpose microprocessor or a special-purpose microprocessor or both, or based on a central processing unit of any other type. Typically, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by or incorporated into a dedicated logic circuit. Typically, a computer will also include one or more large-capacity storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operably coupled to one or more large-capacity storage devices for storing data, to receive data from the large-capacity storage device or to transmit data to the large-capacity storage device or to both receive and transmit data. However, a computer does not need to have such a device. In addition, a computer can be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), etc.

[0163] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0164] To provide for interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form, including auditory, voice, or tactile input. Additionally, a computer may interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on a user's device in response to a request received from the web browser. Furthermore, a computer may interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user.

[0165] The data processing apparatus for implementing machine learning models may also include, for example, dedicated hardware accelerator units for handling common compute-intensive portions of machine learning training or production (i.e., inference) workloads.

[0166] Machine learning models can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework).

[0167] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface, a web browser, or an application program through which a user can interact with an implementation of the subject matter described in this specification), or includes any combination of one or more such back-end components, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs) (e.g., the Internet).

[0168] A computing system may include a client and a server. The client and the server are typically remote from each other and typically interact via a communication network. The relationship between a client and a server arises from computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, a server transmits data (e.g., an HTML page) to a user device acting as a client, for example, to display data to a user interacting with the device and to receive user input from the user. Data generated at the user device (e.g., the result of a user interaction) can be received from the device at the server.

[0169] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Specific features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in a certain combination and even initially claimed as such, in some cases one or more features from the claimed combination may be deleted from that combination, and the claimed combination may involve subcombinations or variations of subcombinations.

[0170] Similarly, although operations are depicted in a particular order in the drawings and recited in a particular order in the claims, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0171] Specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying figures do not necessarily require the specific order shown or sequential order to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method comprising: receiving data defining an amino acid sequence of a capsid monomer; processing data defining the amino acid sequence of the capsid monomer using the embedded neural network according to values ​​of a set of embedded neural network parameters to generate an embedding of the capsid monomer; processing the capsid monomer embeddings using a fitness prediction machine learning model to generate a fitness score that characterizes the predicted quality of the capsid; as well as Output is a fitness score that characterizes the predicted quality of the capsid.

2. The method of claim 1, further comprising: The representation of the capsid monomer is processed using an error prediction machine learning model to generate a fitness error, wherein the fitness error is an estimate of the error in the fitness score generated by the fitness prediction machine learning model.

3. The method of claim 2, further comprising: The fitness score is updated by combining: (i) the fitness score generated by the fitness prediction machine learning model, and (ii) the fitness error generated by the error prediction machine learning model.

4. The method according to claim 3, wherein: Updating the fitness score includes: The fitness score generated by the fitness prediction machine learning model and the fitness error generated by the error prediction machine learning model are summed.

5. The method according to any one of claims 2 to 4, wherein The fitness prediction machine learning model has a different model architecture than the error prediction machine learning model.

6. The method according to any one of claims 2 to 5, wherein Processing the representation of the capsid monomer using the error prediction machine learning model includes: The error prediction machine learning model is used to process a representation of the amino acid sequence of the capsid monomer.

7. The method according to claim 6, wherein: The error prediction machine learning model processes different model inputs than the fitness prediction machine learning model.

8. The method of claim 1, wherein: The embedded neural network was used to process the data defining the amino acid sequence of the capsid: Instantiate the corresponding amino acid embedding corresponding to each position in the amino acid sequence of the capsid monomer; updating each amino acid embedding by processing the amino acid embeddings through one or more neural network layers of the embedded neural network; and The insertion of the capsid protein monomer is generated based on one or more of these updated amino acid insertions.

9. The method of claim 8, wherein: Generating the insertion of the capsid protein monomer based on one or more of these updated amino acid insertions includes: The insertion of the capsid protein monomer is generated by combining these updated amino acid insertions.

10. The method according to any one of claims 8 to 9, wherein Updating each amino acid embedding by processing the amino acid embeddings through one or more neural network layers of the embedded neural network includes: These amino acid embeddings are processed using one or more self-attention neural network layers of the embedding neural network.

11. The method according to any one of claims 1 to 7, wherein Processing the data defining the amino acid sequence of the capsid monomer using the embedded neural network includes: Instantiating an amino acid embedding sequence comprising an embedding of a corresponding amino acid corresponding to each position in the amino acid sequence of the capsid monomer; and sequentially processing each amino acid embedding in the amino acid embedding sequence using a recurrent neural network to generate an updated hidden state of the recurrent neural network; and The embedding of the capsid monomer is determined based on the updated hidden state of the recurrent neural network.

12. A method as claimed in any preceding claim, wherein The embedded neural network has been trained on a set of training proteins to perform the token unmasking task.

13. The method of claim 12, wherein: Training the embedded neural network to perform the token unmasking task involves, for each training protein: generating a masked amino acid sequence by masking one or more positions in the amino acid sequence of the training protein; processing the data defining the masked amino acid sequence using the embedded neural network to determine, for each masked position in the amino acid sequence of the training protein, a corresponding probability distribution over the set of possible amino acids; and The embedded neural network is trained to optimize an objective function that measures, for each masked position in the amino acid sequence of the training protein, the error between: (i) a probability distribution generated for the masked position, and (ii) the identity of the amino acid at the masked position in the training protein.

14. The method according to any one of claims 1 to 11, wherein The embedded neural network has been trained on a set of training proteins to perform the next token prediction task.

15. The method of claim 14, wherein: Training this embedded neural network to perform the next token prediction task involves, for each training protein: Using the embedded neural network, generating a corresponding probability distribution over a set of possible amino acids for each of one or more target positions in the amino acid sequence of the training protein, wherein, for each target position, the probability distribution generated for the target position depends only on the amino acid at the previous position in the amino acid sequence of the training protein; and The embedded neural network is trained to optimize an objective function that measures, for each target position in the amino acid sequence of the training protein, the error between: (i) a probability distribution generated for the target position, and (ii) the identity of the amino acid at the target position in the training protein.

16. A method as claimed in any preceding claim, wherein The fitness prediction neural network has been trained on a set of training examples, where each training example corresponds to a corresponding training monomer of a capsid and includes: (i) an embedding of the training monomer of the capsid, and (ii) a target fitness score for the training monomer of the capsid.

17. The method of claim 16, wherein: Training the fitness prediction neural network on the set of training examples includes rebalancing the set of training examples based on target fitness scores specified by the training examples.

18. A method as claimed in any preceding claim, wherein The fitness prediction machine learning model is a parametric machine learning model.

19. The method of claim 18, wherein: The fitness prediction machine learning model includes one or more of the following: a neural network model, a random forest model, or a support vector regression model.

20. The method according to any one of claims 1 to 17, wherein The fitness prediction machine learning model is a non-parametric model.

21. The method of claim 20, wherein: The fitness prediction machine learning model includes a k-nearest neighbor model.

22. The method according to any one of claims 1 to 21, wherein The predicted quality of the capsid is indicative of the predicted manufacturability of the capsid.

23. The method of claim 22, wherein: The manufacturability of the capsid is defined based on the number of capsid instances generated when a plasmid encoding the capsid is transfected into one or more cells.

24. The method of claim 23, wherein: The number of capsid instances generated when a plasmid encoding that capsid is transfected into one or more cells is normalized by the number of wild-type sequences generated.

25. The method of any one of claims 1 to 21, wherein The predicted quality characterization of the capsid includes the predicted ability of the capsid to prevent viral neutralization.

26. The method of any one of claims 1 to 21, wherein The predicted quality of the capsid is indicative of the predicted immunoreactivity of the capsid.

27. The method of any one of claims 1 to 21, wherein The predicted quality characterization of the capsid includes the predicted ability of the capsid to penetrate the target tissue.

28. The method of any one of claims 1 to 21, wherein The predicted mass of the capsid is indicative of the packaging capacity of the capsid.

29. The method of any one of claims 1 to 21, wherein The predicted quality characterization of the capsid includes the predicted ability of the capsid to integrate virally into the host genome.

30. A method as claimed in any preceding claim, wherein This capsid corresponds to that of adeno-associated virus (AAV).

31. The method of any preceding claim, further comprising: using the fitness prediction machine learning model to select a capsid monomer from a set of possible capsid monomers; as well as One or more viruses are physically generated, the one or more viruses comprising proteins having selected capsid protein monomers.

32. The method of claim 31, wherein Using this fitness prediction machine learning model to select capsid monomers from a set of possible capsid monomers includes: generating a corresponding fitness score for each capsid monomer in the set of possible capsid monomers using the fitness prediction machine learning model; and Capsid monomers are selected based on these fitness scores.

33. The method of claim 32, wherein: Capsid monomer selection based on these fitness scores includes: The capsid monomer with the highest fitness score is selected from among the capsid monomers included in the set of possible capsid monomers.

34. The method of any one of claims 31 to 33, further comprising: The generated virus is applied to a subject to achieve a therapeutic effect on the subject.

35. A system comprising: one or more computers; as well as One or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method of any one of claims 1 to 34.

36. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method of any one of claims 1 to 34.