Genetic engineering of proteins and protein design methods

WO2026019876A3PCT designated stage Publication Date: 2026-02-19MASSACHUSETTS INST OF TECH +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/037842
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-20
Filing Date
2025-07-16
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing protein language models (PLMs) struggle to significantly improve protein activity due to limited generalization and training data, requiring extensive experimental testing for de novo-designed variants, and previous methods fail to generalize well beyond proof-of-concept demonstrations.

Method used

EVOLVE-Pro, an ensemble model combining a foundational protein language model with a top-layer discrimination model, uses active learning to iteratively nominate high-activity protein variants through a few-shot learning framework, enabling efficient and generalizable protein design.

Benefits of technology

EVOLVE-Pro achieves 2- to 515-fold improvements in protein activity with minimal experimental testing by learning a protein's functional landscape, outperforming existing methods in identifying high-activity variants across diverse protein classes.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This disclosure provides protein mutants generated using the large language program EVOLVE-Pro, which is designed for a wide range of applications. Examples of protein mutants include, without limitation, prime editor and antigen binding monoclonal antibodies. EVOLVE-Pro substantially enhances the efficiency and effectiveness of in silico protein evolution, surpassing current state-of-the-art methods and yielding proteins with more than 100-fold improvement of desired properties. EVOLVE-Pro demonstrates the necessity for protein engineering models to zoom in on desired functional properties rather than predicted fitness, paving the way for broader applications of AI-guided protein engineering in biology and medicine.
Need to check novelty before this filing date? Find Prior Art

Description

GENETIC ENGINEERING OF PROTEINS AND PROTEIN DESIGN METHODSCross Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial Nos. 63 / 672,398 and 63 / 722,658 filed respectively on July 17, 2024, and November 20, 2024. The entire content of the above-referenced patent applications is incorporated by reference herein.Sequence Listing

[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on July 7, 2025, is named 76707 l_083474-042PC3_SL.xml and is 484,945 bytes in size.Background

[0003] Billions of years of evolutionary pressures have shaped protein diversity, filtering potential design space for diverse biological activities. There is emerging evidence that these sequences embody a fundamental language of biology that can be modeled with deep learning and offer unique insights into the evolutionary processes that have sculpted life on our planet. Protein language models (PLMs) learn the grammar of protein diversity by training to complete masked amino acids in a large protein sequence database, resulting in novel representations of biology. PLMs, with their internal representations of the protein evolutionary landscape n, have been used to nominate variants with improved activity with limited success. Generative PLMs, such as ESM3, ProtGPT2, and ProGen, can design novel proteins, but these de novo-designed variants typically only reach wild-type level activity after extensive rounds of experimental testing. This inability of PLMs to drastically improve upon protein activity in zero-shot is partially driven by their inability to generalize to new contexts due to evolutionary constraints as well as limited training data. Active learning methods that utilize context-specific data in combination with deep learning models, including machine learning-directed protein evolution (MLDE) methods, have effectively improved diverse proteins, but typically require extensive efforts to experimentally evaluate many variants. Previous attempts combined protein representation models with active learning to simplify the evolution process, but these approaches have not generalized well beyond proof-of-concept demonstrations like GFP engineering.

[0003] Accordingly, there exists a need for novel protein variants nominated by means of PLM and methods of using same for biological applications.Summary

[0004] In one aspect, the disclosure provides a prime editor comprising an amino acid sequence at least 70% identical to any one of the amino acid sequences from Table 3.

[0005] In another embodiment, the disclosure provides a nucleic acid molecule encoding the prime editor of claim 1.

[0006] In another embodiment, the disclosure provides a vector comprising the nucleic acid molecule of claim 2.

[0007] In another embodiment, the disclosure provides a cell comprising the vector of claim 3.

[0008] In another embodiment, the disclosure provides a method of site-specific integration of a nucleic acid into a genome of a cell, the method comprising (a) incorporating an integration site at a desired location in the genome by introducing into the cell: a DNA binding nuclease linked to a reverse transcriptase, wherein the DNA binding nuclease comprises a nickase activity; and a guide RNA (gRNA) comprising a primer binding sequence linked to an integration sequence, wherein the gRNA interacts with the DNA binding nuclease and targets the desired location in the genome, wherein the DNA binding nuclease nicks a strand of the genome and the reverse transcriptase incorporates the integration sequence of the gRNA into the nicked site, thereby providing the integration site at the desired location of the genome; and (b) integrating the nucleic acid into the genome by introducing into the cell: a DNA or RNA strand comprising the nucleic acid linked to a sequence that is complementary or associated to the integration site; and an integration enzyme, wherein the integration enzyme incorporates the nucleic acid into the genome at the integration site by integration, recombination, or reverse transcription of the sequence that is complementary or associated to the integration site, thereby introducing the nucleic acid into the desired location of the cell genome of the cell wherein the DNA binding nuclease linked to the reverse transcriptase comprises the prime editor of claim 1.

[0009] In another embodiment, the disclosure provides Cl 43 antibody comprising an amino acid sequence at least 70% identical to any one of the amino acid sequences from Table 5.

[0010] In another embodiment, the disclosure provides a C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 relative to a wild-type C143 antibody sequence of SEQ ID NO. 135.

[0011] In another embodiment, the disclosure provides a C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 and / or 40 and / or at least one mutation at heavy chain position 39 relative to a wild- type Cl 43 antibody sequence of SEQ ID NO. 135.

[0012] In another embodiment, the disclosure provides a C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28, 40, 50, 45 and / or 14 and / or at least one mutation at heavy chain position 39, 63, and / or 89 relative to a wild-type C143 antibody sequence of SEQ ID NO. 135.

[0013] In another embodiment, the disclosure provides a C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 and / or 14 and / or at least one mutation at heavy chain position 33, 39, and / or 58 relative to a wild-type C143 antibody sequence of SEQ ID NO. 135.

[0014] In another embodiment, the disclosure provides a C143 antibody comprising an amino acid sequence including a mutation at light chain position N28R / Q40K and a mutation at heavy chain position R39K relative to a wild-type Cl 43 antibody sequence of SEQ ID NO. 135.

[0015] In another embodiment, the disclosure provides a C143 antibody comprising an amino acid sequence including a mutation at light chain position N28K relative to a wild-type Cl 43 antibody sequence of SEQ ID NO. 135.

[0016] In another embodiment, the disclosure provides a method for treating a disease or a disorder in a subject, the method comprising administering to the subject in need thereof the antibody of claim 6, thereby treating the disease or the disorder in the subject.

[0017] In another embodiment, the disclosure provides a aCD71 antibody comprising an amino acid sequence at least 70% identical to any one of the amino acid sequences from Table 6.

[0018] In another embodiment, the disclosure provides a aCD71 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 and / or 40and / or at least one mutation at heavy chain position 39 relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

[0019] In another embodiment, the disclosure provides a aCD71 antibody comprising an amino acid sequence including at least one mutation at heavy chain position 70 and / or 92 and / or at least one mutation at light chain position 38 relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

[0020] In another embodiment, the disclosure provides a aCD71 antibody comprising an amino acid sequence including at least one mutation at position S92A relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

[0021] In another embodiment, the disclosure provides a aCD71 antibody comprising an amino acid sequence including mutations at T70A_S92V relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

[0022] In another embodiment, the disclosure provides a aCD71 antibody comprising an amino acid sequence including at least one mutation at heavy chain position 39, 63, and / or 89 and / or at least one mutation at light chain position 14, 40, 50 and / or 45 relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

[0023] In another embodiment, the disclosure provides a method for treating a disease or a disorder in a subject, the method comprising administering to the subject in need thereof the antibody of claim 14, thereby treating the disease or the disorder in the subject.Brief Description of the Drawings

[0024] Aspects, features, benefits, and advantages of the embodiments described herein will be apparent with regard to the following description, appended claims, and accompanying drawings.

[0025] Fig. l.(A) Developing and benchmarking EVOLVE-Pro for protein language model- guided engineering. Schematic describing the EVOLVE-Pro method. Proteins of interest go through iterative rounds of low-N screening. A foundational PLM generates embeddings for all mutants of a protein and the average embedding by pooling across all residues is used as input for the top layer model. Each mutant’s activity is experimentally determined and used to train a domain expert top layer model with PLM embedding as input. The top layer model then nominates the top-N mutants for the next round of testing and the weights are updatediteratively in an active learning format. Fig. l.(A) discloses SEQ ID NOS 324, 325, 326, 325, 326 and 324 respectively, in order of appearance. (B) Benchmarking of foundational models across a panel of 12 comprehensive deep mutational scanning (DMS) datasets. Each point is a unique protein and its DMS data. ESM2-15B has the highest average percent success in high activity variants prediction. (C) Comparison between EVOLVE-Pro in active learning format, in zero-shot pretraining format, and an existing zero-shot prediction method using protein language model (30) across 12 DMS datasets. Each point is a unique protein using its DMS data. (D) Performance over 10 rounds of EVOLVE-Pro with 16 mutants per round, compared to two different non-language model encoding schemes (One-hot encoding and integer encoding). Model performance is benchmarked on four datasets(16, 21, 25, 60) and compared to zero-shot ESM2 nomination success rate and background random sampling (30). Error bar represents the standard deviation for n=10 random simulations. (E) Engineering of REGN10987 over five rounds of EVOLVE-Pro. Data shows cumulative top 10 mutants’ fold improvement over wild-type binding affinity to the target antigen across 5 evolution rounds. Percentages show the percent of mutants that have higher activity than wild-type REGN 10987 each round. (F) Mapping of the top mutations on the structure of REGN 10987 (PDB: 6XDG).

[0026] Fig. 2(A) Evolution of two monoclonal antibodies with EVOLVEpro. Schematic of the evolution strategy with EVOLVEpro for engineering two monoclonal antibodies across two parameters (binding affinity and antibody expression). (B) Engineering of the Cl 43 antibody over five rounds of EVOLVEpro. Data shows cumulative top 10 mutants’ fold improvement over wild-type binding affinity to the target antigen across 5 evolution rounds. (C) IC50 value estimated from ELISA binding data for the WT C143 antibody, the best single mutant (LC N28K) and the best multi-mutant (LC N28R / Q40K+HC R39K). Error bars represent standard error of mean with n=3 technical replicates. A one way ANOVA was run between the three groups (*, p<0.05, **, p<0.01). (D) Scatter plot showing each individual mutant’s expression fold improvement versus binding affinity improvement for the C143 antibody. The best mutant in each round is highlighted with a larger circle. (E) Engineering of the aCD71 over five rounds of EVOLVEpro. Data shows cumulative top 10 mutants’ fold improvement over wild-type binding affinity to the target antigen across 5 evolution rounds. (F) IC50 value estimated from ELISA binding data for the WT anti-CD71 antibody, the best single mutant (S92A) and the best multi-mutant (T70A_S92V). Y axis is shown on log 10 scale. Error bars represent standard error of mean with n=3 technical replicates. A one wayANOVA was run between the three groups (*, p<0.05, **, p<0.01). (G) Scatter plot showing each individual mutant’s expression fold improvement versus binding affinity improvement for the aCD71 antibody. The best mutant in each round is highlighted with a larger circle. (H- I) Mapping of the top mutations on the predicted structure of Cl 43 (H) and anti-CD71 (I) respectively (AF3). (J) Scatter plot comparing the predicted naive ESM-2 C143 protein fitness (predicted masked marginal score) and scaled tested activity of nominated mutants across evolution. Scatter points are colored by rounds in evolution. The correlation and linear regression line are shown in red and the R square of the correlation is reported. (K) Comparison of the Cl 43 embedding latent space with either predicted naive ESM-2 protein fitness landscape or EVOLVEpro protein activity landscape. Yellow rhombus denotes wildtype sequence.

[0027] Fig. 3(A). Evolution of prime editor with EVOLVE-Pro. Schematic of the evolution strategy with EVOLVE-Pro for engineering a prime editor to be more efficient in attB insertion. (B) Engineering of the prime editor PE2 with twinPE guides over seven rounds of EVOLVE-Pro. Data shows cumulative top 10 mutants from current and preceding rounds, as measured by fold improvement of prime editing activity to install a 46 bp AttB site at the murine NOLC1 genomic locus. (C) Validation of 4 evolved prime editors in the installation of attB sites at three different endogenous sites in either mouse or human genomes. A two- sided Student’s t-test was run between WT and each evolved prime editor (ns, nonsignificant, *, p<0.05, **, p<0.01, ***, p<0.001, ****, p<0.0001). Fold change over wildtype PE2 is shown for the best mutant on each genomic locus. Error bars represent standard deviation with n=3 biological replicates. (D) Mapping of the top mutations on the AlphaFold3 model of M-MLV RT. The RT active site is indicated by a red circle. (E) Heatmap showing most common PE2 mutations explored by EVOLVE-Pro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. (F) Scatter plot comparing the predicted naive ESM-2 protein fitness (predicted masked marginal score) and scaled tested activity of nominated mutants across evolution, scatter points are colored by rounds in evolution. (G-H) Comparison of the PE2 embedding latent space with either predicted naive ESM-2 protein fitness landscape or EVOLVE-Pro protein function landscape. (I) A kernel density estimate plot of protein fitness as predicted by ESM-2 versus protein function as predicted by EVOLVE-Pro. The correlation and linear regression line are shown in red and the R square of correlation is reported.

[0028] Fig. 4(A). EVOLVE-Pro model optimization. Summary of parameter grid searches for EVOLVE-Pro with an ESM-2 15B foundational model, showing the random forest regressor combined with the top 10 active learning selection strategy returned the highest average binary top fitness success rate across 12 DMS datasets. (B) The optimized EVOLVE- Pro model with n=l 6 mutants per round shows both a higher median protein activity score and max protein activity score as the model progresses into later rounds of evolution, showing the utility of active learning. Error bars represent standard deviation with n=10 simulations. C) Comparison of the number of nominated mutants from n=10 to n=l 00 per round on the impact of EVOLVE-Pro evolution. Each line graph depicts the percent high fitness for one of the 12 DMS datasets and the error bar represents the standard error of the mean for 10 random simulations.

[0029] Fig. 5. (A) Additional Evolve-Pro model characterization. Performance over 10 rounds of EVOLVE-Pro with 16 mutants per round, compared to two different non- language model encoding schemes (One-hot encoding, integer encoding). Model performance is benchmarked on 8 additional datasets and compared to both the zero-shot ESM2 nomination success rate and background random sampling (1). Error bars represent the standard deviations for 10 random simulations.

[0030] Fig. 6(A): Additional characterization of the evolution of C143 antibody with EVOLVEpro. IC50 of individual nominated mutants from EVOLVEpro across 5 rounds of evolution relative to WT. Error bars represent standard deviation with n=3 technical replicates. B) Heatmap showing the most common Cl 43 antibody mutations explored by EVOLVEpro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. C) A kernel density estimate of protein fitness as predicted by ESM-2 versus protein activity as predicted by EVOLVEpro. The correlation and linear regression line are shown in red and the R square of correlation is reported. D) EVOLVEpro ’s mutational trajectory from round 1 to round 4 on Cl 43 antibody.

[0031] Fig. 7(A): Additional characterization of evolution of anti-CD71 antibody with EVOLVEpro IC50 of individual nominated mutants from EVOLVEpro across 5 rounds of evolution relative to WT. Error bars represent standard deviation with n=3 technical replicates. B) Heatmap showing the most common anti-CD71 antibody mutations explored by EVOLVEpro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. C) Scatter plot comparing the predicted ESM-2 protein fitness score versus experimentally measured anti-CD71 antibody binding affinity scaled foldimprovement across evolution rounds. The correlation and linear regression line are shown in the plot. D) Comparison of the anti-CD71 antibody latent space with either predicted ESM-2 protein fitness (masked marginal score) or EVOLVEpro protein activity fold improvement. Y ellow rhombus denotes wild-type sequence. E) A kernel density estimate of protein fitness as predicted by ESM-2 versus protein activity as predicted by EVOLVEpro. The correlation and linear regression line are shown in red and the R square of correlation is reported. F) EVOLVEpro ’s mutational trajectory from round 1 to round 4 on anti-CD71 antibody.

[0032] Fig. 8. (A) Additional characterization of evolution of REGN10987 with EVOLVE- Pro. Fold improvement of individual nominated REGN 10987 mutants from EVOLVE-Pro across 5 rounds of evolution relative to WT. Error bars represent standard deviation of three biological replicates. (B) Percent success as defined by IC50 lower than WT for the mutants in each round of evolution. (C) Heatmap showing the most common REGN 10987 mutations explored by EVOLVE-Pro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. (D) Scatter plot comparing the predicted ESM-2 protein fitness score versus experimentally measured REGN 10987 binding affinity scaled fold improvement across evolution rounds. The correlation and linear regression line are shown in the plot. (E-F) Comparison of the REGN 10987 antibody latent space with either predicted ESM-2 protein fitness (masked marginal score) or EVOLVE-Pro protein activity fold improvement. (G) A kernel density estimate of protein fitness as predicted by ESM-2 versus protein function as predicted by EVOLVE-Pro. The correlation and linear regression line are shown in red and the R square of the correlation is reported.

[0033] Fig. 9. Additional characterization of the evolution of the prime editor PE2 with EVOLVE-Pro. Bar chart of each Individual PE2 mutant’s fold improvement for murine genomic NOLC1 attB insertion frequency across evolution rounds. Error bars represent standard deviation of three biological replicates.Detailed Description

[0034] Evolve-pro

[0035] EVOLVE-Pro is a frontier multi-modal protein design model. EVOLVE-Pro evolves high-activity protein variants with few-shot learning and minimal experimental testing, achieving accurate prediction of sequence-to-function for general properties. This performance comes from an ensemble approach, combining evolutionary-scale proteinfoundation models with a top-layer discrimination model to learn a protein’s functional landscape and guide the directed evolution process in silico. By applying EVOLVE-Pro in a few-shot, active learning framework, protein sequences with significantly higher activity can be efficiently nominated in a generalizable fashion with minimal effort. The modularity of the EVOLVE-Pro architecture allows this framework to scale with larger parameter PLMs. Moreover, EVOLVE-Pro prompting only requires protein sequences to be evolved without any structural information, expert knowledge, or prior data. As EVOLVE-Pro is multi-modal, multiple protein features of any type or data class can be simultaneously engineered, opening up vast possibilities for its use in biology and medicine.

[0036] We benchmark EVOLVE-Pro in silico across a panel of different proteins, showing state-of-the-art performance, and then apply the final model for different applications with proteins: 1) a monoclonal COVID antibody and 2) a prime editor. EVOLVE-Pro yields mutants with 2- to 515-fold improvement over initial proteins. Analyzing nominated mutations, EVOLVE-Pro explores disparate sites, and the learned functional landscape is separate, and often negatively correlated with, the fitness inferred by the underlying protein language model. Lastly, we showcase EVOLVE-Pro ’s utility in nominating multi-mutant protein designs out of a vast sequence space that enables final mutants that are much more active than naturally observed proteins. EVOLVE-Pro establishes the capabilities of few-shot active learning with protein language models for optimizing proteins for diverse activities.

[0037] Development and benchmarking of the EVOLVE-Pro model

[0038] To establish EVOLVE-Pro, we designed an ensemble model that involves: 1) a foundational protein language model to encode protein sequences into an information-rich latent space, and 2) a top-layer discrimination model to learn protein functional grammar in this evolutionary landscape and rank protein sequences according to a designed policy framework, and 3) an active learning framework using top layer discrimination model to nominate the next set of protein variants for experimental evaluation. This cycle is performed iteratively to evolve defined protein activities until they reach desired levels (Fig. 1 A).

[0039] We optimized EVOLVE-Pro across five parameters: 1) the strategy employed for the first round mutant selection, 2) the top layer discrimination model that learns the fitness landscape, 3) the active learning strategy for selecting mutants for the next round, 4) the evolution policy, and 5) the embedding vector transformation (Table 1). To perform a grid search across this space, we curated a panel of twelve unique deep mutagenesis scanning(DMS) datasets for in silico validation (16-26) (table 2). These twelve proteins represent diverse functions, including viral spike proteins, RNA-guided nucleases, lactases, and kinases, ensuring that the resulting model will be as generalizable as possible for learning diverse protein activity landscapes in PLM latent space. Table 1. Summary of parameters grid searchTable 2. Description for 12 DMS datasets

[0040] We first focused on the ESM-2 protein language model because of its large training data and available model size of >200M proteins and 15B parameters, respectively. Using the ESM-2 15B parameter model, our grid search found the optimal strategy was: 1) selecting a random set of first-round variants, 2) employing a random forest regressor discriminatorymodel to predict protein function, 3) using residue pooled average embeddings, and 4) using a top-N selection strategy in each round of evolution (Fig. 4A). This policy nominated a high frequency of gain-of-function protein variants in only 5 rounds (Fig. 4A). Since we focused on percent of activity passing a threshold as our evaluation metric in the grid search, we next checked for increasing function during in silico evolution. We found that both the median activity and the activity of the nominated top mutant increased monotonically from round to round across all DMS datasets, further validating the model’s performance in this low-N active learning setting (Fig. 4B).

[0041] In general, 16 mutants per round of evolution for 10 rounds identified top mutants with fitness in the 50th percentile for eleven of the twelve DMS datasets. To understand how the number of variants per round affected performance, we tested between 10 and 100 variants per round, finding that larger rounds increased prediction accuracy (Fig. 4C). This performance trade off indicates that EVOLVE-Pro can be used for both extremely low-N evolution (<20 mutants per round) for rapid and cheap experimental characterization and medium-N (—100 mutants per round) for quicker and more efficient evolution with fewer rounds.

[0042] After optimizing the top layer model and learning strategies, we optimized the PLM, comparing ESM-2 15B to a panel of foundational models. Using the optimal parameters from the grid search, we benchmarked performance against smaller versions of ESM-2 and ESM- 1(27), UniRep(75, 28), ProtT5(29), ProteinBERT(4), Ankh(3), one-hot encoding, and integer encoded protein representations. ESM-2 15B parameter model outperformed all the other models for identifying the highest fitness proteins for all datasets except two, confirming its final selection for the EVOLVE-Pro latent space model (Fig. IB, Data S3). Importantly, large parameter PLMs showed a significant boost in prediction accuracy compared to nonlanguage model-based architectures, indicative of the powerful feature extraction present in transformer-based models (Fig. IB).

[0043] We next benchmarked EVOLVE-Pro ’s performance relative to other PLM -based engineering approaches. As many methods require pre-training a discriminatory model on thousands of variants, we tested versions of EVOLVE-Pro augmented with various amounts of pre-training (Fig. 1C). Reinforcement learning drastically reduced the overall number of mutants required: EVOLVE-Pro with only 5 rounds of evolution (16 mutants per round) was equivalent in performance to EVOLVE-Pro pre-trained with 160 mutants, while 10 rounds of evolution (16 mutants per round) was equivalent to pre-training with 500 mutants. Moreover,EVOLVE-Pro significantly outperformed zero-shot prediction methods (30). This comparison confirms that the few-shot nature of EVOLVE-Pro allows for efficient directed evolution with minimal effort and low-N testing per round (Fig. 1C).

[0044] Lastly, we analyzed the per-round evolution improvement for EVOLVE-Pro compared to one-hot and integer encoding and zero-shot prediction, finding that by round 5 variants with significantly enhanced fitness could universally be found (at 16 mutations per round) (Fig. ID, Fig. 5). Moreover, in many cases, the one-hot and integer encoding frameworks saturated much earlier in the evolution process and never reached the fitness levels achieved by EVOLVE-Pro. Interestingly, for some proteins we observe a non-linear increase in protein fitness after round 3, suggesting greater gains in mapping the protein fitness landscape as EVOLVE-Pro evolution proceeds.

[0045] Antibody optimization with EVOLVE-Pro (REGN 10987)

[0046] As a first test of EVOLVE-Pro, we optimized the binding interaction of the REGN10987 antibody to the extracellular epitope of the SARS-CoV-2 spike protein. REGN 10987, a component of an approved COVID- 19 therapy (37), is engineered for neutralization of the SARS-CoV-2 spike protein and is refractory to previous in silico optimization (30). Given the challenge of improving REGN 10987, it is an ideal first test of EVOLVE-Pro’s capabilities. See Table 4 for list of mutant COVID REGN 10987 antibodies. Mutagenizing the heavy chain variable region with EVOLVE-Pro, we compared wild-type REGN 10987 to 11 nominated mutants each round with an enzyme-linked immunosorbent assay (ELISA) against the SP6 stabilizing variants of SARS-CoV-2 spike protein(32). We found significant fold improvement (FI) in binding from only 1 round of EVOLVE-Pro, and five rounds of mutagenesis yielded variants showed a 63% improvement by ELISA with an IC50 of 11.9 nM (Fig. IE, Fig. 8A). We observed that the success rate of EVOLVE-Pro increased each round, with success increasing from 9.1% in the first random round to 54.5% in the last round. This improvement demonstrates EVOLVE-Pro’s top layer model learns a functional grammar distinct from the fitness grammar captured initially by the underlying foundational model (Fig. IE, Fig. 8B).

[0047] In the structure of REGN 10987 bound to the spike epitope, one high-performing mutation (D63K) lies inside the 2nd complementarity-determining region (CDR), with other mutations clustered around the CDR3 in the framework region (HFR3), suggesting a role in enhanced antigen binding (Fig. IF). To understand EVOLVE-Pro’s mutational trajectory, werepresented the model’s attention to particular residues as cumulative frequency of individual residues being explored by the model and found that multiple residues are repeatedly explored by the model including D63, D91, and N32 (Fig. 8C). Across all nominated mutations, successive rounds of training focused on regions around CDR2, emphasizing the increased attention of the EVOLVE-Pro to specific regions of the protein. We next analyzed each mutant’s observed activity versus the PLM-predicted fitness. We calculated the mutant fitness as predicted marginal masked score within the ESM2 embeddings (pMMS) and found that the function of EVOLVE-Pro variants did not correlate with fitness (Fig. 8D). To extrapolate this finding across the entire REGN 10987 mutational landscape, we projected the base layer PLM fitness score and top layer random forest predicted FI (pFI) in the latent space for every possible single mutation variant (Fig. 8E-F). There was relatively little overlap between the two distributions, with a negative correlation of -0.22 between predicted fitness and predicted function (Fig. 8G). Collectively, this data shows that PLMs do not learn protein function in isolation, highlighting the importance of few-shot learning with PLMs.

[0048] Antibody optimization with EVOLVEpro (C143. SARS-COV-2 and aCD7 l )

[0049] As another test of EVOLVE-Pro, we used EVOLVEpro to optimize two therapeutically relevant monoclonal antibodies: C143, an antibody against the SARS-COV-2 spike protein, and aCD71 , an antibody against the human transferrin receptor used for delivery of drugs and siRNA to muscle and cardiac cells in vivo (15, 48, 49). aCD71 can be used to treat diseases and disorders involving tissue cancer, such as glioblastoma and other disorders involving iron metabolism. aCD71 has more than 90% sequence homology to Delpacibart, a phase II clinical stage therapy for Myotonic dystrophy. Both antibodies have low nanomolar affinities against their cognate antigen, presenting a challenge for further improvement by EVOLVEpro. We designed a multi-objective optimization with EVOLVEpro on antibody expression levels and binding affinity to the target antigen.Optimization over multiple features shows that the model can jointly evolve antibody binding and yield, as antibody mutations often affect multiple antibody features, including expression, stability, solubility, half-life, or immunogenicity. Critically, zero-shot developability optimization is difficult with protein language or structure-based inverse folding models, as the relationship between sequence and developability or other non-binding features is not directly captured by evolutionary sequence or structure data (15, 20). In our multi-objective directed evolution scheme, we weighed the binding affinity at four times the expression levels (i.e., developability score) to prioritize variants that bind with higher affinity (Fig. 2A).

[0050] For C143 evolution, we quantified binding affinity with an enzyme- linked immunosorbent assay (ELISA) against the SP6 stabilizing variants of the SARS-CoV-2 spike (Wuhan strain) protein (50). We saw improved binding after 3 rounds of EVOLVEpro evolution, surpassing previous zero-shot approaches(15) (Fig. 2B, Fig. 6A). At round 4, we found significant improvement with a light chain mutant (N28K), with an IC50 of 0.19 nM (Fig. 2C). Using our 4 rounds of single mutations, we had EVOLVEpro design multi-mutant combinations for a fifth round. The best multi-mutant (light chain N28R / Q40K with heavy chain R39K) bound to the SP6 spike antigen with an IC50 of 60 pM (Fig. 2C), likely due to the synergistic interaction between N28R on the light chain and R39K on the heavy chain. We stopped at one round of multi-mutant evolution due to the substantial improvements we observed, but in practice, multiple rounds of multi-mutant testing are likely needed to reach convergence on desired properties. We found that many improved binders compromised yields (Fig. 2D), a tradeoff due to the bias towards binding during the multi-objective design. Despite this tradeoff, a subset of C143 mutants, such as R39K, had both an increase in affinity and protein expression, showing that developability can be co-optimized alongside binding affinity.

[0051] We explored the likelihood of the top EVOLVEpro nominated mutations relative to training data and known antibody variants observed in nature. We analyzed the top 10 mutations for both occurrence in hotspot regions and deviation from the germline sequence. We found none of the top 10 mutations are mutated back to the germline unmutated common ancestor sequence (UCA). The UCA sequence at light chain N28 is a serine (S). However, the affinity-enhancing mutation recommended by EVOLVEpro is either a lysine or arginine, both with likelihoods less than 0.05 when compared to the Uniprot training input, reinforcing the notion that EVOLVEpro mutations are rare and novel. Moreover, this observation highlights the utility of top layer regression model to explore rare mutations not seen in the training input of PLM by promoting exploration into unknown regions of the protein fitness landscape. Furthermore, we found the top single mutant N28R (light chain) happens in the complementarity determining region (CDR), but the majority of the top affinity-enhancing mutations (7 out of the top 10 mutations) are located in the framework region. This highlights the de novo exploration of EVOLVEpro on the entire variable region of the antibody to find affinity-enhancing mutations that might not seem likely or intuitive in the framework region. To further understand EVOLVEpro’s mutational trajectory, we represented the model’s attention to particular residues as a cumulative frequency and found residues like K33, R39,and D58 on the heavy chain and S 14 and N28 on the light chain are repeatedly explored (Fig. 6B).

[0052] For using EVOLVEpro to in silico evolve an anti-CD71 antibody, we measure the target binding affinity using enzyme-linked immunosorbent assay (ELISA) against the human TfR protein and measured antibody expression with an anti-IgG. We saw improvement in binding after just rounds of evolution with EVOLVEpro (Fig. 2E, Fig. 7A). We acquired 10 mutants using the efficient evolution algorithm to benchmark against EVOLVEpro and found that EVOLVEpro nominated mutants 35 -fold better than wild-type whereas efficient evolution’s best mutant is only 8-fold better (15). At round 4, we found the best single mutant heavy chain S92A to bind to the antigen with an IC50 of 29 pM, significantly higher than that of the WT at 551 pM (Fig. 2F). We also asked the model to rank multi-mutants based on the single mutant data from the first 4 rounds, and performed one round of multi-mutant testing. We improved binding and expression in the multi-mutant round 5 with heavy chain T70A / S92V mutant. The multi-mutant binds to the hTfr protein with an IC50 of 19 pM (Fig. 2F). Interestingly, most of the mutants nominated after round 1 showed a marked increase in expression profile and binding affinity, showing that EVOLVEpro simultaneously engineered the developability and binding (Fig. 2G). This finding contrasts with the results from the Cl 43 antibody, implying the Pareto frontier of binding and expression likely differs between the two antibodies, with anti-CD71 ’s wild-type sequence easier to engineer across multiple properties than C143’s wild-type sequence. Future work examining this trade-off between multiple properties for additional antibodies or proteins will enable EVOLVEpro to better traverse Pareto frontiers more efficiently.

[0053] Upon analyzing the novelty of the aCD71 antibody mutations, we found only one of the top 10 mutations is mutated back to the germline unmutated common ancestor sequence (UCA) at position V73. The UCA sequence at the site of the best mutation in heavy chain S92 is a threonine (T). However, the affinity- enhancing mutation recommended by EVOLVEpro is either an alanine or valine. S92V mutation has a mutation likelihood of less than 0.05 when compared to the Uniprot training input, highlighting its rarity. This shows EVOLVEpro ’s ability to insightfully choose novel mutations not seen in the training input of the PLM base layer. Furthermore, we found all top 10 affinity-enhancing mutations are located in the framework region rather than in the CDR that is commonly thought to determine binding affinity. Lastly, to understand EVOLVEpro ’s mutational trajectory on anti-CD71, we represented the model’s attention to particular residues as the cumulativefrequency of individual residues being explored by the model and found that multiple residues are repeatedly explored by the model including T70 and S92 on the heavy chain and Q38 on the light chain (Fig. 7B).

[0054] We used AlphaFold 3 to model the structure of the anti-CD71 and C143 antibodies (Fig. 2H-I). We found two major clusters of exploration by EVOLVEpro on C143 antibody in the framework region with light chain mutations S14, Q40, L50, and K45 co-located and R39, S63, and E89 in close proximity on the heavy chain. These mutations likely alter binding through structural changes in the variable region. Additionally, there was a CDR mutation, N28, on the light chain located in the CDR-L1 region that likely directly alters the interaction between the Cl 43 antibody and the antigen, which is not possible to model with AF3 due to a low confidence score of the complex (Fig. 2H). For the anti-CD71 antibody, we found all the best mutations clustered around one region in the heavy chain domain. As they are all in the framework region, they likely alter the binding affinity indirectly, a hypothesis supported by the increase in expression relative to the wild-type sequence (Fig. 21).

[0055] Lastly, we analyzed each mutant’s observed activity versus the PLM-predicted fitness landscape. We calculated the mutant fitness as a predicted marginal masked score within the ESM2 embeddings (pMMS) and found that the activities of EVOLVEpro variants did not correlate with predicted ESM2 fitness (Fig. 2J, Fig. 7C). To extrapolate this finding across the entire C143 and anti-CD71 mutational landscape, we projected the base layer PLM fitness score and top-layer random forest predicted fold improvement (pFI) in the latent space for every possible single mutant variant, generating EVOLVEpro determined protein activity landscape (Fig. 6C, Fig. 7E). There was relatively little overlap between the two distributions, with a negative correlation of -0.16 for C143 antibody and 0.01 for anti-CD71 antibody between predicted fitness and predicted activity, further highlighting ESM2’s lack of understanding of protein activity (Fig. 6C, Fig. 7E). We projected the individual mutants onto the PCA space of the ESM2 embedding and found two opposing directions between higher fitness and higher function (Fig. 2K, Fig. 7D). Analyzing the evolution trajectory from round 1 to the final round by calculating the geometric midpoint of each round revealed the top layer model directed the evolutionary process towards the higher side of PCA1 for C143 antibody and the higher side of PCA2 for anti-CD7, which we observe to correlate with higher protein function (Fig. 6D, Fig. 7F).

[0056] Engineering improved prime editors with EVOLVE-Pro

[0057] Many molecular tools, such as next-generation genome editing proteins, function as multiple enzymes acting in concert. Prime editing, which uses an RNA-templated reverse transcriptase to programmably install diverse genome edits, is the fusion of a SpCas9 nicking mutant (nCas9) with an engineered Moloney Murine Leukemia Virus Reverse Transcriptase (M-MLV RT) [D200N, L603W, T306K, W313F, T330P] (termed PE2). We surmised that EVOLVE-Pro could improve upon these rational mutations, given further optimizations discovered on M-MLV RT by other directed evolution approachcs(4 / ). As PE-based insertion has difficulty installing longer (>40 nt) edits, we focused on editing outcomes with longer (46 bp) insertions, which have particular utility for programmable gene insertion methods, such as PASTE(42). We set up the evolution policy by using a previously described twinPE approach, where two overlapping pegRNAs are used in combination to install a 46bp attB site in the NOLC1 loci in murine hepatocyte cell line (Hepal-6). Editing was quantified at NOLC1 loci using amplicon sequencing and NGS readout and the top layer EVOLVE-Pro model was trained to predict the insertion efficiency (Fig. 3A).

[0058] Over successive rounds of optimization, we found that EVOLVE-Pro progressively learned the activity landscape of the RT of PE2, yielding improved variants after the initial random selection round and substantially improving upon PE2 -based editing by round 4 (Fig. 3B, Fig. 9). To check for bias against these single loci in the genome, we took the top 4 performing variants (A660S, L670C, L670K, and L671R) and surveyed their editing efficiency at two additional genomic loci (human AAVS1 and mouse Factor IX) in two different cell lines and found statistically significant improvements for L670C and A660S in all three sites (Fig. 3C). These results point to the general protein activity improvement by our model, delivering an additional set of RT mutations specifically for larger edits.

[0059] Projecting the top mutations onto the Alphafold3 -predicted structure of the RT reveals that most of them are clustered in the C-terminal RNase H domain (Fig. 3D), which is a surprising result since most PE evolution focuses on RT mutagenesis. We hypothesize that these mutations could increase the cleavage of the template DNA in the RNA-DNA heteroduplex by the RNaseH domain, facilitating the completion of the prime editing reaction, a route that has not been explored by traditional engineering of prime editors. We then further analyzed EVOLVE-Pro ’s residue site preference during evolution and observed significant attention in residues like L670, L671, and A660, suggesting it was learning that these positions could be quite beneficial for improving activity (Fig. 3E). Analysis of predicted fitness (pMMS) scored by the bottom layer PLM again showed a divergencebetween fitness and activity for the prime editor (Fig. 3F), allowing Evolve-Pro to successfully use the top discrimination layer to navigate to the higher activity variants in later evolution rounds (Fig. 3F).

[0060] Lastly, we try to understand the global mutational trajectory by projecting the activity landscape learned by the random forest regressor and base layer ESM2’s protein fitness landscape onto the first two PC As of the embedding (Fig. 3G-H). We found almost no convergence between the two distributions with a negative correlation of 0.08 (Fig. 31). This analysis points again to the divergence between the mutational landscape of a protein’s activity and the commonly used fitness landscape learned during a foundational model’s training on all protein sequences.

[0061] EVOLVE-Pro for biological and drug development applications

[0062] We demonstrate EVOLVE-Pro as an ensemble model for few-shot learning to evolve proteins. Over consecutive rounds of improvement, EVOLVE-Pro yields variants with 2- to 515- fold improvements in desired properties, including binding, catalytic efficiency, and immunogenic byproducts. Using evolutionary scale protein language models, EVOLVE-Pro learns general rules about protein function and can reason protein designs that are much more active using a discriminatory interpretation model layer, resulting in the rapid evolution of diverse proteins. Moreover, because of the rich latent space generated by PLM and powerful feature selections present in the top layer module, EVOLVE-Pro evolution is a low-N learning approach that requires minimal intervention and experimentation by humans. We benchmarked EVOLVE-Pro across different DMS datasets and different protein classes, showing its superiority in the low-N evolution setting. In this benchmarking work, we evaluated all known embedding-based PLMs and performed a large parameter grid search examining this approach over different discriminatory and active learning selection strategies, as well as different normalization techniques towards the embeddings and fitness measurements. We find that PLMs are essential and high-quality representations of protein sequences as dimensionality reduction of the embedding space through PCA did not improve performance (see methods and Data SI). Because of the modular design of EVOLVE-Pro, we anticipate that further improvements in autoregressive PLMs can be readily integrated into our approach, scaling up EVOLVE-Pro even more towards de novo generation of highly active mutants. In this work, we present the first comprehensive evaluation of Al protein engineering models across tough to engineer proteins, where the wild-type sequence is already highly efficient. These proteins have low correlation between activity and fitness, andin some cases require multiple properties to be optimized simultaneously, rendering them hard to evolve with existing MLDE approaches. In addition, none of the assays used for measuring protein activity in this work would be compatible with pooled screening, making typical directed evolution strategies impossible. Our work compliments previous efforts to establish proper computational metrics for evaluating sequence designs in the setting of high- activity mutant noinination(47). Therefore, EVOLVE-Pro represents a significant step forward for the Al-guided rapid design of “gain-of- function” mutants that are SOTA for downstream biological use cases.

[0063] Evolutionary scale PLMs have been powerful for annotating proteins, clustering families, predicting stability, and simulating protein structures, but prediction of protein sequences with superior function has been elusive. This is a direct manifestation of a PLM’s training objective function, which is to learn the evolutionary representation of proteins. As natural evolution does not necessarily select for maximizing protein function, the PLM’s learned evolutionary landscape will often not be correlated with a protein’s functional landscape. When it is correlated, such as in the case of antibodies due to biophysical considerations, PLM protein evolution may work with some success(30, 48), but optimizing enzymes with PLMs is much more challenging. Unlike antibodies that depend mainly on the biophysical properties of their variable regions for antigen binding, enzymes need enhancements in both biophysical and biochemical properties(5), which are complex and often poorly understood. Enzymes also exhibit significant diversity in their classes and catalytic mechanisms, each necessitating unique engineering approaches. Consequently, the effectiveness of using zero-shot PLMs trained on all protein sequences or fine-tuned on a subset of sequences for selecting the best enzyme variants remains uncertain. It has been shown that PLMs can scale with increasing parameters just like large language models, but recent analysis have shown saturating scaling effects of PLMs with limited input training dataset (Uniref) on larger models 49-51). Thus, it is likely that simply increasing the parameters of these PLMs will not enable better prediction on protein activities and other downstream tasks.

[0064] On the generative side, many attempts at de-novo generation of functional proteins like GPP and CRISPR nucleases have been performed(6, 9). However, most protein sequences designed in this scheme are non- functional and the final candidate typically has lower activity than nearest natural protein neighbors or engineered SOTA variants, limiting utility for general protein engineering efforts. In some design scenarios, generative PLMs donot outperform non-deep learning methods of sequence diversification, such as Gaussian Process models(52) or ancestral sequence reconstruction(55). With these limitations in mind, we do expect generative PLMs to perform better in the future with better architectures and training data; in fact, an end-to-end design pipeline by combining EVOLVE-Pro with zeroshot generative models might emerge as a powerful de novo protein design framework, in which low activity-high fitness variants designed by models like ESM3(6) or ProGen(S) can be rapidly optimized toward an objective property, thereby completing the design cycle.

[0065] EVOLVE-Pro functions in a “high p, low N” paradigm, meaning that although it leverages protein embeddings with dimensions between hundreds to thousands, each experimental evolution round only adds tens of new data points. This framework makes use of a random forest regressor, which imposes regularization and is structured to account for covariation between independent variables, which is particularly useful for this context. Further optimization of a vanilla random forest regressor and other emerging computational approaches that apply these algorithms to high dimensional data(54, 55) could further improve our approach. The active learning field is still in its infancy, and selection strategies in lab-in-a-loop-based approaches in biological contexts remain elementary(56). During the active-learning evolution, we found that the top-n selection of variants is the most effective, a strategy that also allows for a user to observe how EVOLVE-Pro leads to round-by -round improvement in real-time. Further Bayesian-driven frameworks that aim to quantify an uncertainty measurement in protein prediction could be tractably employed to select variants on which the model is least informed, but this would blind the user from observing improvements until a final round that assays the top predicted variants(57).

[0066] The true test of EVOLVE-Pro is whether it can perform out-of-distribution protein engineering to SOTA levels outside of benchmark DMS datasets. Every protein we attempted to improve was able to be engineered to higher levels of desired function using novel residue mutations that were not seen before. The success level of mutagenesis was high with 100% of mutations in later rounds being more active, highlighting the reasoning and insightfulness capabilities of EVOLVE-Pro. We engineered state-of-the-art variants of 5 new proteins representing antibodies, genome-editing enzymes, and polymerases that, as far as we know, have never been described before. In validation experiments, EVOLVE-Pro variants achieved SOTA performance compared to either existing wild-type versions or engineered variants. Upon analysis of top mutations, structural analysis highlighted many of the mechanisms ofactivity improvement in some cases suggesting very insightful and novel findings that can suggest future directions of improvement for enzymes.

[0067] In the context of protein design, EVOLVE-Pro is a highly capable protein engineering model. Whereas generative protein design models, such as ProGen and ESM-3, have relatively low rates of success and require high-quality initial prompting / protein backbones to properly hallucinate the full functional protein, EVOLVE-Pro 1) has high rates of success, 2) requires no special knowledge about the protein, 3) can be used for multi-objective function optimization, and 4) is multi-modal, allowing for any property with a quantifiable assay to be used as an input. Moreover, we explored the multi-mutant landscape of protein function and observed that EVOLVE-Pro is able to select highly active single mutants out of more than 16,000 possible sequences and multi-mutants from more than 780 billion possible sequences. While there is currently no protein function dataset with enough data points to cover even a fraction of the 23Npossible sequences in the complete protein design space, in the future, we expect an “ImageNet” like moment in biology where high throughput data collections on certain proteins start to yield enough data to truly enable de novo design of SOTA proteins(5S), similar to the development of AlphaFold(59). Until such a moment, we present EVOLVE-Pro as a reasoning model for general protein engineering that will be useful for the biology and drug development communities, allowing any user to engineer a protein with minimal effort and cost.

[0068] SequencesTable 3: Prime editor MLV mutants.Table 4: COVID REGN 10987 antibody mutants.Table 5: Cl 43 antibody mutants.Table 6: CD71 antibody mutants.EXAMPLESExample 1. Benchmarking on 12 DMS datasets

[0069] For model benchmarking, we took 9 existing deep mutational scanning (DMS) datasets which were employed in a previous zero-shot high fitness prediction approach. From this work, we leveraged a pre-determined cutoff for high-fitness variants for each dataset to select variants that were low and high fitness. To augment the use of these datasets, we also selected three additional DMS datasets: an AsCasl2f compact genome editor, Cov2 viral spike receptor-binding domain, and Zika virus envelope protein. For these datasets, cutoff values were set based on the general distribution of high-activity variants. To facilitate downstream work, a tabular format CSV file and a fasta file of all available mutant sequences with activity measurements were generated from each dataset.Example 2. Extraction of PLM embeddings

[0070] From the set of all available protein variants in the 12 DMS datasets, protein language model embeddings were generated. We optimized the use of ESM, ProtTrans, ProteinBERT, Ankh, and UniRep methodologies to operate on fasta files of all available mutants. These were used with default parameters. In the case of ESM embeddings, mean embeddings were generated, for esmlb, esmlv, and esm2 models of varying size. For ProtTrans, a per-protein embedding was generated. For ProteinBERT, global representations were generated. For Ankh, the mean of the last hidden state was taken for both the base and large models. For UniRep, the h_avg was generated. In all of these cases, a vector-based representation of each protein variant was the final objective. In addition to this, one-hot encodings and integer encodings of all variants were generated. In order to preserve the single n-dimensional vector embedding for each protein, one-hot encodings were concatenated across all positions.Example 3. EVOLVE-Pro Parameter Grid search

[0071] We conducted an extensive grid search to evaluate various strategies for optimizing fitness in a low number of rounds. The grid search explored the following parameters:1. Fitness measurement: Raw fitness values from each dataset or min-max normalized fitness.2. First-round strategy: Random selection of variants or diverse selection using K- medoids clustering on protein language model embeddings.3. Learning strategies: We compared several strategies for selecting variants in subsequent rounds, including: a. Random selection b. Top n predicted fitness variants c. Top n / 2 and bottom n / 2 predicted fitness variants d. Maximizing Euclidean embedding distance from previously selected variants4. Embedding types: We compared different embedding representations, including raw embeddings, normalized embeddings across residues, and PCA-reduced (10 PCs) embeddings to account for the fact that this was a high p, low n paradigm. This was entirely done on the largest (15B parameter) ESM2 model.5. Regression types: We evaluated various regression models for fitness prediction, including ridge regression, lasso regression, elastic net, linear regression, neuralnetworks with a linear last layer, random forest regression, and gradient boosting regression. These were largely used with default parameters.

[0072] For each combination of parameters, we ran three simulations (to vary the first round of selected variants) using 16 variants per round to account for stochastic variability.Performance was assessed using the proportion of high-fitness variants out of the top 16 variants that the updated model would predict. We quantified the overall effectiveness of each parameter value by counting the number of datasets for which it achieved the highest mean fitness binary percentage. This "winning strategy" count provided a simple yet informative summary of which approaches were most successful across diverse protein systems. The “winning strategy” was: random first round, raw fitness, top-n selection, random forest regression, and raw embeddings.

[0073] PLM benchmarking: Lastly, in order to understand if the largest ESM2 (15B parameter) model was optimal for our EvolvePro model, we sought to understand if our “winning strategy” would be aided by using a different underlying PLM. For each PLM, we ran ten simulations of our winning strategy and assessed this using the proportion of high- fitness variants out of the top 16 variants that the updated model would predict over 10 rounds.Example 4. EVOLVE-Pro Model

[0074] In EVOLVE-Pro, we utilize a Random Forest regressor as our top layer model to learn the functional grammar of proteins based on information-rich latent space embeddings generated by a protein language model in an active learning setting. Let xtdenote the amino acid sequence of the i-th protein variant. The protein language model embedding transformation transforms xtinto a d dimensional vector representation:

[0075] To reduce the number of features fed into the random forest regressor in the low-N setting, we obtain the average embedding vector for all variants by computing the mean of the embeddings across all residues for a given protein at length n:

[0076] The Random Forest model f then operates on these reduced embeddings to predict the fitness value yt of each protein variant. Each decision tree htwithin the Random Forest makes a prediction based on recursive binary splits of the input latent space features based on a defined threshold 0 and the composite function average across T individual decision trees.

[0077] Then this top layer domain expert model is trained to minimize the Mean Squared Error (MSE) between the predicted and actual activity values in a round-to-round active learning fashion where N varies from 10 to 100:Example 5. Comparison with zero-shots model

[0078] To assess whether our model would perform better than other one-shot strategies, we also compared 5 and 10 rounds of evolution through EvolvePro to a one-shot round of pretraining on 16, 80, 160, 500, and 1000 mutants. In this approach, we essentially conducted a single random round of this size, and again assessed using the proportion of high-fitness variants out of 16 variants predicted in this one-shot format, across 10 simulations of the initial round. We also compared a default implementation of a zero-shot high fitness prediction approach(7), where we generated variants for the 3 additional datasets not assessed in this paper for full benchmarking. Given that this one-shot approach was quite narrow, in selecting less than 32 variants, and that ultimately across 5 rounds of our strategy, we were pulling 80 variants, we manipulated parameters in this zero-shot approach to attempt to select more than 80 variants for better benchmarking.Example 6. Comparison of few shots versus many shots

[0079] We logically assumed that EvolvePro would improve if more rounds and more variants per round were provided to the underlying grid searches, given that with more training data our model would improve. Therefore, we did not vary these parameters in the initial grid search. However, to observe how the varying round size would affect the saturation of selecting only high- fitness variants, we tested 10, 20, 30, 40, 50, 100, 200, and 500 variants per round using the “winning strategy” across all 12 datasets. Given the size of certain DMS datasets, this was capped at 100 and 200 variants per round for certain datasets for which no more than 1000 or 2000 variants were available. All in-silico optimization of EvolvePro is available at https: / / github.com / matlOd / EvolveProExample 7. Experimental use of EvolvePro

[0080] Given the optimization of EvolvePro done across the 12 DMS datasets, we sought to apply our “winning strategy” across a range of protein fitness and function optimization tasks with experimental output. In this work, the “minimal strategy” was put into use, with the caveat that raw fitness was not usable given the round-by-round experimental fitness variation. During the experimental setup, the top 10 to top 12 mutants were selected for testing to better suit the high throughput setup.Example 8. SM2 fitness calculation

[0081] We use the previously established EMS2’ masked marginal scoring function(7):where M are the masked residues where mutations occur, x;mutis the mutant-type residue at position i, and x;WTis the wild-type residue at position i. This function was shown to perform best.

[0082] Embeddings are computed with ESM2-15B and then the individual mutant’s fitness score is calculated by the function above. The code for performing this calculation is available within the GitHub repositories.Example 9. Measurement of luciferase activity

[0083] Media containing secreted or intracellular luciferase was harvested 48 hours after transfection unless otherwise noted. 20pL of media is used to measure secreted luciferase activity using Targeting Systems Cypridinia and Targeting systems Gaussia luciferase assay kits (Targeting Systems) on a Biotek Synergy 4 plate reader with an injection protocol. All replicates were performed as biological replicates. Intracellular Nanoluc and firefly luciferase were measured by lysing the cell in the luciferase assay mix (Promega) according to the manufacturer’s protocol. 5 minutes after lysis at room temperature, the signal is read out using a Biotek Synergy 4 plate reader.Example 10. Quantification of protein expression

[0084] Two days after the transfection of HEK293FT or BJ Fibroblast cells, the Nano-Gio HiBiT Lytic Detection System (Promega) was used for the quantification of the HiBiT tags, in cell lysates. For the preparation of the Nano-Gio HiBiT Lytic Reagent, the Nano-Gio HiBit Lytic Buffer (Promega) was mixed with Nano-Gio HiBiT Lytic Substrate (Promega) and the LgBiT Protein (Promega) according to the manufacturer’s protocol. The volume of Nano-Gio HiBiT Lytic Reagent added was equal to the culture medium present in each well, and the samples were placed on an orbital shaker at 600 rpm for 3 minutes. After incubation of 10 minutes at room temperature, the readout took place with 125 gain and 2 seconds integration time using a plate reader (Biotek Synergy Neo 2). The control background was subtracted from the final measurements.Example 11. Harvest of total RNA and quantitative PCR

[0085] For gene expression experiments in mammalian cells, cell harvesting and reverse transcription for cDNA generation were performed using a previously described modification of the commercial Cells-to-Ct kit (Thermo Fisher Scientific) 48 h after transfection. Transcript expression was then quantified with qPCR using Fast Advanced Master Mix (Thermo Fisher Scientific) and TaqMan qPCR probes (Thermo Fisher Scientific) with GAPDH control probes (Thermo Fisher Scientific). All qPCR reactions were performed in 10-j.il reactions with two technical replicates in a 384-well format and read out using a LightCycler 480 Instrument II (Roche). For multiplexed targeting reactions, readout of different targets was performed in separate wells. Expression levels were calculated bysubtracting housekeeping control (GAPDH) cycle threshold (Ct) values from target Ct values to normalize for total input, resulting in ACt levels. Relative transcript abundance was computed as 2-ACt. All replicates were performed as biological replicates.Example 12. AAV production and purification

[0086] Recombinant AAV2 / 8 was produced by transient HEK293 cell transfection and CsCl sedimentation by the University of Massachusetts Medical School Viral Vector Core, as previously described68. Vector preparations were monitored by ddPCR, and purity was assessed by 4%— 12% SDS-acrylamide gel electrophoresis and silver staining (Invitrogen).Example 13. In vivo Flue mRNA delivery and comparative in vivo bioluminescence

[0087] Prior to bioluminescence imaging, 8 to 10-week-old Albino B6 were anesthetized with 3% isoflurane and injected with 5 pg of synthesized mRNA via retro-orbital injection using homemade lipid nanoparticles. At the indicated time points post-injection, the mice were anesthetized again with 3% isoflurane and immediately administered 200 pl of 15 mg / mL D-luciferin (PerkinElmer) for imaging. Ventral bioluminescence images were acquired using an IVIS Spectrum In Vivo Imaging System (PerkinElmer). The following conditions were used for image acquisition: exposure time = 60 sec, binning = medium: 4, field of view = 15 x 15 cm, and f / stop = 1. Bioluminescent images were analyzed using Living Image 4.3 software (PerkinElmer) and normalized radiance (photons / s) was reported.Example 14. Circular RNA production

[0088] The construction of the plasmid template for circular mRNA synthesis has been previously described (Chen et al., 2022). The plasmid was digested with Notl restriction enzyme (Thermo Fisher, FD0593), and the resulting DNA product, serving as the transcriptional template for circular RNA, was column purified using a MinElute PCR Purification Kit (Qiagen, 28006). For each 20 pl IVT reaction, 500 ng of the purified transcriptional template was used. The IVT reactions were incubated at 37°C for 12 hours, followed by degradation of the DNA template with 2 pl of DNase I per 500 ng of transcriptional template for 30 minutes at 37°C. The remaining RNA was column purified before further enzymatic processing.

[0089] A separate circularization step was performed to promote covalently closed circular RNA formation from the un-circularized IVT product. This step included IX T4 RNA ligase I buffer (NEB, B0216L), 2 mM GTP (NEB, N0450L), and 0.5 units of RNase inhibitor (NEB, M0314L) incorporated into 45 pg of silica column-purified RNA in a 50 pl reaction. The reaction mixture was heated at 55°C for 8 minutes and then silica column purified (NEB, T2040L). To isolate circular RNAs, the column-purified RNA was digested with 4 units of RNase R (Abeam, ab286929) per microgram of RNA for 15 minutes at 37°C. The samples were subsequently column purified, quantified using a Nanodrop One Microvolume UV spectrophotometer, and verified for complete digestion using the E-Gel Electrophoresis System or an Agilent TapeStation, following the manufacturer's instructions.Example 15. Antibody and antigen production

[0090] For high throughput antibody production, we transfected HEK293FT cells with 50 ng of heavy chain encoding plasmid and 50 ng of light chain encoding plasmid 24 hours after seeding at 20,000 cells per well in a 96-well plate. Antibodies are harvested 72 hours after transfection by centrifuging at 4,000x g for 10 minutes and supernatants were directly used for downstream ELISA. Here, we chose the previously established REGN10987 heavy chain only for mutation(7).

[0091] For antigen purification, SP6 stabilized antigen was His-tagged and purified using HisPur Ni-NTA resin (Thermo Fisher Scientific, 88222). Cell supernatants were diluted with 1 / 3 volume of wash buffer (20 mM imidazole, 20 mM 4-(2-hydroxyethyl)-l- piperazineethanesulfonic acid (HEPES) pH 7.4, 150 mM sodium chloride (NaCl) or 20 mM imidazole, 1 x PBS), and the Ni-NTA resin was added to diluted cell supernatants.Resin / supernatant mixtures were added to chromatography columns for gravity flow purification. The resin in the column was washed with wash buffer (20 mM imidazole, 20 mM HEPES pH 7.4, 150 mM NaCl or 20 mM imidazole, 1 x PBS), and the proteins were eluted with 250 mM imidazole, 20 mM HEPES pH 7.4, 150 mM NaCl or 20 mM imidazole, l x PBS. Column elutions were concentrated using centrifugal concentrators at 10-kDa, 50- kDa, or 100-kDa cutoffs, and stored in a storage buffer at -20°C.Example 16. High throughput Antibody concentration determination and ELISA

[0092] We adopted a previously developed high-throughput ELISA where we coated protein binding titer plates with 50uL of polyclonal goat-anti-human IgG at 2.5 pg / inl in 1 x PBSovernight at 4°C(S). After overnight coating, we removed the coating solution by flicking the plate and tapping the residual liquid on paper towels. We washed each well of the ELISA plate six times with 200 pl of washing buffer and removed the remaining fluids by flicking the plate and blocking the plate with 200 pl of blocking buffer per well for at least 45 min at 37 °C. We then washed each well of the ELISA plate six times with 200 pl of washing buffer and flicked the remaining fluids. A standard human myeloma IgG 1 kappa at 4 pg / ml in 1 x PBS was used to create a serial dilution of standards and added in with dilutions of produced antibodies for 45 minutes at room temperature. We wash each well of the ELISA plate six times with 200 pl of washing buffer and flick the remaining fluids. We then diluted HRP- conjugated secondary antibody 1 : 1,000 in a blocking buffer and added 50uL of the secondary antibodies to each well. Plates were incubated at room temperature for 45 minutes. After incubation, plates were washed again six times with wash buffer, and 5 Opl of ABTS solution at room temperature was added per well. 5 minutes after addition, 2M sulfuric acid is added to stop the reaction and OD450 is read out using a Biotek plate reader. For the binding activity of the antibody, the same protocol is used except coating the plate with lOOnM of purified antigen at 4°C overnight.Example 17. dsRNA ELISA

[0093] The dsRNA byproduct from in vitro transcription was detected using a dsRNA sandwich enzyme-linked immunosorbent assay (ELISA) that selectively identifies multispecies dsRNA molecules larger than 30-40 bp. This assay was performed with a multispecies dsRNA ELISA Kit (Novus Biologicals, NBP3-11368), using antibodies as previously described by Schonbom, J., et al. The KI (IgG2a) mouse monoclonal antibody was immobilized on 96-well Immulon 2 HB plates (Thermo Fisher Scientific, 3455) overnight at 4°C, then blocked with 1% BSA in PBS at 37°C for 2 hours. After triple washing, mRNA samples and Poly (EC) dsRNA standards were added to the plates and incubated for 1 hour at 37°C. Following another triple wash, the plates were incubated with monoclonal antibody K2 (IgM) at 37°C for 1 hour. After triple washing again, the plates were exposed to horseradish peroxidase (HRP)-conjugated F(ab’)2 fragment of goat anti-mouse secondary antibody at 37°C for 1 hour, followed by a final wash. Subsequently, 100 pL of TMB (3, 3', 5, 5' tetramethylbenzidine) substrate solution was added to each well, followed by 100 pL of 2M H2SO4. The absorbance was measured at 450 nm using a BioTek Synergy Neo2 Hybrid Multimode Reader (BioTek, BTNEO2).Example 18. mRNA analysis by Egel system and Tapestation

[0094] Linear RNA or isolated circular RNA was column purified and quantified using a NanoDrop One spectrophotometer. For RNA characterization via the E-gel system, 1200 ng of RNA samples and 4 pL ssRNA Ladder (NEB, N0362S) were denatured by a 1 : 1 dilution with formamide (Sigma, F7503-250ML). The samples were then loaded onto 2% E-Gel™ EX Agarose Gels with SYBR-GOLD II and run on the E-Gel Power Snap Plus Electrophoresis System (Thermo Fisher, G9101) using the settings: E-gel Category “11 wells,” E-gel Type “E-Gel™ EX 2%,” and Time “12 minutes” at room temperature. Images were captured using a Bio-Rad ChemiDoc Imaging System with the “SYBR-Gold” settings. For the characterization of circular RNA using the Agilent TapeStation, 2 pL of 10 ng / pL mRNA samples and RNA ladder were mixed with 1 pL of High Sensitivity RNA Sample Buffer. The samples were denatured at 72°C for 3 minutes and then cooled at 4°C for 2 minutes. They were subsequently loaded into the 4200 TapeStation instrument (Agilent, G2991BA) and analyzed using the 4200 TapeStation Controller Software, following the manufacturer’s instructions.Example 19. LNP production

[0095] LNPs were prepared using a vortex mixing method(9). ALC-0315, DSPC, cholesterol, and DMG-PEG-2000 were dissolved in ethanol at 75, 10, 10, and 10 mg / ml, respectively, and mixed at a molar ratio of 50: 10:38.5: 1.5. The final volume was brought to 30 pl with ethanol. Separately, 10 pg of purified mRNA was diluted in 10 mM citrate buffer (pH 4) to a final volume of 90 pl. The lipid mixture (30 pl) was rapidly pipetted into the vortexing RNA solution (90 pl) at a 1 :3 volume ratio, and vortexed for an additional 20 seconds. The mixture was left at room temperature for 10-15 minutes, then dialyzed against PBS at 4°C overnight to remove ethanol and residual lipids, resulting in the final LNP formulation. RNA encapsulation efficiency was measured using the Thermo Fisher RiboGreen assay. LNP samples were treated with Tris-EDTA or Tris-EDTA + 1% Triton-X. Free RNA was detected using an RNA-binding fluorescent dye, RiboGreen, and fluorescence was measured with a plate reader. Encapsulation efficiency was quantified by comparing the fluorescence of treated samples to RNA standards.

[0096] Alternatively, LNPs were synthesized using the microfluidic organic-aqueous precipitation method. The organic / ethanol phase consisted of lipids Dlin-MC3-DMA, DSPC,Cholesterol, and DMG-PEG2k at a molar ratio of 50: 10:38.5: 1.5. The aqueous phase was prepared by diluting the RNA payload in a 10 mM citrate buffer at pH 3.0. The two phases were prepared at an ethanokaqueous volume ratio of 1 :2, and at an N:P ratio of 5 : 1. The phases were mixed through the NxGen microfluidic cartridge by the NanoAssemblr Ignite (Precision Nanosystems). The Ignite was set to: volume ratio- 2: 1; flow rate- 12 ml / min; waste volume- 0 mL. The resulting LNPs were dialyzed against PBS using 20K MWCO Slide- A-Lyzer™ MINI Dialysis cassettes (ThermoFisher Scientific) at 25°C for 90 min, with an exchange of the buffer reservoir at 45 min.Example 20. Mammalian cell culture and transfection

[0097] Mammalian cell culture experiments were performed in the HEK293FT (Thermo Fisher Scientific), BJ Fibroblast (ATCC CRL-2522), and Hepa 1-6 (ATCC CRL-1830) cell lines, grown in Dulbecco’s Modified Eagle Medium with high glucose, sodium pyruvate, and GlutaMAX (Thermo Fisher Scientific), and supplemented with 1 x penicillin-streptomycin (Thermo Fisher Scientific) and 10% fetal bovine serum (VWR Seradigm). All cells were maintained at confluency below 80%. All transfections were performed with Lipofectamine 3000 (Thermo Fisher Scientific). Cells were plated 16-20 hours prior to transfection to ensure 90% confluency at the time of transfection. For 96-well plates, cells were plated at 2x 104cells / well. For each well on the plate, transfection plasmids were combined with OptiMEM I Reduced Serum Medium (Thermo Fisher Scientific) to a final volume of 10 pL.Example 21. Mammalian genome editing

[0098] To measure genome editing activity, 50 ng of protein expression construct, 50 ng of the corresponding guide construct, and optionally 20 ng of luciferase reporter were transfected in one well of a 96-well plate, using Lipofectamine 3000. After 72 h, the cells were washed once with 1 x DPBS (Sigma Aldrich). The cells were resuspended in 50 pL QuickExtract DNA Extraction Solution (Lucigen) and cycled at 65°C for 15 min, 68°C for 15 min, and then at 95°C for 10 min for lysis. A 2.5 pL aliquot of the cell lysate was used as input for each PCR reaction. For library amplification, target reporter regions were amplified with a 12-cycle PCR using NEBNext High Fidelity 2 x PCR Master Mix (NEB) with an annealing temperature of 63°C for 15 s, followed by a second 20-cycle round of PCR to add Illumina adapters and barcodes. The libraries were gel extracted and subject to paired-endsequencing on an Illumina MiSeq with Read 1 220 cycles, Index 1 8 cycles, Index 2 8 cycles, and Read 2 80 cycles. Insertion / deletion (indel) frequency was analyzed using CRISPResso2(70). All guide sequences are available in Data S5.Example 22. Cloning of twinPE pegRNA

[0099] pegRNA were cloned by Golden Gate assembly of PCR products. Guide products were amplified by PCR (KAPA HiFi HotStart DNA polymerase, Roche) off of the Cas9 single guide RNA scaffold, with the forward primer containing spacer sequences and the reverse primer containing desired PBS, RT, and attB insertion sequences, in the case of the pegRNA. PCR products were purified by gel extraction (Monarch gel extraction kit, NEB) and assembled in a Golden Gate assembly containing 6.25 ng of pU6-atgRNA-GG-acceptor (Addgene, 132777), purified PCR product (approximately two- to four fold molar excess), 0.125 pl ofFermentas Eco31I (Thermo Fisher Scientific), 0.0625 pl ofT7 DNA ligase (Enzymatics), 0.0625 pl of 20 mg ml-1 bovine serum albumin (NEB), 2x reaction ligation buffer (Enzymatics) and water, for a 6.25-pl total reaction volume. Reactions were incubated between 37 °C and 20 °C for 5 min each for a total of 15 cycles. Two micro liters of assembled reactions were transformed into 20 pl of competent Stbl3 cells generated by Mix and Go! competency kit (Zymo) and plated on agar plates supplemented with appropriate antibiotics. After overnight growth at 37 °C, colonies were picked into Terrific Broth (TB) medium (Thermo Fisher Scientific) and incubated with shaking at 37 °C for 24 h. Cultures were collected using a QIAprep Spin Miniprep kit (Qiagen) according to the manufacturer’s instructions.Example 23. High throughput Cloning of mutants

[0100] Expression constructs for antibody and prime editor were cloned for mammalian expression via Gibson cloning using Hifi Assembly mix (NEB) according to the manufacturer’s instructions. Overlapping reverse and forward primer- carrying mutations for desired amino acids are used to amplify the plasmid around the globe with 18bp of homology. Then DPNI is used to clean up the plasmid from PCR reactions followed by column cleanup. 50ng of the cleaned-up PCR product is then used to perform Gibson reactions according to the manufacture’s protocol. For all Gibson clonings, 2 pl of assembled reactions were transformed into 20 pl of competent Stbl3 cells generated by Mix and Go!competency kit (Zymo) and plated on agar plates supplemented with appropriate antibiotics. After growth overnight at 37 °C, colonies were picked into TB medium (Thermo Fisher Scientific) and incubated with shaking at 37 °C for 24 h. Cultures were collected using a QIAprep Spin Miniprep kit (Qiagen) according to the manufacturer’s instructions.

[0101] References:1. B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275-283 (2024).2. T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCasl2f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920-4935. e23 (2023).3. A. J. Greaney, T. N. Starr, C. O. Barnes, Y. Weisblum, F. Schmidt, M. Caskey, C. Gaebler, A. Cho, M. Agudelo, S. Finkin, Z. Wang, D. Poston, F. Muecksch, T. Hatziioannou, P. D. Bieniasz, D. F. Robbiani, M. C. Nussenzweig, P. J. Bjorkman, J. D. Bloom, Mapping mutations to the SARS-CoV-2 RBD that escape binding by different classes of antibodies. Nat. Commun. 12, 4196 (2021).4. M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol. 93 (2019).5. A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.6. E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315-1322 (2019).7. J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zeroshot prediction of the effects of mutations on protein function, bioRxiv (2021) p. 2021.07.09.450648.8. L. Gieselmann, C. Kreer, M. S. Ercanoglu, N. Lehnen, M. Zehner, P. Schommers, J. Potthoff, H. Gruell, F. Klein, Effective high-throughput isolation of fully human antibodies targeting infectious pathogens. Nat. Protoc. 16, 3639-3671 (2021).9. X. Wang, S. Liu, Y. Sun, X. Yu, S. M. Lee, Q. Cheng, T. Wei, J. Gong, J. Robinson, D. Zhang, X. Lian, P. Basak, D. J. Siegwart, Preparation of selective organ-targeting (SORT) lipid nanoparticles (LNPs) using multiple technical methods for tissue-specific mRNA delivery. Nat. Protoc. 18, 265-291 (2023).10. K. Clement, H. Rees, M. C. Canver, J. M. Gehrke, R. Farouni, J. Y. Hsu, M. A. Cole, D. R. Liu, J. K. Joung, D. E. Bauer, L. Pinello, CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226 (2019).11. J. Schymkowitz, J. Borg, F. Stricher, R. Nys, F. Rousseau, L. Serrano, The FoldX web server: an online force field. Nucleic Acids Res. 33, W382-8 (2005).12. S. Bae, J. Park, J.-S. Kim, Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473— 1475 (2014).13. Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, N. Smetanin, R. Verkuil, O. Kabeli, Y. Shmueli, A. Dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido, A. Rives, Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023).14. M. Heinzinger, K. Weissenow, J. G. Sanchez, A. Henkel, M. Mirdita, M. Steinegger, B. Rost, Bilingual Language Model for Protein Sequence and Structure, bioRxiv (2024)p. 2023.07.23.550085.15. A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.16. N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, M. Linial, ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics 38, 2102-2110 (2022).17. Y. He, X. Zhou, C. Chang, G. Chen, W. Liu, G. Li, X. Fan, M. Sun, C. Miao, Q. Huang, Y. Ma, F. Yuan, X. Chang, Protein language models-assisted optimization of an uracil- N-glycosylase variant enables programmable T-to-G and T-to-C base editing. Mol. Cell 84, 1257-1270. e6 (2024).18. T. Hayes, R. Rao, H. Akin, N. J. Sofroniew, D. Oktay, Z. Lin, R. Verkuil, V. Q. Tran, J. Deaton, M. Wiggert, R. Badkundri, I. Shafkat, J. Gong, A. Derry, R. S. Molina, N. Thomas, Y. A. Khan, C. Mishra, C. Kim, L. J. Bartie, M. Nemeth, P. D. Hsu, T. Sercu, S. Candido, A. Rives, Simulating 500 million years of evolution with a language model, bioRxiv (2024)p. 2024.07.01.600583.19. N. Ferruz, S. Schmidt, B. Hocker, ProtGPT2 is a deep unsupervised language model for protein design. Nat. Commun. 13, 4348 (2022). 0. A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos Jr, C. Xiong, Z. Z. Sun, R. Socher, J. S. Fraser, N. Naik, Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099— 1106 (2023). 1. J. A. Ruffolo, S. Nayfach, J. Gallagher, A. Bhatnagar, J. Beazer, R. Hussain, J. Russ, J. Yip, E. Hill, M. Pacesa, A. J. Meeske, P. Cameron, A. Madani, Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences, bioRxiv (2024)p. 2024.04.22.590591. 2. K. K. Yang, Z. Wu, F. H. Arnold, Machine-leaming-guided directed evolution for protein engineering. Nat. Methods 16, 687-694 (2019). 3. H. Lu, D. J. Diaz, N. J. Czarnecki, C. Zhu, W. Kim, R. Shroff, D. J. Acosta, B. R. Alexander, H. O. Cole, Y. Zhang, N. A. Lynd, A. D. Ellington, H. S. Alper, Machine learning-aided engineering of hydrolases for PET depolymerization. Nature 604, 662- 667 (2022). 4. N. Thomas, D. Belanger, C. Xu, H. Lee, K. Hirano, K. Iwai, V. Polic, K. D. Nyberg, K. G. Hoff, L. Frenz, C. A. Emrich, J. W. Kim, M. Chavarha, A. Ramanan, J. J. Agresti, L.J. Colwell, Engineering of highly active and diverse nuclease enzymes by combiningmachine learning and ultra-high-throughput screening, bioRxiv (2024)p. 2024.03.21.585615.25. Z. Wu, S. B. J. Kan, R. D. Lewis, B. J. Wittmann, F. H. Arnold, Machine learning- assisted directed protein evolution with combinatorial libraries. Proc. Natl. Acad. Sci. U. S. A. 116, 8852-8858 (2019).26. B. J. Wittmann, Y. Yue, F. H. Arnold, Informed training set design enables efficient machine learning-assisted directed protein evolution. Cell Syst 12, 1026-1045. e7 (2021).27. S. Biswas, G. Khimulya, E. C. Alley, K. M. Esvelt, G. M. Church, Low-N protein engineering with data-efficient deep learning. Nat. Methods 18, 389-396 (2021).28. L. Brenan, A. Andreev, O. Cohen, S. Pantel, A. Kamburov, D. Cacchiarelli, N. S. Persky, C. Zhu, M. Bagul, E. M. Goetz, A. B. Burgin, L. A. Garraway, G. Getz, T. S. Mikkelsen, F. Piccioni, D. E. Root, C. M. Johannessen, Phenotypic Characterization of a Comprehensive Set of MAPK1 / ERK2 Missense Mutants. Cell Rep. 17, 1171-1183 (2016).29. P. Notin, A. W. Kollasch, D. Ritter, L. van Niekerk, S. Paul, H. Spinner, N. Rollins, A. Shaw, R. Weitzman, J. Frazer, M. Dias, D. Franceschi, R. Orenbuch, Y. Gal, D. S. Marks, ProteinGym: Large-Scale Benchmarks for Protein Design and Fitness Prediction. bioRxiv, doi: 10.1101 / 2023.12.07.570727 (2023).30. T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCasl2f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920M935.c23 (2023).31. H. K. Haddox, A. S. Dingens, J. D. Bloom, Experimental Estimation of the Effects of All Amino-Acid Mutations to HIV’s Envelope Protein on Viral Replication in Cell Culture. PLoS Pathog. 12, el006114 (2016).32. E. D. Kelsic, H. Chung, N. Cohen, J. Park, H. H. Wang, R. Kishony, RNA Structural Determinants of Optimal Codons Revealed by MAGE-Seq. Cell Syst 3, 563-57 l.e6(2016).33. M. A. Stiffler, D. R. Hekstra, R. Ranganathan, Evolvability as a function of purifying selection in TEM- 1 p-lactamase. Cell 160, 882-892 (2015).34. C. J. Markin, D. A. Mokhtari, F. Sunden, M. J. Appel, E. Akiva, S. A. Longwell, C. Sabatti, D. Herschlag, P. M. Fordyce, Revealing enzyme functional architecture via high-throughput microfluidic enzyme kinetics. Science 373 (2021).35. A. O. Giacomelli, X. Yang, R. E. Lintner, J. M. McFarland, M. Duby, J. Kim, T. P. Howard, D. Y. Takeda, S. H. Ly, E. Kim, H. S. Gannon, B. Hurhula, T. Sharpe, A. Goodale, B. Fritchman, S. Steelman, F. Vazquez, A. Tshemiak, A. J. Aguirre, J. G. Doench, F. Piccioni, C. W. M. Roberts, M. Meyerson, G. Getz, C. M. Johannessen, D. E. Root, W. C. Hahn, Mutational processes shape the landscape of TP53 mutations in human cancer. Nat. Genet. 50, 1381-1387 (2018).36. E. M. Jones, N. B. Lubock, A. J. Venkatakrishnan, J. Wang, A. M. Tseng, J. M. Paggi, N. R. Latorraca, D. Cancilla, M. Satyadi, J. E. Davis, M. M. Babu, R. O. Dror, S. Kosuri, Structural and functional characterization of G protein-coupled receptors with deep mutational scanning. Elife 9 (2020).37. M. B. Doud, J. D. Bloom, Accurate Measurement of the Effects of All Amino-Acid Mutations on Influenza Hemagglutinin. Viruses 8 (2016).38. J. M. Lee, J. Huddleston, M. B. Doud, K. A. Hooper, N. C. Wu, T. Bedford, J. D. Bloom, Deep mutational scanning of hemagglutinin helps predict evolutionary fates of human H3N2 influenza variants. Proc. Natl. Acad. Sci. U. S. A. 115, E8276-E8285 (2018).39. A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma, R. Fergus, Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. U. S. A. 1 18 (2021).40. E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315-1322 (2019).41. A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y. Wang, L. Jones, T. Gibbs, T.Feher, C. Angerer, M. Steinegger, D. Bhowmik, B. Rost, ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE Trans. Pattern Anal. Mach. Intell. 44, 7112-7127 (2022). B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275-283 (2024). A. Baum, D. Ajithdoss, R. Copin, A. Zhou, K. Lanza, N. Negron, M. Ni, Y. Wei, K. Mohammadi, B. Musser, G. S. Atwal, A. Oyejide, Y. Goez-Gazi, J. Dutton, E. Clemmons, H. M. Staples, C. Bartley, B. Klaffke, K. Alfson, M. Gazi, O. Gonzalez, E. Dick Jr, R. Carrion Jr, L. Pessaint, M. Porto, A. Cook, R. Brown, V. Ali, J. Greenhouse, T. Taylor, H. Andersen, M. G. Lewis, N. Stahl, A. J. Murphy, G. D. Yancopoulos, C. A. Kyratsous, REGN-COV2 antibodies prevent and treat SARS-CoV-2 infection in rhesus macaques and hamsters. Science 370, 1110-1115 (2020). C.-L. Hsieh, J. A. Goldsmith, J. M. Schaub, A. M. DiVenere, H.-C. Kuo, K. Javanmardi, K. C. Le, D. Wrapp, A. G. Lee, Y. Liu, C.-W. Chou, P. O. Byrne, C. K. Hjorth, N. V. Johnson, J. Ludes-Meyers, A. W. Nguyen, J. Park, N. Wang, D. Amengor, J. J. Lavinder, G. C. Ippolito, J. A. Maynard, I. J. Finkelstein, J. S. McLellan, Structure-based design of prefiision-stabilized SARS-CoV-2 spikes. Science 369, 1501-1505 (2020). C. Xin, J. Yin, S. Yuan, L. Ou, M. Liu, W. Zhang, J. Hu, Comprehensive assessment of miniature CRISPR-Casl2f nucleases for gene disruption. Nat. Commun. 13, 5623 (2022). Z. Wu, Y. Zhang, H. Yu, D. Pan, Y. Wang, Y. Wang, F. Li, C. Liu, H. Nan, W. Chen, Q. Ji, Programmed genome editing by a miniature CRISPR-Casl2f nuclease. Nat. Chem. Biol. 17, 1132-1138 (2021). X. Xu, A. Chemparathy, L. Zeng, H. R. Kempton, S. Shang, M. Nakamura, L. S. Qi, Engineered miniature CRISPR-Cas system for mammalian genome regulation and editing. Mol. Cell 81, 4333-4345.e4 (2021). B. P. Kleinstiver, A. A. Sousa, R. T. Walton, Y. E. Tak, J. Y. Hsu, K. Clement, M. M. Welch, J. E. Homg, J. Malagon-Lopez, I. Scarfo, M. V. Maus, L. Pinello, M. J. Aryee, J. K. Joung, Engineered CRISPR-Cas 12a variants with increased activities and improvedtargeting ranges for gene, epigenetic and base editing. Nat. Biotechnol. 37, 276-282 (2019).49. X. Kong, H. Zhang, G. Li, Z. Wang, X. Kong, L. Wang, M. Xue, W. Zhang, Y. Wang, J. Lin, J. Zhou, X. Shen, Y. Wei, N. Zhong, W. Bai, Y. Yuan, L. Shi, Y. Zhou, H. Yang, Engineered CRISPR-OsCasl2fl and RhCasl2fl with robust activities and expanded target range for genome editing. Nat. Commun. 14, 2046 (2023).50. L. Zhang, J. A. Zuris, R. Viswanathan, J. N. Edelstein, R. Turk, B. Thommandru, H. T. Rube, S. E. Glenn, M. A. Collingwood, N. M. Bode, S. F. Beaudoin, S. Lele, S. N. Scott, K. M. Wasko, S. Sexton, C. M. Borges, M. S. Schubert, G. L. Kurgan, M. S. McNeill, C. A. Fernandez, V. E. Myer, R. A. Morgan, M. A. Behlke, C. A. Vakulskas, AsCasl2a ultra nuclease facilitates the rapid generation of therapeutic cell medicines. Nat. Commun. 12, 3908 (2021).51. D. Y. Kim, J. M. Lee, S. B. Moon, H. J. Chin, S. Park, Y. Lim, D. Kim, T. Koo, J.-H.Ko, Y.-S. Kim, Efficient CRISPR editing with a hypercompact Casl2fl and engineered guide RNAs delivered by adeno-associated virus. Nat. Biotechnol. 40, 94-102 (2022).52. P. Pausch, B. Al-Shayeb, E. Bisom-Rapp, C. A. Tsuchida, Z. Li, B. F. Cress, G. J. Knott, S. E. Jacobsen, J. F. Banfield, J. A. Doudna, CRISPR-CasO from huge phages is a hypercompact genome editor. Science 369, 333-337 (2020).53. J. L. Doman, S. Pandey, M. E. Neugebauer, M. An, J. R. Davis, P. B. Randolph, A. McElroy, X. D. Gao, A. Raguram, M. F. Richter, K. A. Everette, S. Banskota, K. Tian, Y. A. Tao, J. Tolar, M. J. Osborn, D. R. Liu, Phage-assisted evolution and protein engineering yield compact, efficient prime editors. Cell 186, 3983M002. e26 (2023).54. M. T. N. Yamall, E. I. loannidi, C. Schmitt-Ulms, R. N. Krajeski, J. Lim, L. Villiger, W. Zhou, K. Jiang, S. K. Garushyants, N. Roberts, L. Zhang, C. A. Vakulskas, J. A. Walker, A. P. Kadina, A. E. Zepeda, K. Holden, H. Ma, J. Xie, G. Gao, L. Foquet, G. Bial, S. K. Donnelly, Y. Miyata, D. R. Radiloff, J. M. Henderson, A. Ujita, O. O. Abudayyeh, J. S. Gootenberg, Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases. Nat. Biotechnol., 1-13 (2022).55. J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zeroshot prediction of the effects of mutations on protein function, bioRxiv (202 l)p.2021.07.09.450648.56. A. Dousis, K. Ravichandran, E. M. Hobert, M. J. Moore, A. E. Rabideau, An engineered T7 RNA polymerase that produces mRNA free of immunostimulatory byproducts. Nat. Biotechnol. 41, 560-568 (2023).57. Z. J. Kartje, H. I. Janis, S. Mukhopadhyay, K. T. Gagnon, Revisiting T7 RNA polymerase transcription in vitro with the Broccoli RNA aptamer as a simplified realtime fluorescent reporter. J. Biol. Chem. 296, 100175 (2021).58. R. Chen, S. K. Wang, J. A. Belk, L. Amaya, Z. Li, A. Cardenas, B. T. Abe, C.-K. Chen, P. A. Wender, H. Y. Chang, Author Correction: Engineering circular RNA for enhanced protein production. Nat. Biotechnol. 41, 293 (2023).59. S. R. Johnson, X. Fu, S. Viknander, C. Goldin, S. Monaco, A. Zelezniak, K. K. Yang, Computational scoring and experimental evaluation of enzymes generated by neural networks. Nat. Biotechnol., doi: 10.1038 / s41587-024-02214-2 (2024).60. V. R. Shanker, T. U. J. Bruun, B. L. Hie, P. S. Kim, Unsupervised evolution of protein and antibody complexes with a structure-informed language model. Science 385, 46-53 (2024).61. Y. Serrano, A. Ciudad, A. Molina, Are Protein Language Models Compute Optimal?, arXiv [q-bio.BM] (2024). http: / / arxiv.org / abs / 2406.07249.62. X. Cheng, B. Chen, P. Li, J. Gong, J. Tang, L. Song, Training Compute-Optimal Protein Language Models, bioRxiv (2024)p. 2024.06.06.597716.63. B. Chen, X. Cheng, P. Li, Y.-A. Geng, J. Gong, S. Li, Z. Bei, X. Tan, B. Wang, X. Zeng, C. Liu, A. Zeng, Y. Dong, J. Tang, L. Song, xTrimoPGLM: Unified lOOB-Scale Pretrained Transformer for Deciphering the Language of Protein, bioRxiv (2024)p.2023.07.05.547496.64. C. N. Bedbrook, K. K. Yang, J. E. Robinson, E. D. Mackey, V. Gradinaru, F. H. Arnold, Machine learning-guided channelrhodopsin engineering enables minimally invasive optogenetics. Nat. Methods 16, 1176-1184 (2019).65. J. W. Thornton, Resurrecting ancient genes: experimental analysis of extinct molecules.Nat. Rev. Genet. 5, 366-375 (2004). D. Ghosh, J. Cabrera, Enriched Random Forest for High Dimensional Genomic Data.IEEE / ACM Trans. Comput. Biol. Bioinform. 19, 2817-2828 (2022). A. Kirjner, J. Yim, R. Samusevich, S. Bracha, T. S. Jaakkola, R. Barzilay, I. R. Fiete, Improving protein optimization with smoothed fitness landscapes (2023). https: / / openreview.net / pdf7idmxlF2Zv8x0. K. Huang, R. Lopez, J.-C. Hutter, T. Kudo, A. Rios, A. Regev, Sequential Optimal Experimental Design of Perturbation Screens Guided by Multi-modal Priors, bioRxiv (2023)p. 2023.12.12.571389. P. M. Groth, M. H. Kerrn, L. Olsen, J. Salomon, W. Boomsma, Protein property prediction with uncertainties, arXiv [q-bio.BM] (2024). http: / / arxiv.org / abs / 2407.00002. J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, “ImageNet: A large-scale hierarchical image database” in 2009 IEEE Conference on Computer Vision and Pattern Recognition (IEEE, 2009), pp. 248-255. J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figumov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Zidek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, D. Hassabis, Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021). M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol. 93 (2019).

Claims

What is claimed:

1. A prime editor comprising an amino acid sequence at least 70% identical to any one of the amino acid sequences from Table 3.

2. A nucleic acid molecule encoding the prime editor of claim 1.

3. A vector comprising the nucleic acid molecule of claim 2.

4. A cell comprising the vector of claim 3.

5. A method of site-specific integration of a nucleic acid into a genome of a cell, the method comprising:(a) incorporating an integration site at a desired location in the genome by introducing into the cell: i. a DNA binding nuclease linked to a reverse transcriptase, wherein the DNA binding nuclease comprises a nickase activity; and ii. a guide RNA (gRNA) comprising a primer binding sequence linked to an integration sequence, wherein the gRNA interacts with the DNA binding nuclease and targets the desired location in the genome, wherein the DNA binding nuclease nicks a strand of the genome and the reverse transcriptase incorporates the integration sequence of the gRNA into the nicked site, thereby providing the integration site at the desired location of the genome; and(b) integrating the nucleic acid into the genome by introducing into the cell: iii. a DNA or RNA strand comprising the nucleic acid linked to a sequence that is complementary or associated to the integration site; and iv. an integration enzyme, wherein the integration enzyme incorporates the nucleic acid into the genome at the integration site by integration, recombination, or reverse transcription of the sequence that is complementary or associated to the integration site, thereby introducing the nucleic acid into the desired location of the cell genome of the cellwherein the DNA binding nuclease linked to the reverse transcriptase comprises the prime editor of claim 1.

6. A C143 antibody comprising an amino acid sequence at least 70% identical to any one of the amino acid sequences from Table 5.

7. A C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 relative to a wild-type Cl 43 antibody sequence of SEQ ID NO. 143.

8. A C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 and / or 40 and / or at least one mutation at heavy chain position 39 relative to a wild-type Cl 43 antibody sequence of SEQ ID NO. 143.

9. A C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28, 40, 50, 45 and / or 14 and / or at least one mutation at heavy chain position 39, 63, and / or 89 relative to a wild-type C143 antibody sequence of SEQ ID NO. 143.

10. A C143 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 and / or 14 and / or at least one mutation at heavy chain position 33, 39, and / or 58 relative to a wild-type C143 antibody sequence of SEQ ID NO. 143.

11. A C 143 antibody comprising an amino acid sequence including a mutation at light chain position N28R / Q40K and a mutation at heavy chain position R39K relative to a wildtype C143 antibody sequence of SEQ ID NO. 135.

12. A C143 antibody comprising an amino acid sequence including a mutation at light chain position N28K relative to a wild-type C143 antibody sequence of SEQ ID NO. 135.

13. A method for treating a disease or a disorder in a subject, the method comprising administering to the subject in need thereof the antibody of claim 6, thereby treating the disease or the disorder in the subject.

14. A aCD71 antibody comprising an amino acid sequence at least 70% identical to any one of the amino acid sequences from Table 6.

15. A aCD71 antibody comprising an amino acid sequence including at least one mutation at light chain position 28 and / or 40 and / or at least one mutation at heavy chain position 39 relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

16. A aCD71 antibody comprising an amino acid sequence including at least one mutation at heavy chain position 70 and / or 92 and / or at least one mutation at light chain position 38 relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

17. A aCD71 antibody comprising an amino acid sequence including at least one mutation at position S92A relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

18. A aCD71 antibody comprising an amino acid sequence including mutations at T70A S92V relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

19. A aCD71 antibody comprising an amino acid sequence including at least one mutation at heavy chain position 39, 63, and / or 89 and / or at least one mutation at light chain position 14, 40, 50 and / or 45 relative to a wild-type CD71 antibody sequence of SEQ ID NO. 222.

20. A method for treating a disease or a disorder in a subject, the method comprising administering to the subject in need thereof the antibody of claim 14, thereby treating the disease or the disorder in the subject.

Citation Information

Patent Citations

  • Prime editing composition with improved editing efficiency

    WO2022169235A1