Systems and methods for protein design
ProseLM addresses the limitations of existing protein design methods by incorporating structural and functional context, resulting in improved engineered antibodies and genome editors with enhanced stability and binding affinity.
Patent Information
- Application Number
- PCT/US2025/040355
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2025-08-01
- Publication Date
- 2026-02-05
AI Technical Summary
Existing protein design methods, such as physics-based and deep learning approaches, are limited by the accuracy and quantity of training data, restricting the design of novel proteins with desired functions.
The use of ProseLM, a protein language model that incorporates structural and functional context, including backbone and nearby molecule interactions, to enhance protein sequence design.
ProseLM improves the design of engineered antibodies by enhancing stability, binding affinity, and reducing immunogenicity, while also improving the performance of genome editors and therapeutic antibodies.
Smart Images

Figure IMGF000035_0001 
Figure IMGF000036_0001 
Figure IMGF000087_0001
Abstract
Description
[0001] SYSTEMS AND METHODS FOR PROTEIN DESIGN
[0002] FIELD
[0003] Provided herein are systems and methods for protein sequence design. Also provided herein are engineered antibody sequences.
[0004] CROSS REFERENCE TO RELATED APPLICATIONS
[0005] This application claims the benefit of U.S. Provisional Application No. 63 / 678,316, filed August 1, 2024, the content of which is herein incorporated by reference in its entirety.
[0006] SEQUENCE LISTING STATEMENT
[0007] The content of the electronic sequence listing titled PROF 43534 601 SequenceListing. xml (Size: 379,167 bytes; and Date of Creation: July 31, 2025) is herein incorporated by reference in its entirety.
[0008] BACKGROUND
[0009] Protein sequence design aims to identify an amino acid sequence that will fold into a desired backbone and carry out a function of interest. The ability to design proteins with novel functions has broad applications in biotechnology and medicine, including the development of therapeutics, vaccines, and industrial enzymes. Physics-based methods, such as Rosetta, approach protein design as an optimization problem, searching for sequences that minimize an energy function. However, these methods are limited by the accuracy of the underlying energy function and are computationally expensive in practice. Recently, deep learning approaches, which learn a mapping from structure to sequence, have emerged as an alternative to physicsbased methods. Generative models for protein design, such as ProteinMPNN, have proven successful across a variety of tasks, including design of protein binders, assemblies, diversified enzymes, and conformational switches. Despite their success, protein design models trained solely on experimentally determined structures are ultimately limited by the quantity and diversity of their training data. While the Protein Data Bank (PDB) contains over 180,000 protein structures, this number is dwarfed by the number of sequences identified through genomic and metagenomic sequencing efforts. To overcome this disparity, predictions from AlphaFold2 have been used to supplement the structures used for training. This approach has been shown to improve the performance of protein design models, but is still limited by the number of structures that can be predicted and the accuracy of the predictions. SUMMARY
[0010] Provided herein are engineered antibodies. In some embodiments, the engineered antibodies are generated using the methods and system for protein sequence design described herein.
[0011] In some embodiments, the engineered antibodies comprise a heavy chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 1 and / or a light chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 2. In some embodiments, the engineered antibodies comprise a heavy chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 3 and / or a light chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 4. The engineered antibodies may have improved stability, increased binding affinity, reduced immunogenicity, decreased conformational diversity or flexibility, or a combination thereof as compared to an antibody lacking the one or more substitutions, deletions, or additions.
[0012] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions as compared to SEQ ID NO: 1 at positions selected from: 3, 5, 11, 13, 16, 19, 21, 23, 27, 30, 32, 37, 43, 52, 58, 59, 80, 88, 89, 97, and 99. In some embodiments, the one or more substitutions are selected from: Q3K; V5Q; VI IL; Q13K; R16E, R16T, or R16G; R19K; D21S; K23A, K23T, K23R, or K23V; I27V, I27L, I27F, or I27M; S30N; S32Y, S32F, S32H, S32K, S32R, S32N, or S32L; V37I; K43Q; W52S; R58K, R58T, or R58E; Y59F, Y59N, or Y59H; F80Y or F80H; A88S or A88T; E89D; A97T; N99T, N99E, N99S, N99H, N99D, or N99Y; and combinations thereof.
[0013] In some embodiments, the one or more substitutions comprises a substitution at position 32, as compared to SEQ ID NO: 1. In some embodiments, the substitution at position 32 is selected from S32Y, S32F, S32H, S32K, S32R, S32N, or S32L.
[0014] In some embodiments, the one or more substitutions comprises a substitution at one or more or all of positions: 11, 21, 23, 58, and 80, as compared to SEQ ID NO: 1. In some embodiments, the one or more substitutions are selected from VI IL; D21S; K23A, K23T, K23R, or K23V; R58K, R58T, or R58E; F80Y or F80H; and combinations thereof.
[0015] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions as compared to SEQ ID NO: 2 at positions selected from: 1, 3, 4, 9, 10, 14, 28, 29, 30, 31, 58, 60, 70, 77, 81, 90, 91, 93, and 97. In some embodiments, the one or more substitutions are selected from: EID or EIQ; V3L or V3Q; L4M; A9S or A9D; T10S; S14K or S14A; S28T; V29I; S30D, S30H or S30T; S3 IT, S3 II, or S31R; I58V; A60D; D70E; S77R; E81D; Q90H; S91R; N93S, N93A, N93T, N93D, or N93Q; T97L; and combinations thereof.
[0016] In some embodiments, the one or more substitutions comprise a substitution at position 93, as compared to SEQ ID NO: 2. In some embodiments, the substitution at position 93 is selected from N93S, N93A, N93T, N93D, or N93Q.
[0017] In some embodiments, the one or more substitutions comprise a substitution at position 1, as compared to SEQ ID NO: 2. In some embodiments, the substitutions at position 1 is selected from EID or EIQ.
[0018] In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 5-99 and / or a light chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 100- 194. In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence of any one of SEQ ID NOs: 5-99 and / or a light chain having an amino acid sequence of any one of SEQ ID NOs: 100-194.
[0019] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions as compared to SEQ ID NO: 3 at positions selected from: 1, 2, 3, 5, 13, 28, 31, 33, 35, 37, 43, 49, 50, 52, 53, 54, 57, 58, 61, 62, 85, 88, 97, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, and 123. In some embodiments, the one or more substitutions are selected from: EIQ; V2Q; Q3S, Q3K, or Q3T; V5Q or V5L; Q13K; T28I; N31 S, N3 ID, or N31R; W33Y or W33A; N35S or N35T; V37I; K43Q; A49S; A50N, A50S, A50Y, or A50D; N52K; Q53E, Q53R, Q53T, Q53N, Q53K, or Q53S; D54N; E57N; K58T, K58E, or K58I; V61A; G62D; S85N; V88A; V97A or V97T; D99H, D99E, D99L, D99S, D99R, D99N, D99W, D99K, D99Y, D99Q, or D99A; Y100S, Y100R, Y100K, Y100T, Y100F, Y100A, Y100L, Y100W, Y100E, Y100G, or Y100V; YIOIH, Y101F, Y101N, Y101R, or Y101M; D102S, D102T, D102Y, D102W, D102R, D102L, D102E, D102F, D102A, or D102K; I103V, Il 03 S, I103T, 1103 Y, I103L, I103D, I103F, or I103R; LI 04V, LI 04 A, L104E, LI 041, L104Y, L104D, L104T, L104F, L104R, or L104M; T105N, T105D, T105S, T105Y, T105V, T105R, T105E, T105A, T105H, T105W, T105Q, T105I, or T105K; D106N, D106Y, D106E, D106A, D106S, D106T, D106W, D106G, or D106Q; Y107R, Y107L, Y107H, Y107D, Y107T, Y107Q, Y107K, Y107E, Y107A, Y107S, Y107N, Y107M, Y107F, Y107G, or Y107W; Y108L, Y108F, Y108H, Y108D, Y108W, Y108T, Y108G, Y108Q, or Y108M; I109V, I109L, I109Y, I109D, I109T, I109G, or I109Q; Hl ION, H110D, H110W, Hl 10R, Hl 10G, Hl 10Y, Hl 10L, Hl 10A, Hl 10T, or Hl 10Q; Y111H, Y11 II, Y11 IE, Y11 IL, Y111R, Y11 ID, Y111G, Y11 IQ, Y11 IV, Y111W, Y11 IM, Y11 IS, Y11 IF, Y11 IN, or Y11 IT; W1 12M, W112Y, W112S, W112D, W112Q, W112A, W112N, W112T, W112G, W112V, or W112L; Y113F, Y113D, Y113G, Y113I, Y113A, Y113V, Y113Q, Y113T, Y113L, or Y113R; Fl 14Y, Fl 14W, Fl 14V, Fl 14T, Fl 14G, Fl 14M, or Fl 14D; DI 15W, DI 15M, DI 15T, or DI 15V; LI 16Y or LI 16T; W117V or W117T; G118T; R119Q, R119A, R119S, or R119V; G120S; T121S; L122M, L122S, L122Q, L122T, or L122A; V123S; and combinations thereof.
[0020] In some embodiments, the one or more substitutions comprise a substitution at one or more positions selected from: 35, 50, 58, 88, 97, 99, 100, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, and 119, as compared to SEQ ID NO: 3. In some embodiments, the one or more substitutions comprise a substitution at one or more positions selected from: 35, 50, 88, 97, 100, 102, 103, 104, 105, 112, 114, and 119, as compared to SEQ ID NO: 3.
[0021] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions as compared to SEQ ID NO: 4 at positions selected from: 1, 2, 3, 4, 7, 9, 10, 12, 13, 24, 29, 30, 32, 33, 34, 44, 54, 59, 61, 78, 94, 97, 101, 104, and 105. In some embodiments, the one or more substitutions are selected from: EID; I2V; V3Q; L4M; S7T; G9A, G9S, or G9D; T10S or T10I; S12A; LBV; R24K; V29I; S30R; S32N, S32R, or S32D; Y33L, Y33N, Y33S, Y33R, Y33F, or Y33H; L34V or L34A; A44S; S54T; I59V; D61A; R78S; S94T; C97T, C97R, C97Y, C97V, C97S, C97I, C97Q, or C97W; Q101G; R104K; L105V; and combinations thereof.
[0022] In some embodiments, the one or more substitutions comprise a substitution at one or more positions selected from: 1, 2, 9, 10, 32, 33, 34, 44, 97, 104, 105, as compared to SEQ ID NO: 4. In some embodiments, the one or more substitutions comprise a substitution at one or more positions selected from: 2, 9, 33, 97, 104, and 105, as compared to SEQ ID NO: 4. In some embodiments, the one or more substitutions comprise a substitution at one or more positions selected from: 1, 9, 10, 32, 33, 34, 44, 97, and 104, as compared to SEQ ID NO: 4.
[0023] In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 195-290 and / or a light chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 291 -386. In some embodiments, the engineered antibodies comprises a heavy chain having an amino acid sequence of any one of SEQ ID NOs: 195-290 and / or a light chain having an amino acid sequence of any one of SEQ ID NOs: 291-38.
[0024] Also provided herein are compositions comprising an engineered antibody as described herein and nucleic acids encoding an engineered antibody as described herein.
[0025] Further provided are methods of treating or preventing a disease or disorder in a subject, comprising administering to the subject an effective amount of an engineered antibody, composition, or nucleic acid, as described herein.
[0026] Additionally provided are methods for generating engineered protein sequences as shown and described in FIG. 1. In some embodiments, the methods comprise a neural network having a series of stacked layers, each layer comprising a self-attention mechanism and feed forward neural network (FFNN) followed by a structural adapter. In some embodiments, the structural adapter incorporates structural and functional context. In some embodiments, the structural and function context comprises backbone interactions, protein-protein interactions, and interactions with non-protein molecules.
[0027] Other aspects and embodiments of the disclosure will be apparent in light of the following detailed description.
[0028] BRIEF DESCRIPTION OF FIGURES
[0029] FIGS. 1A-1E show the design of protein sequences for diverse backbones. FIG. 1A is a diagram of proseLM architecture. Structural adapter layers are placed after each layer of the pretrained language model. In each adapter layer, structural context from the causal encoder is combined with language model embeddings to condition sequence generation. Shown is SEQ ID NO: 389. FIG. IB is a graph of the perplexity of ProGen2 pre-trained language models and single-chain proseLM models on the CATH 4.2 test set collected by Ingraham et al. (Advances in neural information processing systems, 32, 2019). FIG. 1C is a graph of the perplexity of proseLM models on the clustered PDB dataset collected by Dauparas et al. (Science, 378(6615): 49-56, 2022). Perplexity is reported as the mean of cluster-averaged values. For all models, perplexity decreases as additional structural and functional context is provided. FIG. ID is a graph of perplexity with respect to total model parameters for proseLM models trained on the PDB with increasing levels of structural and functional context provided. FIG. IE is a graph of the native sequence recovery for fully designed sequences from proseLM models. Sequence recovery is reported as the median of cluster-averaged values. For all models, native sequence recovery increases when structural and functional context is provided.
[0030] FIGS. 2A-2D show modeling protein functional context and fitness. FIG. 2 A shows median native sequence recovery for residues near nucleic acids, small-molecule ligands, and ions, as indicated. FIG. 2B shows the change in native sequence recovery with and without structural and functional context as a function of distance from provided context. Smaller models show the largest increases in recovery near the provided context, while all models converge towards lower overall increases in recovery as distance increases. FIG. 2C shows the sequence recovery as a function of residue burial for protein complex targets. When designed with backbone only (single chain) or with protein context, recovery is largely determined by degree of burial. FIG. 2D shows the evaluation of fitness prediction, measured as normalized discounted cumulative gain, for proseLM models. Overall performance is averaged over landscape categories: stability, binding, activity, expression, and organismal fitness.
[0031] FIGS. 3A-3I show optimization of genome editors with proseLM. FIG. 3A is a schematic of exemplary methodology for design of optimized SpCas9 nuclease sequences. Sequences are generated using ProseLM-Base with conditioning on binary and catalytic states from PDB IDs 4ZT0 and 7Z4J, respectively. Sequences are generated with the PAM-interacting domain residues fixed to maintain compatibility with SpCas9 target sites. Shown are SEQ ID NOs: 387 and 388. FIG. 3B shows the editing efficiency of seven designed Cas9 variants across three target sites, with comparisons to parental SpCas9. Five of seven designs showed some activity across at least two guides. FIG. 3C shows the distribution of mutations across SpCas9 domains for two high-activity nuclease designs, with 102 and 59 total mutations. FIG. 3D shows the structure of adenosine base editor (PDB ID 6VPC), with deaminase active site highlighted. FIG. 3E shows the maximum A-to-G editing efficiency for deaminases with design focused on active site or non-active site residues. The parental deaminase activity is indicated by open circles, while active and inactive designs are indicated by blue and gray circles, respectively. The editing efficiencies of three deaminases designed through directed evolution are indicated by purple markers. FIG. 3F is a graph of the number of mutations from parental deaminase for experimentally tested active site designs. Fraction of designs with observable and improved are indicated by light and dark blue bars, respectively. FIG. 3G is a graph of the number of mutations from parental deaminase for experimentally tested non-active site designs. Fraction of designs with observable and improved are indicated by light and dark blue bars, respectively. FIG. 3H is a structural model of active site six mutations that resulted in highest A:G editing efficiency. FIG. 31 shows the positions of 19 mutations at non-active site positions that resulted in highest A:G editing efficiency.
[0032] FIGS. 4A-4H show design of therapeutic antibodies with proseLM. (a-d) All perplexity and sequence recovery values are reported as the mean and median, respectively, of cluster- averaged values. FIG. 4A is a graph of perplexity of proseLM models on the antibody dataset collected from SAbDab. For all models, perplexity decreases when antigen context is provided. FIG. 4B is a graph of perplexity with respect to total model parameters for proseLM models with and without antigen context provided. Blue lines and purple points represent performance of PDB-trained (proseLM [PDB]) and SAbDab-trained (proseLM-Ab) models, respectively. FIG. 4C shows native sequence recovery for fully designed heavy and light chain variable fragment sequences, with and without antigen context. FIG. 4D is a graph of sequence recovery for designed heavy and light chain variable fragments by structural region. FIG. 4E is a graph of the percentage of experimentally tested nivolumab variants that retained PD-1 binding for each design strategy. FIG. 4F is a graph of binding affinity (-log(KD), higher better) for nivolumab variants that retained PD-1 binding. Horizontal line indicates the binding affinity of parental nivolumab antibody. FIG. 4G is a structural model of mutations for highest affinity nivolumab variant from CDR-directed optimization strategy. The locations of the two mutations in the Fv are shown as red spheres, with potential novel interactions highlighted. FIG. 4H shows the position of mutations for highly diversified secukinumab variants, with 31 positions mutated across the heavy and light chains.
[0033] FIGS. 5A-5D show visualization of protein and atomic graph components. FIG. 5A shows each protein residue represented as a rigid body frame with the Caatom at the origin and the frame's orientation defined by the positions of the N and C atoms. FIG. 5B shows nonprotein atoms represented as rigid body frames with the primary atom at the origin and the two closest atoms defining the frame's orientation. FIG. 5C is a schematic showing protein and atomic graphs are processed by alternating MPNN and IPMP layers. MPNN layers operate on the graph nodes and edges, while IPMP layers additionally incorporate frame-based geometric features. FIG. 5D shows the edge connectivity for the protein-only residue graph (left), atomic- only graph (upper right), and protein-atomic graph (lower right). Black lines illustrate the connectivity between nodes in the graph.
[0034] FIGS. 6A-6D show performance of proseLM models on CATH 4.2 benchmark. Evaluation of proseLM model performance on the CATH 4.2 test set curated by Ingraham et al. (Advances in neural information processing systems, 32, 2019). Aggregate sequence recovery values are reported as the median all proteins (or subsets when appropriate). FIG. 6A is a graph of the recovery of native sequence residues for proseLM models. FIG. 6B is a graph of the recovery of native sequence residues for proseLM models binned by residue burial, calculated as the average Cp distance to the nearest eight neighbors (lower is more buried). All models achieve high rates of native sequence recovery among buried residues and reduced recovery at less- buried surface positions. FIG. 6C is a plot of the recovery of native sequence residues for proseLM models, binned by sequence length. Longer sequences exhibit higher rates of native sequence recovery, with larger proseLM models having particularly high recovery for large proteins. FIG. 6D shows the relationship between perplexity and sequence recovery for proseLM models. Spearman correlation coefficients are reported for each model. Larger models show less correlation between perplexity and sequence recovery.
[0035] FIGS. 7A-7D show impact of coordinate noising on single-sequence structure prediction. Evaluation of proseLM model performance on the CATH 4.2 test set with and without Gaussian noise added to coordinates. Structure prediction accuracy and confidence are for single-sequence predictions using AlphaFold2. FIG. 7A is a graph of the fraction of designed sequences that are predicted to successfully recapitulate the input structure (1DDT > 90). Models trained with 0.1 A of Gaussian noise added to coordinates achieve higher structure prediction success. Larger models approach the lower level of structure prediction success of the native sequences (Horizontal line) FIG. 7B is a graph of the structure prediction success rates for noised models relative to unnoised for several 1DDT thresholds. Larger models show less sensitivity to coordinate noising (lower relative success rates). FIG. 7C is a graph of the fraction of designed sequences that yield highly confident structure predictions (pLDDT > 90). Models trained with coordinate noising yield high-confidence structures more frequently. Larger models approach the lower level of confident structure prediction rates for the native sequences (Horizontal line). FIG. 7D is a graph of the confidence in structure prediction for noised models relative to unnoised for several pLDDT thresholds. Larger models show less sensitivity to coordinate noising (lower relative rates of confident predictions).
[0036] FIG. 8 shows the perplexity of masked structural spans. Evaluation of proseLM model perplexity on contiguous structurally masked spans of twenty residues. Perplexity for structurally masked residues (dashed lines) increases significantly over the same residues without masking (solid lines). Compared to the respective ProGen2 models, proseLM models achieve lower perplexity for structurally masked residues, indicating that the surrounding context is effectively incorporated when predicting for masked positions.
[0037] FIG. 9 shows the fitness prediction measured by normalized discounted cumulative gain. Comparison of proseLM and recent structure-conditioned sequence design models (ESM-IF1 and ProteinMPNN) for prediction of mutational fitness landscapes. Performance is reported as normalized discounted cumulative gain for the top ten percent of samples according to experimental fitness (NDCG10%). The metric is relevant for protein design tasks, where it is important to accurately prioritize the highest fitness sequences for experimental characterization.
[0038] FIG. 10 shows the fitness prediction measured by Spearman's rank correlation coefficient. Comparison of proseLM and recent structure-conditioned sequence design models (ESM-IF1 and ProteinMPNN) for prediction of mutational fitness landscapes. Performance is reported as Spearman's rank correlation coefficients. This metric indicates how well model scores rank sequences according to fitness across the entire dataset.
[0039] FIG. 11 shows model scores for SpCas9 variants. Mutational hamming distance from SpCas9 and proseLM model perplexities for designed SpCas9 variants. Mutation distances and model scores are shown for generations directly from proseLM (light gray), generations from proseLM augmented with position-specific residue propensities from deep mutational scans and multiple-sequence alignments (dark gray), and the subset selected for experimental validation (blue). Model scores for SpCas9 are indicated by vertical black lines.
[0040] FIGS. 12A-12D show the model scores for adenine base editor variants. Mutational hamming distance from parental deaminase and proseLM model perplexities for designed deaminase variants. Mutation distances and model scores for the complete set of generations and the subset selected for experimental validation are shown in gray and blue, respectively. Model scores for the parental deaminase are indicated by vertical black lines. FIG. 12A shows the positions of all mutated residues among selected active site variants (spheres). FIG. 12B shows the mutation and score distributions for active site variants. FIG. 12C shows the positions of all mutated residues among selected non-active site variants (spheres). FIG. 12D shows the mutation and score distributions for non-active site variants.
[0041] FIG. 13 shows the model scores for nivolumab CDR variants. Mutational hamming distance from parental nivolumab and proseLM model perplexities for designed nivolumab CDR loop variants. Mutation distances and model scores for the complete set of generations and the subset selected for experimental validation are shown in gray and blue, respectively. Model scores for nivolumab are indicated by vertical black lines. Y-axis is shown on a log scale.
[0042] FIG. 14 shows model scores for nivolumab framework variants. Mutational hamming distance from parental nivolumab and proseLM model perplexities for designed nivolumab framework variants. Mutation distances and model scores for the complete set of generations and the subset selected for experimental validation are shown in gray and blue, respectively. Model scores for nivolumab are indicated by vertical black lines. Y-axis is shown on a log scale.
[0043] FIG. 15 shows model scores for diversified secukinumab variants. Mutational hamming distance from parental secukinumab and proseLM model perplexities for designed secukinumab variants. Mutation distances and model scores for the complete set of generations and the subset selected for experimental validation are shown in gray and blue, respectively. Model scores for secukinumab are indicated by vertical black lines. Y-axis is shown on a log scale.
[0044] DETAILED DESCRIPTION
[0045] Protein language models present an alternative means of modeling protein sequencefunction relationships. These models learn directly from sequences through self-supervised training objectives, such as masked residue prediction and next-residue prediction. With increasing numbers of parameters, protein language models have been shown to capture properties including structure and function. For design tasks, protein language models are typically steered towards a particular functional family through fine-tuning on curated sequence datasets. However, this dependence on natural examples for fine-tuning restricts the scope of design with protein language models to known protein families and functions. Further, it does not offer a straightforward means of incorporating atomistic constraints on the design space, instead relying on the model to implicitly learn these constraints from the fine-tuning data.
[0046] The present disclosure provides automated systems and methods for protein sequence design based on adaptation of protein language models to incorporate structural and functional context, including the backbone-of-interest and nearby molecule context (proteins, nucleic acids, ligands, ions, etc.), herein referred to as ProseLM. ProseLM effectively incorporates this context while benefiting from the scaling trends of the underlying language models, enabling high rates of native sequence recovery. The addition of non-protein context, nucleic acids, ligands, and ions improves sequence design across model scales.
[0047] ProseLM functionally improved genome editors and therapeutic antibodies and can be used to identify functionally relevant mutations in complex protein-nucleic acid systems. ProseLM improved binding affinity of highly optimized therapeutic antibodies through two distinct strategies (CDR-directed and framework-directed optimization), as well as to diversify the binding region of a structurally complex antibody, retained activity, and in some cases improved on-target editing efficiency of nucleases, and redesigned the active site of a deaminase yielding variants with significantly increased activity.
[0048] Section headings as used in this section and the entire disclosure herein are merely for organizational purposes and are not intended to be limiting.
[0049] Definitions
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In case of conflict, the present document, including definitions, will control. Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in practice or testing of the present disclosure. All publications, patent applications, patents and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and not intended to be limiting.
[0051] The terms “comprise(s),” “include(s),” “having,” “has,” “can,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. As used herein, comprising a certain sequence or a certain SEQ ID NO usually implies that at least one copy of said sequence is present in recited peptide or polynucleotide. However, two or more copies are also contemplated. The singular forms “a,” “and,” and “the” include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments “comprising,” “consisting of,” and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.
[0052] As used herein, terms and phrases such as “having,” “may have,” “include,” or “may include” a feature (such as a number, function, operation, or component, such as a component) indicate the presence of that feature, and do not preclude the presence of other features. Further, as used herein, the phrase “a or B,” “at least one of a and / or B,” or “one or more of a and / or B” may include all possible combinations of a and B. For example, “a or B,” “at least one of a and B,” and “at least one of a or B” may indicate all of the following: (1) comprises at least one A, (2) comprises at least one B, or (3) comprises at least one A and at least one B. Furthermore, as used herein, the terms “first” and “second” may modify various components without regard to importance, and do not limit the components. These terms are only used to distinguish one component from another. For example, the first user device and the second user device may indicate user devices that are different from each other regardless of the order or importance of the devices. A first component may be termed a second component, and vice-versa, without departing from the scope of the present disclosure.
[0053] It will be understood that when an element (such as a first element) is referred to as being (operatively or communicatively) “coupled / coupled” or “connected / connected” to another element (such as a second element), it can be directly coupled or connected / coupled or connected to the other element (such as the second element) or via a third element. Conversely, it will be understood that when an element (such as a first element) is referred to as being “directly coupled” / ” directly coupled to” or “directly connected” / ” directly connected” to another element (such as a second element), there is no other element (such as a third element) intervening between the element and the other element.
[0054] As used herein, the phrase “configured (or set) to” may be used interchangeably with the phrases “adapted to”, “having .. . capability”, “designed to”, “adapted to”, “made to”, or “capable”, as the case may be. The phrase “configured (or set) to” does not substantially mean “specially designed in hardware”. Rather, the phrase “configured to” may indicate that a device is capable of performing an operation with another device or component. For example, the phrase “a processor configured (or arranged) to perform A, B and C” may refer to a general- purpose processor (such as a CPU or an application processor) or a special-purpose processor (such as an embedded processor) that may perform operations by executing one or more software programs stored in a memory device.
[0055] The various functions described below may be implemented or supported by one or more computer programs, each formed from computer-readable program code and embodied in a computer-readable medium. The terms “application” and “program” refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in suitable computer readable program code.
[0056] As used herein, the term “computer” refers to a machine, apparatus, or device that is capable of accepting and performing logic operations from software code. The term “application”, “software”, “software code” or “computer software” refers to any set of instructions operable to cause a computer to perform an operation. Software code may be operated on by a “rules engine” or “processor.” Thus, in some embodiments, the methods and systems of the present invention may be performed by a computer or computing device having a processor based on instructions received by computer applications and software.
[0057] The term “computer readable medium” as used herein refers to any medium that participates in providing instructions to the processor for execution. A computer readable medium may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical, magnetic disks, and magneto-optical disks, such as the hard disk or the removable media drive. Volatile media includes dynamic memory, such as the main memory. Transmission media includes coaxial cables, copper wire and fiber optics, including the wires that make up the bus. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. Non-transitory computer readable media includes all computer readable media, with the sole exception being a transitory, propagating signal per se.
[0058] As used herein the term “data network” or “network” shall mean an infrastructure capable of connecting two or more computers such as client devices either using wires or wirelessly allowing them to transmit and receive data. Non-limiting examples of data networks may include the Internet or wireless networks which may include Wi-Fi and cellular networks. For example, a network may include a local area network (LAN), a wide area network (WAN) (e.g., the Internet), a mobile relay network, a metropolitan area network (MAN), an ad hoc network, a telephone network (e g., a Public Switched Telephone Network (PSTN)), a cellular network, a Zigby network, or a voice-over-IP (VoIP) network.
[0059] As used herein, the term “database” shall generally mean a digital collection of data or information. For the purposes of the present disclosure, a database may be stored on a remote server and accessed by a client device (e.g., through the Internet) or alternatively in some embodiments the database may be stored on the client device or remote computer itself.
[0060] For the recitation of numeric ranges herein, each intervening number there between with the same degree of precision is explicitly contemplated. For example, for the range of 6-9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0-7.0, the number 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly contemplated.
[0061] The terms and phrases used herein are used only to describe some embodiments of the present disclosure and do not limit the scope of other embodiments of the present disclosure. It is to be understood that the singular includes plural referents unless the context clearly dictates otherwise. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as commonly understood by one of ordinary skill in the art to which embodiments of the present disclosure belong. It will be further understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some instances, the terms and phrases defined herein may be construed to exclude embodiments of the disclosure.
[0062] As used herein, “nucleic acid” or “nucleic acid sequence” refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (See Albert L. Lehninger, Principles of Biochemistry, 793-800 (Worth Pub. 1982)). The present technology contemplates any deoxyribonucleotide, ribonucleotide, or nucleoprotein component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogenous or homogenous in composition and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states. In some embodiments, a nucleic acid or nucleic acid sequence comprises other kinds of nucleic acid structures such as, for instance, a DNA / RNA helix, peptide nucleic acid (PNA), morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14): 4503-4510 (2002) and U.S. Pat. No. 5,034,506), locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97: 5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122: 8595-8602 (2000)), and / or a ribozyme. Hence, the term “nucleic acid” or “nucleic acid sequence” may also encompass a chain comprising non-natural nucleotides, modified nucleotides, and / or non- nucleotide building blocks that can exhibit the same function as natural nucleotides (e.g., “nucleotide analogs”); further, the term “nucleic acid sequence” as used herein refers to an oligonucleotide, nucleotide or polynucleotide, and fragments or portions thereof, and to DNA or RNA of genomic or synthetic origin, which may be single or double-stranded, and represent the sense or antisense strand. The terms “nucleic acid,” “polynucleotide,” “nucleotide sequence,” and “oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.
[0063] As used herein, “peptide,” “polypeptide,” or “protein” refer to a sequence of two or more amino acids linked by peptide bonds. The polypeptide can be natural, synthetic, or a modification or combination of natural and synthetic. The peptide or polypeptide may be modified by the addition of sugars, lipids or other moieties not included in the amino acid chain. The terms “polypeptide,” “oligopeptide,” and “peptide” are used interchangeably herein. The peptide(s) may be produced by recombinant genetic technology or chemical synthesis. The peptide(s) may be isolated and purified by any number of standard methods including, but not limited to, differential solubility (e.g., precipitation), centrifugation, chromatography (e.g., affinity, ion exchange, and size exclusion), or by any other standard techniques known in the art.
[0064] The term “amino acid” or “any amino acid” as used here refers to any and all amino acids, including naturally occurring amino acids (e.g., a-amino acids), unnatural amino acids, modified amino acids, and non-natural amino acids. It includes both D- and L-amino acids. Natural amino acids include those found in nature, such as, e.g., the 23 amino acids that combine into peptide chains to form the building-blocks of a vast array of proteins. These are primarily L stereoisomers, although a few D-amino acids occur in bacterial envelopes and some antibiotics. For the most part, the names of naturally occurring and non-naturally occurring aminoacyl residues used herein follow the naming conventions suggested by the IUPAC Commission on the Nomenclature of Organic Chemistry and the IUPAC-IUB Commission on Biochemical Nomenclature as set out in “Nomenclature of a-Amino Acids (Recommendations, 1974)” Biochemistry, 14(2), (1975). To the extent that the names and abbreviations of amino acids and aminoacyl residues employed in this specification and appended claims differ from those suggestions, they will be made clear to the reader. Throughout the present specification, unless naturally occurring amino acids are referred to by their full name (e.g., alanine, arginine, etc.), they are designated by their conventional three-letter or single-letter abbreviations (e.g., Ala or A for alanine, Arg or R for arginine, etc.). The term “L-amino acid,” as used herein, refers to the “L” isomeric form of a peptide, and conversely the term “D-amino acid” refers to the “D” isomeric form of a peptide (e.g., Dphe, (D)Phe, D-Phe, orDF for the D isomeric form of Phenylalanine). Amino acid residues in the D isomeric form can be substituted for any L-amino acid residue, as long as the desired function is retained by the peptide.
[0065] Nucleic acid or amino acid sequence “identity,” as described herein, can be determined by comparing a nucleic acid or amino acid sequence of interest to a reference nucleic acid or amino acid sequence. A number of mathematical algorithms for obtaining the optimal alignment and calculating identity between two or more sequences are known and incorporated into a number of available software programs. Examples of such programs include CLUSTAL-W, T- Coffee, and ALIGN (for alignment of nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST 2.1, BL2SEQ, and later versions thereof) and FASTA programs (e.g., FASTA3x, FAS™, and SSEARCH for sequence alignment and sequence similarity searches).
[0066] The terms “non-naturally occurring,” “engineered,” and “synthetic” are used interchangeably and indicate the involvement of the hand of man. The terms, when referring to nucleic acid molecules or polypeptides mean that the nucleic acid molecule or the polypeptide is at least substantially free from at least one other component with which it is naturally associated in nature and as found in nature, and / or the nucleic acid molecule or the polypeptide is associated with at least one other component with which it is not naturally associated in nature and / or that there is one or more changes in nucleic acid or amino acid sequence as compared with such sequence as it is found in nature and / or that the nucleic acid or polypeptide sequence was generated de novo, e.g., not based on or derived from any naturally occurring sequence.
[0067] “Antibody” and “antibodies” as used herein refers to monoclonal antibodies, monospecific antibodies (e.g., which can either be monoclonal, or may also be produced by other means than producing them from a common germ cell), multi-specific antibodies, human antibodies, humanized antibodies (fully or partially humanized), animal antibodies such as, but not limited to, a bird (for example, a duck or a goose), a shark, a whale, and a mammal, including a non-primate (for example, a cow, a pig, a camel, a llama, a horse, a goat, a rabbit, a sheep, a hamster, a guinea pig, a cat, a dog, a rat, a mouse, etc.) or a non-human primate (for example, a monkey, a chimpanzee, etc.), recombinant antibodies, chimeric antibodies, singlechain Fvs (“scFv”), single chain antibodies, single domain antibodies sdAbs that are naturally occurring, e.g., as in cartilaginous fishes and camelid, or which are synthetic, e.g., nanobodies, VHH, or other domain structure), and functionally active epitope-binding fragments of any of the above, such as Fab fragments, F(ab’) fragments, F(ab’)2 fragments, disulfide-linked Fvs (“sdFv”), and anti -idiotypic (“anti-Id”) antibodies, dual-domain antibodies, dual variable domain (DVD) or triple variable domain (TVD) antibodies (dual-variable domain immunoglobulins and methods for making them are described in Wu, C., et al., Nature Biotechnology, 25(11 ): 1290- 1297 (2007) and PCT International Application WO 2001 / 058956, the contents of each of which are herein incorporated by reference), or domain antibodies (dAbs) (e.g., such as described in Holt et al., Trends in Biotechnology 21 :484-490 (2014)), and including single domain antibodies sdAbs that are naturally occurring, e.g., as in cartilaginous fishes and camelid, or which are synthetic, e.g., nanobodies, VHH, or other domain structure), and functionally active epitopebinding fragments of any of the above. In particular, antibodies include immunoglobulin molecules and immunologically active fragments of immunoglobulin molecules, namely, molecules that contain an analyte-binding site. Immunoglobulin molecules can be of any type (for example, IgG, IgE, IgM, IgD, IgA, and IgY), class (for example, IgGl, IgG2, IgG3, IgG4, IgAl, and IgA2), or subclass.
[0068] “Antibody fragment” as used herein refers to a portion of an intact antibody that retain the ability to specifically bind to an antigen (e.g., comprises the antigen-binding site or variable region). Any antigen-binding fragment of the antibody described herein is within the scope of the present disclosure. The antibody may not include the constant heavy chain domains (e.g., CH2, CH3, or CH4, depending on the antibody isotype) of the Fc region of the intact antibody. Examples of antibody fragments include, but are not limited to, Fab fragments, Fab’ fragments, Fab’-SH fragments, F(ab’)2 fragments, Fd fragments, Fv fragments, diabodies, single-chain Fv (scFv) molecules, single-chain polypeptides containing only one light chain variable domain, single-chain polypeptides containing the three CDRs of the light-chain variable domain, singlechain polypeptides containing only one heavy chain variable region, and single-chain polypeptides containing the three CDRs of the heavy chain variable region.
[0069] Typically, an immunoglobulin or antibody is a protein that comprises at least one complementarity determining region (CDR). The CDRs form the “hypervariable region” of an antibody, which is responsible for antigen binding. “CDR” is used herein to refer to the “complementarity determining region” within an antibody variable sequence. There are three CDRs in each of the variable regions of the heavy chain and the light chain. Proceeding from the N-terminus of a heavy or light chain, these regions are denoted “CDR1,” “CDR2,” and “CDR3,” for each of the variable regions. The term “CDR set” as used herein refers to a group of three CDRs that occur in a single variable region that binds the antigen. An antigen-binding site, therefore, may include six CDRs, comprising the CDR set from each of a heavy and a light chain variable region. A polypeptide comprising a single CDR, (e.g., a CDR1, CDR2, or CDR3) may be referred to as a “molecular recognition unit.” Crystallographic analyses of antigen-antibody complexes have demonstrated that the amino acid residues of CDRs form extensive contact with bound antigen, wherein the most extensive antigen contact is with the heavy chain CDR3. Thus, the molecular recognition units may be primarily responsible for the specificity of an antigenbinding site. In general, the CDR residues are directly and most substantially involved in influencing antigen binding.
[0070] A whole antibody typically consists of four polypeptides: two identical copies of a heavy (H) chain polypeptide and two identical copies of a light (L) chain polypeptide. Each of the heavy chains contains one N-terminal variable (VH) region and three C-terminal constant (CHI, CH2, and CH3) regions, and each light chain contains one N-terminal variable (VL) region and one C-terminal constant (CL) region. The light chains of antibodies can be assigned to one of two distinct types, either kappa (K) or lambda (X), based upon the amino acid sequences of their constant domains. In a typical antibody, each light chain is linked to a heavy chain by disulfide bonds, and the two heavy chains are linked to each other by disulfide bonds. The light chain variable region is aligned with the variable region of the heavy chain, and the light chain constant region is aligned with the first constant region of the heavy chain. The remaining constant regions of the heavy chains are aligned with each other. The variable regions of each pair of light and heavy chains form the antigen binding site of an antibody. The VH and VL regions have the same general structure, with each region comprising four framework (FW or FR) regions. The term “framework region,” as used herein, refers to the relatively conserved amino acid sequences within the variable region which are located between the CDRs. There are four framework regions in each variable domain, which are designated FR1, FR2, FR3, and FR4. The framework regions form the 0 sheets that provide the structural framework of the variable region.
[0071] “CDR” is used herein to refer to the “complementarity determining region” within an antibody variable sequence. There are three CDRs in each of the variable regions of the heavy chain and the light chain. Proceeding from the N-terminus of a heavy or light chain, these regions are denoted “CDR1,” “CDR2,” and “CDR3,” for each of the variable regions. The term “CDR set” as used herein refers to a group of three CDRs that occur in a single variable region that binds the antigen. An antigen-binding site, therefore, may include six CDRs, comprising the CDR set from each of a heavy and a light chain variable region. A polypeptide comprising a single CDR, (e.g., a CDR1, CDR2, or CDR3) may be referred to as a “molecular recognition unit.” Crystallographic analyses of antigen-antibody complexes have demonstrated that the amino acid residues of CDRs form extensive contact with bound antigen, wherein the most extensive antigen contact is with the heavy chain CDR3. Thus, the molecular recognition units may be primarily responsible for the specificity of an antigen-binding site. In general, the CDR residues are directly and most substantially involved in influencing antigen binding.
[0072] The exact boundaries of these CDRs have been defined differently according to different systems. The system described by Kabat (Kabat et al., Sequences of Proteins of Immunological Interest (National Institutes of Health, Bethesda, Md. (1987) and (1991)) not only provides an unambiguous residue numbering system applicable to any variable region of an antibody, but also provides precise residue boundaries defining the three CDRs. These CDRs may be referred to as “Kabat CDRs”. Chothia and coworkers (Chothia and Lesk, J. Mol. Biol, 196: 901-917 (1987); and Chothia et al., Nature, 342: 877-883 (1989)) found that certain sub-portions within Kabat CDRs adopt nearly identical peptide backbone conformations, despite having great diversity at the level of amino acid sequence. These sub-portions were designated as “LI,” “L2,” and “L3,” or “Hl,” “H2,” and “H3,” where the “L” and the “H” designate the light chain and the heavy chain regions, respectively. These regions may be referred to as “Chothia CDRs,” which have boundaries that overlap with Kabat CDRs. Other boundaries defining CDRs overlapping with the Kabat CDRs have been described by Padlan, FASEB J., 9: 133-139 (1995), and MacCallum, J. Mol. Biol., 262(5): 732-745 (1996). Still other CDR boundary definitions may not strictly follow one of the herein systems, but will nonetheless overlap with the Kabat CDRs, although they may be shortened or lengthened in light of prediction or experimental findings that particular residues or groups of residues or even entire CDRs do not significantly impact antigen binding. The methods used herein may utilize CDRs defined according to any of these systems, although certain aspects use Kabat- or Chothia-defmed CDRs.
[0073] “Humanized” forms of non-human (e.g., rodent) antibodies are chimeric antibodies that contain minimal sequence derived from the non-human antibody. For the most part, humanized antibodies are human immunoglobulins (recipient antibody) in which residues from a hypervariable region of the recipient are replaced by residues from a hypervariable region of a non-human species (donor antibody) such as mouse, rat, rabbit, or non-human primate having the desired antibody specificity, affinity, and capability. In some instances, framework region (FR) residues of the human immunoglobulin are replaced by corresponding non-human residues. Furthermore, humanized antibodies can comprise residues that are not found in the recipient antibody or in the donor antibody. These modifications are made to further refine antibody performance. In general, the humanized antibody will comprise substantially all of at least one, and typically two, variable domains, in which all or substantially all of the hypervariable loops correspond to those of a nonhuman immunoglobulin and all or substantially all of the FRs are those of a human immunoglobulin sequence. The humanized antibody optionally also will comprise at least a portion of an immunoglobulin constant region (Fc), typically that of a human immunoglobulin. For further details, see Jones et al., Nature 321 :522-525 (1986); Riechmann et al., Nature 332:323-329 (1988); and Presta, Curr. Op. Struct. Biol. 2:593-596 (1992).
[0074] As used herein, when an antibody or other entity (e.g., antigen binding domain) “specifically recognizes” or “specifically binds” an antigen or epitope, it preferentially recognizes the antigen in a complex mixture of proteins and / or macromolecules, and binds the antigen or epitope with affinity which is substantially higher than to other entities not displaying the antigen or epitope. In this regard, “affinity which is substantially higher” means affinity that is high enough to enable detection of an antigen or epitope which is distinguished from entities using a desired assay or measurement apparatus. Typically, it means binding affinity having a binding constant (Ka) of at least 107M1(e.g., >107M'1, >108M’1, >109M’1, >1O10M’1, >10uM’ \ >1012M’1, >1013M’1, etc.). In certain such embodiments, an antibody is capable of binding different antigens so long as the different antigens comprise that particular epitope. In certain instances, for example, homologous proteins from different species may comprise the same epitope.
[0075] “Affinity” refers to the strength of the sum total of noncovalent interactions between a single binding site of a molecule (e.g., an antibody) and its binding partner (e.g., an antigen). Unless indicated otherwise, as used herein, “binding affinity” refers to intrinsic binding affinity which reflects a 1 : 1 interaction between members of a binding pair (e.g., antibody and antigen). The affinity of a molecule X for its partner Y can generally be represented by the dissociation constant (KD or KD). Affinity can be measured by common methods known in the art, including those described herein. Specific illustrative and exemplary embodiments for measuring binding affinity are described in the following.
[0076] The term “monoclonal antibody,” as used herein, refers to an antibody produced by a single clone of B lymphocytes that is directed against a single epitope on an antigen. Monoclonal antibodies typically are produced using hybridoma technology. Monoclonal antibodies may also be produced using recombinant DNA methods, isolated from phage display antibody libraries, or produced from transgenic mice carrying a fully human immunoglobulin system. In contrast, “polyclonal” antibodies are antibodies that are secreted by different B cell lineages within an animal. Polyclonal antibodies are a collection of immunoglobulin molecules that recognize multiple epitopes on the same antigen.
[0077] The term “monospecific” antibody as used herein denotes an antibody that has one or more binding sites each of which bind to the same epitope of the same antigen.
[0078] The term “bispecific” antibody as used herein denotes an antibody that has at least two binding sites each of which bind to different epitopes of the same antigen or a different antigen.
[0079] The term “multi-specific” antibody as used herein denotes an antibody that has binding specificities for at least two different sites.
[0080] A “parent sequence” as used herein means a sequence comprising a region or residue that is unmodified in relationship to the position being modified in a variant or engineered sequence.
[0081] The terms “immunogen” and “antigen” are used interchangeably herein and refer to any molecule, compound, or substance that induces an immune response in an animal (e.g., a mammal). An “immune response” can entail, for example, antibody production and / or the activation of immune effector cells. An antigen in the context of the disclosure can comprise any subunit, fragment, or epitope of any proteinaceous or non-proteinaceous (e.g., carbohydrate or lipid) molecule that provokes an immune response in a mammal. By “epitope” is meant a sequence of an antigen that is recognized by an antibody or an antigen receptor. Epitopes also are referred to in the art as “antigenic determinants.” In certain embodiments, an epitope is a region of an antigen that is specifically bound by an antibody. In certain embodiments, an epitope may include chemically active surface groupings of molecules such as amino acids, sugar side chains, phosphoryl, or sulfonyl groups. In certain embodiments, an epitope may have specific three- dimensional structural characteristics (e.g., a “conformational” epitope) and / or specific charge characteristics. The antigen can be a protein or peptide of viral, bacterial, parasitic, fungal, protozoan, prion, cellular, or extracellular origin, which provokes an immune response in a mammal, preferably leading to protective immunity.
[0082] As used herein, the terms “providing,” “administering,” and “introducing” are used interchangeably herein and refer to the placement into a subject by a method or route which results in at least partial localization to a desired site.
[0083] A “subject” or “patient” may be human or non-human and may include, for example, animal strains or species used as “model systems” for research purposes, such a mouse model as described herein. Likewise, patient may include either adults or juveniles (e.g., children). Moreover, patient may mean any living organism, preferably a mammal (e.g., human or non- human) that may benefit from the administration of compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the Mammalian class: humans, non- human primates such as chimpanzees, and other apes and monkey species; farm animals such as cattle, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice and guinea pigs, and the like. Examples of nonmammals include, but are not limited to, birds, fish, and the like. In one embodiment, the mammal is a human.
[0084] As used herein, “treat,” “treating,” and the like means a slowing, stopping, or reversing of progression of a disease or disorder when provided a bispecific binding molecule or composition described herein to an appropriate subject. The term also includes a reversing of the progression of such a disease or disorder to a point of eliminating or greatly reducing the disease. As such, “treating” means an application or administration of the bispecific binding molecule or compositions described herein to a subject, where the subject has a disease or a symptom of a disease, where the purpose is to cure, heal, alleviate, relieve, alter, remedy, ameliorate, improve, or affect the disease or symptoms of the disease.
[0085] Definitions for other specific words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior, as well as future uses of such defined words and phrases.
[0086] None of the description in this application should be read as implying that any particular element, step, or function is an essential element which must be included in the claim scope. The scope of patented subject matter is defined only by the claims. Any other term used in the claims, including, but not limited to, “mechanism,” “module,” “device,” “unit,” “assembly,” “element,” “member,” “apparatus,” “machine,” “system,” “processor,” or “controller,” is understood by the applicants to refer to structures known to those of ordinary skill in the relevant art.
[0087] Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in practice or testing of the present disclosure. All publications, patent applications, patents and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and not intended to be limiting.
[0088] Antibodies
[0089] Disclosed herein are engineered antibodies, compositions comprising thereof, and methods for using the engineered antibodies. In some embodiments, the engineered antibodies comprise one or more substitutions, deletions or additions, as compared to a known parent antibody sequence (e.g., heavy or light chain sequence). The one or more substitutions, deletions or additions in the engineered antibodies may modulate of binding affinity, stability, immunogenicity, conformational diversity or flexibility, or a combination thereof
[0090] In some embodiments, the engineered antibodies comprise one or more substitutions in a framework region. In some embodiments, the one or more substitutions are in residues within 10 A (e.g., within 10 A, within 9 A, within 8 A, within 7 A, within 6 A, within 6 A, within 4 A, within 3 A, within 2 A, within 1 A, or less) of the antigen, e.g., when the antigen is bound. In some embodiments, the one or more substitutions stabilize the framework region, or the antibody as a whole. While the framework region does not directly interact with antigen, substitutions in this region may modulate the binding affinity and / or specificity of the CDRs with the target antigen.
[0091] In some embodiments, the engineered antibodies comprise one or more substitutions in one or more of the CDRs. For example, the CDRs are shown in bold in SEQ ID NOs: 1-4 in Tables 1 and 2 below: SEQ ID NO: 1 (heavy chain) CDR1 comprises positions 31 to 35, CDR2 comprises positions 50 to 66, CDR3 comprises positions 99 to 102; SEQ ID NO: 2 (light chain) CDR1 comprises positions 24 to 34, CDR2 comprises positions 50 to 56, CDR3 comprises positions 89 to 97; SEQ ID NO: 3 (heavy chain) CDR1 comprises positions 26 to 35, CDR2 comprises positions 50 to 66, CDR3 comprises positions 96 to 118; and SEQ ID NO: 4 (light chain) CDR1 comprises positions 24 to 35, CDR2 comprises positions 51 to 57, CDR3 comprises positions 90 to 98.
[0092] Thus, engineered antibodies comprising one or more substitutions in any of the above listed or indicated CDRs are also described herein. For example, as described below, the engineered antibody may comprise a mutation at position 32 in SEQ ID NO: 1, equivalent to heavy chain CDR1. Thus, provided is an engineered antibody having a heavy chain comprising a substitution of S32 to S32Y, S32F, S32H, S32K, S32R, S32N, or S32L in heavy chain CDR1, with CDR2 and CDR3 as provided in Table 1 or individually with one or more substitutions. Also, as described below, the engineered antibody may comprise a mutation at position 93 in SEQ ID NO: 2, equivalent to light chain CDR3. Thus, provided is an engineered antibody having a heavy chain comprising a substitution of S32 in heavy chain CDR1 and a light chain comprising a substitution of N93 in light chain CDR3.
[0093] The engineered antibodies may comprise one or more substitutions in a framework region and one or more substitutions in one or more of the CDRs.
[0094] In some embodiments, the antibodies target PD-1. In some embodiments, the engineered antibodies comprise a heavy chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 1 and / or a light chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 2.
[0095] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions to SEQ ID NO: 1. In some embodiments, the one or more substitutions are at positions selected from: 3, 5, 11, 13, 16, 19, 21, 23, 27, 30, 32, 37, 43, 52, 58, 59, 80, 88, 89, 97, 99, and combinations thereof, in reference to SEQ ID NO: 1 . In some embodiments, the one or more substitutions are selected from: Q3K; V5Q; VI IL; Q13K; R16E, R16T, or R16G; R19K; D21S; K23A, K23T, K23R, or K23V; I27V, I27L, I27F, or I27M; S30N; S32Y, S32F, S32H, S32K, S32R, S32N, or S32L; V37I; K43Q; W52S; R58K, R58T, or R58E; Y59F, Y59N, or Y59H; F80Y or F80H; A88S or A88T; E89D; A97T; N99T, N99E, N99S, N99H, N99D, or N99Y, and combinations thereof, in reference to SEQ ID NO: 1.
[0096] In some embodiments, the engineered antibodies comprise a heavy chain having a substitution at position 32, in reference to SEQ ID NO: 1. In some embodiments, the engineered antibodies comprise a heavy chain having a S32Y, S32F, S32H, S32K, S32R, S32N, or S32L substitution.
[0097] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions to SEQ ID NO: 1 at positions selected from: 11, 21, 23, 32, 58, 80, and combinations thereof. In some embodiments, the one or more substitutions are selected from: VI IL; D21S; K23A, K23T, K23R, or K23V; S32Y, S32F, S32H, S32K, S32R, S32N, or S32L; R58K, R58T, or R58E; F80Y or F80H, and combinations thereof.
[0098] In some embodiments, the engineered antibodies comprise a heavy chain having substitutions at one or more or all of positions 11, 21, 23, 58, and 80, in reference to SEQ ID NO: 1. In some embodiments, the engineered antibodies comprise a heavy chain having the following substitutions: VI IL; D21S; K23A, K23T, K23R, or K23V; R58K, R58T, or R58E; and F80Y or F80H.
[0099] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions to SEQ ID NO:2. In some embodiments, the one or more substitutions are at positions selected from: 1, 3, 4, 9, 10, 14, 28, 29, 30, 31, 58, 60, 70, 77, 81, 90, 91, 93, 97, and combinations thereof, in reference to SEQ ID NO: 2. In some embodiments, the one or more substitutions are selected from: EID or EIQ; V3L or V3Q; L4M; A9S or A9D; T10S; S14K or S14A; S28T; V29I; S30D, S30H or S30T; S3 IT, S3 II, or S31R; I58V; A60D; D70E; S77R; E81D; Q90H; S91R; N93S, N93A, N93T, N93D, or N93Q; T97L; and combinations thereof.
[0100] In some embodiments, the engineered antibodies comprise a light chain having a substitution at position 93, in reference to SEQ ID NO: 2. In some embodiments, the engineered antibodies comprise a light chain having a N93S, N93A, N93T, N93D, or N93Q substitution. In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions to SEQ ID NO: 2 at positions selected from: 1, 93, and combinations thereof. In some embodiments, the one or more substitutions are selected from: EID or EIQ; N93S, N93A, N93T, N93D, or N93Q; and combinations thereof.
[0101] In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence with at least 75% (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least
[0102] 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least
[0103] 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any one of SEQ ID NOs: 5-99 and / or a light chain having an amino acid sequence with at least 75% (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any one of SEQ ID NOs: 1 GO-
[0104] 194. In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence of one of SEQ ID NOs: 5-99 and / or a light chain of one of SEQ ID NOs: 100-194. In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence with at least 75% identity those sequences in Table 1.
[0105] In some embodiments, the engineered antibodies target IL-17A. In some embodiments, the engineered antibodies comprise a heavy chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 3 and / or a light chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 4. In some embodiments, the engineered antibodies comprise one or more substitutions within the CDR loops of SEQ ID NOs: 3 and 4, particularly, the CDR loops of the heavy chain, e.g., CDR H3.
[0106] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions to SEQ ID NO: 3. In some embodiments, the one or more substitutions are at positions selected from: 1, 2, 3, 5, 13, 28, 31, 33, 35, 37, 43, 49, 50, 52, 53, 54, 57, 58, 61, 62, 85, 88, 97, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, and combinations thereof, in reference to SEQ ID NO: 3. In some embodiments, the one or more substitutions are selected from: EIQ; V2Q; Q3S, Q3K, or Q3T; V5Q or V5L; Q13K; T28I; N31 S, N3 ID, or N31R; W33 Y or W33A; N35S or N35T; V37I; K43Q; A49S; A50N, A50S, A50Y, or A50D; N52K; Q53E, Q53R, Q53T, Q53N, Q53K, or Q53S; D54N; E57N; K58T, K58E, or K58I; V61 A; G62D; S85N; V88A; V97A or V97T; D99H, D99E, D99L, D99S, D99R, D99N, D99W, D99K, D99Y, D99Q, or D99A; Y100S, Y100R, Y100K, Y100T, Y100F, Y100A, Y100L, Y100W, Y100E, Y100G, or Y100V; Y101H, Y101F, Y101N, Y101R, or Y101M; D102S, D102T, D102Y, D102W, D102R, D102L, D102E, D102F, D102A, or D102K; I1O3V, I1O3S, I1O3T, 1103 Y, I103L, I103D, I103F, or I103R; LI 04V, LI 04 A, L104E, LI 041, L104Y, L104D, L104T, L104F, L104R, or L104M; T105N, T105D, T105S, T105Y, T105V, T105R, T105E, T105A, T105H, T105W, T105Q, T105I, or T105K; D106N, D106Y, D106E, D106A, D106S, D106T, D106W, D106G, or D106Q; Y107R, Y107L, Y107H, Y107D, Y107T, Y107Q, Y107K, Y107E, Y107A, Y107S, Y107N, Y107M, Y107F, Y107G, or Y107W; Y108L, Y108F, Y108H, Y108D, Y108W, Y108T, Y108G, Y108Q, or Y108M; I109V, I109L, I109Y, I109D, I109T, I109G, or I109Q; Hl ION, H110D, H110W, H110R, H110G, H110Y, H110L, H110A, H110T, or HUOQ; Y111H, Y111I, Y111E, Y111L, Y111R, Y11 ID, Y111G, Y11 IQ, Y11 IV, Y111W, Y11 IM, Y11 IS, Y11 IF, Y11 IN, or Y11 IT; W1 12M, W112Y, W112S, W112D, W112Q, W112A, W112N, W112T, W112G, W112V, or W1 12L; Y113F, Y113D, Y113G, Y1131, Y113A, Y113V, Y113Q, Y113T, Y113L, or Y113R; F114Y, F114W, Fl 14V, F114T, F114G, F114M, or Fl 14D; DI 15W, DI 15M, DI 15T, or DI 15V; LI 16Y or LI 16T; W117V or W117T; G118T; R119Q, R119A, R119S, or R119V; G120S; T121S; L122M, L122S, L122Q, L122T, or L122A; V123S; and combinations thereof.
[0107] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions to SEQ ID NO: 3 at positions selected from: 35, 50, 58, 88, 97, 99, 100, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 119, and combinations thereof In some embodiments, the one or more substitutions are selected from: N35S or N35T; A50N, A50S, A50Y, or A50D; K58T, K58E, or K58I; V88A; V97A or V97T; D99H, D99E, D99L, D99S, D99R, D99N, D99W, D99K, D99Y, D99Q, or D99A; Y100S, Y100R, Y100K, Y100T, Y100F, Y100A, Y100L, Y100W, Y100E, Y100G, or Y100V; D102S, D102T, D102Y, D102W, D102R, D102L, D102E, D102F, D102A, or D102K; I103V, I103S, I103T, 1103 Y, I103L, I103D, I103F, or I103R; L104V, L104A, L104E, L104I, L104Y, L104D, L104T, L104F, L104R, or L104M; T105N, T105D, T105S, T105Y, T105V, T105R, T105E, T105A, T105H, T105W, T105Q, T105I, or T105K; D106N, D106Y, D106E, D106A, D106S, D106T, D106W, D106G, or D106Q; Y107R, Y107L, Y107H, Y107D, Y107T, Y107Q, Y107K, Y107E, Y107A, Y107S, Y107N, Y107M, Y107F, Y107G, or Y107W; Y108L, Y108F, Y108H, Y108D, Y108W, Y108T, Y108G, Y108Q, or Y108M; I109V, I109L, I109Y, I109D, I109T, I109G, or I109Q; Hl ION, Hl 10D, Hl 10W, Hl 10R, Hl 10G, Hl 10Y, Hl 10L, Hl 10A, Hl 1OT, or Hl 1OQ; Y111H, Y1111, Y11 IE, Y11 IL, Y111R, Y11 ID, Y111G, Y11 IQ, Y11 IV, Y111W, Y11 IM, Y11 IS, Y11 IF, Y11 IN, or Y11 IT; W112M, W112Y, W112S, W112D, W112Q, W112A, W112N, W112T, W1 12G, W112V, or W112L; Y113F, Y113D, Y113G, Y1131, Y113A, Y113V, Y113Q, Y113T, Y113L, or Y113R; F114Y, F114W, F114V, F114T, F114G, F114M, or F114D; D115W, DI 15M, DI 15T, or DI 15V; R119Q, R119A, R119S, or R119V; and combinations thereof.
[0108] In some embodiments, the engineered antibodies comprise a heavy chain having one or more substitutions to SEQ ID NO: 3 at one or more or all of positions 35, 50, 88, 97, 100, 102, 103, 104, 105, 112, 114, 119, in reference to SEQ ID NO: 3. In some embodiments, the one or more substitutions are selected from: N35S or N35T; A50N, A50S, A50Y, or A50D; V88A; V97A or V97T; Y100S, Y100R, Y100K, Y100T, Y100F, Y100A, Y100L, Y100W, Y100E, Y100G, or Y100V; D102S, D102T, D102Y, D102W, D102R, D102L, D102E, D102F, D102A, or D102K; H03V, I103S, H03T, I103Y, H03L, H03D, H03F, or H03R; L104V, L104A, L104E, L104I, L104Y, L104D, L104T, L104F, L104R, or L104M; T105N, T105D, T105S, T105Y, T105V, T105R, T105E, T105A, T105H, T105W, T105Q, T105I, or T105K; W112M, W1 12Y, W112S, W112D, W112Q, W112A, W112N, W112T, W112G, W112V, or W112L; Fl 14Y, Fl 14W, Fl 14V, Fl 14T, Fl 14G, Fl 14M, or Fl 14D; R119Q, R119A, R119S, or R119V; and combinations thereof.
[0109] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions to SEQ ID NO: 4. In some embodiments, the one or more substitutions are at positions selected from: 1, 2, 3, 4, 7, 9, 10, 12, 13, 24, 29, 30, 32, 33, 34, 44, 54, 59, 61, 78, 94, 97, 101, 104, 105, and combinations thereof, in reference to SEQ ID NO: 4. In some embodiments, the one or more substitutions are selected from: EID; I2V; V3Q; L4M; S7T; G9A, G9S, or G9D; T10S or T10I; S12A; L13V; R24K; V29I; S30R; S32N, S32R, or S32D; Y33L, Y33N, Y33S, Y33R, Y33F, or Y33H; L34V or L34A; A44S; S54T; I59V; D61A; R78S; S94T; C97T, C97R, C97Y, C97V, C97S, C97I, C97Q, or C97W; Q101G; R104K; L105V; and combinations thereof.
[0110] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions to SEQ ID NO: 4 at positions selected from: 1, 2, 9, 10, 32, 33, 34, 44, 97, 104, 105, and combinations thereof. In some embodiments, the one or more substitutions are selected from: EID; I2V; G9A, G9S, or G9D; T10S or T10I; S32N, S32R, or S32D; Y33L, Y33N, Y33S, Y33R, Y33F, or Y33H; L34V or L34A; A44S; C97T, C97R, C97Y, C97V, C97S, C97I, C97Q, or C97W; R104K; LI 05V; and combinations thereof.
[0111] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions to SEQ ID NO: 4 at one or more or all of positions 2, 9, 33, 97, 104, and 105, in reference to SEQ ID NO: 4. In some embodiments, the one or more substitutions are selected from: I2V; G9A, G9S, or G9D; Y33L, Y33N, Y33S, Y33R, Y33F, or Y33H; C97T, C97R, C97Y, C97V, C97S, C97I, C97Q, or C97W; R104K; LI 05V; and combinations thereof.
[0112] In some embodiments, the engineered antibodies comprise a light chain having one or more substitutions to SEQ ID NO: 4 at one or more or all of positions 1, 9, 10, 32, 33, 34, 44, 97, and 104, in reference to SEQ ID NO: 4. In some embodiments, the one or more substitutions are selected from: EID; G9A, G9S, or G9D; T10S or T10I; S32N, S32R, or S32D; Y33L, Y33N, Y33S, Y33R, Y33F, or Y33H; L34V or L34A; A44S; C97T, C97R, C97Y, C97V, C97S, C97I, C97Q, or C97W; R104K; and combinations thereof.
[0113] In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence with at least 75% (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least
[0114] 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least
[0115] 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any one of SEQ ID NOs: 195-290 and / or a light chain having an amino acid sequence with at least 75% (e.g., at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any one of SEQ ID NOs: 291- 386. In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence of one of SEQ ID NOs: 195-290 and / or a light chain of one of SEQ ID NOs: 291- 386. In some embodiments, the engineered antibodies comprise a heavy chain having an amino acid sequence with at least 75% identity those sequences in Table 2.
[0116] Any of the engineered antibodies described herein may comprise one or more (e g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or more, etc.) amino acid substitutions as compared to any of the recited sequences (e.g., sequences in Tables 1 and 2). An amino acid “replacement” or “substitution” refers to the replacement of one amino acid at a given position or residue by another amino acid at the same position or residue within a polypeptide sequence. Amino acids are broadly grouped as “aromatic” or “aliphatic”. An aromatic amino acid includes an aromatic ring. Examples of aromatic amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y or Tyr), and tryptophan (W or Trp). Non-aromatic amino acids are broadly grouped as aliphatic. Examples of aliphatic amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Vai), leucine (L or Leu), isoleucine (I or He ), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine (C or Cys), proline (P or Pro), glutamic acid (E or Glu), aspartic acid (D or Asp), asparagine (N or Asn), glutamine (Q or Gin), lysine (K or Lys), and arginine (R or Arg).
[0117] The amino acid replacement or substitution can be conservative, semi-conservative, or non-conservative. The phrase “conservative amino acid substitution” or “conservative mutation” refers to the replacement of one amino acid by another amino acid with a common property. A functional way to define common properties between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (Schulz and Schirmer, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such analyses, groups of amino acids may be defined where amino acids within a group exchange preferentially with each other and therefore resemble each other most in their impact on the overall protein structure (Schulz and Schirmer). Examples of conservative amino acid substitutions include substitutions of amino acids within the sub-groups described above, for example, lysine for arginine and vice versa such that a positive charge may be maintained, glutamic acid for aspartic acid and vice versa such that a negative charge may be maintained, serine for threonine such that a free -OH can be maintained, and glutamine for asparagine such that a free -NH2 can be maintained. “Semi-conservative mutations” include amino acid substitutions of amino acids within the same groups listed above, but not within the same sub-group. For example, the substitution of aspartic acid for asparagine, or asparagine for lysine, involves amino acids within the same group, but different sub-groups. “Non-conservative mutations” involve amino acid substitutions between different groups, for example, lysine for tryptophan, or phenylalanine for serine, etc. In some embodiments, the engineered antibody is a monoclonal antibody, a humanized antibody, a chimeric antibody, a recombinant antibody, a monospecific antibody, a bispecific antibody, or a multi-specific antibody.
[0118] Nucleic acids or vectors may be used to propagate the engineered antibodies disclosed herein in an appropriate cell and / or to allow expression from the segment (e.g., an expression vector). The person of ordinary skill in the art would be aware of the various vectors available for propagation and expression of a nucleic acid sequence. The vector(s) and nucleic acid(s) can be introduced into a cell that is capable of expressing the polypeptide encoded thereby, including any suitable prokaryotic or eukaryotic cell.
[0119] In one embodiment, a DNA segment encoding the engineered antibodies disclosed herein is contained in a plasmid vector that allows expression of the protein and subsequent isolation and purification of the protein produced by the recombinant vector. Accordingly, the engineered antibodies disclosed herein can be purified following expression, obtained by chemical synthesis, or obtained by recombinant methods.
[0120] To construct cells that express the engineered antibodies disclosed herein, expression vectors for stable or transient expression of the engineered antibodies may be constructed via methods as described herein or known in the art and introduced into cells. For example, nucleic acids encoding an engineered antibody as described herein may be cloned into a suitable expression vector, such as a plasmid or a viral vector in operable linkage to a suitable promoter.
[0121] Vectors according to the present disclosure can be transformed, transfected, or otherwise introduced into a wide variety of host cells. Transfection refers to the taking up of a vector by a host cell whether or not any coding sequences are in fact expressed. Numerous methods of transfection are known to the ordinarily skilled artisan, for example, lipofectamine, calcium phosphate co-precipitation, electroporation, DEAE-dextran treatment, microinjection, viral infection, and other methods known in the art.
[0122] Further disclosed herein are compositions comprising an engineered antibody described herein. The compositions may comprise excipients or pharmaceutically acceptable carriers. The choice of excipients or pharmaceutically acceptable carriers will depend on factors including, but not limited to, the particular mode of administration, the effect of the excipient on solubility and stability, and the nature of the dosage form. The compositions of the present invention will be readily apparent to those skilled in the art. Techniques and formulations may be found, for example, in Remington’s Pharmaceutical Sciences, 19th Edition (Mack Publishing Company, 1995).
[0123] The term “pharmaceutically acceptable carrier,” as used herein, means a non-toxic, inert solid, semi-solid or liquid fdler, diluent, encapsulating material, surfactant, cyclodextrins or formulation auxiliary of any type. Some examples of materials which can serve as pharmaceutically acceptable carriers are sugars such as, but not limited to, lactose, glucose and sucrose; starches such as, but not limited to, corn starch and potato starch; cellulose and its derivatives such as, but not limited to, sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients such as, but not limited to, cocoa butter and suppository waxes; oils such as, but not limited to, peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, com oil and soybean oil; surfactants such as, but not limited to, cremophor EL, cremophor RH 60, Solutol HS 15 and polysorbate 80; cyclodextrins such as, but not limited to, alpha-CD, beta-CD, gamma-CD, HP-beta-CD, SBE-beta-CD; glycols; such as propylene glycol; esters such as, but not limited to, ethyl oleate and ethyl laurate; agar; buffering agents such as, but not limited to, magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer’s solution; ethyl alcohol, and phosphate buffer solutions, as well as other non-toxic compatible lubricants such as, but not limited to, sodium lauryl sulfate and magnesium stearate, as well as releasing agents, coating agents, preservatives and antioxidants can also be present in the composition, according to the judgment of the formulator.
[0124] The route by which the disclosed engineered antibodies are administered and the form of the composition will dictate the type of carrier to be used. The composition may be in a variety of forms, suitable, for example, for systemic administration (e.g., oral, rectal, nasal, sublingual, buccal, implants, or parenteral injections) or topical administration (e.g., dermal, pulmonary, nasal, aural, ocular, liposome delivery systems, or iontophoresis).
[0125] In some embodiments, the composition comprises a buffering agent. The buffer serves to maintain a physiologically suitable pH. In addition, the buffer can serve to enhance isotonicity and chemical stability of the composition. Generally, the composition has a physiologically suitable pH. In some embodiments, the composition has a pH of about 5 to about 7, about 5.5 to about 6.5, preferably about 6.0 to about 6.5. In select embodiments, the composition has a pH of about 6. Ranges intermediate to the above recited pH levels, for example, about pH 5.2 to about pH 6.3, preferably 6.0 or pH 6.2, are also encompassed. The pH may be adjusted as necessary. For example, HC1 may be added as necessary to adjust the pH to desired levels.
[0126] In some embodiments, the composition further comprises a tonicity agent, an antioxidant, a stabilizer, or a combination thereof. A tonicity agent contributes to maintaining the isotonicity of the composition and preserving the level, ratio, or proportion of the engineered antibody. An antioxidant preserves the composition by preventing oxidation. A stabilizer interacts and stabilizes biological molecules and / or general pharmaceutical excipients in a composition. In some embodiments, the compositions disclosed herein further comprise an adjuvant. An adjuvant refers to one or more substances that cause stimulation of the immune system.
[0127] Any of the above compositions or formulations disclosed herein may further comprise at least one additional therapeutic agent.
[0128] The disclosed engineered antibodies may be useful in a variety of methods. For example, the disclosed engineered antibodies may be utilized in methods directed to the localization and / or quantitation of a target molecule (e.g., for use in measuring levels of a target molecule (e.g., an antigen, a receptor, a ligand, a substrate etc.) within appropriate physiological samples) in diagnostic or imaging methods. The disclosed engineered antibodies may be utilized in methods directed to isolating, purifying, or detecting a target molecule in techniques such as affinity chromatography, immunofluorescence, flow cytometry, immunohistochemistry, or immunoprecipitation. The disclosed engineered antibodies may be particularly useful in prophylactic or therapeutic methods (e.g., to treat or prevent a disease or disorder).
[0129] Further disclosed herein are methods of treating disease or disorder, comprising administering to a subject an effective amount of an engineered antibodies, or a nucleic acid encoding thereof, or composition disclosed herein. The subject may be suffering from, diagnosed as having, or at risk for developing the disease or disorder. In some embodiments, the subject is human.
[0130] Protein structure-encoded language model
[0131] Disclosed herein are systems and methods for protein design. In particular, the present disclosure provides systems and methods to design protein sequences using methods which include adapters to incorporate structural and functional context to the protein sequence. The structural and functional context comprises backbone interactions, protein-protein interactions, and interactions with arbitrary non-protein molecules (e.g., nucleic acids, ligands, ions, etc.). The method for protein sequence design, proseLM, leverages pre-trained protein language models for structure-conditioned design. ProGen2 language models are used as a foundation and a conditional adapter is used for parameter-efficient incorporation of structural and functional context.
[0132] The methods include adapters, which introduce a set of bottlenecked operations that modify the outputs of each model layer. Adapter layers update the outputs of the simultaneous attention and feed-forward operations of each ProGen2 layer (FIG. 1A). With sufficiently reduced bottleneck dimension, the parameters of adapter layers are miniscule with respect to the pre-trained model. To inject structural information into the ProGen2 language models, an adapter architecture conditions in the low-rank representation using a multi-layer perceptron (MLP). In this way, the adapter layers maintain parameter efficiency while incorporating structural context throughout the depth of the language model.
[0133] Structural context for conditioning is obtained from a pre-trained causal encoder, which is trained in a similar fashion to prior encoder-decoder protein design models. The causal encoder architecture consists of alternating invariant-point message-passing (IPMP) and message-passing (MPNN) encoder layers (eight total) followed causally masked IPMP and MPNN layers (four total). IPMP layers modify standard MPNN layers by adding frame-based inter-residue features. Following ProteinMPNN, the causal encoder is trained to decode randomly permuted sequences given a fixed backbone. Different from the natural N-to-C fixed decoding order, this formulation enables conditioning on later residues for design tasks where some of the sequence should remain constant. The causal encoder is additionally trained with masked structural spans, following the span sampling scheme used for ESM-IF1. To enable span masking with IPMP layers, which depend on residue structure, updates on masked residues to the MPNN layers are restricted. When trained on the CATH 4.2 dataset, the causal encoder achieves a median sequence recovery of 47.24% and perplexity of 4.76 (Table 5) on the test set. Taken together, this performance is similar to ProteinMPNN (trained without noise), which achieves a median sequence recovery of 45.96% and perplexity of 4.61 on the same proteins.
[0134] To implement the method, the pre-trained causal encoder is combined with pre-trained ProGen2 language models using parameter-efficient adapter layers. Standard adapter layers are typically composed of a feed-forward network with a bottleneck dimension, expressed as: where h is the language model embeddings, / is a non-linear activation function (typically ReLU), and Rdown and PKupare weight matrices for the down- and up-projections, respectively. The conditional adapter architecture incorporates structural context into the language model embeddings using an MLP. The conditional adapter layers are expressed as: where ALM and ACE are the language model and causal encoder embeddings, respectively. The causal encoder embeddings are taken from the last decoder layer, just prior to amino acid prediction. By concatenating the causal encoder embeddings with the low-rank language model embeddings, the conditional adapter maintains the parameter-efficiency of standard adapters while incorporating conditioning information. Conditional adapter layers are placed after each of the simultaneous attention and feed-forward layers of ProGen2, with separately trained weights for each adapter layer. During training, the parameters of the causal encoder and conditional adapter layers are updated, while the parameters of the language model are frozen.
[0135] To test the effectiveness of the conditional adapters for structure-conditioned sequence modeling, the perplexity of models trained on the CATH 4.2 dataset are compared to their constitutive models. For ProGen2 models, perplexity scaled with model size, with ProGen2- xlarge (6.4B parameters) approaching the performance of the structure-aware causal encoder (FIG. IB). For proseLM models, perplexity further improved over the causal encoder, following the same scaling trend as the underlying language models. The native sequence recovery of proseLM models were compared to the causal encoder. Native sequence recovery is defined as the percentage of designed residues that match the native sequence for a particular structure. Native sequence recovery increased with proseLM model scale, with proseLM-XL achieving a 3.59% higher median recovery rate than the causal encoder (FIG. 6A). This increase in sequence discovery was distributed across surface and core residues (FIG. 6B) and was most dramatic for larger proteins (FIG. 6C).
[0136] Training structure-conditioned sequence design methods with Gaussian noise applied to the coordinates has been shown to increase robustness of single-sequence AlphaFol d2 predictions, which can serve as a proxy for design quality. To assess the impact of training with coordinate noise on proseLM, an additional set of causal encoder and proseLM models were trained with 0.1 A Gaussian noise added to the backbone coordinates. As reported for ProteinMPNN, all proseLM models trained with coordinate noise achieved higher rates of single-sequence prediction structure prediction success with AlphaFold2 and yielded more confident structures (FIG. 7). Interestingly, these improvements were most pronounced for the causal encoder and smaller proseLM models, suggesting that the larger models are more robust to structural noise.
[0137] Functional context constrains design space
[0138] Protein function depends on interactions with other molecules, ranging from recognition of specific DNA sequences to coordination of metal ions. The methods disclosed herein incorporate functional context beyond the backbone to include not only protein-protein interactions, but also interactions with arbitrary non-protein molecules (e.g., nucleic acids, ligands, ions, etc.). The causal encoder is extended to consider protein complexes, which is achieved by expanding the existing protein graph to include multiple chains.
[0139] Non-protein context is represented as an atomistic graph, with each atom represented as a node with an associated reference frame. Each atomic frame is constructed using the atom-of- interest and its two nearest neighbors. Nodes are initialized with the types of each atom encoded as node features, along with the distances between the central atom and its neighbors. To incorporate information from the non-protein context graph into the causal encoder, an IPMP layer is introduced for updating the atomistic features, followed by a cross-graph IPMP layer for updating the protein residue features. The cross-graph IPMP layer is similar to the standard IPMP layer, but uses a different set of weights for protein and atomistic nodes, and does not update the atomistic features. The context-aware causal encoder and corresponding proseLM models were trained on the multi -chain dataset used for ProteinMPNN in a similar manner as the single-chain versions, except the causal encoder parameters were frozen during proseLM training to reduce memory requirements.
[0140] The models were evaluated on the subset of the PDB test set that contained protein complexes and interactions with some non-protein entity (Table 7). For each protein chain, the perplexity is computed given increasing levels of context: backbone only, with other protein chains, and with full context. In each scenario, perplexity decreased as more context was provided (FIG. 1C). This trend held across model scales, suggesting that even the largest language models benefit from explicit functional context (FIG. ID). The capabilities for fullsequence design are evaluated on the test set structures that had any context (protein or non- protein). The median native sequence recovery for designs given the backbone only and full context show increases in sequence recovery from scaling models and providing functional context were largely additive, with the disclosed model achieving a consistent 3-5% increase in sequence recovery with full context (FIG. IE).
[0141] The most significant increases in native sequence recovery were for residues within 5 A of nucleic acids, ligands, and ions (FIG. 2A). The causal encoder achieves sequence recovery (with / without context) of 44.44% / 54.90%, 50.00% / 64.00%, and 46.15% / 66.67% for residues near nucleic acids, ligands, and ions, respectively.
[0142] With the addition of language model components, proseLM models achieve further increases in native sequence recovery in the vicinity of functional context (Table 7). For ligand and ion interactions, these increases were mostly localized to residues within 5 A of the provided context (FIG. 2B). For nucleic acid interactions, smaller models showed the largest increases in sequence recovery near the provided context, while larger models showed more consistent improvements at greater distances. This suggests that while smaller models benefit from the direct influence of nucleic acid context for local sequence recovery, larger models may be leveraging evolutionary information from pre-training to improve sequence recovery at greater distances. For protein-protein interactions, the effects of context were observable up to 11 A away, largely due to changes in burial of residues across the designed protein chains (FIG. 2C).
[0143] Structure improves modeling of protein fitness
[0144] Protein language models trained on evolutionarily diverse sequences are strong zero-shot predictors (i.e., without labeled data) of protein function, with performance on some landscapes benefitting from increased model scale. ProseLM models were evaluated on a set of 201 deep mutational scan (DMS) datasets curated for ProteinGym. Analysis was to datasets where the protein sequence length was less than 1024 residues, which was the context length used for training ProGen2 models. Log likelihoods were computed for proseLM models trained on the CATH 4.2 (backbone-only) and PDB (full context) datasets using AlphaFol d2-predicted structures from ProteinGym. Datasets are coarsely categorized as stability, binding, activity, expression, or organismal fitness, according to the property measured. An overall performance as the average of performance over each category to avoid imbalances in the number of datasets available for each fitness type. Fitness prediction performance is quantified according to the normalized discounted cumulative gain (NDCG) at ten percent (FIG. 9) and the Spearman's rank correlation coefficient (FIG. 10). Although Spearman values are more commonly reported for performance across entire fitness landscapes, the NDCG metric, which focuses on a model's ability to prioritize the highest- fitness seqeunces, is more aligned with practical protein engineering settings. Overall, proseLM models achieve higher NDCG values than the causal encoder and ProGen2 models (FIG. 2D). Interestingly, the largest improvements in NDCG are for smaller language models, while larger proseLM models often performed comparable to the underlying ProGen2 models.
[0145] For stability landscapes, opposing trends were observed between structure-conditioned and standard language models. The causal encoder was a better predictor of high-stability proteins, while proseLM models showed degraded performance with increased scale. For binding-related fitness landscapes, proseLM models outperformed the causal encoder and underlying ProGen2 models. Interestingly, the backbone-only proseLM models trained on the CATH 4.2 dataset consistently outperformed those trained on the full PDB with additional context. This discrepancy highlights a tradeoff between models that learn to implicitly model functional context and those that explicitly incorporate given context. While the proseLM models trained on the PDB are able to effectively leverage binding partners for prediction, they appear less capable of inferring binding from protein structure alone.
[0146] In some embodiments, the technology described herein is associated with a programmable machine designed to perform a sequence of arithmetic or logical operations as provided by the methods described herein. For example, some embodiments of the technology are associated with (e.g., implemented in) computer software and / or computer hardware. In one aspect, the technology relates to a computer comprising a form of memory, an element for performing arithmetic and logical operations, and a processing element (e.g., a microprocessor) for executing a series of instructions (e.g., a method as provided herein) to read, manipulate, and store data.
[0147] In some embodiments, the various embodiments of the present disclosure are associated with a plurality of programmable devices that operate in concert to perform a method as described herein. For example, in some embodiments, a plurality of computers (e.g., connected by a network) may work in parallel to collect and process data, e.g., in an implementation of cluster computing or grid computing or some other distributed computer architecture that relies on complete computers (with onboard CPUs, storage, power supplies, network interfaces, etc.) connected to a network (private, public, or the internet) by a conventional network interface, such as Ethernet, fiber optic, or by a wireless network technology.
[0148] For example, some embodiments provide a computer that includes a computer-readable medium. The embodiment includes a random access memory (RAM) coupled to a processor. The processor executes computer-executable program instructions stored in memory. Such processors may include a microprocessor, an ASIC, a state machine, or other processor, and can be any of a number of computer processors, such as processors from Intel Corporation of Santa Clara, California and Motorola Corporation of Schaumburg, Illinois. Such processors include, or may be in communication with, media, for example computer-readable media, which stores instructions that, when executed by the processor, cause the processor to perform the steps described herein.
[0149] Computers are connected in some embodiments to a network. Computers may also include a number of external or internal devices such as a mouse, a CD-ROM, DVD, a keyboard, a display, or other input or output devices. Examples of computers are personal computers, digital assistants, personal digital assistants, cellular phones, mobile phones, smart phones, pagers, digital tablets, laptop computers, internet appliances, and other processor-based devices. In general, the computers related to aspects of the technology provided herein may be any type of processor-based platform that operates on any operating system, such as Microsoft Windows, Linux, UNIX, Mac OS X, etc., capable of supporting one or more programs comprising the technology provided herein. Some embodiments comprise a personal computer executing other application programs (e.g., applications). The applications can be contained in memory and can include, for example, a word processing application, a spreadsheet application, an email application, an instant messenger application, a presentation application, an Internet browser application, a calendar / organizer application, and any other application capable of being executed by a client device. All such components, computers, and systems described herein as associated with the technology may be logical or virtual.
[0150] Examples
[0151] Example 1 Genome editor design Genome editing technologies repurposed from bacterial anti-phage defense systems have revolutionized life science research and are being actively developed for agricultural and therapeutic applications. In particular, the Cas9 protein from Streptococcus pyogenes (SpCas9) has formed the foundation of several downstream editing technologies, including the targeted editing of base pairs in the genome. As an RNA-guided endonuclease, the functional activity of SpCas9 is highly dependent on interactions with its guide RNA and the target DNA.
[0152] To design variants of SpCas9 with proseLM while maintaining functional activity, a multi-state conditioning strategy was utilized (FIG. 3A) that included both the binary and catalytic states of the protein in complex with guide RNA (and DNA for catalytic state). Evolutionary and functional information was incorporated in the form of residue-wise conservation patterns from aligned natural sequences and experimental mutation-scanning data, respectively. When combined with multi-state conditioning, this design strategy enabled sampling of low-perplexity sequences within 200 mutations of the wild type sequence (FIG. 11). A set of seven designs were tested for genome editing in HEK293T cells via co-transfection of the designed proteins and single-guide RNAs (sgRNA) targeting one of three previously characterized target sites. Across all three sites, a wide range of editing efficiencies were observed, with a subset of variants showing activity on-par or higher than SpCas9 (FIG. 3B). The most active variant, OPT-Cas9-2, showed a significant increase in editing compared to wild-type SpCas9 at two of three target sites. Notably, this variant contained considerable mutational load in the REC-1 and HNH domains (FIG. 3C), which may facilitate increased on-target editing through reduced specificity and increased nuclease activity.
[0153] Next considered was the design of base editors, fusions of a deaminase domain to a Cas9 nickase scaffold (FIG. 3D). Adenine base editors enable the targeted conversion of A:T base pairs to G:C base pairs in the genome, and have been used to correct pathogenic mutations in human cells. As a testbed for optimization with proseLM, a deaminase domain previously designed using protein language models with editing activity on par with early base editors derived from the natural E. coli TadA protein was selected. A structural model of the base editor functional state was created by aligning an AlphaFold2 prediction of the dimeric deaminase to the previously solved structure of ABE8e. Using proseLM, design was focused on active site residues (within 5 A of the bound adenine) or non-active site residues (outside 5 A and not in the deaminase dimer interface). 40 designs from each strategy were tested for A-to-G editing efficiency in HEK293T cells and both strategies yielded base editors with significantly higher editing efficiency than the parental deaminase (FIG. 3E). Among the active site designs, improvements in editing efficiency were achieved with as few as three mutations, while the best design featured six mutations (FIG. 3F). For non-active site designs, which ranged from 13 to 27 mutations, fewer designs retained editing efficiency and only one showed improvment over the parental sequence (FIG. 3G). If FIG. 3H, the set of six active-site mutations yielding a 50% relative improvement in editing efficiency are depicted on the predicted structure of the parental deaminase. While it is difficult to ascertain the precise contribution of each mutation, the design contains several non-conservative mutations, representing a significant reworking of functionally important residues. Meanwhile, the non-active site design with the highest editing efficiency contained 19 mutations distributed across the surface and core of the deaminase (FIG. 31), which may facilitate increased stability or solubility of the protein.
[0154] Example 2 Therapeutic antibody design
[0155] Antibodies are a class of immune proteins that have been developed for a wide range of research and clinical applications. The design of specific and biophysically well-behaved antibodies has been a long-standing challenge, due in large part to the complexity and sensitivity of protein-protein interactions typical of antibody-antigen complexes. Recently, protein language models (including antibody-specific models) have been used for targeted optimization of particular antibody attributes, such as stability or immunogenicity. However, prior approaches have primarily focused on sequence-based optimization, ignoring the structural context of antibody binding, and focusing on mutations to the framework region.
[0156] Although proseLM models effectively recovered native residues at protein-protein interfaces, neither the causal encoder nor the ProGen2 models were exposed to significant numbers of antibody sequences during training. To address this limitation, proseLM-Ab (based on adaptation of the ProGen2-OAS model) was trained on a set of antibody structures from the Structural Antibody Database (SAbDab). The perplexity of ProGen2 and proseLM models was compared on a heldout set of antibodies and proseLM-Ab achieved lower perplexity than proseLM models trained on the PDB, including those with significantly more parameters (FIG. 4A). These improvements came despite the underlying ProGen2-OAS model assigning relatively high perplexity to the heldout antibodies, likely due to an under-representation of light chains in its training corpus. When provided antigen context, all proseLM models assigned lower perplexity to the antibody sequences (FIG. 4B) and typically achieved higher rates of native sequence recovery (FIG. 4C). For heavy and light chain framework regions, a scaling trend was observed in sequence recovery, with larger models achieving higher recovery rates (FIG. 4D). A similar trend was observed for CDR regions, although the recovery rates were generally lower than for the framework regions. ProseLM-Ab frequently achieved the highest rates of recovery for CDR loops, but showed the largest improvements in the framework regions, demonstrating the utility of adapting antibody -specific language models trained on diverse antibody sequences.
[0157] To test the design capabilities of proseLM models, nivolumab, which targets the PD-1 antigen, was used in affinity optimization. Mutations were focused in either the complementarity-determining regions (CDRs) or the framework regions. For CDR-directed designs, all single and double mutations (excluding mutations to or from cysteine or proline) at positions within 8 A of the antigen were considered (FIG. 13). For framework variants, the entire heavy and light chain variable fragments were redesigned, conditioned on the residues within 8 A of the antigen (FIG. 14). Designs from both strategies were scored using an ensemble of proseLM models and the best 55 CDR-directed and 40 framework-directed variants were selected for experimental characterization. Binding to PD-1 was tested via surface plasmon resonance (SPR) and KD values for 25.4% of CDR-directed designs and 92.5% of framework- directed designs (FIG. 4E). Among these, a wide range of binding affinities were observed, with the most improved variants from each strategy achieving a 2.5-fold increase in binding affinity relative to nivolumab (FIG. 4F). The most improved CDR-directed variant contained two mutations near the antigen interface, which likely facilitated tighter binding through improved rigidification of the paratope (FIG. 4G), despite one mutation (HC S32H) ablating binding and the other (LC N93S) only moderately improving affinity in isolation. The most improved framework-directed variant contained seven mutations that may have indirectly improved binding through stabilization of the framework.
[0158] Secukinumab, which binds the IL-17A cytokine through contacts mediated by an extended 18-residue CDR H3 loop, was selected for a redesign of the entire heavy and light chain variable fragments (FIG. 15). Due to the relative ease of recovering framework residues, the majority of variation was focused within the CDR loops (particularly CDRH3). 96 designs were tested for binding to IL-17A via SPR and found two variants that retained binding (FIG. 4H). These variants contained 18 and 31 mutations and bound with 135 nM and 102 nM affinity, respectively. These results demonstrate the potential of proseLM for the diversification of structurally challenging therapeutic antibodies.
[0159] Table 1.
[0160]
[0161]
[0162]
[0163]
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179] Table 2.
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200] Materials and Methods
[0201] Protein structure datasets Three datasets were assembled for training proseLM models: single-chain proteins, protein complexes with non-protein context, and antibody complexes. For training on diverse single-chain proteins, the CATH 4.2 dataset constructed by Ingraham et al., which has been widely adopted as a benchmark for single-chain protein design methods, was used. For training on protein complexes with non-protein context, the dataset was adapted to train multi-chain variants of ProteinMPNN. This dataset was originally constructed by clustering the entire PDB, as of August 2, 2021, at 40% sequence identity and identifying clusters for testing that did not include proteins that co-occurred in biological assemblies with proteins used for training. This dataset was extended by extracting non-protein context within 5 A of the protein chains. For training on antibody complexes, all antibody structures were collected from SAbDab, as of July 1, 2023, and performed clustering at 80% identity on the concatenated heavy and light chain variable fragment sequences with MMseqs2. These clusters were used to divide the dataset into training, validation, and test splits, such that roughly 80%, 10%, and 10% of clusters were allocated to each split, respectively.
[0202] Model architecture
[0203] Structure featurization Protein structures were formulated as nearest-neighbor graphs, with residues as nodes and inter-residue relationships as edges. The graph adjacency matrix is defined by the promiximity of residues in sequential and three-dimensional space. For each residue, the six closest residues along the sequence were taken then the next thirty closest residues according to Cadistance, for a total of 36 neighbors. The nodes are featurized only with a binary indicator of whether the residue has a backbone structure (e.g., N, Ca, and C atoms). The edges are featurized by the inter-residue distances between pairs of backbone atoms (N, Ca, C, O, virtual Cp) and a relative positional embedding indicating the relative position of neighboring residues along the sequence, up to a maximum of 32 positions in either direction. For inter-chain residue pairs, the relative positional embedding was set to a constant value indicating no sequential relationship, but all other features remain the same.
[0204] For atomic-level representation of non-protein context, a similar graph representation was adopted, with individual atoms as nodes and inter-atomic relationships as edges. The adjacency matrix for the atomic graph is constructed by selecting the ten nearest atoms in three-dimensional space. Each atom is represented as a node with an associated reference frame formed by the atom-of-interest and its two nearest neighbors. The nodes are initialized with the types of each atom, along with the distances between the primary atom and its neighbors. The edges of the atomic graph are featurized by the inter-atomic distances between these triplets of atoms. Finally, to incorporate information from the atomic graph into the protein graph, a cross-graph adjacency matrix was construct between each protein residue and the nearest thirty atoms (up to 12 A away) in the atomic graph. The edges for the protein-atomic graph are featurized by the pairwise distances between the backbone atoms of the protein residues and the atom triplets of the atomic graph.
[0205] Causal encoder architecture The causal encoder takes as input the protein and atomic context graphs featurized as described above. All discrete node and edge features e.g., (binary structural indicator, relative positional embedding, etc.) are one-hot encoded. All distance-based edge features are encoded by sets of sixteen Gaussian radial basis functions (RBFs) equally spaced between 0 and 20 A. The respective features for nodes and edges are concatenated then processed by two-layer MLPs to bring them to appropriate dimensionality. The causal encoder is composed of a series of sequence-agnostic encoder layers followed by a series of causally masked decoder layers that are ultimately used to predict the amino acid sequence (Algorithm 1). Complete hyperparameters for the causal encoder are provided in Table 3.
[0206] Algorithm 1 Causal encoder
[0207] Require: -> Protein node and edge embeddings
[0208] Require: o Atomic node and edge embeddings
[0209] Require: > Protein-atomic edge embeddings
[0210] Require: i> Rigid transforms for protein and atomic nodes
[0211] Require: > Topology of protein, atomic, and protein-atomic graphs
[0212] Require: > Amino acid swjaence
[0213] Table 3. Casual encoder hyperparameters
[0214] The encoder and decoder layers are parameterized by graph neural networks (GNNs), specifically taking the form of message-passing neural network (MPNN) layers and invariant point message-passing (IPMP) layers. The instantiation of MPNN and IPMP layers differ only in the features used to form messages for updating node and edge embeddings. For MPNN layers, messages are formed by concatenating node and edge embeddings for neighbors in the graph topology (Algorithm 2). For IPMP layers, these embeddings are added to the set of five invariant components proposed by Randolph et al. (Proteins: Structure, Function, and
[0215] Bioinformatics, 2024) (Algorithm 3). To obtain these components, a set of invariant points in the local frame of each node is predicted, then each of the following are computed for a pair of neighboring nodes in the graph (Algorithm 4):
[0216] 1. Node z's invariant points in node z's local frame
[0217] 2. Distances from node z's invariants points to the origin of node z's local frame
[0218] 3. Node / s invariant points in node z's local frame 4. Distances from node j's invariant points to the origin of node z's local frame
[0219] 5. Distances between node z's invariant points and node / s invariant points in the global frame Algorithm 2 MPA!N layer
[0220] Require: u;, ey i> Node and edge embeddings
[0221] Require: .M u Topology of graph specifying neighbors of each node return n;, ey
[0222] Algorithm 3 1PMP layer
[0223] Require: iq, ey > Node and edge embeddings
[0224] Require: T, i> Rigid transforms
[0225] Require: Ar•> Topology of graph specifying neighbors of each node return n}. ey
[0226] Algorithm 4 Invariant point message features
[0227] Require: n., ey > Node and edge embeddings
[0228] Require: T-- Rigid transforms
[0229] Require: A’ > Topology of graph specifying neighbors of each node
[0230] After assembling the message features p;yfor MPNN or IPMP layers, those features are passed through a set of common operations to update node then edge embeddings. For the node embeddings a set of messages m / ;for each neighboring node j E N was computed using a three- layer MLP and aggregate them by taking the mean across all neighbors. These messages are then added to the original node embeddings and passed through a layer normalization. The updated node embeddings are further processed by a two-layer MLP, with the outputs added back to the previously updated node embeddings with layer normalization. For the edge embeddings, a similar procedure is followed omitting the aggregation step and instead directly updating the using the messages.
[0231] To enable the decoder layers to predict the amino acid sequence residue-by-residue at inference time, the ground truth (or presently decoded) sequence was provided through causally masked edge embeddings. Specifically, for neighboring nodes i > j, the protein sequence s?rotis embedded and added to the protein edges. For node and edge updates, messages were formed in a causally consistent manner for neighboring nodes i and j by selectively using encoder node embeddings for node j when i < j $ and decoder node embeddings when / ' > j. The final node embeddings are used to predict the amino acid sequence through a linear layer.
[0232] Causal encoder models were trained for 80 epochs on the CATH 4.2 and PDB datasets with an effective batch size of 32 using the Adam optimizer. The learning rate was increased linearly over 4,000 warmup steps to a maximum value of le-4, then decayed according to an inverse-square-root schedule. An additional set of models were trained with 0.1 A Gaussian noise added to the protein backbone coordinates. The final models were chosen according to validation set loss.
[0233] ProseLM architecture For proseLM models, a pre-trained causal encoder is combined with a pre-trained ProGen2 protein language model using a set of parameter-efficient conditional adapters (Algorithm 5). The causal encoder is used to encode the protein and atomic structures, as well as the causally masked amino acid sequence, into a set of node embeddings n?rot■ These embeddings are then used to condition the outputs of the simultaneous attention and feedforward layers of the language model through a set of MLP adapters (Algorithm 6). The node embeddings are shifted by one position, such that the language model is conditioned on the structural information of the next residue to be predicted, rather than the current residue encoded in the sequence. The adapter layers first down-project the language model embeddings to a reduced dimensionality, then condition on the node embeddings through a two-layer MLP, and finally up-project the conditioned embeddings back to the original dimensionality. Importantly, the linear layer responsible for up-projecting the embeddings is initialized such that the entire adapter layer approximates an identity operation. The conditioned embeddings are ultimately added back to the original language model embeddings and passed through the pre-trained layer normalization of the language model. Hyperparameters for conditional adapters for each proseLM model are provided in Table 4.
[0234] ProseLM models were trained for 5 epochs on the CATH 4.2 dataset and 15 epochs on the PDB dataset with an effective batch size of 64 using the Adam optimizer. The learning rate was set to a fixed value of 2e-4. An additional set of models were trained with 0.1 A Gaussian noise added to the protein backbone coordinates. The final models were chosen according to validation set loss.
[0235] Algorithm 5 proseLM
[0236] Require: > Protein node and edge embeddings Require: t> Atomic node and edge embeddings itequi ♦re: > Protein-atomic edge embeddings
[0237] Require: > Rigid transforms for protein and atomic nodes
[0238] Require: i> Topology of protein, atomic, and protein- atomic graphs
[0239] Require: & Amino acid sequence
[0240] > Embed protein sequence aneous attention mid feedforward layers o Condition on next structure residue
[0241] Sj <--■ Linear(h<) > Predict amino acid sequence return s;
[0242] Algorithm 6 MLP adapter
[0243] Require: h-,pus> Language model embeddings
[0244] Require: up>< 1> Protein node embeddings i> Down-project language model embeddings
[0245] <> Down-project node embeddings
[0246] > Condition on next structure residue
[0247] ■> Up-project language model embeddings return hf Table 4. Prose LM hyperparameters Comparison to LM-Design The most conceptually similar method to proseLM is LM- Design, which adapts masked language models for structure-conditioned sequence design. While proseLM generates sequences autoregressively for a given backbone structure, LM-Design adopts a strategy more akin to sequence refinement using language models given an initial guess. Architecturally, the conditional adapter of proseLM uses significantly few parameters per layer (80K-300K) than LM-Design (5M per layer). This discrepancy enables proseLM to maintain parameter-efficiency while incorporating adapters after every layer of the language model, whereas LM-Design only incorporates adapters after the final layer. To assess the impact of these architectural differences, proseLM was compared to LM-Design on the CATH 4.2 test set (Table 5). The most apt comparison to the reported LM-Design performance is proseLM-M, which contains a similar number of pre-trained language model parameters. ProseLM-M achieves lower perplexity, while LM-Design achieves higher native sequence recovery. This discrepancy likely arises from the different modeling objectives, with proseLM sampling sequences autoregressively and LM-Design iteratively updating sequences by sampling from position-wise marginal distributions. Autoregressive modeling directly captures the co-evolution between pairs of residues, which is critical for protein function, but may yield designs with lower native sequence recovery.
[0248] Table 5. CATH 4.2 benchmark. Comparison of recent protein design models on the CATH 4.2 test set. Perplexity is reported over all residues in the test set. f indicates results cited from Gao, et al. (In The Eleventh International Conference on Learning Representations, 2022). J indicates results cited from Diederik P Kingma and Jimmy Ba. (Adam: A method for stochastic optimization. arXiv preprint arXiv: 1412,6980, 2014, ),
[0249] Table 6. PDB benchmark. Comparison of proseLM modesl on PDB test set derived from Daupraras, et al. (Science, 378(6615):49— 56, 2022). Metrics are for proteins that are part of a complex or have interactions with some non-protein entity. Performance is reported for both backbone-only and full-context inputs (backbone / full). For protein complexes, all other chains are provided as context for the target chain.
[0250] Table 7. Sequence recovery near non-protein context. Comparison of proseLM models on PDB test set derived from Daupraras, et al. (Science, 378(6615):49-56, 2022). Metrics are for residues within 5 A of nucleic acids, ligands, or ions. Performance is reported for both backbone-only and full-context inputs(backbone / full).
[0251] Table 8. SabDab antibody benchmark. Comparison of proseLM models on antibody test set from SabDab. Sequence recover is reported as the median of cluster-averaged values for all antigen-bound heavy (n=361) and light (n=300) chains in the test set. Performance is reported for both antibody-only and full-complex inputs (antibody / complex),
[0252] Table 9. SabDab bound antibody benchmark. Comparison of proseLM models on antibody test set from SAbDab. Sequence recovery is reported as the median of cluster-averaged values for all antigen-bound heavy (n=361) and light (n=300) chains in the test set. Performance is reported for roth antibody-only and full-complex inputs (antibody / complex).
[0253] Design of genome editors
[0254] To design optimized SpCas9 nucleases with proseLM, proseLM-Base was trained by adapting the ProGen2-base model and training on the PDB dataset. This modeled the complete 1,368-residue SpCas9 sequence, which extends beyond the maximum length supported by other ProGen2 and proesLM models. For design, a multi-state conditioning strategy was utilized. Two structures were selected for conditioning that represented the binary (PDB ID 4ZT0) and catalytic (PDB ID 7Z4J) states of SpCas9. To account for missing residues in the structures, AlphaFold2 was used to predict the structures with each state provided as a template. The predicted structures were aligned with sub-A RMSD into the experimental complexes, yielding structurally complete representations of the binary and catalytic states. Using these structures, 1,600 sequences (7=1.0) were designed. All sequences had the PAM-interacting domain and known catalytic residues fixed to their natural identities.
[0255] To generate sequences closer to the natural SpCas9 sequence, positional biases derived from evolutionary and experimental data were introduced. Evolutionary data was obtained from a multiple-sequence alignment of phylogenetically related Cas9 proteins, which were used to create a position-specific scoring matrix (PSSM). Experimental data was obtained from a deep mutational scanning study of SpCas9, which measured the on- and off-target impacts of single amino acid mutations. The PSSM and normalized DMS data were combined to create a positional bias for each residue in the SpCas9 sequence. This bias was added to the logits of proseLM-Base and was used to generate an additional 4,400 sequences (7=0.1). Seven designs were selected from this set for experimental characterization according to the criteria used for OpenCRISPR.
[0256] For base editor optimization, a deaminase domain previously designed using protein language models fine-tuned on the TadA family was selected. The structure of the deaminase dimer was predicted using the ABE8e structure (PDB ID 6VPC) as a template, then the predicted structure was aligned into the functional state of the ABE8e complex with sub-A RMSD. The designs were divided across two strategies, focusing separately on active site and non-functional scaffolding residues. For active site designs, all residues within 5 A of the single-stranded DNA or Cas9 nuclease were selected, excluding positions at the termini or within 5 A of the dimeric interface, and all other residues were kept fixed. For non-active site designs, all residues greater than 5 A from the single-stranded DNA or Cas9 nuclease and greater than 5 A from the dimeric interface selected, and all other residues were kept fixed. When generating designs, the fixed residues were provided to the causal encoder as context by reordering the residue indices, effectively conditioning proseLM on future positions. 300 active site designs (7=1.0) and 200 non-active site designs (7=0.5) were generated with proseLM-S. The 40 best designs from each strategy according to perplexity were selected for experimental characterization.
[0257] Design of antibodies
[0258] To optimize the binding affinity of the therapeutic antibody nivolumab, the crystal structure of nivolumab bound to PD-1 (PDB ID 5WT9) was used. The designs were divided across two strategies, focusing separately on the complementarity-determining regions (CDRs) and the framework regions. For CDR-directed designs, all possible single- and double-mutations to residues within 8 A of the antigen were enumerated, excluding mutations to or from cysteine or proline. In total, this set included 414,477 variants with mutations across 54 positions. For framework-directed designs, designs were generated with conditioning on the CDRs by fixing residues within 6 A of the antigen. 2,000 designs each were generated from proseLM-Ab and proseLM-Base (7=1.0). Half of sequences had the heavy chain designed first, with the light chain successively designed based on the designed heavy, and the other half used the reverse chain order. 55 CDR-directed and 40 framework-directed designs were selected according to an ensemble of proseLM models (causal encoder, proseLM-S, proseLM-base, proseLM-XL, proseLM-Ab), using the product of perplexities as a selection criterion. For CDR-directed designs, 15 single mutations and 40 double mutations were selected. For framework-directed designs, 20 designs with 1-10 mutations and 20 designs with 11-20 mutations were selected. For each strategy, no particular mutation appeared more than ten times in the final set. In total, 95 designs were selected for experimental characterization.
[0259] For diversification of the therapeutic antibody secukinumab, the crystal structure of secukinumab bound to IL-17A (PDB ID 6WIO) was used. Full heavy and light chain variable fragments were generated with 2,000 designs each from proseLM-Ab and proseLM-Base (7=1.0). Half of sequences had the heavy chain designed first, with the light chain successively designed based on the designed heavy, and the other half used the reverse chain order. 48 designs were selected according to an ensemble of proseLM models (causal encoder, proseLM-S, proseLM-base, proseLM-XL, proseLM-Ab), using the product of perplexities as a selection criterion. From each model's designs, 16 designs with 1 to 24 mutations, 16 designs with 25 to 29 mutations, and 16 designs with 30 to 34 mutations were selected. In total, 96 designs were selected for experimental characterization.
[0260] Characterization of genome editors
[0261] DNA oligonucleotides and plasmid assembly Oligonucleotides used in this study were synthesized by IDT with standard desalting. All natural and Al-generated nuclease and deaminase sequences were purchased as synthetic gene fragments (Twist Bioscience), human codon-optimized using the Twist codon optimization tool, and cloned into CMV-driven expression plasmids using HiFi DNA Assembly (New England Biolabs). Single-guide RNA (sgRNA) sequences were cloned into a human U6 (hU6)-driven expression plasmid that also contains a CMV-driven GFP transfection reporter using HiFi DNA Assembly (New England Biolabs). All plasmids were sequence-verified by whole plasmid Nanopore sequencing (Primordium) prior to downstream applications.
[0262] HEK293T cell culture and transient transfection HEK293T cells (ATCC) were cultured at 37°C and 5% (v / v) CO2 in high glucose DMEM with 4 mM L-glutamine, 1 mM sodium pyruvate and phenol red pH indicator (Gibco), supplemented with 10% FBS and IX penicillinstreptomycin. 24 hours prior to transfection, cells were seeded at a density of IxlO3cells / well in 96-well tissue culture-treated plates (Nunc™ Edge™, Thermo Fisher Scientific).
[0263] For each transfection well, 50 ng of sgRNA plasmid and 50 ng of nuclease or base editorexpressing plasmid were added to 5 pL of Opti-MEM (Gibco). 0.2 pL of TransIT®-2020 transfection reagent (Minis Bio) was diluted into 4 pL of Opti-MEM. Plasmid and TransIT®- 2020 mixtures were combined, incubated for 15-30 min at room temperature, and added to HEK293T cells in a dropwise manner. Plates were gently rocked to mix and incubated for 72 hours.
[0264] Targeted-amplicon secpiencing and analysis Cell lysates were generated by washing the cells with IX PBS and adding 25 pL of lysis buffer (100 mM Tris-HCl, pH 7.5; 0.05% SDS; 25 pg / mL Proteinase K) per well. Plates with lysis buffer were incubated at 37°C for 1 hour, and then 25 pL of nuclease-free water was added per well. The lysates were transferred to 96-well PCR plates and boiled at 98°C for 15 min. Locus-specific primers were used to amplify regions of interest from cell lysates by PCR (Q5® High-Fidelity DNA Polymerase). After PCR, the resulting amplicons were purified (Mag-Bind® RxnPure Plus, Omega Bio-tek), DNA yields were quantified (QuantiFluor® kit, Promega), and DNA concentrations were normalized to 2 ng / 100 bp of amplicon length and submitted for Sanger sequencing with the appropriate forward PCR primer. The activity of base editors and nucleases was quantified using the software tools BEAT and Synthego ICE vl.2.0, respectively, with default parameters.
[0265] Characterization of antibodies
[0266] Antibody production Antibody heavy and light chain variable fragment sequences were synthesized as gene fragments and cloned into the pTwist CMV vectors. Nivolumab variants were cloned into IgG4 (heavy chain) and IgK (light chain) vectors. Secukinumab variants were cloned into IgGl (heavy chain) and IgK (light chain) vectors. Antibodies were expressed in ImL HEK293 mammalian cultures and the supernatant was used for binding affinity measurements. Gene synthesis and expression was performed by Twist Bioscience.
[0267] Binding affinity measurements Binding studies were performed in HBSTE running buffer (lOmM HEPES, 150mM NaCl, 3mM EDTA, 0.0% Tween-20) at 25°C. Six-point antigen dilution series were prepared in running buffer starting at 200 nM with a 2.5-fold serial dilution (200-2. OnM). Antibodies were first captured on a Carterra HC30M sensor chip with immobilized goat anti-human Fc pAb. Following 8-10 buffer injection cycles, increasing antigen concentrations were injected over the Ab-captured surfaces with a 5 min of association phase and a 10 min of dissociation phase. The chip surface was finally regenerated and antibodies were recaptured for subsequent antigen binding studies. Double reference subtracted data containing antigen binding at varying concentrations were globally fit using 1 : 1 binding model. Binding affinity measurements and analyses were performed by Twist Bioscience.
[0268] REFERENCES:
[0269] 1. Andrew Leaver-Fay, Michael Tyka, Steven M Lewis, Oliver F Lange, James Thompson, Ron Jacak, Kristian W Kaufman, P Douglas Renfrew, Colin A Smith, Will Sheffler, et al. Rosetta3: an object-oriented software suite for the simulation and design of macromolecules. In Methods in enzymology, volume 487, pages 545-574. Elsevier, 2011.
[0270] 2. John Ingraham, Vikas Garg, Regina Barzilay, and Tommi Jaakkola. Generative models for graph-based protein design. Advances in neural information processing systems, 32, 2019.
[0271] 3. Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learning-based protein sequence design using proteinmpnn. Science, 378(6615):49-56, 2022. 4. Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. In International Conference on Machine Learning, pages 8946-8970. PMLR, 2022.
[0272] 5. BIM Wicky, LF Milles, A Courbet, RJ Ragotte, J Dauparas, E Kinfu, S Tipps, RD Kibler, M Baek, F DiMaio, et al. Hallucinating symmetric protein assemblies. Science, 378(6615):56— 61 , 2022.
[0273] 6. Kiera H Sumida, Reyes Nunez-Franco, Indrek Kalvet, Samuel J Pellock, Basile IM Wicky, Lukas F Milles, Justas Dauparas, Jue Wang, Yakov Kipnis, Noel Jameson, et al. Improving protein expression, stability, and function with proteinmpnn. Journal of the American Chemical Society, 146(3):2054-2061, 2024.
[0274] 7. Florian Praetorius, Philip JY Leung, Maxx H Tessmer, Adam Broerman, Cullen Demakis, Acacia F Dishman, Arvind Pillai, Abbas Idris, David Juergens, Justas Dauparas, et al. Design ofstimulus-responsive two-state hinge proteins. bioRxiv, pages 2023-01, 2023.
[0275] 8. Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne. The protein data bank. Nucleic acids research, 28(1): 235-242, 2000.
[0276] 9. John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figumov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zidek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583-589, 2021.
[0277] 10. Roshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov, and Alexander Rives. Transformer protein language models are unsupervised structure learners. In International Conference on Learning Representations, 2020.
[0278] 11. Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637): 1123-1130, 2023.
[0279] 12. Daniel Hesslow, Niccolo Zanichelli, Pascal Notin, lacopo Poli, and Debora Marks. Rita: a study on scaling up generative protein sequence models. arXiv preprint arXiv:2205.05789, 2022.
[0280] 13. Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models. Cell Systems, 2023.
[0281] 14. Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos Jr, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. Large language models generate functional protein sequences across diverse families. Nature Biotechnology, pages 1-8, 2023.
[0282] 15. Geraldene Munsamy, Ramiro Illanes-Vicioso, Silvia Funcillo, Ioanna T Nakou, Sebastian Lindner, Gavin Ayres, Lesley S Sheehan, Steven Moss, Ulrich Eckhard, Philipp Lorenz, et al. Conditional language models enable the efficient design of proficient enzymes. bioRxiv, pages 2024-05, 2024.
[0283] 16. Jeffrey A Ruffolo, Stephen Nayfach, Joseph Gallagher, Aadyot Bhatnagar, Joel Beazer, Riffat Hussain, Jordan Russ, Jennifer Yip, Emily Hill, Martin Pacesa, et al. Design of highly functional genome editors by modeling the universe of crispr-cas sequences. bioRxiv, pages 2024-04, 2024. 17. Jonas Pfeiffer, Aishwarya Kamath, Andreas Ruckle, Kyunghyun Cho, and Iryna Gurevych. Adapterfusion: Non-destructive task composition for transfer learning. arXiv preprint arXiv:2005.00247, 2020.
[0284] 18. Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Geliy. Parameter-efficient transfer learning for nip. In International Conference on Machine Learning, pages 2790-2799. PMLR, 2019.
[0285] 19. Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
[0286] 20. Nicholas Z Randolph and Brian Kuhlman. Invariant point message passing for protein side chain packing. Proteins: Structure, Function, and Bioinformatics, 2024.
[0287] 21. Christine A Orengo, Alex D Michie, Susan Jones, David T Jones, Mark B Swindells, and Janet M Thornton. Cath-a hierarchic classification of protein domain structures. Structure, 5(8): 1093-1109, 1997.
[0288] 22. Zhangyang Gao, Cheng Tan, and Stan Z Li. Pifold: Toward effective and efficient protein inverse folding. In The Eleventh International Conference on Learning Representations, 2022.
[0289] 23. Justas Dauparas, Gyu Rie Lee, Robert Pecoraro, Linna An, Ivan Anishchenko, Cameron Glasscock, and David Baker. Atomic context-conditioned protein sequence design using ligandmpnn. Biorxiv, pages 2023-12, 2023.
[0290] 24. Lucien Krapp, Femado Meireles, Luciano Abriata, and Matteo Dal Peraro. Context-aware geometric deep learning for protein sequence design. bioRxiv, pages 2023-06, 2023.
[0291] 25. Pascal Notin, Aaron W Kollasch, Daniel Ritter, Lood van Niekerk, Steffanie Paul, Hansen Spinner, Nathan Rollins, Ada Shaw, Ruben Weitzman, Jonathan Frazer, et al. Proteingym: Large-scale benchmarks for protein design and fitness prediction. bioRxiv, pages 2023-12, 2023.
[0292] 26. Nicole M Gaudelli, Alexis C Komor, Holly A Rees, Michael S Packer, Ahmed H Badran, David I Bryson, and David R Liu. Programmable base editing of a* t to g* c in genomic dna without dna cleavage. Nature, 551(7681):464-471, 2017.
[0293] 27. Nicole M Gaudelli, Dieter K Lam, Holly A Rees, Noris M Sola-Esteves, Luis A Barrera, David A Born, Aaron Edwards, Jason M Gehrke, Seung-Joo Lee, Alexander J Liquori, et al. Directed evolution of adenine base editors with increased activity and therapeutic application. Nature biotechnology, 38(7):892-900, 2020.
[0294] 28. Audrone Lapinaite, Gavin J Knott, Cody M Palumbo, Enrique Lin-Shiao, Michelle F Richter, Kevin T Zhao, Peter A Beal, David R Liu, and Jennifer A Doudna. Dna capture by a crispr-cas9- guided adenine base editor. Science, 369(6503):566-571, 2020.
[0295] 29. Brian L Hie, Varun R Shanker, Duo Xu, Theodora UJ Bruun, Payton A Weidenbacher, Shaogeng Tang, Wesley Wu, John E Pak, and Peter S Kim. Efficient evolution of human antibodies from general protein language models. Nature Biotechnology, 2023.
[0296] 30. David Prihoda, Jad Maamary, Andrew Waight, Veronica Juan, Laurence Fayadat-Dilman, Daniel Svozil, and Danny A Bitton. Biophi: a platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning. Tn MAbs, volume 14, page 2020203. Taylor & Francis, 2022.
[0297] 31. James Dunbar, Konrad Krawczyk, Jinwoo Leem, Terry Baker, Angelika Fuchs, Guy Georges, Jiye Shi, and Charlotte M Deane. Sabdab: the structural antibody database. Nucleic acids research, 42 (D1):D1 140-D1146, 2014.
[0298] 32. Richard W Shuai, Jeffrey A Ruffolo, and Jeffrey J Gray. Iglm: Infilling language modeling for antibody sequence design. Cell Systems, 14(11):979— 989, 2023.
[0299] 33. Yeqing Lin, Minji Lee, Zhao Zhang, and Mohammed AlQuraishi. Out of many, one: Designing and scaffolding proteins at the scale of the structural universe with genie 2. arXiv preprint arXiv:2405.15489, 2024.
[0300] 34. Deniz Akpinaroglu, Kosuke Seki, Amy Guo, Eleanor Zhu, Mark JS Kelly, and Tanja Kortemme. Structure-conditioned masked language models for protein sequence design generalize beyond the native sequence space. bioRxiv, pages 2023-12, 2023.
[0301] 35. Martin Steinegger and Johannes Sbding. Clustering huge protein sequence sets in linear time. Nature communications, 9(1): 1-8, 2018.
[0302] 36. Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv: 1412.6980, 2014.
[0303] 37. Zaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou, Fei Ye, and Quanquan Gu. Structure- informed language models are protein designers. bioRxiv, pages 2023-02, 2023.
[0304] 38. Jeffrey M Spencer and Xiaoliu Zhang. Deep mutational scanning of s. pyogenes cas9 reveals important functional domains. Scientific reports, 7(1): 16836, 2017.
[0305] 39. Li Xu, Yakun Liu, and Renzhi Han. Beat: a python program to quantify base editing from sanger sequencing. The CRISPR journal, 2(4):223-229, 2019.
[0306] 40. David Conant, Tim Hsiau, Nicholas Rossi, Jennifer Oki, Travis Maures, Kelsey Waite, Joyce Yang, Sahil Joshi, Reed Kelso, Kevin Holden, et al. Inference of crispr edits from sanger trace data. The CRISPR journal, 5(1): 123-130, 2022.
[0307] 41. Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. In International Conference on Learning Representations, 2020.
Claims
1. CLAIMSWe claim:
1. An engineered antibody comprising: i) a heavy chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 1 and / or a light chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 2; or ii) a heavy chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 3 and / or a light chain comprising one or more substitutions, deletions, or additions as compared to SEQ ID NO: 4.
2. The engineered antibody of claim 1, wherein the engineered antibody comprises a heavy chain having one or more substitutions as compared to SEQ ID NO: 1 at positions selected from:
3. 5, 11, 13, 16, 19, 21, 23, 27, 30, 32, 37, 43, 52, 58, 59, 80, 88, 89, 97, and 99.
3. The engineered antibody of claim 2, wherein the one or more substitutions are selected from: Q3K; V5Q; VI IL; Q13K; R16E, R16T, or R16G; R19K; D21S; K23A, K23T, K23R, or K23V; I27V, I27L, I27F, or I27M; S30N; S32Y, S32F, S32H, S32K, S32R, S32N, or S32L; V37I; K43Q; W52S; R58K, R58T, or R58E; Y59F, Y59N, or Y59H; F80Y or F80H; A88S or A88T; E89D; A97T; N99T, N99E, N99S, N99H, N99D, or N99Y; and combinations thereof.
4. The engineered antibody of claim 2 or 3, wherein the one or more substitutions comprise a substitution at position 32, as compared to SEQ ID NO: 1.
5. The engineered antibody of claim 4, wherein the substitution at position 32 is selected from S32Y, S32F, S32H, S32K, S32R, S32N, or S32L.
6. The engineered antibody of any one of claims 2-5, wherein the one or more substitutions comprises a substitution at one or more or all of positions: 11, 21, 23, 58, and 80, as compared to SEQ ID NO: 1.
7. The engineered antibody of claim 6, wherein the one or more substitutions are selected from VI IL; D21S; K23A, K23T, K23R, or K23V; R58K, R58T, or R58E; F80Y or F80H; and combinations thereof.
8. The engineered antibody of any one of claims 1-7, wherein the engineered antibody comprises a light chain having one or more substitutions as compared to SEQ ID NO: 2 at positions selected from: 1, 3, 4, 9, 10, 14, 28, 29, 30, 31, 58, 60, 70, 77, 81, 90, 91, 93, and 97.
9. The engineered antibody of claim 8, wherein the one or more substitutions are selected from: EID or EIQ; V3L or V3Q; L4M; A9S or A9D; T10S; S14K or S14A; S28T; V29I; S30D, S30H or S30T; S31T, S31I, or S31R; I58V; A60D; D70E; S77R; E81D; Q90H; S91R; N93S, N93A, N93T, N93D, or N93Q; T97L; and combinations thereof.
10. The engineered antibody of claim 8 or 9, wherein the one or more substitutions comprise a substitution at position 93, as compared to SEQ ID NO: 2.
11. The engineered antibody of claim 10, wherein the substitution at position 93 is selected from N93S, N93A, N93T, N93D, or N93Q.
12. The engineered antibody of any one of claims 8-11, wherein the one or more substitutions comprise a substitution at position 1, as compared to SEQ ID NO: 2.
13. The engineered antibody of claim 12, wherein the substitution at position 1 is selected from EID or EIQ.
14. The engineered antibody of any one of claims 1-13, wherein the engineered antibody comprises a heavy chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 5-99 and / or a light chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 100-194.
15. The engineered antibody of any one of claims 1-14, wherein the engineered antibody comprises a heavy chain having an amino acid sequence of any one of SEQ ID NOs: 5-99 and / or a light chain having an amino acid sequence of any one of SEQ ID NOs: 100-194.
16. The engineered antibody of claim 1, wherein the engineered antibody comprises a heavy chain having one or more substitutions as compared to SEQ ID NO: 3 at positions selected from: 1, 2, 3, 5, 13, 28, 31, 33, 35, 37, 43, 49, 50, 52, 53, 54, 57, 58, 61, 62, 85, 88, 97, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, and 123.
17. The engineered antibody of claim 16, wherein the one or more substitutions are selected from: EIQ; V2Q; Q3S, Q3K, or Q3T; V5Q or V5L; Q13K; T28I; N31S, N31D, or N31R; W33Y or W33A; N35S or N35T; V37I; K43Q; A49S; A50N, A50S, A50Y, or A50D; N52K; Q53E, Q53R, Q53T, Q53N, Q53K, or Q53S; D54N; E57N; K58T, K58E, or K58I; V61A; G62D; S85N; V88A; V97A or V97T; D99H, D99E, D99L, D99S, D99R, D99N, D99W, D99K, D99Y, D99Q, or D99A; Y100S, Y100R, Y100K, Y100T, Y100F, Y100A, Y100L, Y100W, Y100E, Y100G, or Y100V; Y101H, Y101F, Y101N, Y101R, or Y101M; D102S, D102T, D102Y, D102W, D102R, D102L, D102E, D102F, D102A, or D102K; I103V, I103S, I103T, 1103 Y, I103L, I103D, I103F, or I103R; L104V, L104A, L104E, L104I, L104Y, L104D, L104T, L104F, L104R, or L104M; T105N, T105D, T105S, T105Y, T105V, T105R, T105E, T105A, T105H, T105W, T105Q, T105I, or T105K; D106N, D106Y, D106E, D106A, D106S, D106T, D106W, D106G, or D106Q; Y107R, Y107L, Y107H, Y107D, Y107T, Y107Q, Y107K, Y107E, Y107A, Y107S, Y107N, Y107M, Y107F, Y107G, or Y107W; Y108L, Y108F, Y108H, Y108D, Y108W, Y108T, Y108G, Y108Q, or Y108M; I109V, I109L, I109Y, I109D, I109T, I109G, or I109Q;Hl ION, Hl 10D, Hl 10W, Hl 10R, Hl 10G, Hl 10Y, Hl 10L, Hl 10A, Hl 10T, or Hl 10Q; Y111H, Y111I, Y111E, Y111L, Y111R, Y111D, Y111G, Y111Q, Y111V, Y111W, Y111M, Y111S, Y11 IF, Y11 IN, or Y11 IT; W112M, W112Y, W112S, W112D, W112Q, W112A, W112N, W1 12T, W112G, W112V, or W112L; Y113F, Y113D, Y113G, Y1131, Y113A, Y113V, Y113Q, Y113T, Y113L, or Y113R; Fl 14Y, Fl 14W, Fl 14V, Fl 14T, Fl 14G, Fl 14M, or Fl 14D; DI 15W, D115M, D115T, or DI 15V; L116Y or LI 16T; W117V or W117T; G118T; R119Q, R119A, R119S, or R119V; G120S; T121S; L122M, L122S, L122Q, L122T, or L122A; V123S; and combinations thereof.
18. The engineered antibody of claim 16 or 17, wherein the one or more substitutions comprise a substitution at one or more positions selected from: 35, 50, 58, 88, 97, 99, 100, 102,103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, and 119, as compared to SEQ ID NO: 3.
19. The engineered antibody of any of claims 16-18, wherein the one or more substitutions comprise a substitution at one or more positions selected from: 35, 50, 88, 97, 100, 102, 103,104, 105, 112, 114, and 119, as compared to SEQ ID NO: 3.
20. The engineered antibody of any one of claims 1 and 16-19, wherein the engineered antibody comprises a light chain having one or more substitutions as compared to SEQ ID NO: 4 at positions selected from: 1, 2, 3, 4, 7, 9, 10, 12, 13, 24, 29, 30, 32, 33, 34, 44, 54, 59, 61, 78, 94, 97, 101, 104, and 105.
21. The engineered antibody of claim 20, wherein the one or more substitutions are selected from: EID; I2V; V3Q; L4M; S7T; G9A, G9S, or G9D; T10S or T10I; S12A; L 13V; R24K; V29I; S30R; S32N, S32R, or S32D; Y33L, Y33N, Y33S, Y33R, Y33F, or Y33H; L34V or L34A; A44S; S54T; I59V; D61A; R78S; S94T; C97T, C97R, C97Y, C97V, C97S, C97I, C97Q, or C97W; Q101G; R104K; LI 05V; and combinations thereof.
22. The engineered antibody of claim 20 or 21, wherein the one or more substitutions comprise a substitution at one or more positions selected from: 1, 2, 9, 10, 32, 33, 34, 44, 97, 104, 105, as compared to SEQ ID NO: 4.
23. The engineered antibody of claim 20 or 21, wherein the one or more substitutions comprise a substitution at one or more positions selected from: 2, 9, 33, 97, 104, and 105, as compared to SEQ ID NO: 4.
24. The engineered antibody of claim 20 or 21, wherein the one or more substitutions comprise a substitution at one or more positions selected from: 1, 9, 10, 32, 33, 34, 44, 97, and 104, as compared to SEQ ID NO: 4.
25. The engineered antibody of any one of claims 16-24, wherein the engineered antibody comprises a heavy chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 195-290 and / or a light chain having an amino acid sequence with at least 90% identity to any one of SEQ ID NOs: 291-386.
26. The engineered antibody of any one of claims 16-25, wherein the engineered antibody comprises a heavy chain having an amino acid sequence of any one of SEQ ID NOs: 195-290 and / or a light chain having an amino acid sequence of any one of SEQ ID NOs: 291-38.
27. The engineered antibody of any one of claims 1-26, wherein the engineered antibody has improved stability, increased binding affinity, reduced immunogenicity, decreasedconformational diversity or flexibility, or a combination thereof as compared to an antibody lacking the one or more substitutions, deletions, or additions.
28. A composition comprising the engineered antibody of any of claims 1-27 and a pharmaceutically acceptable carrier.
29. A nucleic acid encoding the engineered antibody of any of claims 1-27.
30. A method of treating or preventing a disease or disorder in a subject, comprising administering to the subject an effective amount of the engineered antibody of any of claims 1- 27, a composition comprising the engineered antibody, or a nucleic acid encoding the engineered antibody.
31. A method for generating engineered protein sequences as shown and described in FIG. 1.
32. The method of claim 31, comprising a neural network having a series of stacked layers, each layer comprising a self-attention mechanism and feed forward neural network (“FFNN”) followed by a structural adapter.
33. The method of claim 32, wherein the structural adapter incorporates structural and functional context.
34. The method of claim 33, wherein the structural and function context comprises backbone interactions, protein-protein interactions, and interactions with non-protein molecules.
Citation Information
Patent Citations
Human Monoclonal Antibodies To Programmed Death 1(PD-1) And Methods For Treating Cancer Using Anti-PD-1 Antibodies Alone or in Combination with Other Immunotherapeutics
US20090217401A1
Anti-transthyretin antibodies
US20160251418A1