Systems and methods for generating protein variants with target properties
A computational method using predictive models for T-cell and B-cell epitope prediction optimizes protein sequences, addressing deimmunization and stability challenges, resulting in improved protein variants for therapeutic use.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2026-04-09
AI Technical Summary
Current methods for protein therapeutics face challenges in evading antibody recognition, manufacturing at scale, and maintaining stability and function, with existing approaches failing to achieve full deimmunization and practical scalability.
A computational method involving predictive models for T-cell and B-cell epitope prediction, combined with protein design, iteratively samples amino acid sequences to identify mutations that enhance target properties while minimizing immunogenicity and stability issues, using machine learning and evolutionary models to optimize protein sequences.
The method effectively generates protein variants with improved stability, reduced immunogenicity, and enhanced manufacturability, achieving high success rates in identifying functional and stable sequences for therapeutic applications.
Smart Images

Figure US2025049060_09042026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: SES-018WO PATENTSYSTEMS AND METHODS FOR GENERATING PROTEIN VARIANTS WITHTARGET PROPERTIESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 701,872, filed October 1, 2024, and U.S. Provisional Application No. 63 / 744,512, filed January 13, 2025, each of which is hereby incorporated by reference in its entirety.FIELD
[0002] Embodiments provided herein relate to methods and systems for generating protein variants with select target properties.BACKGROUND
[0003] A protein therapeutic must evade recognition by existing antibodies and avoid eliciting new anti-drug antibodies via T cell-dependent B cell activation to be effective. Sufficient modification or mutation to the surface of a drug protein can greatly reduce pre-existing antidrug antibody (pADA) binding. Current efforts at reducing T-cell respons have aimed to eliminate presentation of T-cell epitopes by MHC-II display. This is partly motivated by strong capabilities in measuring and predicting MHC-II display. However, no approach has yet achieved full protein deimmunization at a scale practical for drug development. Another challenge is to manufacture the protein at scale. Barriers to manufacturing include low physical stability, resulting in denaturation or aggregation under temperature or pH stress, and low chemical stability, resulting in post-translational modifications (PTMs) that may obstruct activity. The key drug properties can be improved by predicting chemical and immunogenic liabilities and targeting these for mutation. The fundamental challenge is to identify the mutations that enhance drug properties while avoiding the many mutations that damage stability and function. The embodiments provided for herein fulfill this need as well as others.SUMMARY
[0004] Disclosed herein are predictive models for T-cell epitope prediction, B-cell epitope prediction, and protein design. Embodiments disclosed herein are incorporated by reference into this section. Accordingly, in some embodiments, provided is a computer-implemented method for generating a protein variant amino acid sequence of a target protein having one or more1IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT modified properties, the method comprising: (a) iteratively sampling an input amino acid sequence of the target protein, the iteratively sampling comprising: (i) mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence; (ii) inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two computational models comprise two or more of: a post- translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T- cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model; (iii) generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the at least two computational models that were inputted with the single residue mutant input amino acid sequence; (iv) combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; and (b) sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences to generate a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.
[0005] In some embodiments, provided is a method for identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes, the method comprising: (a) inputting a plurality of subsequences of the target protein into a computational model to determine a plurality of scores representing whether the plurality of subsequences are likely to be presented by a plurality of MHC alleles; (b) for each subsequence, transforming a corresponding score to generate a population-wide measure of antigen presentation; and (c) identifying a subset of the plurality of subsequences as candidate T-cell epitopes based on their population-wide measures of antigen presentation that indicate that the subset of the plurality of subsequences are likely bound by one or more of the plurality of MHC alleles.2IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTBRIEF DESCRIPTION OF THE DRAWINGS
[0006] These and other features, aspects, and advantages of the present disclosure will become better understood with regard to the following description and accompanying drawings.
[0007] FIG. 1 A is an example system overview for predicting target properties of a protein sequence.
[0008] FIG. IB depicts an example block diagram for predicting target properties of a protein sequence.
[0009] FIG. 2 depict a process for predicting target properties of a protein sequence.
[0010] FIGs. 3A-3B depict a process for iteratively sampling an input amino acid sequence of a target protein.
[0011] FIG. 4 illustrates an example computer for implementing the predictive models, methods, and systems disclosed herein.
[0012] FIGs. 5A-5F are a schematic of modular multi-objective design methods for engineering a protein into a potential therapeutic.
[0013] FIGs. 6A-6D show unsupervised sequence models trained on natural repertoires yield functional and stable sequences.
[0014] FIGs. 7A and 7B show multi-objective design of protease results in a high success rate of identifying fully drug design sequences.
[0015] FIGs. 8A-8C show a chematic of protein efficacy, manufacturability, and tolerable immunogenicity.
[0016] FIGs. 9A-9F show machine learning (ML) generated predictive mutations for design drug properties.
[0017] FIGs. 10A-10E show that ML models in combination with structure-based and data- driven rational design can introduce novel activity and specificity.
[0018] FIGs. 11A-11C are a schematic of a general method for enhancing the therapeutic properties of natural proteins. (A) To get a protein drug to the clinic, it must be efficacious, manufacturable, and have tolerable immunogenicity. There are many natural proteins of interest to medicine, but possess post-translational chemical modifications (PTMs), anti-drug antibody (ADA) binding, and T-cell mediated immunogenicity standing in the way of therapeutic value. (B) The solutions provided herein are to develop computational models of each property, where biochemical parameters are used to predict PTMs and ADA probability, and machine learning3IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT models to predict T-cell epitopes and high fitness mutations. (C) Together, the models define the sequence landscape of therapeutic potential.
[0019] FIGs. 12A-12I show that redesigned IdeS has improved drug properties and minimal immunogenicity .
[0020] FIGs. 13A-13F show that accurate T-cell epitope prediction enables deimmunization relying on protein sequence alone.
[0021] FIGs. 14A-14F show that multi-optimization generates drug quality proteins in a single round.
[0022] FIG. 15 depict a process for antibody design.DETAILED DESCRIPTION
[0023] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the ait. The mention of techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the ail. While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.
[0024] As used herein and in the appended claims, the singular forms “a”, “an” and “the” include plural reference unless the context clearly dictates otherwise.
[0025] As used herein, the term “about” means that the numerical value is approximate and small variations would not significantly affect the practice of the disclosed embodiments. Where a numerical limitation is used, unless indicated otherwise by the context, “about” means the numerical value can vary by ±5% and remain within the scope of the disclosed embodiments. Thus, about 100 means 95 to 105.
[0026] It should be understood that the term “at least one of’ includes individually each of the recited objects after the expression and the various combinations of two or more of the recited objects unless otherwise understood from the context and use. The term “and / or” in connection with three or more recited objects should be understood to have the same meaning unless otherwise understood from the context.
[0027] As used herein, the terms “comprising” (and any form of comprising, such as “comprise”, “comprises”, and “comprised”), “having” (and any form of having, such as “have” and “has”),4IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT“including” (and any form of including, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”), arc inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Any composition or method that recites the term “comprising” should also be understood to also describe such compositions as consisting, consisting of, or consisting essentially of the recited components or elements.
[0028] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the present invention remain operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0029] At various places in the present specification, variable or parameters are disclosed in groups or in ranges. It is specifically intended that the description include each and every individual subcombination of the members of such groups and ranges. For example, an integer in the range of 0 to 5 is specifically intended to individually disclose 0, 1, 2, 3, 4, 5, and an integer in the range of 1 to 3 is specifically intended to individually disclose 1, 2, and 3.
[0030] The use of any and all examples, or exemplary language herein, for example, “such as” or “including,” is intended merely to illustrate better the present disclosure and does not pose a limitation on the scope of any invention(s) unless claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of that provided by the present disclosure.
[0031] As used herein, the terms “protein” and “polypeptide” are used interchangeably and generally refer to a macromolecule that includes one or more linked chains of amino acid residues, which can be natural amino acids, unnatural amino acids or both.
[0032] As used herein, the term “epitope” refers to a region of a protein that is specifically recognized by a binding partner, such as an antibody or another binding protein. The epitope may generally span a portion of the protein. Often, proteins may have multiple such regions where binding partners can attach. Epitopes typically fall into two classes: continuous epitopes (also known as linear epitopes), which are epitopes defined by linear sequences of consecutive amino acids, and discontinuous epitopes (also known as conformational epitopes), which are epitopes defined by discontinuous amino acids that are brought together into spatial proximity when a protein is in its folded state.5IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0033] As used herein, the term "paratope" refers to the specific region of a binding molecule (e.g., binding protein) that recognizes and binds an epitope of a target molecule. The paratope typically comprises 5-20 amino acids.
[0034] As used herein, the phrase “structural feature(s)” is generally used in the context of amino acids, e.g., an amino acid residue present in a polypeptide.
[0035] The term “subject” encompasses a cell, tissue, or organism, human or non-human, whether in vivo, ex vivo, or in vitro, male or female.
[0036] The term “mammal” encompasses both humans and non-humans and includes but is not limited to humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.
[0037] The term “sample” can include a single cell or multiple cells or fragments of cells or an aliquot of body fluid, such as a blood sample, taken from a subject, by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, or intervention or other means known in the art. Examples of an aliquot of body fluid include amniotic fluid, aqueous humor, bile, lymph, breast milk, interstitial fluid, blood, blood plasma, cerumen (earwax), Cowper’s fluid (pre-ejaculatory fluid), chyle, chyme, female ejaculate, menses, mucus, saliva, urine, vomit, tears, vaginal lubrication, sweat, serum, semen, sebum, pus, pleural fluid, cerebrospinal fluid, synovial fluid, intracellular fluid, and vitreous humour.
[0038] The terms “marker,” “markers,” “biomarker,” and “biomarkers” encompass, without limitation, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids, genes, and oligonucleotides, together with their related complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other analytes or sample-derived measures. A marker can also include mutated proteins, mutated nucleic acids, variations in copy numbers, and / or transcript variants, in circumstances in which such mutations, variations in copy number and / or transcript valiants are useful for generating a predictive model, or are useful in predictive models developed using related markers (e.g., non-mutated versions of the proteins or nucleic acids, alternative transcripts, etc.).
[0039] The term "antibody" is used in the broadest sense and specifically covers monoclonal antibodies (including full length monoclonal antibodies), polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments that are antigen-binding so long6IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT as they exhibit the desired biological activity, e.g., an antibody or an antigen-binding fragment thereof.
[0040] "Antibody fragment", and all grammatical variants thereof, as used herein are defined as a portion of an intact antibody comprising the antigen binding site or variable region of the intact antibody, wherein the portion is free of the constant heavy chain domains (i.e. CH2, CH3, and CH4, depending on antibody isotype) of the Fc region of the intact antibody. Examples of antibody fragments include Fab, Fab', Fab'-SH, F(ab')2, and Fv fragments; diabodies; any antibody fragment that is a polypeptide having a primary structure consisting of one uninterrupted sequence of contiguous amino acid residues (referred to herein as a "single-chain antibody fragment" or "single chain polypeptide").
[0041] Predictive models, as disclosed herein, are useful for identifying mutations to an amino acid sequence that improve target properties, such as IMM, pADAs, PTMs, FIT, and others.
[0042] Disclosed herein are methods, systems, and non-transitory computer readable media for identifying mutations to an amino acid sequence that improve target properties, such as IMM, pADAs, PTMs, FIT, and others.
[0043] FIG. 1A is an exemplary system overview for identifying mutations to an amino acid sequence of a protein that improve target properties of the protein, such as IMM, pADAs, PTMs, FIT. The system overview includes a prediction system 130 and one or more third parly entities 110A and / or 110B in communication with one another through a network 120. FIG. 1 A depicts one embodiment of the overall system environment. In other embodiments, additional or fewer third party entities 110 in communication with the prediction system 130 can be included. Generally, the prediction system 130 implements methods disclosed herein for identifying mutations to an amino acid sequence of a protein that improve target properties of the protein, such as IMM, pADAs, PTMs, FIT.
[0044] As referenced herein, methods for identifying mutations to an amino acid sequence of a protein that improve target properties of the protein, such as IMM, pADAs, PTMs, FIT include one or more of: iteratively sampling an input amino acid sequence of the target protein; sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences; and generating a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more7IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT amino acid mutations of the single residue mutant input amino acid sequences. Tn some embodiments, methods for identifying mutations to an amino acid sequence of a protein that improve target properties of the protein, such as IMM, pADAs, PTMs, FIT include one or more of: iteratively sampling an input amino acid sequence of the target protein; mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence; inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two trained computational models comprises two or more of: a post-translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T-cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model; generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the trained computational machine learning models that were inputted with the single residue mutant input amino acid sequence; combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences; and generating a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.
[0045] The third party entities 110 communicate with the prediction system 130 for purposes associated with using the characterized one or more structural features of the protein sequence e.g., for rational drug discovery.
[0046] In various embodiments, the third party entity 110 represents a partner entity of the prediction system 130 that operates either upstream or downstream of the prediction system 130. As one example, the third party entity 110 operates upstream of the prediction system 130 and provide information to the prediction system 130 to enable the characterization of one or more structural features of a protein sequence. In this scenario, the prediction system 130 receives data from the third party entity 110, examples of which protein sequence(s) and / or characteristics8IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT of protein sequence(s) (e.g., a protein interface of a protein sequence). The prediction system 130 performs methods disclosed herein to characterize one or more features of the protein sequence.
[0047] Referring to the network 120, any suitable network 120 may be implemented that enables connection between the prediction system 130 and third party entities 110. The network 120 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 120 uses standard communications technologies and / or protocols. For example, the network 120 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of networking protocols used for communicating via the network 704 include multiprotocol label switching (MPLS), transmission control protocol / Intemet protocol (TCP / IP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Data exchanged over the network 704 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML). In some embodiments, all or some of the communication links of the network 120 may be encrypted using any suitable technique or techniques.Predictive Model
[0048] Evolutionary sequence models are powerful tools for predicting the relative fitness of proteins because they are trained on data from large repertoires of natural sequences. These models, regardless of model architecture, number of parameters, and depth of training data, all have one thing in common: they learn features of natural sequences that contribute to the overall fitness of a molecule. This fitness can be reflected in proteins in several ways, including expression or yield, thermostability, and functional activity and can be predictive of disease variants. This is because natural sequences have evolved over time to be stable and functional, often a prerequisite for an organism’s survival. Thus, in some embodiments, a computational model may be trained on evolutionarily related sequence families.
[0049] Machine learning models for the prediction of 3D folds as well as rapid advancements in experimental 3D structure determination has opened up the possibility of using structure-based information to aid in protein engineering. For example, structure models can be used to identify9IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT positions in a structure that are more likely to break the 3D fold and thus should be avoided or to identify positions that arc likely to affect interaction with a binding partner and thus should be sampled with some bias. For the latter example, rational design perspectives can also be coded in to a simplistic function or mask to influence design models to positively or negatively weight specific positions and mutations as desired. Furthermore, as experimental methods for screening constructs continue to improve and data is readily generated, early data points can be used to guide design. For example, experimental screening of point mutants can identify the influence of specific residues or mutations on design goals. If sufficient data are available, a supervised machine learning model may be constructed and directly leveraged within a multi-objective design framework, or employed indirectly by identifying key sequence parameters that may be cast as a set of rules for design considerations. In some embodiments, a structure-based rational design may be used to inform which positions in the enzyme are masked to prevent mutations potentially detrimental to activity.
[0050] In some embodiments, the computational model is an independent position specific scoring matrix model (PSSM). In some embodiments, the computational model an evolutionary couplings based hamiltonian model (EVH). In some embodiments, the computational model a variational auto encoder (VAE) (EVE). In some embodiments, each of the PSSM, EVH, and / or VAE models are trained on an alignment of homologous sequences, but are parametrized in different ways. In some embodiments, the PSSM model learns parameters of each column (position) in the alignment independently by modeling the probability of a sequence x where the energy (log probability) of a sequence is given by the following equation:where hi are the learned column- specific bias terms.
[0051] In some embodiments, the EVH model expands on the PSSM model by also learning pairwise coupling terms, Jij and calculating energy as follows:
[0052] In some embodiments, the EVE model can be used to calculate the energy of a sequence derived from the ELBO as follows:Ffitix) = log^(A ) = Er / [logp(x|z,0)] - O^(2k^)||p(z))10IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0053] In some embodiments, the computational models provided herein may be used to rank amino acid sequences and / or individual mutations based on their predicted scores. In some embodiments, amino acid sequences and / or individual mutations are ranked based on their fitness properties.
[0054] The challenge of protein drug development is to minimize risk-factors such as post- translational modifications (PTMs), pre-existing anti-drug antibody (ADA) binding, and T-cell mediated immunogenicity (IMM), while maintaining essential activity and stability (FIT).
[0055] In some embodiments, computational models may be used to predict PTMs, pADA, IMM, and / or FIT.
[0056] In some embodiments, the computational model takes a target protein amino acid sequence as input and outputs: (1) the positions of predicted liabilities (PTMs, pADAs, and T cell epitopes), and (2) the predicted effects on liabilities and fitness for all possible amino acid substitutions. In some embodiments, mutations that are predicted to both remove / reduce liabilities and to improve or maintain evolutionary fitness are retained as opportunities to improve the drug. In some embodiments, the predictions are quantified by four score functions (IMM, ADA, PTM, FIT), provided a target protein’s sequence x and structure s, which is understood to be computationally predicted from sequence if an experimentally solved structure is unavailable. Jointly, accordingly to the formula: P = (x, s').
[0057] 1.1 FIT P): In some embodiments, evolutionary fitness models are trained on databases of natural protein sequences, and / or on both sequences and structure. In some embodiments, fitness models can be used to generate mutant sequences with predicted improvements in fitness, or to output probability of an input sequence. In some embodiments, the probability of a mutated sequence is predictive of the mutation’s effect on evolutionarily conserved functions (such as enzymatic activity or binding to a target protein) and / or thermostability (as quantified by melting temperature (Tm) or half-life). In some embodiments, prior art in designing drug-like properties of proteins has also made use of fitness models or physics-based score functions, such as but not limited to Rosetta, to quantify these properties. In some embodiments, the fitness score FIT P) is calculated using the log probabilities log p(x) from an EVcouplings model trained on a multiple sequence alignment of sequence homologs to the target protein of interest.
[0058] 1.2 IMM P): Without wishing to be bound by a particular theory, a key factor in immunogenicity of biologies is B-cell activation via T-cell recognition of MHC -bound peptides11IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT derived from the therapeutic protein. Thus, in some embodiments, provided herein is a computational model for predicting the display of T cell epitopes on MHC-II alleles across a population representing a target clinical cohort. In some embodiments, the model uses the netMHCIIpan4.1 model, which distinguishes true T cell epitopes displayed by a given MHC-II allele from random peptides with high accuracy. In some embodiments, to capture epitopes displayed across a given population, sequences are scored in terms of their predicted binding to a set of MHC-II alleles representing broad HL A haplotype diversity. In some embodiments, to create a single score function EPlpop, accounting for the full patient population, a sigmoid transformation is applied to the netMHCIIpan percentile rank score, and average scores for all alleles weighted by their frequency in the world population according to the following formula:
[0059] In some embodiments, to prioritize mutations for deimmunization, the EPIpopscores are summed for all subsequences of the mutant sequence according to the following formula:
[0060] In some embodiments, the IMM computational model yields a single score i MM to identify mutations that eliminate epitope probability across all patient genotypes. In some embodiments, substitutions with negative values within an epitope ( i MM ~ —EPIpOp(.xcore) + 0.05) typically reduce netMHCIIpan score to an insignificant value (%Rank > 20%) simultaneously for all MHC-II genotypes.
[0061] 1.3 ADA(P In some embodiments, ADA is defined as the negative hamming distance (x„rat, xwt) between the parent protein and the variant. In some embodiments, ADA can be modified by predictors of antibody -binding surfaces, or substituted for a model that predicts mutations most likely to escape antibody binding.
[0062] 1.4 PPM(P) In some embodiments, methods for PTMs prediction may analyze a candidate protein for regions that match PTM sequence motifs and exceed solvent accessibility thresholds in the corresponding 3D structure. In some embodiments, the PTM computational model considers deamidation, isomerization, oxidation, DP-clipping, and glycosylation as the target liabilities. In some embodiments, short sequence patterns that are enriched for measured PTM sites are used as PTM motifs. In some embodiments, the resulting score is the tally of all12IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT predicted PTMs throughout the sequence (0 is the solvent-accessibility threshold, and is 8 the Dirac-delta function):
[0063] In some embodiments, an approach based on sequence motifs, solvent accessibility, and secondary structure to identify deamidation, isomerization, oxidation, DP-clipping, and glycosylation sites may be used for PTM prediction. In some embodiments, this information is combined into a score per sequence that counts the total predicted PTM sites as follows:where 0 is the solvent accessibility threshold for the PTM site in a 8sasa Dirac-delta function and X is the secondary structure of a loop region in a 8ss Kronecker-delta function. In some embodiments, a crystal structure of a protein is used to calculate solvent accessibility and secondary structure for chemical liability prediction.
[0064] B-cell epitope prediction: Without wishing to be bound by a particular theory, a preexisting ADA binding may be detrimental to a drug’s efficacy and safety because it can cause the drug to be quickly cleared from the body and induce immune response. In some embodiments, the computational method provided for herein for pADA binding removal may use: a Hamming distance between the starting sequence and design, a measure of the uniformity of the mutations made at the surface of the protein, and / or a count of the number of mutations made at positions that are strongly predicted to belong to immunodominant B-cell epitope hotspots based on a proprietary B-cell epitope predictor.
[0065] Generative Model: In some embodiments, a general computational framework allows any given input protein sequence to be designed for immunogenicity, manufacturability, and fitness simultaneously.
[0066] In some embodiments, a score function that incorporates quantitative predictions of each target property is as according to the following formula:13IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT where SltS2, ■■■ ■ > Snare score functions for each individual target property, and wltw2, . . . , wnare weights quantifying the relative contribution of each property to the overall score. In some embodiments, the score function for four properties is as according to the following formula:
[0067] In some embodiments, the score function may be used to frame the protein design objective as identifying local maxima of S(P . In some embodiments, a random sequence is sampled. In some embodiments, the random sequence is then designed in iterations. In some embodiments, each iteration comprises a single substitution sampled from the score distribution of all possible point mutations such that the sampled substitution improves the score S(P). In some embodiments, a final set of designed sequences predicted to exhibit high fitness, low immunogenicity, few PTMs, and low pADA binding is obtained.
[0068] In some embodiments, the T-cell or B-cell epitope prediction task results in a small number of annotated epitope residues out of hundreds to thousands of residues constituting a given antigen. In some embodiments, not annotated residues may indeed belong to an epitope for an antibody whose interaction with the antigen has yet to be characterized.Training the Predictive Model
[0069] In some embodiments, the predictive model comprises the step of obtaining or having obtained the input amino acid sequence of a candidate protein. In some cmbodimcns, the candidate protein may be any protein, such as any parental protein, anitibody, and the like. In some embodiments, the predictive model comprises the step of obtaining or having obtained the input amino acid sequence of a candidate protein prior to iteratively sampling the input amino acid sequence.
[0070] In some embodiments, the predictive model provided herein is trained on a databased of cell epitopes derived from a database comprising protein, antibody, and / or antigen structures.
[0071] In various embodiments, the predictive model achieves a particular AUC performance metric. In various embodiments, the predictive model achieves an AUC of at least 0.60, at least0.61, at least 0.62, at least 0.63, at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least0.68, at least 0.69, at least 0.70, at least 0.71, at least 0.72, at least 0.73, at least 0.74, at least0.75, at least 0.76, at least 0.77, at least 0.78, at least 0.79, at least 0.80, at least 0.81, at least14IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT0.82, at least 0.83, at least 0.84, at least 0.85, at least 0.86, at least 0.87, at least 0.88, at least 0.89, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, at least 0.96, at least 0.97, at least 0.98, or at least 0.99. In various embodiments, the predictive model achieves an AUC of at least 0.60. In various embodiments, the predictive model achieves an AUC of at least 0.61. In various embodiments, the predictive model achieves an AUC of at least 0.62. In various embodiments, the predictive model achieves an AUC of at least 0.63. In various embodiments, the predictive model achieves an AUC of at least 0.64. In various embodiments, the predictive model achieves an AUC of at least 0.65. In various embodiments, the predictive model achieves an AUC of at least 0.66. In various embodiments, the predictive model achieves an AUC of at least 0.67. In various embodiments, the predictive model achieves an AUC of at least 0.68. In various embodiments, the predictive model achieves an AUC of at least 0.69. In various embodiments, the predictive model achieves an AUC of at least 0.70. In various embodiments, the predictive model achieves an AUC of at least 0.71. In various embodiments, the predictive model achieves an AUC of at least 0.72. In various embodiments, the predictive model achieves an AUC of at least 0.73. In various embodiments, the predictive model achieves an AUC of at least 0.74. In various embodiments, the predictive model achieves an AUC of at least 0.75. In various embodiments, the predictive model achieves an AUC of at least 0.76. In various embodiments, the predictive model achieves an AUC of at least 0.77. In various embodiments, the predictive model achieves an AUC of at least 0.78. In various embodiments, the predictive model achieves an AUC of at least 0.79. In various embodiments, the predictive model achieves an AUC of at least 0.80. In various embodiments, the predictive model achieves an AUC of at least 0.81. In various embodiments, the predictive model achieves an AUC of at least 0.82. In various embodiments, the predictive model achieves an AUC of at least 0.83. In various embodiments, the predictive model achieves an AUC of at least 0.84. In various embodiments, the predictive model achieves an AUC of at least 0.85. In various embodiments, the predictive model achieves an AUC of at least 0.86. In various embodiments, the predictive model achieves an AUC of at least 0.87. In various embodiments, the predictive model achieves an AUC of at least 0.88. In various embodiments, the predictive model achieves an AUC of at least 0.89. In various embodiments, the predictive model achieves an AUC of at least 0.90. In various embodiments, the predictive model achieves an AUC of at least 0.91. In various embodiments, the predictive model achieves an AUC of at least 0.92. In various embodiments,15IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT the predictive model achieves an AUC of at least 0.93. In various embodiments, the predictive model achieves an AUC of at least 0.94. In various embodiments, the predictive model achieves an AUC of at least 0.95. In various embodiments, the predictive model achieves an AUC of at least 0.96. In various embodiments, the predictive model achieves an AUC of at least 0.97. In various embodiments, the predictive model achieves an AUC of at least 0.98. In various embodiments, the predictive module achieves an AUC of at least 0.99.Deploying the Predictive Model
[0072] In some embodiments, the predictive model comprises the step of obtaining or having obtained the input amino acid sequence of a candidate protein. In some embodimens, the candidate protein may be any protein, such as any parental protein, anitibody, and the like. In some embodiments, the predictive model comprises the step of obtaining or having obtained the input amino acid sequence of a candidate protein prior to iteratively sampling the input amino acid sequence.
[0073] Reference is made to FIG. 2, which provides a flowchart with steps for identifying mutations to an amino acid sequence of a protein that improve target properties of the protein, in accordance with an embodiment.
[0074] With reference to FIG. 2, step 160 comprises iteratively sampling an input amino acid sequence of the target protein. As used herein, the term “iteratively” means for every amino acid, every amino acid sequence, or every variant amino acid sequence. As used herein, the term “sampling” means applying repeatedly steps 160 and 170. Thus, “iteratively sampling” as used herein, means that each mutation sampled in each step 160 and 170 is then used as the input to step 160. Therefore, mutations are iteratively layered and updated, with each iteration resulting in a sequence composite of one or more mutations combined from each prior iteration. In some embodiments, the iterations can continue until a specified number of iterations elapse, or can terminate upon a property score or combined score achieving a specified criteria value. In some embodiments, a sequence of interest can be selected from those resulting from any iteration. For example, the sequence of interest can be selected according to the combined score.
[0075] In some embodiments, step 160 comprises step 161 (FIG. 3 A) comprising mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence. In some embodiments, step 160 comprises step 162 (FIG. 3 A) comprising16IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein. In some embodiments, step 160 comprises step 163 (FIG. 3 A) comprising generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the trained computational models that were inputted with the single residue mutant input amino acid sequence. In some embodiments, step 160 comprises step 164 (FIG. 3A) comprising combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence.
[0076] In some embodiments, step 162 (FIG. 3 A) comprises step 162-1 (FIG. 3B) comprising a post-translational modification (PTM) model. In some embodiments, step 162 (FIG. 3A) comprises step 162-2 (FIG. 3B) comprising a pre-existing anti-drug antibody (pADA) model. In some embodiments, step 162 (FIG. 3A) comprises step 162-3 (FIG. 3B) comprising a T-cell mediated immunogenicity (IMM) model. In some embodiments, step 162 (FIG. 3A) comprises step 162-4 (FIG. 3B) comprising an activity and stability (FIT) model. In some embodiments, step 162 (FIG. 3A) comprises step(s) 162-1, 162-2, 162-3, or 162-4, or any combination thereof. In some embodiments, step 162 (FIG. 3A) comprises step(s) 162-1, 162-2, 162-3, and 162-4.
[0077] In some embodiments, step 170 comprises repeating steps 160 through 170 on the single residue mutant input amino acid sequence to generate the protein variant amino acid sequence of the target protein with one or more modified properties with a combined protein score exceeding a set threshold. In some embodiments, the threshold is a score threshold. In some embodiments, the threshold is a target property threshold. In some embodiments, the threshold is at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% improvement in a target property. In some embodiments, the threshold is an iteration threshold. In some embodiments, the iteration threshold is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least17IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT38, at least 39, at least 40, at least 41 , at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, or more iterations.
[0078] In some embodiments, step 170 may also comprise interpolation used to interrogate the latent space of the VAE model. In some embodiments, in order to test the nature of regions of latent space outside of the loci of natural sequences, the encoded Z- vector of the wild-type starting sequence of the protein and the mean Z- vector of the learned VAE distribution may be interpolated. In some embodiments, incremental linear interpolation steps are taken between the start and end points and decode the interpolated vectors through the VAE decoder.
[0079] In refernce to FIG. 15, the design pipeline 500 provided herein uses structure and sequence based modeling to balance objectives of maintaining activity while minimizing the presence of sequences that have higher chances of immunogenicity and developability issues. Thus, in some embodiments, the the design pipeline 500 may be used for antibody design.
[0080] With reference to FIG. 15, step 501 comprises inverse folding and / or generating amino acid sequences from parental antibody structures, controlling for design candidate CDRs.
[0081] In some embodiments, step 501 comprises use of an antibody-specific protein design model, such as, without limitation, FvHallucinator (see Mahajan, S. P., Ruffolo, J. A., Frick, R., & Gray, J. J. (2022). Hallucinating structure-conditioned antibody libraries for target-specific binders. Frontiers in Immunology, 13, 999034. https: / / doi.org / 10.3389 / fimmu.2022.999034). In some embodiments, the antibody- specific protein design model is built on a deep learning model, sue has but not limited to, DeepAb that translates sequence to structure (inverse folding). In some embodiments, the antibody -specific protein design model uses both structure and fixed sequence to create the design subsequence, yielding CDR libraries conditioned on retaining the conformation of the input structure. In some embodiments, step 501 comprises the inputting the structure of the protein backbone and creating a representative graph. In some embodiments, step 501 comprises the use of ProteinMPNN2 (see Dauparas, J., Anishchenko, I., Bennett, N., Bai,18IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTH., Ragotte, R. J., Milles, L. F., Wicky, B. T. M., Courbet, A., de Haas, R. J., Bethel, N., Leung, P. J. Y., Huddy, T. F., Pollock, S., Tischcr, D., Chan, F., Kocpnick, B., Nguyen, H., Kang, A., Sankaran, B., ... Baker, D. (2022). Robust deep learning-based protein sequence design using ProteinMPNN. Science, 378(6615), 49-56. https: / / doi.Org / 10.l 126 / science.add2187). In some embodiments, Step 501 comprises use of any inverse folding model, such as but not limited to, FvHallucinator, ProteinMPNN, ESM-IF1, and the like. In some embodiments, step 501 further comprises use of a GNN to process the structure, capturing local and global relationships between residues. In some embodiments, step 501 further comprises predicting amino acids most likely to respect the backbone structural constraints.
[0082] With reference to FIG. 15, step 502 comprises filtering out chemical liability motifs. In some embodiments, the chemical liability is any chemical liability, such as those provided herein. In some embodiments, the chemical liabilty is a post-translational modification (PTM). Step 503 comprises generating homology model inverse folded sequences. Step 504 comprises predicting B-cell epitope for each design. Step 505 comprises predicting a sequence similarity score for each design. In some embodiment, the sequence similarity score may be predicted using any protein language model, such as but not limited to Progen, or ESM-2. In some embodiments, Progen is an LLM trained on about 280 million protein sequences from more than about 19,000 families (Madani, A., Krause, B., Greene, E. R., Subramanian, S., Mohr, B. P., Holton, J. M., Olmos, J. L., Xiong, C., Sun, Z. Z., Socher, R., Fraser, J. S., & Naik, N. (2023). Large language models generate functional protein sequences across diverse families. Nature Biotechnology, 41(8), 1099-1106. https: / / doi.org / 10.1038 / s41587-022-01618-2).
[0083] With reference to FIG. 15, step 506 comprises combining the sequence similarity score and the B-cell epitope scores into an overall score.Implementation of the Predictive Model for Protease Deimmunization
[0084] The immunoglobulin G-degrading protease from Streptococcus pyogenes, IdeS, is used as a therapeutic for HLA desensitization of patients toward tissue transplants by eliminating anti- HLA antibodies. It has potential to treat a spectrum of diseases mediated by a patient’s antibodies, such as anti-glomerular basement vasculitis, myasthenia gravis, and rheumatoid arthritis. Likewise, IdeS may prove useful in combination with protein or AAV therapies otherwise hindered by pADAs. However, repeat dosing of IdeS is currently restricted due to the19IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT presence of pre-existing anti-IdeS antibodies in patients with prior Strep infections, as well as the development of induced anti-IdeS antibodies within one week of treatment, leading to toxicity concerns and limiting its applicability for many clinical indications.
[0085] In some embodiments, the prediction system 130 may be applied to redesign IdeS for drug development. Thus, in some embodiments, the amino acid sequence of IdeS, a multiple sequence alignment of cysteine proteases with homology to IdeS (for training the EVcouplings model), and a crystal structure of the IdeS monomer to compute solvent accessibility for PTM prediction may be used in step 160. In each screening round of step 160, a library of mutated sequences may be designed by combining mutations predicted by the predictive model to improve target properties, such as those provided herein.
[0086] In some embodiments, steps 170 and 180 comprise a score function that incorporates quantitative predictions of each property according to the formula:where S1, S2, ... Snare score functions for each individual target property, and vq, w2, . . . , wnare weights quantifying the relative contribution of each property to the overall score. In some embodiments, steps 170 and 180 comprise score functions for four properties:S(P) = wfitFIT(P) + wimmIMM(P) + wptmPTM(P) + wadaADA(P
[0087] Thus, in some embodiments, a random sequence is sampled, and then designed in iterations. In each iteration of step 160, a single substitution is sampled from the score distribution of all possible point mutations such that the sampled substitution improves the score S(P). In some embodiments, by iterating this process, a final set is obtained of designed sequences predicted to exhibit target properties, such as high fitness, low immunogenicity, few PTMs, and low pADA binding.
[0088] In some embodiments, sampling of step 170 produces a combined protein score with improved scores of at least three or all of PTM, pADA, IMM, and FIT.
[0089] In some embodiments, the improved score for PTM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in PTMs. In some embodiments, selecting for improved PTM results in at most few PTM sites in the protein variant. In some embodiments, selecting for PTM results in no identified PTMs in the protein variant.20IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0090] In some embodiments, selecting for pADA binding results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in predicted pADA binding. In some embodiments, selecting for pADA results in a prediction of low pADA binding. In some embodiments, selecting for pADA results in a prediction of no pADA binding.
[0091] In some embodiments, selecting for IMM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in immunogenicity. In some embodiments, selecting for IMM results in predicted low immunogenicity. In some embodiments, selecting for IMM results in a prediction of no immunogenicity. In some embodiments, selecting for FIT results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% increase in FIT. In some embodiments, selecting for FIT improves predicted FIT. In some embodiments, selecting for FIT results in high FIT.
[0092] In some embodiments, the PTM model comprises instructions for searching the single residue mutant input amino acid sequence for regions that match PTM motifs and exceed solvent accessibility thresholds in a corresponding 3D structure of the protein. In some embodiments, the PTM motifs comprise short sequence patterns that are enriched for measured PTM sites. In some embodiments, the PTM motifs comprise deamidation sites (e.g., N[DNPTGSC]), DP clipping (DP), N-glycosylation sites (e.g., N[AP][ST]), integrin binding motif (RGD), oxidation sites (MW), and isomerization sites (e.g., D[DCSAG]). In some embodiments, the PTM model generates a PTM score comprising a tally of all predicted PTMs throughout the single residue mutant input amino acid sequence.
[0093] In some embodiments, the pADA model comprises generating a negative hamming distance score between the single residue mutant input amino acid sequence and the input amino acid sequence.
[0094] In some embodiments, the IMM model comprises identifying one or more sets of amino acids representing binding cores of the input amino acid sequence; determining a set of MHC alleles representative of MHC alleles of a human population; determining scores for pairs of the one or more sets of amino acids and MHC alleles; combining the determined scores to generate a combined score for each set of amino acids representing a binding core; and selecting one or more sets of amino acids representing a binding core as T-cell epitopes.21IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0095] In some embodiments, the selected one or more sets of amino acids have higher combined scores compared to non-sclcctcd one or more sets of amino acids.
[0096] In some embodiments, the MHC alleles comprise class II MHC alleles. In some embodiments, the MHC alleles comprise class I MHC alleles. In some embodiments, the MHC alleles are selected from DR1 / 15, DR7, DR4, DR15, DR13, DR8 / 11, DR8, DR3, or any combination thereof. In some embodiments, the MHC alleles are selected from DRB 1*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, or any combination thereof. In some embodiments, the MHC alleles are selected from HLA-DPAl*02:02 / DPBl*26:01, HLA-DPAl*02:02 / DPB 1*13:01, HLA- DPA 1 *02:02 / DPB 1 *03 :01 , HLA-DPA 1 *01 :03 / DPB 1 *03 :01 , HLA-DPA 1 *01 :03 / DPB 1*11:01, HLA-DPAl*02:02 / DPB 1*09:01, HLA-DPAl*01:03 / DPBl*17:01, HLA- DPA1*O2:O1 / DPB1*18:O1, HLA-DPAl*01:03 / DPBl*18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPA l*02:02 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*02:01, HLA- DPAl*01:03 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA-DPA l*01:03 / DPB 1*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA- DPA1*O1:O3 / DPB1*19:O1, HLA-DPA l*01:03 / DPB 1*05:01, or any combination threreof. In some embodiments, the MHC alleles are selected from HLA-DQA1 *05 :01 / DQB 1*04:02, HLA- DQAl*01:02 / DQB 1*05:01, HLA-DQA 1 *05 :01 / DQB 1*05:01, HLA- DQA 1 *05 :01 / DQB 1*02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA- DQA 1*01 :02 / DQB 1*03:01, HLA-DQAl*01:02 / DQBl*02:01, HLA- DQAl*01:03 / DQBl*06:03, HLA-DQAl*01:02 / DQBl*06:02, HLA-DQA 1*05:01 / DQB 1*03:01, HLA-DQAl*05:01 / DQBl*03:03, or any combination thereof. In some embodiments, the MHC alleles are selected from DRBl*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6,22IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTDRB 1 * 13:02, DRB3*01 :01 , DRB3*02:02, DRB3*03:01 , DRB4*01 :01 , DRB4*01 :03, DRB5*01:01, HLA-DPA l*02:02 / DPB 1*26:01, HLA-DPAl*02:02 / DPBl*13:01, HLA- DPA 1 *02:02 / DPB 1 *03:01 , HLA-DPA 1 *01 :03 / DPB 1 *03:01 , HLA-DPA 1 *01 :03 / DPB 1*11:01, HLA-DPAl*02:02 / DPB 1*09:01, HLA-DPAl*01:03 / DPBl*17:01, HLA-DPAl*02:01 / DPBl*18:01, HLA-DPAl*01:03 / DPBl*18:01, HLA-DPAl*02:02 / DPBl*15:01,HLA-DPA l*02:02 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*02:01, HLA- DPAl*01:03 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA-DPA l*01:03 / DPB 1*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPAl*01:03 / DPBl*19:01, HLA-DPAl*01:03 / DPBl*05:01, HLA-DQAl*05:01 / DQBl*04:02, HLA-DQA1*O1:O2 / DQB1*O5:O1, HLA-DQA1*O5:O1 / DQB1*O5:O1, HLA-DQA 1 *05 :01 / DQB 1*02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA- DQAl*01:02 / DQB 1*03:01, HLA-DQAl*01:02 / DQB 1*02:01, HLA- DQAl*01:03 / DQBl*06:03, HLA-DQA1 *01 :02 / DQB 1*06:02, HLA- DQA1*O5:O1 / DQB1*O3:O1, HLA-DQAl*05:01 / DQBl*03:03, or any combination thereof.
[0097] In some embodiments, the IMM model comprises a NetMHC model, or a variant thereof. In some embodiments, the FIT model comprises an EVcouplings model, or a variant thereof. In some embodiments, the FIT model generates a fitness score using log probabilities of improved fitness.
[0098] In some embodiments, the improved fitness comprises improvement to enzymatic activity, binding to a target protein, thermostability, melting temperature, and / or half-life.Implementation of the Predictive Model for Identifying One or More T-cell Epitopes
[0099] In some embodiments, the prediction system 130 may be applied to identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes. In some embodiment, the methds comprises the steps of: inputting a plurality of subsequences of the target protein into a computational model to determine a plurality of scores representing whether the plurality of subsequences are likely to be presented by a plurality of MHC alleles; for each subsequence, transforming a corresponding score to generate a population-wide measure of antigen presentation; and identifying a subset of the plurality of subsequences as candidate T-cell epitopes based on their population-wide measures of antigen presentation that indicate that the subset of the plurality of subsequences are likely bound by one23IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT or more of the plurality of MHC alleles. In some embodiments, the method further comprises the steps of introducing one or more amino acid mutations into a candidate T-cell epitope of the input amino acid sequence to generate a variant amino acid sequence; and evaluating the impact of the introduced one or more amino acid mutations on T-cell IMM of a protein therapeutic comprising the variant amino acid sequence in comparison to the protein comprising the input amino acid sequence.
[0100] In some embodiment, the methods disclosed herein comprise the steps of: inputting a plurality of subsequences of the target protein into a computational model to determine a plurality of scores representing whether the plurality of subsequences are likely to be presented by a plurality of MHC alleles; for each subsequence, transforming a corresponding score to generate a population-wide measure of antigen presentation; identifying a subset of the plurality of subsequences as candidate T-cell epitopes based on their population-wide measures of antigen presentation that indicate that the subset of the plurality of subsequences are likely bound by one or more of the plurality of MHC alleles; introducing one or more amino acid mutations into a candidate T-cell epitope of the input amino acid sequence to generate a variant amino acid sequence; and evaluating the impact of the introduced one or more amino acid mutations on T- cell IMM of a protein therapeutic comprising the variant amino acid sequence in comparison to the protein comprising the input amino acid sequence.
[0101] In some embodiments, the computational model is the pan-specific binding of peptides to MHC class I proteins of known sequence (NetMHC pan model).
[0102] In some embodiments, the step of evaluating the impact of the introduced one or more amino acid mutations on T-cell IMM of a protein therapeutic comprising the variant amino acid sequence in comparison to the protein comprising the input amino acid sequence comprises summating population-wide measures of presentation of a plurality of subsequences of the variant amino acid sequence to generate a IMM score for the variant amino acid; and comparing the IMM score for the variant amino acid to a IMM score of the input amino acid sequence of the protein.
[0103] In some embodiments, a larger difference between the IMM score for the variant amino acid and the IMM score of the input amino acid sequence of the protein is indicative of a larger impact of the introduced one or more amino acid mutations.24IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0104] Tn some embodiments, provided herein are methods for identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes presentation by class I MHC. In some embodiments, the MHC alleles comprise class I MHC alleles. In some embodiments, the plurality of MHC alleles comprise class MHC-I alleles. Designing protein sequences for reduced T-cell epitope presentation by class I MHC, particularly in context of gene delivery or intracellular drug delivery.
[0105] In some embodiments, provided herein are methods for identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes presentation by class I MHC for gene delivery or intracellular drug delivery. In some embodiments, the MHC alleles comprise class I MHC alleles. In some embodiments, the plurality of MHC alleles comprise class MHC-I alleles.
[0106] In some embodiments, provided herein are methods for identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes presentation by class I MHC for development of a vaccine. In some embodiments, the MHC alleles comprise class I MHC alleles. In some embodiments, the plurality of MHC alleles comprise class MHC-I alleles. In some embodiments, selecting for IMM results in predicted high immunization.
[0107] In some embodiments, the input amino acid sequence is greater than 30 amino acids in length, and wherein the plurality of subsequences is less than 20 amino acids in length. In some embodiments, the plurality of subsequences are about 15 amino acids in length. In some embodiments, the plurality of MHC alleles is a population- wide representation of MHC alleles. In some embodiments, the population-wide representation of MHC alleles comprises at least 50 different MHC alleles. In some embodiments, the plurality of MHC alleles comprise class MHC- II alleles. In some embodiments, transforming a corresponding score to generate a populationwide measure of presentation comprises weighting the corresponding score by population-wide frequencies of the plurality of MHC alleles. In some embodiments, the MHC alleles are selected from DR1 / 15, DR7, DR4, DR15, DR13, DR8 / 11, DR8, DR3, or any combination thereof. In some embodiments, the MHC alleles are selected from DRB 1*04:03, DRB 1*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01,25IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTDRB1* 13:03, DRB 1 * 14:07, DRB1 * 14:01 , DRB1 * 13:01 , DRB1 * 11 :04, DRB1 * 14:06, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, or any combination thereof. In some embodiments, the MHC alleles are selected from HLA-DPAl*02:02 / DPB 1*26:01, HLA-DPAl*02:02 / DPB 1*13:01, HLA- DPAl*02:02 / DPBl*03:01, HLA-DPAl*01:03 / DPBl*03:01, HLA-DPAl*01:03 / DPBl*l l:01, HLA-DPA l*02:02 / DPB 1*09:01, HLA-DPAl*01:03 / DPBl*17:01, HLA-DPA 1 *02:01 / DPB 1*18:01, HLA-DPA 1 *01 :03 / DPB 1*18:01, HLA-DPA 1 *02:02 / DPB 1*15:01, HLA-DPAl*02:02 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*02:01, HLA-DPA 1*01 :03 / DPB 1*04:01, HLA-DPA l*02:02 / DPB 1*04:01, HLA-DPA l*02:02 / DPB 1*01:01, HLA-DPA l*01:03 / DPB 1*01:01, HLA-DPA 1 *02 :02 / DPB 1*05:01, HLA- DPA1*O1:O3 / DPB1*19:O1, HLA-DPA l*01:03 / DPB 1*05:01, or any combination threreof. In some embodiments, the MHC alleles are selected from HLA-DQAl*05:01 / DQBl*04:02, HLA- DQAl*01:02 / DQB 1*05:01, HLA-DQAl*05:01 / DQBl*05:01, HLA- DQA 1 *05 :01 / DQB 1*02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA- DQAl*01:02 / DQB 1*03:01, HLA-DQAl*01:02 / DQBl*02:01, HLA- DQAl*01:03 / DQBl*06:03, HLA-DQA1 *01 :02 / DQB 1*06:02, HLA-DQA 1 *05 :01 / DQB 1*03:01, HLA-DQA1*O5:O1 / DQB1*O3:O3, or any combination thereof. In some embodiments, the MHC alleles are selected from DRBl*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, HLA-DPA l*02:02 / DPB 1*26:01, HLA-DPAl*02:02 / DPBl*13:01, HLA- DPA 1 *02:02 / DPB 1 *03 :01 , HLA-DPA 1 *01 :03 / DPB 1 *03 :01 , HLA-DPA 1 *01 :03 / DPB 1*11:01, HLA-DPAl*02:02 / DPBl*09:01, HLA-DPAl*01:03 / DPBl*17:01, HLA- DPAl*02:01 / DPBl*18:01, HLA-DPAl*01:03 / DPBl*18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPA l*02:02 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*02:01, HLA-DPA 1*01 :03 / DPB 1*04:01, HLA-DPA l*02:02 / DPB 1*04:01, HLA-DPA l*02:02 / DPB 1*01:01, HLA-DPA l*01:03 / DPB 1*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPA 1 *01 :03 / DPB 1*19:01, HLA-DPA 1 *01 :03 / DPB 1 *05 :01 , HLA-DQA 1 *05 :01 / DQB 1 *04:02,26IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTHLA-DQ A 1 *OEO2 / DQB1 *05 :01 , HLA-DQ A 1 *05:01 / DQB 1 *05:01 , HLA- DQAl*05:01 / DQBl*02:01, HLA-DQA1*OEO2 / DQB 1*06:04, HLA- DQAl*01:02 / DQB 1*03:01, HLA-DQA1*OEO2 / DQB 1*02:01, HLA- DQA1*OEO3 / DQB1*O6:O3, HLA-DQA1*OEO2 / DQB1*O6:O2, HLA- DQA 1 *05 :01 / DQB 1*03:01, HLA-DQAl*05:01 / DQBl*03:03, or any combination thereofComputer Implementation
[0108] The methods disclosed herein, such as the methods of identifying mutations to an amino acid sequence of a protein that improve target properties of the protein, are, in some embodiments, performed on one or more computers. For example, the building and deployment of a predictive model to analyze protein structural data, and database storage can be implemented in hardware or software, or a combination of both. In one embodiment, a machine-readable storage medium is provided, the medium comprising a data storage material encoded with machine readable data which, when using a machine programmed with instructions for using said data, is capable of displaying any of the datasets and execution and results of a predictive model. Such data can be used for a variety of purposes, such as patient monitoring, treatment considerations, and the like. Methods disclosed herein can be implemented in computer programs executing on programmable computers, comprising a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), a graphics adapter, a pointing device, a network adapter, at least one input device, and at least one output device. Program code may be applied to input data to perform the functions described above and generate output information. The output information is applied to one or more output devices, in known fashion. The computer can be, for example, a personal computer, microcomputer, or workstation of conventional design.
[0109] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also27IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[0110] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern information. The databases as described herein can be recorded on computer readable media, e.g. any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the ait can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a recording of the present database information. "Recorded" refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g. word processing text file, database format, etc.
[0111] FIG. 4 illustrates an example computer 600 for implementing the predictive models, methods, systems, and data described herein. The computer 600 includes at least one processor 602 coupled to a chipset 604. The chipset 604 includes a memory controller hub 620 and an input / output (VO) controller hub 622. A memory 606 and a graphics adapter 612 are coupled to the memory controller hub 620, and a display 618 is coupled to the graphics adapter 612. A storage device 608, an input device 614, and network adapter 616 are coupled to the VO controller hub 622. Other embodiments of the computer 600 have different architectures.
[0112] The storage device 608 is a non-transitory computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 506 holds instructions and data used by the processor 602. The input device 614 is a touch-screen interface, a mouse, track ball, or other type of pointing device, a keyboard, or some combination thereof, and is used to input data into the computer 600. In some embodiments, the computer 600 may be configured to receive input (e.g., commands) from the input device 614 via gestures from the user. The graphics adapter 612 displays images and other28IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT information on the display 618. The network adapter 616 couples the computer 600 to one or more computer networks.
[0113] The computer 600 is adapted to execute computer program modules for providing functionality described herein. As used herein, the term “module” refers to computer program logic used to provide the specified functionality. Thus, a module can be implemented in hardware, firmware, and / or software. In one embodiment, program modules are stored on the storage device 608, loaded into the memory 606, and executed by the processor 602.
[0114] The types of computers 600 can vary depending upon the embodiment and the processing power required by the entity. For example, the can run in a single computer 600 or multiple computers 600 communicating with each other through a network such as in a server farm. The computers 600 can lack some of the components described above, such as graphics adapters 612, and displays 618.
[0115] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[0116] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern information of the present invention. The databases of the present invention can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the ail can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a29IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT recording of the present database information. "Recorded" refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g., word processing text file, database format, etc.
[0117] In some embodiments, provided herein is a system comprising a non-transitory computer-readable storage medium and a processor, wherein the non-transitory computer- readable storage medium comprises: (a) iteratively sampling an input amino acid sequence of the target protein, the iteratively sampling comprising: (i) mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence; (ii) inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two trained computational models comprises two or more of: a post- translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T- cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model; (iii) generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the trained computational models that were inputted with the single residue mutant input amino acid sequence; (iv) combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; and (b) sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences to generate a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.ENUMERATED EMBODIMENTS1. A computer-implemented method for generating a protein variant amino acid sequence of a target protein having one or more modified properties, the method comprising:30IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT(a) iteratively sampling an input amino acid sequence of the target protein, the iteratively sampling comprising:(i) mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence;(ii) inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two trained computational models comprises two or more of: a post-translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T-cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model;(iii) generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the trained computational models that were inputted with the single residue mutant input amino acid sequence;(iv) combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; and(b) sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences to generate a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.2. The method of embodiment 1, wherein sampling of step (b) comprising repeating steps (a)-(b) on the single residue mutant input amino acid sequence to generate the protein variant amino acid sequence of the target protein with one or more modified properties with a combined protein score exceeding a set threshold.31IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT3. The method of embodiment 2, wherein the threshold is an improvement in target property, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% improvement in a target property.4. The method of embodiment 2, wherein the threshold is and iteration threshold, such as at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, or more iterations.5. The method of embodiments 1 or 2, the method further comprising obtaining or having obtained the input amino acid sequence of the candidate protein prior to iteratively sampling the input amino acid sequence.6. The method of any one of embodiments 1-5, wherein the one or more amino acid mutations that, when included in the input amino acid sequence, improves at least one target property of the candidate protein7. The method of any one of embodiments 1-6, wherein (iv) combining the at least one weighted relative contribution of the single residue mutant to the at least one target property to generate an individual protein score comprises determining the weighted summation of the at least one weighted relative contribution of the single residue mutant to the at least one target property.8. The method of any one of embodiments 1-7, wherein the (b) sampling the at least one weighted relative contribution of the single residue mutant comprises sampling at least one weighted relative contribution of at least one single residue mutant.32IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT9. The method of any one of embodiments 1 -8, wherein step the (b) sampling produces a combined protein score with improved scores of at least three or all of PTM, pADA, IMM, and FIT.10. The method of embodiment 9, wherein the improved score for PTM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in PTMs.11. The method of embodiments 9 or 10, wherein selecting for improved PTM results in at most few PTM sites in the protein variant.12. The method of embodiment 11, wherein selecting for PTM results in no identified PTMs in the protein variant.13. The method of embodiment 9, wherein selecting for pADA binding results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in predicted pADA binding.14. The method of embodiment 13, wherein selecting for pADA results in a prediction of low pADA binding.15. The method of embodiment 14, wherein selecting for pADA results in a prediction of no pADA binding.16. The method of embodiment 9, wherein selecting for IMM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in immunogenicity.17. The method of embodiment 16, wherein selecting for IMM results in predicted low immunogenicity.18. The method of embodiment 16, wherein selecting for IMM results in a prediction of no immunogenicity .19. The method of embodiment 9, wherein selecting for FIT results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% increase in FIT.20. The method of embodiment 19, wherein selecting for FIT improves predicted FIT.21. The method of embodiment 19, wherein selecting for FIT results in high FIT.33IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT22. The method of embodiment 1 , wherein the PTM model comprises instructions for searching the single residue mutant input amino acid sequence for regions that match PTM motifs and exceed solvent accessibility thresholds in a corresponding 3D structure of the protein.23. The method of embodiment 22, wherein the PTM motifs comprise short sequence patterns that are enriched for measured PTM sites.24. The method of embodiment 22, wherein the PTM motifs comprise deamidation sites (e.g., N[DNPTGSC]), DP clipping (DP), N-glycosylation sites (e.g., N[AP][ST]), integrin binding motif (RGD), oxidation sites (MW), and isomerization sites (e.g., D[DCSAG]).25. The method of any one of embodiments 22-24, wherein the PTM model generates a PTM score comprising a tally of all predicted PTMs throughout the single residue mutant input amino acid sequence.26. The method of embodiment 1, wherein the pADA model comprises generating a negative hamming distance score between the single residue mutant input amino acid sequence and the input amino acid sequence.27. The method of embodiment 1, wherein the IMM model comprises: identifying one or more sets of amino acids representing binding cores of the input amino acid sequence; determining a set of MHC alleles representative of MHC alleles of a human population; determining scores for pairs of the one or more sets of amino acids and MHC alleles; combining the determined scores to generate a combined score for each set of amino acids representing a binding core; and selecting one or more sets of amino acids representing a binding core as T-cell epitopes.28. The method of embodiment 27, wherein the selected one or more sets of amino acids have higher combined scores compared to non-selected one or more sets of amino acids.29. The method of embodiment 27, wherein the MHC alleles comprise class II MHC alleles or class I MHC alleles.30. The method of embodiment 27, wherein the MHC alleles are selected from DRB 1*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4,34IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTDRB 1 * 14:06, DRB 1 * 13:02, DRB3*01 :01 , DRB3*02:02, DRB3*03:01 , DRB4*01 :01 ,DRB4*01:03, DRB5*01:01, HLA-DPAl*02:02 / DPBl*26:01, HLA-DPAl*02:02 / DPBl*13:01, HLA-DPAl*02:02 / DPBl*03:01, HLA-DPA l*01:03 / DPB 1*03:01, HLA-DPA 1 *01 :03 / DPB 1*11:01, HLA-DPA 1 *02:02 / DPB 1 *09:01 , HLA-DPA 1 *01 :03 / DPB 1*17:01, HLA-DPA1*O2:O1 / DPB1*18:O1, HLA-DPA1 *01 :03 / DPB 1*18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPAl*02:02 / DPBl*19:01, HLA-DPAl*01:03 / DPBl*02:01,HLA-DPA l*01:03 / DPB 1*04:01, HLA-DPA 1 *02 :02 / DPB 1*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA-DPAl*01:03 / DPBl*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPA 1 *01 :03 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*05:01, HLA-DQA1 *05 :01 / DQB 1*04:02, HLA-DQA1*O1:O2 / DQB1*O5:O1, HLA-DQA 1 *05 :01 / DQB 1*05:01, HLA-DQA1 *05 :01 / DQB 1*02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA-DQAl*01:02 / DQB 1*03:01, HLA-DQAl*01:02 / DQBl*02:01, HLA-DQA1 *01 :03 / DQB 1*06:03, HLA- DQAl*01:02 / DQBl*06:02, HLA-DQA 1 *05 :01 / DQB 1*03:01, HLA- DQA 1 *05 :01 / DQB 1*03:03, or any combination thereof.31. The method of embodiment 1, wherein the IMM model comprises a NetMHC model.32. The method of embodiment 1, wherein the FIT model comprises an EVcouplings model.33. The method of embodiment 32, wherein the FIT model generates a fitness score using log probabilities of improved fitness.34. The method of embodiment 33, wherein the improved fitness comprises improvement to enzymatic activity, binding to a target protein, thermostability, melting temperature, and / or halflife.35. The method of any one of embodiments 1-34, wherein the input amino acid sequence of the target protein is a wild-type sequence.36. The method of any one of embodiments 1-35, wherein the target protein is a bacterial enzyme.37. The method of embodiment 36, wherein the bacterial enzyme is a protease or a glycosidase.38. The method of embodiment 37, wherein the protease is an immunoglobulin-degrading protease, such as IgG, IgE, or an IgM protease.35IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT39. The method of embodiment 38, wherein the protease is selected from deS, IdeSsuis, IdcZ, IdcE, IdcE2, IdcZ2, Idc85, or IdcC, or a variant thereof.40. The method of any one of embodiments 1-39, wherein the (i) mutating an amino acid of the input amino acid sequence to generate a single residue mutant input amino acid sequence comprises introducing any of a substitution, insertion, or deletion to the input amino acid sequence.41. A method for identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes, the method comprising:(a) inputting a plurality of subsequences of the target protein into a computational model to determine a plurality of scores representing whether the plurality of subsequences are likely to be presented by a plurality of MHC alleles;(b) for each subsequence, transforming a corresponding score to generate a populationwide measure of antigen presentation; and(c) identifying a subset of the plurality of subsequences as candidate T-cell epitopes based on their population-wide measures of antigen presentation that indicate that the subset of the plurality of subsequences are likely bound by one or more of the plurality of MHC alleles.42. The method of embodiment 41, the method further comprising:(d) introducing one or more amino acid mutations into a candidate T-cell epitope of the input amino acid sequence to generate a variant amino acid sequence; and(e) evaluating the impact of the introduced one or more amino acid mutations on T-cell IMM of a protein therapeutic comprising the variant amino acid sequence in comparison to the protein comprising the input amino acid sequence.43. The method of embodiment 42, wherein step (e) comprises: summating population-wide measures of presentation of a plurality of subsequences of the variant amino acid sequence to generate a IMM score for the variant amino acid; and comparing the IMM score for the valiant amino acid to a IMM score of the input amino acid sequence of the protein.44. The method of embodiment 43, wherein a larger difference between the IMM score for the variant amino acid and the IMM score of the input amino acid sequence of the protein is indicative of a larger impact of the introduced one or more amino acid mutations.36IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT45. The method of embodiment 41 or 42, wherein the computational model comprises a NctMHC pan model.46. The method of any one of embodiments 41-45, wherein the input amino acid sequence is greater than 30 amino acids in length, and wherein the plurality of subsequences is less than 20 amino acids in length.47. The method of embodiment 46, wherein the plurality of subsequences are about 15 amino acids in length.48. The method of any one of embodiments 41-47, wherein the plurality of MHC alleles is a population-wide representation of MHC alleles.49. The method of embodiment 48, wherein the population- wide representation of MHC alleles comprises at least 50 different MHC alleles.50. The method of any one of embodiments 41-49, wherein the plurality of MHC alleles comprise class II MHC alleles or class I MHC alleles.51. The method of any one of embodiments 41-50, wherein transforming a corresponding score to generate a population- wide measure of presentation comprises weighting the corresponding score by population- wide frequencies of the plurality of MHC alleles.52. The method of embodiment 41, wherein the MHC alleles are selected from DRB 1*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, HLA-DPAl*02:02 / DPBl*26:01, HLA-DPAl*02:02 / DPB 1*13:01, HLA-DPAl*02:02 / DPBl*03:01, HLA-DPAl*01:03 / DPB 1*03:01, HLA-DPA 1 *01 :03 / DPB 1*11:01, HLA-DPA 1 *02:02 / DPB 1 *09:01 , HLA-DPA 1 *01 :03 / DPB 1*17:01, HLA-DPA1*O2:O1 / DPB1*18:O1, HLA-DPA l*01:03 / DPBl* 18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPAl*02:02 / DPBl*19:01, HLA-DPAl*01:03 / DPBl*02:01, HLA-DPA l*01:03 / DPB 1*04:01, HLA-DPAl*02:02 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA-DPAl*01:03 / DPBl*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPA 1 *01 :03 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*05:01, HLA-DQA1 *05 :01 / DQB 1*04:02, HLA-DQAl*01:02 / DQB 1*05:01, HLA-37IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTDQA 1*05:01 / DQB 1*05:01, HLA-DQ A 1 *05:01 / DQB 1 *02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA-DQA1*O1:O2 / DQB1*O3:O1, HLA-DQAl*01:02 / DQB 1*02:01, HLA-DQAl*01:03 / DQBl*06:03, HLA- DQAl*01:02 / DQBl*06:02, HLA-DQA 1*05:01 / DQB 1*03:01, HLA- DQA 1*05:01 / DQB 1*03:03, or any combination thereof.53. A system for designing a sequence of a protein, the system comprising: at least one computer processor; and at least one non-transitory computer-readable storage medium storing processorexecutable instructions, that when executed by the at least one computer processor, cause the at least one computer processor to perform:(a) iteratively sampling an input amino acid sequence of the target protein, the iteratively sampling comprising:(i) mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence;(ii) inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two trained computational models comprises two or more of: a post-translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T-cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model;(iii) generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the trained computational models that were inputted with the single residue mutant input amino acid sequence;(iv) combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; and38IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT(b) sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences to generate a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.54. The system of any one of the preceding embodiments, wherein sampling of step (b) comprising repeating steps (a)-(b) on the single residue mutant input amino acid sequence to generate the protein variant amino acid sequence of the target protein with one or more modified properties with a combined protein score exceeding a set threshold.55. The system of any one of the preceding embodiments, wherein the threshold is an improvement in target property, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% improvement in a target property.56. The system of any one of the preceding embodiments, wherein the threshold is and iteration threshold, such as at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, or more iterations.57. The system of any one of the preceding embodiments, the method further comprising obtaining or having obtained the input amino acid sequence of the candidate protein prior to iteratively sampling the input amino acid sequence.39IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT58. The system of any one of the preceding embodiments, wherein the one or more amino acid mutations that, when included in the input amino acid sequence, improves at least one target property of the candidate protein59. The system of any one of the preceding embodiments, wherein (iv) combining the at least one weighted relative contribution of the single residue mutant to the at least one target property to generate an individual protein score comprises determining the weighted summation of the at least one weighted relative contribution of the single residue mutant to the at least one target property.60. The system of any one of the preceding embodiments, wherein the (b) sampling the at least one weighted relative contribution of the single residue mutant comprises sampling at least one weighted relative contribution of at least one single residue mutant.61. The system of any one of the preceding embodiments, wherein step the (b) sampling produces a combined protein score with improved scores of at least three or all of PTM, pADA, IMM, and FIT.62. The system of any one of the preceding embodiments, wherein the improved score for PTM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in PTMs.63. The system of any one of the preceding embodiments, wherein selecting for improved PTM results in at most few PTM sites in the protein variant.64. The system of any one of the preceding embodiments, wherein selecting for PTM results in no identified PTMs in the protein variant.65. The system of any one of the preceding embodiments, wherein selecting for pADA binding results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in predicted pADA binding.66. The system of any one of the preceding embodiments, wherein selecting for pADA results in a prediction of low pADA binding.67. The system of any one of the preceding embodiments, wherein selecting for pADA results in a prediction of no pADA binding.68. The system of any one of the preceding embodiments, wherein selecting for IMM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in immunogenicity.40IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT69. The system of any one of the preceding embodiments, wherein selecting for IMM results in predicted low immunogenicity.70. The system of any one of the preceding embodiments, wherein selecting for IMM results in a prediction of no immunogenicity.71. The system of any one of the preceding embodiments, wherein selecting for FIT results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% increase in FIT.72. The system of any one of the preceding embodiments, wherein selecting for FIT improves predicted FIT.73. The system of any one of the preceding embodiments, wherein selecting for FIT results in high FIT.74. The system of any one of the preceding embodiments, wherein the PTM model comprises instructions for searching the single residue mutant input amino acid sequence for regions that match PTM motifs and exceed solvent accessibility thresholds in a corresponding 3D structure of the protein.75. The system of any one of the preceding embodiments, wherein the PTM motifs comprise short sequence patterns that are enriched for measured PTM sites.76. The system of any one of the preceding embodiments, wherein the PTM motifs comprise deamidation sites (e.g., NfDNPTGSC]), DP clipping (DP), N-glycosylation sites (e.g., N[AP][ST]), integrin binding motif (RGD), oxidation sites (MW), and isomerization sites (e.g., D[DCSAG]).77. The system of any one of the preceding embodiments, wherein the PTM model generates a PTM score comprising a tally of all predicted PTMs throughout the single residue mutant input amino acid sequence.78. The system of any one of the preceding embodiments, wherein the pADA model comprises generating a negative hamming distance score between the single residue mutant input amino acid sequence and the input amino acid sequence.79. The system of any one of the preceding embodiments, wherein the IMM model comprises: identifying one or more sets of amino acids representing binding cores of the input amino acid sequence;41IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT determining a set of MHC alleles representative of MHC alleles of a human population; determining scores for pairs of the one or more sets of amino acids and MHC alleles; combining the determined scores to generate a combined score for each set of amino acids representing a binding core; and selecting one or more sets of amino acids representing a binding core as T-cell epitopes.80. The system of any one of the preceding embodiments, wherein the selected one or more sets of amino acids have higher combined scores compared to non- selected one or more sets of amino acids.81. The system of any one of the preceding embodiments, wherein the MHC alleles comprise class II MHC alleles or class I MHC alleles.82. The system of any one of the preceding embodiments, wherein the MHC alleles are selected from DRBl*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01,DRB1*16:O1, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02,DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03,DRBl*08:07, DRB1*14:O2, DRBl*ll:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1,DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02,DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, HLA-DPA l*02:02 / DPB 1*26:01, HLA-DPAl*02:02 / DPBl*13:01, HLA-DPA 1 *02 :02 / DPB 1*03:01, HLA-DPAl*01:03 / DPBl*03:01, HLA-DPAl*01:03 / DPBl*ll:01, HLA-DPAl*02:02 / DPBl*09:01, HLA-DPA1*O1:O3 / DPB1*17:O1, HLA-DPA1*O2:O1 / DPB1*18:O1, HLA-DPAl*01:03 / DPBl*18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPAl*02:02 / DPBl*19:01,HLA-DPA l*01:03 / DPB 1*02:01, HLA-DPAl*01:03 / DPB 1*04:01, HLA-DPA 1 *02:02 / DPB 1 *04:01, HLA-DPA l*02:02 / DPB 1*01:01, HLA-DPA 1*01 :03 / DPB 1*01:01,HLA-DPAl*02:02 / DPB 1*05:01, HLA-DPAl*01:03 / DPBl*19:01, HLA-DPA 1*01 :03 / DPB 1*05:01, HLA-DQAl*05:01 / DQBl*04:02, HLA-DQAl*01:02 / DQB 1*05:01, HLA-DQA 1 *05 :01 / DQB 1*05:01, HLA-DQA 1 *05 :01 / DQB 1*02:01, HLA-DQA1 *01 :02 / DQB 1*06:04, HLA-DQA 1*01 :02 / DQB 1*03:01, HLA-DQA l*01:02 / DQB 1*02:01, HLA-DQAl*01:03 / DQBl*06:03, HLA-DQA1 *01 :02 / DQB 1*06:02, HLA-DQA1*O5:O1 / DQB1*O3:O1, HLA-DQAl*05:01 / DQBl*03:03, or any combination thereof.42IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT83. The system of any one of the preceding embodiments, wherein the IMM model comprises a NctMHC model.84. The system of any one of the preceding embodiments, wherein the FIT model comprises an EVcouplings model.85. The system of any one of the preceding embodiments, wherein the FIT model generates a fitness score using log probabilities of improved fitness.86. The system of any one of the preceding embodiments, wherein the improved fitness comprises improvement to enzymatic activity, binding to a target protein, thermostability, melting temperature, and / or half-life.87. The system of any one of the preceding embodiments, wherein the input amino acid sequence of the target protein is a wild-type sequence.88. The system of any one of the preceding embodiments, wherein the target protein is a bacterial enzyme.89. The system of any one of the preceding embodiments, wherein the bacterial enzyme is a protease or a glycosidase.90. The system of any one of the preceding embodiments, wherein the protease is an immunoglobulin-degrading protease, such as IgG, IgE, or an IgM protease.91. The system of any one of the preceding embodiments, wherein the protease is selected from deS, IdeSsuis, IdeZ, IdeE, IdeE2, IdeZ2, Ide85, or IdeC, or a variant thereof.92. The system of any one of the preceding embodiments, wherein the (i) mutating an amino acid of the input amino acid sequence to generate a single residue mutant input amino acid sequence comprises introducing any of a substitution, insertion, or deletion to the input amino acid sequence.93. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform:(a) iteratively sampling an input amino acid sequence of the target protein, the iteratively sampling comprising:(i) mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence;(ii) inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two43IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two trained computational models comprises two or more of: a post-translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T-cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model;(iii) generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the trained computational models that were inputted with the single residue mutant input amino acid sequence;(iv) combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; and(b) sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences to generate a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.94. The non-transitory computer readable medium of any one of the preceding embodiments, wherein sampling of step (b) comprising repeating steps (a)-(b) on the single residue mutant input amino acid sequence to generate the protein variant amino acid sequence of the target protein with one or more modified properties with a combined protein score exceeding a set threshold.95. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the threshold is an improvement in target property, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% improvement in a target property.44IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT96. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the threshold is and iteration threshold, such as at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, or more iterations.97. The non-transitory computer readable medium of any one of the preceding embodiments, the method further comprising obtaining or having obtained the input amino acid sequence of the candidate protein prior to iteratively sampling the input amino acid sequence.98. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the one or more amino acid mutations that, when included in the input amino acid sequence, improves at least one target property of the candidate protein99. The non-transitory computer readable medium of any one of the preceding embodiments, wherein (iv) combining the at least one weighted relative contribution of the single residue mutant to the at least one target property to generate an individual protein score comprises determining the weighted summation of the at least one weighted relative contribution of the single residue mutant to the at least one target property.100. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the (b) sampling the at least one weighted relative contribution of the single residue mutant comprises sampling at least one weighted relative contribution of at least one single residue mutant.45IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT101 . The non-transitory computer readable medium of any one of the preceding embodiments, wherein step the (b) sampling produces a combined protein score with improved scores of at least three or all of PTM, pADA, IMM, and FIT.102. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the improved score for PTM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in PTMs.103. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for improved PTM results in at most few PTM sites in the protein variant.104. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for PTM results in no identified PTMs in the protein variant.105. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for pADA binding results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in predicted pADA binding.106. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for pADA results in a prediction of low pADA binding.107. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for pADA results in a prediction of no pADA binding.108. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for IMM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in immunogenicity .109. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for IMM results in predicted low immunogenicity.110. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for IMM results in a prediction of no immunogenicity.111. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for FIT results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% increase in FIT.46IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT112. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for FIT improves predicted FIT.113. The non-transitory computer readable medium of any one of the preceding embodiments, wherein selecting for FIT results in high FIT.114. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the PTM model comprises instructions for searching the single residue mutant input amino acid sequence for regions that match PTM motifs and exceed solvent accessibility thresholds in a corresponding 3D structure of the protein.115. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the PTM motifs comprise short sequence patterns that are enriched for measured PTM sites.116. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the PTM motifs comprise deamidation sites (e.g., N[DNPTGSC]), DP clipping (DP), N- glycosylation sites (e.g., N[AP][ST]), integrin binding motif (RGD), oxidation sites (MW), and isomerization sites (e.g., D[DCSAG]).117. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the PTM model generates a PTM score comprising a tally of all predicted PTMs throughout the single residue mutant input amino acid sequence.118. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the pADA model comprises generating a negative hamming distance score between the single residue mutant input amino acid sequence and the input amino acid sequence.119. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the IMM model comprises: identifying one or more sets of amino acids representing binding cores of the input amino acid sequence; determining a set of MHC alleles representative of MHC alleles of a human population; determining scores for pairs of the one or more sets of amino acids and MHC alleles; combining the determined scores to generate a combined score for each set of amino acids representing a binding core; and selecting one or more sets of amino acids representing a binding core as T-cell epitopes.47IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT120. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the selected one or more sets of amino acids have higher combined scores compared to non-selected one or more sets of amino acids.1 E The non-transitory computer readable medium of any one of the preceding embodiments, wherein the MHC alleles comprise class II MHC alleles or class I MHC alleles.122. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the MHC alleles are selected from DRBl*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*ll:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, HLA- DPAl*02:02 / DPBl*26:01, HLA-DPAl*02:02 / DPBl*13:01, HLA-DPAl*02:02 / DPBl*03:01, HLA-DPAl*01:03 / DPBl*03:01, HLA-DPAl*01:03 / DPBl* 11:01, HLA- DPAl*02:02 / DPBl*09:01, HLA-DPAl*01:03 / DPBl*17:01, HLA-DPAl*02:01 / DPBl*18:01, HLA-DPAl*01:03 / DPBl*18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPAl*02:02 / DPBl*19:01, HLA-DPAl*01:03 / DPBl*02:01, HLA-DPAl*01:03 / DPBl*04:01, HLA-DPAl*02:02 / DPB 1*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA- DPAl*01:03 / DPBl*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPAl*01:03 / DPBl*19:01, HLA-DPAl*01:03 / DPB 1*05:01, HLA-DQA1 *05 :01 / DQB 1*04:02, HLA- DQAl*01:02 / DQB 1*05:01, HLA-DQAl*05:01 / DQBl*05:01, HLA-DQA 1 *05 :01 / DQB 1*02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA- DQAl*01:02 / DQB 1*03:01, HLA-DQAl*01:02 / DQB 1*02:01, HLA- DQAl*01:03 / DQBl*06:03, HLA-DQA1 *01 :02 / DQB 1*06:02, HLA- DQA1*O5:O1 / DQB1*O3:O1, HLA-DQAl*05:01 / DQBl*03:03, or any combination thereof.123. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the IMM model comprises a NetMHC model.124. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the FIT model comprises an EVcouplings model.125. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the FIT model generates a fitness score using log probabilities of improved fitness.48IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT126. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the improved fitness comprises improvement to enzymatic activity, binding to a target protein, thermostability, melting temperature, and / or half-life.127. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the input amino acid sequence of the target protein is a wild-type sequence.128. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the target protein is a bacterial enzyme.129. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the bacterial enzyme is a protease or a glycosidase.130. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the protease is an immunoglobulin-degrading protease, such as IgG, IgE, or an IgM protease.131. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the protease is selected from deS, IdeSsuis, IdeZ, IdeE, IdeE2, IdeZ2, Ide85, or IdeC, or a variant thereof.132. The non-transitory computer readable medium of any one of the preceding embodiments, wherein the (i) mutating an amino acid of the input amino acid sequence to generate a single residue mutant input amino acid sequence comprises introducing any of a substitution, insertion, or deletion to the input amino acid sequence.133. A designed protein comprising one or more amino acid mutations relative to a protein, the designed protein exhibiting improved at least one target property in comparison to the protein, wherein the designed protein is identified by performing the method of any one of embodiments 1-52.EXAMPLES
[0118] Below are examples of specific embodiments. The examples are offered for illustrative purposes only and are not intended to limit the scope. Efforts have been made to ensure accuracy with respect to numbers used, but some experimental error and deviation should be allowed for.49IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTExample 1: Development of multi-objective methods for simultaneous design of both function and drug-like properties.Unsupervised generative sequence models
[0119] Evolutionary sequence models are powerful tools for predicting the relative fitness of proteins because they are trained on data from large repertoires of natural sequences. These models, regardless of model architecture, number of parameters, and depth of training data, all have one thing in common: they learn features of natural sequences that contribute to the overall fitness of a molecule. This fitness can be reflected in proteins in several ways, including expression or yield, thermostability, and functional activity and can be predictive of disease variants. This is because natural sequences have evolved over time to be stable and functional, often a prerequisite for an organism’s survival. Here, three models that are trained on evolutionarily related sequence families: an independent postion specific scoring matrix model (PSSM), an evolutionary couplings based hamiltonian model (EVH), and a variational auto encoder (VAE) were used (FIG. 5A). FIG. 5 is a modular multi-objective design methods for engineering a protein into a potential therapeutic. FIG. 5A, shows enerative models of protein sequences can be trained on alignments of natural sequences and learn the site-wise position parameters hi, pairwise interaction parameters Jij and latent space representations Z. FIG. 5B, shows developability considerations for designing a potential therapeutic. FIG. 5C shows identifying local maxima for both fitness and developability. FIG. 5D, shows schema for generation methods of simulated annealing, conditional sampling, and interpolation sampling from left to right. FIG. 5E, shows a matrix of choices for trainign data, model architecture, generation schemes, and developability constraints. FIG. 5F, shows experimental design of design categories for production and characterization, broken down by the modular pieces of the multi-objective design methods
[0120] Each of these models was trained on an alignment of homologous sequences, but was parametrized in different ways. The PSSM model learned parameters of each column (position) in the alignment independently by modeling the probability of a sequence x where the energy (log probability) of a sequence was given by the following equation:where hi were the learned column- specific bias terms.50IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0121] The EVH model expanded on the PSSM model hy also learning pairwise coupling terms, Jij and calculating energy as follows:
[0122] The EVE model was used to calculate the energy of a sequence derived from the ELBO as follows:Ffit(x) = log p(x|0) = Eq [log p(x|z, 0)] - DKL(q(z|x, <|>)||p(z))
[0123] Using these predicted energies, full sequences and individual mutations were ranked based on their predicted scores, thereby prioritizing designs that have the best fitness properties.Structure-based and data-driven rational design
[0124] Machine learning models for the prediction of 3D folds as well as rapid advancements in experimental 3D structure determination have opened up the possibility of using structure-based information to aid in protein engineering. For example, structure models can be used to identify positions in a structure that are more likely to break the 3D fold and thus should be avoided or to identify positions that arc likely to affect interaction with a binding partner and thus should be sampled with some bias. Rational design perspectives can also be coded in to a simplistic function or mask to influence design models to positively or negatively weight specific positions and mutations as desired. Furthermore, as experimental methods for screening constructs continue to improve and data is readily generated, early data points can be used to guide design. For example, experimental screening of point mutants can identify the influence of specific residues or mutations on design goals. If sufficient data are available, a simple supervised machine learning model can be constructed and directly leveraged within a multi-objective design framework, or employed indirectly by identifying key sequence parameters that may be cast as a set of rules for design considerations. Here, structure-based rational design was used to inform which positions in the enzyme should be masked to prevent mutations potentially detrimental to activity. Additionally, experimentally tested mutations were also identified and confirmed to be harmful to activity and masked those mutations.Chemical liability prediction51IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0125] Chemical liabilities can contribute to the instability and poor manufacturability of a molecule. Because of this, it is important to remove these post-translational modification (PTM) sites when possible. There are various predictive models for individual liabilities. Here, an approach based on sequence motifs, solvent accessibility, and secondary structure to identify deamidation, isomerization, oxidation, DP-clipping, and glycosylation sites was used. This information was combined into a score per sequence that counted the total predicted PTM sites as follows:where 0 was the solvent accessibility threshold for the PTM site in a Osasa Dirac-delta function and X was the secondary structure of a loop region in a 5ss Kronecker-delta function. For chemical liability prediction, a crystal structure of a protease was used to calculate solvent accessibility and secondary structure.B-cell epitope prediction
[0126] Pre-existing ADA binding is detrimental to a drug’s efficacy and safety because it can cause the drug to be quickly cleared from the body and induce immune response. It has been shown previously that pADA binding is inversely correlated with the number of substitutions made to an immunogenic protein. Models maximizing the mutations made at the surface can be highly effective for preventing recognition by pADAs. For pADA binding removal, three methods were used and included: a simple Hamming distance between the starting sequence and design, a measure of the uniformity of the mutations made at the surface of the protein, and a count of the number of mutations made at positions that are strongly predicted to belong to immunodominant B-cell epitope hotspots based on a proprietary B-cell epitope predictor.T-cell epitope prediction
[0127] To reduce T-cell response through protein engineering, epitopes that arc displayed on antigen-presenting cells via binding to MHC-II molecules were eliminated. To this end the netMHCIIpan-4.1 model was used. The score was transformed using a sigmoid function on the output percentile rank score for eluted ligand and binding affinity prediction from netMHCIIpan. In order to get a single score for each sequence, a weighted average of scores per possible52IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT epitope peptides (9-mers) for all MHC-TI alleles according to the human population frequencies was calcuatcd and then summc across all peptides according to the following formula:Multi-objective design models
[0128] Given these models of fitness and drug-like properties, new protein sequences were generated according to predicted probabilities, harnessing the power of the unsupervised generative models trained on natural sequence repertoires. Using the probabilities output by these models, the fitness landscape was searched for simultaneously designed sequences (FIG. 5C) by summing up weighted scores of the individual properties. Here, there were three generation schemes developed and tested experimentally: 1) simulated annealing, 2) conditional sampling, and 3) interpolation.
[0129] Simulated annealing was used to frame the protein design objective as a stochastic design problem to identify local maxima (FIG. 5D). In each iteration, a single substitution was sampled from the score distribution of all possible point mutations such that the sampled substitution improved the total probability score. By iterating this process, a final set of sequences designed for desired developability constraints was obtained.
[0130] At each iteration, the stalling sequence was encoded via the VAE encoder into the latent space and sequences were randomly sampled by sampling from the Z-space according to the learned distribution and decoded through the VAE decoder. These sequences were then scored through the predictive models for developability along with the fitness score derived from the ELBO. A new starting sequence was randomly sampled according to these probabilities while ensuring that the new sequence had a higher predicted score than the previous sequence.Through iteration, a highly designed sequence was obtained.
[0131] Interpolation was used to interrogate the latent space of the VAE model. The latent space of the VAE embeded discrete natural protein sequences within a smooth manifold. In order to test the nature of regions of latent space outside of the loci of natural sequences, an interpolation was conducted between the encoded Z- vector of the wild-type starting sequence of protease and the mean Z- vector of the learned VAE distribution. To sample sequences,53IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT incremental linear interpolation steps were taken between these start and end points and decode the interpolated vectors through the VAE decoder.Protein design and experimental methodology
[0132] A set of approximately 400 proteases was carefully designed for experimental testing in order to compare and contrast these multi-objective design methods. A combination of different training datasets (b05: 3,021 sequences; b06: 946 sequences; b08: 179 sequences), generative model architectures, generation schemes, and subsets of developability constraints (FIG. 5E) was used. Each designed construct was expressed in E. coli and screened for production purity of monomeric protease via a SEC (% POI), melting temperature (Tm) and onset temperature (Ton) of unfolding, and activity as measured by both production of the full cleavage product Fab’2 and remainder of the uncleaved product IgGl.Comparison of model architectures
[0133] In order to compare the model architectures of the unsupervised generative models trained on natural sequences, a set of constructs was designed with 20-30 examples generated via simulated annealing for each of the generative models, PSSM (FIGs. 6A-6C), EVH (FIGs. 6A- 6C), and EVE (FIGs. 6A-6C) without including any supervised models for developability constraints. FIG. 6 shows unsupervised sequence models trained on natural repertoires yield functional and stable sequences. FIG. 6A, shows percent monomer distributions for each model and generation method. FIG. 6B, shows activity distributions as measured by both full cleavage (% Fab’2 production) and remainder of intact uncleaved material for each model and generation method. FIG. 6C, shows Tm and Ton distributions for each model and generation method. FIG. 6D, shows results for simulated annealing method with three model architectures, broken down by the number of sequences in the training datasets. The PSSM model produced the most designs with high monomeric percentage (% POI, FIG. 6A). The EVE and EVH models produced the most designs that retained full IgG cleavage activity, whereas almost half of the constructs produced by the PSSM model had strongly reduced cleavage activity (FIG. 6B). Finally, all three models produced designs that had a range of thermostability measurements (Tm) between 40 and 65°C and onset melting temperatures (Ton) between 30 and 55°C. The EVE model did produce fewer constructs with especially low Ton, which is important for manufacturability and stability54IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT at body temperature. These results suggest that while the PSSM model may be able to generate proteins that arc highly monomeric and generally stable, it lacks the complexity needed to capture functional constraints for preserving full cleavage activity.Comparison of generation methods
[0134] For the variational autoencoder generations, three generation schemes were compared and included: simulated annealing (FIGs. 6A-6C), conditional sampling (FIGs. 6A-6C), and interpolation (FIGs. 6A-6C). The interpolation sampling constructs yielded the 196 largest fraction of constructs with high % POI (FIG. 6A). The group of sequences that had the better purity had also reduced activity. All the methods that sampled from latent space generated sequences with higher Tm and Ton than simulated annealing.Effect of depth of training data
[0135] For a set of simulated annealing designs, sequences were generated using datasets of three different alignment sizes, from a broad alignment b05 of 3,021 sequences to a medium alignment b06 of 946 sequences to a narrow alignment b08 of 179 sequences (FIG. 6D). From the experimental design, the alignment depth had a particularly strong effect on the PSSM model. EVH and EVE produced the best designs with narrower alignments, reflecting the same trend seen with the PSSM model. The VAE model trained on the broad alignments produced the highest thermostability constructs with designs that less than 50% homologous to the starting sequence. Constructs that were designed with more constraints generally had lower stability than the wild-type parental construct. This was seen in both Tm and Ton. FIG. 7 A, shows the breakdown of constructs from each category that had like- wild- type or better properties versus worse properties for purity, activity, and thermostability as well as the number of predicted B- cell epitopes, T-cell epitopes, and chemical liabilities removed from each construct in the design categories. FIG. 7 is a multi-objective design of protease results in a high success rate of identifying fully drug designed sequences. FIG. 7A, shows for each design group described in the upset plot, there is a corresponding bar for each of the experimental characterizations that reflects the fraction of sequences with like or better than wild-type and worse than wild-type and a corresponding stripplot for the distributions of the predicted number of T-cell epitopes and chemical liabilities removed for the sequences in the group. A corresponding stripplot of number55IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT of mutations is included as a proxy for the predicted degree of B-cell epitopes removed in each of the constructs. FIG. 7B, shows that examining the extent of epitope and PTM removal, it can be observed that FIT+ADA+IMM reduces T cell epitopes scores, and FIT+ADA+IMM+PTM reduces T cell epitope scores and eliminates PTM motifs. The top construct eliminated all PTMs and 6 of 7 T cell epitopes, and achieves near' wildtype melting temperature, with minor decrease in enzyme potency. This was achieved stalling from the wildtype sequence, and assaying only 37 designs across all sets, and only 14 sequences in the full drug design set
[0136] For designs generated with all developability constraints, the VAE produced the most viable constructs for activity, purity, and in particular for thermostability. Introducing more chemical liability removal mutations and T-cell epitope removal mutations had an inverse relationship with thermostability. Strikingly, for each combination of unsupervised model, generation method and developability constraints, at least one construct with Tm > 50°C, Ton > 40°C, and full cleavage activity similar to wild-type function was identified, demonstrating the success and modularity of the multi-objective design methods.High hit rates for multi-objective design
[0137] The designs output of the EVH model with simulated annealing with was analyzed three sets of developability constraints in addition to the fitness score provided by the EVH model itself: 1) B-cell epitope removal (FIT+ADA), 2) B-cell and T-cell epitope removal (FIT+ADA+IMM), and 3) B-cell and T-cell epitope and chemical liability removal (FIT+ADA+IMM+PTM) (FIG. 7B). A sequence was considered successful when the in silico objectives were achieved, the protein expressed with sufficient yield (>60ng / pE), and the measured Tm and cleavage reaction completion were sufficiently near to wildtype (> 50°C, > 50%C Fc product).
[0138] Out of a total of 37 sequences across design categories, a success rate of 91% was observed for B-cell epitope removal alone, 67% for deimmunization by B-cell and T-cell epitope removal, and 74% for full drug design with B-cell and T-cell epitope and chemical liability removal. All strategies also sampled multiple high activity sequences with enhanced stability (Tm > 50°C) (FIG. 7B).56IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTExample 2: Generalizable machine learning models applied to design an IgG-degrading protease for therapeutic developability.
[0139] Since the FDA approved Humulin as the first recombinant protein therapeutic in 1982, development of protein-based drugs has accelerated rapidly into a market of nearly $400 billion annually, with hundreds of candidates approved and in clinical trials. Antibodies have led the way, empowered by methods such as high-throughput library panning and humanization that enable a rapid path to the clinic. But lacking these conveniences, design of non-antibody biologies remains challenging and slow. Many hurdles stand in the way of a protein and the clinic, including: therapeutic efficacy, safety, manufacturability, and immunogenicity (FIG. 8). Having not naturally evolved to exhibit drug-like properties, few non-human and non-antibody proteins have ever made it to the clinic, and new protein design approaches must be explored to unlock the therapeutic potential of all protein families.
[0140] In particular, the intrinsic immunogenicity of non-human proteins presents a major challenge to their use as medicines. Anti-drug antibodies (AD As) elicited against the therapeutic protein can neutralize activity and induce toxicity. Indeed, various proteins from human-invading pathogens have activities promising for medicine, but a combination of pre-existing ADAs (pADAs) and novel ADA responses reduces their therapeutic efficacy and limits the feasibility of repeat dosing. Another challenge is to manufacture the protein at scale. Barriers include poor physical stability, resulting in denaturation or aggregation under temperature or pH stress, poor storage stability (required for commercial viability), and poor chemical stability, resulting in post-translational modifications (PTMs) that may obstruct activity. To address these developability constraints, provided herein is a variety of machine learning (ML) models to introduce mutations to simultaneously achieve high expression, thermostability, chemical stability, and low immunogenicity within a potential therapeutic protein, all while maintaining high therapeutic activity and with minimal need for experimental iteration.Computational and experimental methodology
[0141] In order to engineer a natural protein to be therapeutically developable, risk-factors such as PTMs, ADA binding, and T-cell mediated immunogenicity (IMM) must be minimized, while maintaining essential activity and stability (FIT) (FIG. 8). To do so, computational models were developed to predict 1) the positions of each liability and 2) the effects of all possible mutations57IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT with respect to each liability (FIG. 9A). FIG. 9 shows that ML models accelerate drug discovery by predicting mutations that design drug properties. FIG. 9A, shows that the models provided for herein take a target sequence input and outputs the positions of predicted liabilities (PTMs, ADAs, and T-cell epitopes) and a table of predicted mutation effects on liabilities and evolutionary fitness. Mutations predicted to both remove liabilities and have high fitness are identified as opportunities to improve the drug. FIG. 9B, shows that applying the method to redesign IdeS enabled rapid drug development in just 3 screening rounds, each phenotyping a library of sequences designed by combining mutations identified by the model. FIG. 9C, shows that in round 1, variants with 5-30 mutations were designed to maximize predicted fitness and minimize ADA binding. The top design was selected among those with sufficient enzymatic activity and enhanced stability (left) with greatest decrease in ADAs (right). FIG. 9D, shows that in round 2, PTMs and T-cell epitopes were predicted, and 48 variants were designed to remove these liabilities. The top design was selected for yield and activity, and minimum predicted T- cell epitopes. FIG. 9E, shows the design was measured in MAPPS and time-course PTM measurements, informing further PTM and T-cell epitope mutations in a library of 20 designs. The final design was selected for yield and activity, and maximum number of T cell epitope and PTM mutations, possessing 34 total mutations.
[0142] Mutations that are predicted to both remove liabilities and to improve or maintain evolutionary fitness were retained as opportunities to improve the drug. The predictions were quantified by four score functions (IMM, ADA, PTM, FIT ), provided a target protein’s sequence x and structure s, which was understood to be computationally predicted from sequence if an experimentally solved structure was unavailable. Biochemical parameters were used to predict PTMs and pADA probability, and machine learning models to predict T-cell epitopes and fitness-improving mutations.
[0143] These models were used to design a protease for drug-like properties. A protease was used as a single-dose therapeutic to decrease tissue transplant risk in sensitized patients by transiently degrading immunoglobulins, including anti-HLA antibodies. More generally, an IgG- degrading protease has the potential to treat a spectrum of autoantibody-mediated diseases, such as anti-glomerular basement disease and myasthenia gravis. However, repeat dosing of a protease is restricted due to its high immunogenicity. Pre-existing anti-protease antibodies in patients with past Strep infections and induced anti-protease antibodies developed within one58IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT week of treatment have caused toxicity concerns and limit the applicability of protease for many clinical indications. To create a therapeutic for chronic autoantibody-mediated indications that would benefit from repeat dosing, a protease was designed to minimize its immunogenicity while retaining therapeutic activity and increasing its stability. Undesired PTMs that arise during protein expression, purification, and throughout long-term storage, were also eliminated to ensure a stable and high-quality drug product.
[0144] The process of redesign took place over 3 rounds of screening (FIG. 9B), in which mutations were systematically introduced into the protease based on predictions from a variety of machine learning models, with the goal of jointly designing its drug-like properties. To construct a machine learning model of protease fitness, a multiple sequence alignment of 2,942 cysteine proteases with homology to the protease was used to construct an EVCouplings model capable of generating fitness-improving mutations. To predict PTMs, chemical liability motif scanning was combined with a crystal structure of the protease monomer to identify solvent accessible liability loci on the protease surface, into which mutations guided by the EVCouplings model that would remove the liability motif were introduced.
[0145] Predictions of T cell epitopes from netMHCIIpan-4.1 were combined with a model of HLA prevalence across the global population to identify the molecular drivers of IdeS immunogenicity. In each screening round, a library of mutated sequences was designed by combining mutations predicted by the EVCouplings model to benefit or minimally impact fitness while removing T cell epitopes and PTM sequence motifs. These mutant proteins were recombinantly expressed and assayed to measure key drug properties, and the best design from each round was selected for further engineering or as a final design. Per round, between 20-94 mutated sequences were screened, leveraging the high precision of the machine learning-guided process. These small, efficient libraries enabled plate-based screening and avoided the need for high-throughput screening methods in the discovery of a therapeutically developable protease.Models to predict effects of mutations on fitness and pADA binding
[0146] Fitness models can be used to generate mutant sequences with predicted improvements in fitness, or to output probability of an input sequence. The probability of a mutated sequence is predictive of the mutation’s effect on evolutionarily conserved functions (such as enzymatic activity or binding to a target protein) and thermostability (as quantified by melting temperature59IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT(Tm) or half-life). Here, the fitness score was calcualted using the log probabilities log p(x) from an EVcouplings model trained on a multiple sequence alignment of sequence homologs to the protease. ADA was defined as the negative hamming distance h(xmut, xwt) between the parent protein and the variant. Variants with lower sequence similarity to the parental protein are less likely to be recognized by pADAs due to antibodies’ high target specificity.Screening of ML generated sequences for fitness and reduced pADA binding
[0147] 30 full sequence designs that each combine 5-30 mutations to maximize predicted fitness and minimize pADA binding were tested. 20 of 30 full sequence designs maintained enzymatic activity comparable to the protease (Fc product > 70%), with equal or greater thermostability (Tm > 53°C). Of those, 11 strongly increased thermostability (Tm > 57°C). Decreased pADA binding was linearly correlated with mutation number (Pearson r = 0.836, p = 9.3e-22). The active design with the greatest decrease in pADAs possessed 24 mutations, had a melting temperature of 57 °C, and exhibited comparable proteolytic activity to the protease (Fc product = 89%) (FIG. 9C).Models to predict chemical liabilities and T-cell epitopes
[0148] Because a key factor in immunogenicity of biologies is B-cell activation via T-cell recognition of MHC -bound peptides derived from the therapeutic protein, a model to predict the display of T cell epitopes on MHC-II alleles across a population representing a target clinical cohort was developed. Predictions from the netMHCIIpan-4.1 model were used, which distinguishes true T cell epitopes displayed by a given MHC-II allele from random peptides with high accuracy (AUC= 0.98). To capture epitopes likely to be displayed across a given population, sequences were scored in terms of their predicted binding to a set of MHC-II alleles representing broad HLA haplotype diversity. To prioritize mutations for deimmunization, a sigmoid transformation was applied to the netMHCIIpan percentile rank score (EPI(xcore)), average scores for all alleles weighted by their frequency in the world population, and sum across all subsequences of the mutant sequence:60IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0149] Methods for PTMs prediction search a candidate protein for regions that match PTM sequence motifs and exceed solvent accessibility thresholds in the corresponding 3D structure. Here, the PTM model considered deamidation, isomerization, oxidation, DP-clipping, and glycosylation as the target liabilities. For PTM motifs, short sequence patterns that were enriched for measured PTM sites, solvent accessibility, and secondary structure were used.Screening of ML designed sequences for T-cell epitope and chemical liability removal
[0150] As starling points for round 2, the 24-mutation top round 1 design, and a 26-mutation chimera combining the top design with a loop from a homolovous sequence were used. Mutations were introduced to remove predicted PTMs and T-cell epitopes. The T-cell epitope model predicted eight population prevalent epitopes in the protease. One epitope was removed by truncating the N-terminus in all designs. For each of the remaining seven epitopes, all substitutions within the 9 amino acid epitope binding core were considered, and those mutations with simultaneous maximal fitness and minimal immunogenicity were selected. 12 PTMs were preicted, and 8 were prioritized for mutation based on their proximity to IgG in the protease-IgG complex 3D structure. For each PTM, all motif-overlapping mutations were considered, and those with simultaneous maximal fitness and minimal predicted PTMs were selected. A library of 48 designs was created by combining subsets of these mutations, expressed, and assayed for activity and stability. The design with maximum predicted decrease in T-cell epitopes and PTM liabilities among those with sufficient activity and stability was selected. This top design possessed 30 mutations relative to the protease, removing 1 PTM and 6 of 8 predicted epitopes (FIG. 9D). The protease and the top round 2 design for T-cell epitope display were asseyed by MHC associated peptide proteomics (MAPPs). MAPPs verified the protease epitope prediction to be 100% accurate, with all 8 epitopes predicted exhibiting MHC-II DRB display across a panel of 15 PBMC donors, and no other epitopes being displayed. Furthermore, MAPPs of the top design showed 6 of 6 mutated epitopes were successfully eliminated. Finally, in round 3, a set of 20 constructs was designed with the purpose of removing the remaining liabilities observed in the MAPPs and PTM assays.
[0151] Final ML designed construct was designed for desirable drug properties. The final design was chosen as that with the fewest remaining T-cell epitopes and PTMs, while achieving high activity (Fc product > 70%) and high monomeric yield (POI > 85%) (FIG. 9E). The final61IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT design increased melting temperature by 2.7°C to 57. 1°C and retained like-wild-type activity. To confirm that liabilities were removed and that therapeutic efficacy was mainted, the final design was compared to wild-type protease in a suite of experimental assays. First, direct pADA binding to Fc-protease and to the design was measured by electrochemical luminescence in the MesoScale Discovery (MSD) immunoassay (FIG. 10F). FIG. 10 shows ML models in combination with structure-based and data-driven rational design can introduce novel activity and specificity. FIG. 10A, shows homology search and clustering by sequence identity to sample representative sequences in the family of immunoglobulin degrading cysteine proteases. These proteases were screened against immunoglobulin isotypes to determine activity and specificity profiles. FIG. 10B, shows electrostatic potentials of IgGl versus substrate shows IgGl interface is much more negative than substrate. FIG. 10C, shows characterization of purity (left), melting temperature (middle), and onest of unfolding temperature (right). FIG. 10D, shows cleavage of substrate (activity, measured by % Fab’2 production) and lack of cleavage of IgGl (specificity, measured by remaining % of intact IgG). FIG. 10E, is a dose response curves of top constructs of each round of design for cleavage of substrate (top) and IgGl (bottom).
[0152] The average pADA signal of the design was 31% that of Fc-IdeS on average across all 50 donors. Next, PTM liabilities were evaluated based on the fraction of modified peptide across timepoints in several month-long time courses and diverse stressed conditions. The final design decreased the significant PTMs from 7 to 2 compared to wild-type protease. As for T-cell epitopes, MAPPs verified that all 8 epitopes measured in the parent are eliminated in the final design. To quantify T-cell response, the parent and design proteins were incubated with PBMCs from 40 healthy donors and measured the extent of CD4+ T-cell proliferation via flow cytometry (FIG. 10G). 90% of donors responded to the parent protein (Fc-protease) whereas 20% responded to the design proteins, a more than 4-fold decrease in immunogenicity. This placed the design within the spectrum of clinically-approved biologies, which typically illicit responses in 0 - 20% of donors.Computational subsampling and experimental screening of Ig-degrading cysteine protease asfamily
[0153] To identify the homologous sequences to the protease, a jackhmmer tool was used to generate a multiple sequence alignment of 1,222 sequences from Uniprot and Mgnify databases.62IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTThe candidates were then sampled by calculating all pairwise sequence identities and clustering hierarchically according to increasing sequence identity. These resulted in 342 representative sequences out of which the domains to express were selected by using the sequence alignment and pfam domain information. The 342 sequences sampled from evolution can be clustered by sequence into 4 major clusters (FIG. 10A). These representative constructs were then synthesized, expressed in E. coli, purified, and screened for cleavage of human IgGl, IgG2, IgG3, IgG4, IgAl, IgA2, IgM and IgE. This screen identified a suite of proteases with diverse specificity profiles across all the clusters,Improving favorable production properties of protease using rational design and machine learning
[0154] To assess the quality of the starting candidate, the thermostability and purity of the protease were characterized. For thermostability, the protease had a measured Tm of 45.4°C. The protease also exhibited a low monomer percentage as measured by aSEC, with just 44.0% peak of interest (POI). These poor properties were prohibitive to further production and / or characterization. In order to continue engineering and developing a protease therapeutic, first the poor stability and production yield of the molecule needed to be addressed. To do so, two parallel engineering efforts were employed: 1) combining rational design and evolutionary sequence models for low-N mutations at or near the interface (R00) and 2) evolutionary sequence models for low-N mutations away from the interface (RO).
[0155] In R00, information from the evolutionary screen of the nearest homologous sequences that had higher measured Tm as well as information from a structure-based homoogy model was used. In the evolutionary screen, there were constructs that were over 90% homologous with the candidate protease, but which had higher Tm and yield. By comparing the sequences, positions that may contribute to stability were identified. AlphaFold2 was used to predict the structure of the candidate protease along with homology modeling to a complex structure of the protease bound to IgGl Fc. 7 constructs with % POI higher than the wild-type starting sequence, 9 constructs with higher Tm, and 14 with higher Ton were identified (FIG. 10C).
[0156] In R0, evolutionary sequence models were used to predict the effects of mutations and screen designs in silico for improved fitness. Two evolutionary sequence models, one built on a larger alignment of nearly 3000 sequences and another narrow alignment of under 10063IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT sequences. The larger alignment captured more information across diverse sequences, but the narrow alignment focused on just the nearest neighbors of the query sequence. All single, double, and triple mutants, and randomly subsampled quadruple and quintuple mutations were screened. From this in silico scan, 13 constructs were selected for experimental testing. Of these, 12 had improved % POI over the wild-type starting sequence, 11 constructs with higher Tm, including 5 constructs with Tm > 50°C, and 12 with higher Ton, including 5 with Ton > 37°C (FIG. IOC).Leveraging machine learning, structure-based and data-driven rational design to increase potency and specificity
[0157] In both R00 and RO, changes in activity and / or specificity through specific mutations were observed, and a data-driven rational design and evolutionary sequence models were employed to expand the mutations and positions tested as well as combine mutation sets from R00 and RO. Through these efforts, a protease that had 100% Fab’2 production (full cleavage) and 100% intact IgGl remaining (no cleavage) at a concentration of 3uM protease (FIG. 10D) was identifed. Additional designs in R1 also had high levels of substrate cleavage with low levels of IgG cleavage, demonstrating great success in introducing novel activity and specificity to a cysteine protease. The top R1 construct has an EC50 of 0.082uM for a substrate with no discernible IgGl cleavage detected even at a high concentration of 11.65uM protease (FIG. 10E).Example 3: Use of Predictive Models for Protease Design.
[0158] 30 full sequence designs that each combine 5-30 mutations to maximize predicted fitness and minimize pADA binding were tested according to the function: argmax(F / T(F) + wada. A DA(F)) p where wada is a negative-valued weight to decrease pADA binding, and optimal sequences are sampled by simulated annealing.
[0159] 64 hypothesis-driven constructs were also tested and included: consensus mutations from the designs, mutations in patches of the 3D surface, and loop regions exchanged with a homologous protease (IdeZ). 20 of 30 full sequence designs maintained enzymatic activity comparable to IdeS (Fc product > 70%), with equal or greater thermostability (Tm> 53°C). Of those, 11 strongly increased thermostability (Tm> 57°C). Decreased pADA binding was linearly64IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT correlated with mutation number (Pearson r = 0.836, p = 9.3e-22). The active design with the greatest decrease in ADAs possessed 24 mutations, had a melting temperature of 57.0°C, and exhibited comparable proteolytic activity to IdeS (Fc product = 89%) (FIG. 9C).
[0160] As stalling points for round 2, the 24-mutation top round 1 design was used, and a 26- mutation chimera combining the top design with the top loop-exchange construct. Mutations were introduced to remove predicted PTMs and T-cell epitopes. The T-cell epitope model described in the above equation predicted eight population-prevalent epitopes in IdeS. For each of the remaining seven epitopes, all substitutions within the 9 amino acid epitope binding core were considered, selecting those mutations with simultaneous maximal fitness and minimal immunogenicity : argmax(F F(F), -IMM P))0.10
[0161] 12 PTMs were predicted, and prioritized 8 for mutation based on their proximity to IgG in the IdeS-IgG complex 3D structure (PDB: 8A47). For each PTM, all motif-overlapping mutations were considered, selecting those with simultaneous maximal fitness and minimal predicted PTMs: argmax(F7T(F);PTXE Pp0.05
[0162] A library of 48 designs was created by combining subsets of these mutations, expressed, and assayed for activity and stability. The 3 high yield (POI > 60%) and high activity (Fc product > 80%) constructs had 5 T-cell epitope mutations atop the chimera. The design with maximum predicted decrease in T-cell epitopes and PTM liabilities among those with sufficient activity and stability was selected. This top design possessed 30 mutations relative to IdeS, removing 1 PTM and 6 of 8 predicted epitopes (FIG. 9D). IdeS and the top round 2 design were assayed for T-cell epitope display by MAPPS. MAPPS verified the IdeS epitope prediction to be 100% accurate, with all 8 epitopes predicted exhibiting MHC-II DRB display across a panel of 15 PBMC donors, and no other epitopes being displayed. Furthermore, MAPPS of the top design showed 6 of 6 mutated epitopes were successfully eliminated.
[0163] A set of 20 constructs was designed with the purpose of removing the remaining liabilities observed in the MAPPS and PTM assays. The final design was chosen as that with the65IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT fewest remaining T-cell epitopes and PTMs, while achieving high activity (Fc product > 70%) and high monomeric yield (POI > 85%) (FIG. 9E). This final design possessed 37 mutations relative to IdeS. The final design, runner-up designs, and controls were carried forward to final validation experiments, including MAPPS, donor T-cell proliferation, time-course PTM measurements, and in vivo PK / PD.Final design with desirable drug properties
[0164] The final drug candidate showed a dramatic improvement in PTMs, T cell epitopes, immunogenicity, and ADA binding relative to IdeS (FIG. 12A). FIG. 12 shows that redesigned IdeS has improved drug properties and minimal immunogenicity. FIG. 12A, shows the final design shows a dramatic improvement in PTMs, T-cell epitopes, immunogenicity, and ADA binding. FIG. 12B, shows enhanced stability, increasing Tm by 3oC. FIG. 12C, shows monodispersity in large-scale production increased by 20%. FIG. 12D, shows in vitro potency slightly decreased, from EC50 26nm to 78nM. FIG. 12E, shows that total IgG cleavage in vivo was closely comparable to Fc-IdeS. Mice were infused with human IVIG as substrate, and then dosed with either Fc-IdeS or the design. At 24 hours post-dose, 99% of IVIG elimination was observed by Fc-IdeS, and 91% by the present design (based on the method provided for herein). FIG. 12F, shows PTM measurements across month-long time courses of various stressed conditions, revealed the positions and significance of PTMs in Fc-IdeS and the design. Values above the dashed line are significant. Fc-IdeS has 7 significant PTMs whereas the design has just 2. FIG. 12G, shows T-cell epitopes measured in the parent and final design by MAPPs. IdeS had epitopes in eight regions, and all are eliminated in the present design. In one donor, a new epitope was observed in the design, unique from those seen in the parent. FIG. 12H, shows ADA binding was measured using serum IgG from 50 donors. Each donor’s binding to the design was normalized to corresponding IdeS signal. The design showed an average pADA signal 31% that of parent, with 70% of donors showing less than 10% signal. FIG. 121 shows immunogenicity was measured by the proliferation of CD4+ T-cells in donor PBMC samples exposed to protein sample. 90% of donors responded to the parent protein whereas 20% responded to the present design, a >4x decrease in immunogenicity.Stability66IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0165] Sufficient thermostability is essential to achieve adequate non-aggregated protein yield, shelf-life, and minimal toxicity. Thermostability was quantified by measuring melting temperature (Tm), observing 54.4°C for IdeS and 57.1°C for the final design, an increase of 2.7°C (FIG. 12B). In contrast to typical drug development that introduces minimal mutations because of heavy costs to activity and stability, the design here showed improved stability despite the heavy mutation load addressing epitopes and PTMs.Therapeutic activity
[0166] Next, the potency of parent protein and the design in vitro was quantified by titrating drug protein, using human plasma samples as substrate and the Fc product of IgG cleavage as the readout. IdeS displayed an EC50 of 26nM while the design protein measured an EC50 of 78nM (FIG. 12D). To measure potency in vivo, mice were injected with human IVIG, dosed with drug protein (t=0), and quantified the remaining intact IVIG. At 24 hours, a 99% IVIG elimination by Fc-IdeS, and 91% by the design protein were observed (FIG. I2E).Chemical liabilities
[0167] PTM liabilities were evaluated based on the fraction of modified peptide across timepoints in several month-long time courses and diverse stressed conditions (FIG. 12F). Some predicted PTMs were observed, but with only low frequency (<5%) or only in certain stressed conditions (37°C and pH 8.5 for 1 week, pH 4.5 for 1 week, or 0.01% H2O2 for 24 hours). Quantifying the therapeutic activity of drug material throughout the time course showed no loss of potency, suggesting that the two remaining PTMs have negligible effect on activity.Highly accurate epitope prediction and removal
[0168] The predictive models provided herein enabled selection of jointly optimal mutations along the epitope-fitness pareto front (FIG. 13A and B). FIG. 13 shows that accurate T-cell epitope prediction enables deimmunization relying on protein sequence alone. FIG. 13A, shows that throughout the protein, the predicted effect of all amino acid substitutions on epitope display and protein fitness must be computed. FIG. 13B, shows that for a given epitope, considering all possible mutations, those that jointly minimize epitope display and maximize fitness were identified. The top choice of mutation was substitutions that result in epitope score < 0.10 with67IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT highest fitness (flagged by an asterisk). FTG. 13D, are epitope scores throughout the protein sequence predict epitope display of regions 38-48, 73-88, 117-135, 155-163, 184-192, 255-266, 271-289, and 296-303. Each point in the plot represents a 9aa binding core starting at the x- coordinate. FIG. 13C, is the 3D placement of targeted T-cell epitopes within the resolved 3D structure of IdeS-IgG complex poses a challenge for mutation owing to high overlap with buried and / or activity-determining positions. FIG. 13E, shows T-cell epitopes measured in the parent and final design by MHC-associated peptide proteomics confirm that the 8 predicted epitopes were displayed, and zero epitopes were unpredicted. The x-axis range of each rectangle shows the region of sequence displayed by MHC-II, with opacity corresponding to the number of donors that display peptides in that region. FIG. 13F, shows that MAPPs verifies that all eight epitopes measured in the parent are eliminated in the final design. However, a single epitope was observed at 238-246, unique from those seen in the parent, owing to a mutation in the design sequence.
[0169] The precision of these models to recommend viable mutations was important, because of the challenging placement of the T-cell epitopes (5 of 8 epitopes had 5-9 of 9 residues either buried in the 3D fold or directly involved in the essential interaction with IgG)(FIG. 13C). T-cell epitopes were measured in the parent and final design by MHC-associated peptide proteomics (FIG. 12G, and FIG. 13E and F). The parent showed epitope display of eight sequence regions: 34-53, 70-93, 107-128, 148-167, 179-197, 253-268, 269-294, and 293-307. The measured IdeS T-cell epitopes corresponded to the eight regions predicted from sequence alone. In the final design, the N-terminus was truncated at amino acid 48, eliminating the first epitope. Mutations were introduced at all 7 other epitopes, each predicted to reduce MHC-II display across a set of 52 MHC alleles most frequently observed in the world population. MAPPs verified that all 8 epitopes measured in the parent are eliminated in the final design. A single epitope was observed at 238-246 (FIG. 13E).Pre-existing AD As
[0170] Because the majority of the world population has been immunized to IdeS via S. pyogenes exposure, pADAs are highly prevalent in patients. One approach to decrease pADA binding here was by introducing mutations throughout the protein surface, creating mismatch to IdeS residues recognized by pADAs. ADA binding signal of IdeS and the design proteins were68IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT measured by electrochemical luminescence in the MesoScale Discovery (MSD) immunoassay (FIG. 12H), normalizing the signal to IdcS. The average pADA signal of the final design was 31% that of the parent on average across all 50 donors, and 70% of donors showed less than 10% signal.Immunogenicity
[0171] To quantify T cell response, the parent and the design proteins were incubated with PBMCs from 40 healthy donors and measured the extent of CD4+ T-cell proliferation via flow cytometry (FIG. 121). 90% of donors responded to the parent protein (Fc-IdeS) whereas 20% responded to the design protein, a more than 4-fold decrease in immunogenicity. Low immunogenicity clinically-approved biologies typically illicit responses in 0-20% of donors.
[0172] Next, it was assessed whether reduced T-cell epitopes and reduced in vitro response amount to a commensuiserate decrease of in vivo immunogenicity, or if unidentified features of the protein might trigger immunogenicity. Groups of black 6 mice were twice dosed with the parental protease and the design proteins, and ADA titers wer emeasured over a 27-day timecourse. A strong reduction of immunogenic responses was observed to design proteins. All mice treated with the parental protease showed high ADA levels (>5ug / mL) or died upon the second dose. By contrast, only 20% of mice treated with the design proteins showed high ADA levels, and all of the design protein treated mice survived redosing. Thus, the design proteins showed reduction in immunogenicity as compared to the parental protease.
[0173] Based on the success of the 3 round approach to deimmunizing IdeS, a general computational framework was developed whereby any given input protein sequence can be designed for immunogenicity, manufacturability, and fitness simultaneously.
[0174] The objective of the redesign process was defined as a score function that incorporated quantitative predictions of each target property:where SltS2, .... ,Snwere score functions for each individual target property, and wltw2, . . . , wnwere weights quantifying the relative contribution of each property to the overall score. The design of IdeS comprised score functions for four properties:69IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT
[0175] The protein design objective was framed as identifying local maxima of S(P), a stochastic design problem, which was approached by simulated annealing (Fig 14A). First, a random sequence was sampled, and then redesigned in iterations. In each iteration, a single substitution was sampled from the score distribution of all possible point mutations such that the sampled substitution improved the score S(P). By iterating this process, a final set of redesigned sequences predicted to exhibit high fitness, low immunogenicity, few PTMs, and low pADA binding was obtained.
[0176] Next, a panel of IdeS variants was designed. 37 designed IdeS variants were designed using the multi-objective simulated annealing method, recombinantly expressed in E. coll, and assayed for melting temperature and IgG cleavage activity. Constructs were designed to better understand the tradeoffs between competing objectives, and to separately design for stabilization (FIT -I- ADA), deimmunization (FIT + ADA -I- IMM), and full drug design (FIT + ADA + IMM + PTM)(FIG. 14B). FIG. 14 shows that multi-optimization generates drug quality proteins in a single round. FIG. 14A, shows that sequence space was searched for designs with optimal therapeutic potential. The multi- annealing algorithm sampled a random sequence and iteratively designed it by mutation. The final sequence had a high total weighted score, resulting in high FIT, low IMM, low PTM, and low ADA scores. FIG. 14B, shows that designs were explored for FIT+ADA, FIT+ADA+IMM, and FIT+ADA+IMM+PTM. The IgG cleavage activity and melting temperature of each design was measured, and a high rate of designs with wildtype-like cleavage activity, and increased melting temperature were observed. FIG. 14C, shows that all design objectives resulted in improvements to stability, though the more constraints added to design the less improvement to stability was observed. FIG. 14D, shows that the in silico score functions predicted a tradeoff between fitness and epitope removal, with FIT+ADA designs (triangles) as higher fitness, whereas FIT+ADA+IMM designs (circles) reduce epitopes at a fitness cost. FIG. 14E, shows measured stability and activity followed the predicted fitnessepitope tradeoff. FIG. 14F, shows that examining the extent of epitope and PTM removal, it was observed that FIT+ADA+IMM reduced T cell epitopes scores, and FIT+ADA+IMM+PTM reduced T cell epitope scores and eliminated PTM motifs. The top construct eliminated all PTMs and 6 of 7 T cell epitopes (an 8thepitope is removed in all constructs by the N terminus deletion), and achieved near wildtype melting temperature, with minor decrease in enzyme potency. This70IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT was achieved, starting from the wildtype sequence, and assaying only 37 designs. Showing that drug-like properties can be achieved in a single 96- well plate of designs.
[0177] Sequences 10-12, 20-22, and 30-32 mutations distant from wildtype IdeS were generated for a total of nine design criteria, assaying 3-5 designs for each.
[0178] A sequence was considered successful when the in silico objectives were achieved, the protein expressed with sufficient yield, and the measured Tm and cleavage reaction completion were sufficiently near to wildtype (>50°C, >50% Fc product). A success rate of 91% was observed for stabilization, 67% for deimmunization, and 74% for full drug design. All strategies also sampled multiple high activity sequences with enhanced stability (Tm > 57°C) (FIG. 14C). For deimmunization, 20 and 30 mutations had notably higher success rates than 10 mutations, likely due to favorable mutations compensating for costly epitope-removal mutations. For full drug design, success rate for 30 mutations was lower than for 10 or 20. The top design removed 7 of 8 epitopes, and 6 of 6 chemical liabilities (FIG. 14F), comparable in mutation number, epitope-removal, and liability-removal to the drug candidate. A single plate of designs produced by multi-objective design can achieve drug-like properties, an early example of highly accelerated and low cost drug discovery.Example 4: Use of Predictive Models for Enzyme Design.
[0179] The general framework for invizibilization via machine learning of Example 1 was used for immunogenic enzyme design. Here, the multi-objective design was enabled by searching for local maxima in a fitness landscape via stochastic optimization for other types of enzymes that are not cysteine proteases like IdeS or variants thereof.. Each individual term in the linear’ combination represented a different property to be optimized (stability, activity, immunogenicity, chemical liabilities, etc.). The models of Example 1 were provided with the enzyme sequence that was not a cysteine protease to predict T cell epitopes, produce invisiblized designs, and predict the success rate per mutant. Applying the predictive model of Example 1 to said enzyme sequence produced an 8 residue epitope prediction and deimmunization fitness probability of each epitope residue. According, the predictive model produced an in sillico output of an invisibilized enzyme allowing for diversification of designs to maximize the chances of potent an non-immunogenic therapeutics. Thus, these results demonstrate that the methods provided for herein can be used with a broad class of proteins and enzymes.71IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTExample 5: Use of Predictive Models for Antibody Design.
[0180] The general framework for invizibilization via machine learning of Example 1 was used for antibody deimmunization. Here, the multi-objective design was enabled by searching for local maxima in a fitness landscape via stochastic optimization for an antibody. Each individual term in the linear combination represented a different property to be optimized (stability, activity, immunogenicity, chemical liabilities, and the like). An antibody predicted to have several liabilities in the HCDR3 and inaccessible to conventional humanization was analyzed using the predictive methods provided herein. Here, population-level T cell epitope prediction method was used to predict epitopes and identify mutations most likely to remove them. A homology model of the antibody was created and applied to a Rosetta cartesian ddg to estimate the effects of deimmunizing mutations on monomer free energy. Next, a protein large language model (ProGen) was used to estimate mutation effects on antibody naturalness (likelihood of stable soluble expression). These predictions were then combined to identify most favorable mutations that also remove the DP clipping motif. The predictive models provided herein identify most favorable mutations that removed 2 strongly predicted T cell epitopes and DP clipping site in HCDR3. In vitro T cell proliferation assay confirmed successful epitope removal and deimmunization.Example 6: Antibody Design Pipeline for Antibody Design
[0181] Screening large libraries of antibody designs in the lab is expensive, time-consuming, and the success rate obtaining binders with desirable function, epitope specificity, and developability properties is low. Incorporation of computational methods in this process can drastically increase efficiency by narrowing libraries to only include designs with predicted desirable properties. For example, large language models (LLMs) trained on large datasets of protein sequences have demonstrated progress in improving the affinity of known antibodies using a small number of designs. To address the problem of predicting epitope specificity, graph neural networks (GNNs) were developed to predict antibody- specific B cell epitopes on general antigens. Combining inverse folding with the B-cell epitope predictor allowed to design valiants around a starting sequence with modulated affinity and developability properties while maintaining epitope specificity and function. This novel method allows to select for variant72IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT sequences with a high likelihood of binding without the need for case-specific input data heyond the sequence of an antibody and its target antigen. The B cell epitope predictor is used to compare enrichment of binders among three ML-based sequence design methods, and this prediction is then validated through a yeast display assay with 10A7 variants.
[0182] In refernce to FIG. 15, the design pipeline 500 provided herein uses structure and sequence based modeling to balance objectives of maintaining activity while minimizing the presence of sequences that have higher chances of immunogenicity and developability issues.
[0183] With reference to FIG. 15, step 501 comprises inverse folding and / or generating amino acid sequences from parental antibody structures, controlling for design candidate CDRs.
[0184] In some embodiments, step 501 comprises use of an antibody- specific protein design model, such as, without limitation, FvHallucinator (see Mahajan, S. P., Ruffolo, J. A., Frick, R., & Gray, 1. 1. (2022). Hallucinating structure-conditioned antibody libraries for target- specific binders. Frontiers in Immunology, 13, 999034. https: / / doi.org / 10.3389 / fimmu.2022.999034). In some embodiments, the antibody -specific protein design model is built on a deep learning model DeepAb that translates sequence to structure (inverse folding). In some embodiments, the antibody-specific protein design model uses both structure and fixed sequence to create the design subsequence, yielding CDR libraries conditioned on retaining the conformation of the input structure. In some embodiments, step 501 comprises the inputting the structure of the protein backbone and creating a representative graph. In some embodiments, step 501 comprises the use of ProteinMPNN2 (see Dauparas, I., Anishchenko, I., Bennett, N., Bai, H., Ragotte, R. I., Milles, L. F., Wicky, B. I. M., Courbet, A., de Haas, R. J., Bethel, N., Leung, P. I. Y., Huddy, T. F., Pellock, S., Tischer, D., Chan, F., Koepnick, B., Nguyen, H., Kang, A., Sankaran, B., ... Baker, D. (2022). Robust deep learning-based protein sequence design using ProteinMPNN. Science, 378(6615), 49-56. https: / / doi.org / 10.1126 / science.add2187). A GNN is then used to process the structure, capturing local and global relationships between residues. Amino acids most likely to respect the backbone structural constraints are predicted.
[0185] With reference to FIG. 15, step 502 comprises filtering out chemical liability motifs. Step 503 comprises generating homology model inverse folded sequences. Step 504 comprises predicting B-cell epitope for each design. Step 505 comprises predicting progen score for each design. In some embodiments, Progen is an LLM trained on about 280 million protein sequences from more than about 19,000 families (Madani, A., Krause, B., Greene, E. R., Subramanian, S.,73IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTMohr, B. P., Holton, J. M., Olmos, J. L., Xiong, C., Sun, Z. Z., Socher, R., Fraser, J. S., & Naik, N. (2023). Large language models generate functional protein sequences across diverse families. Nature Biotechnology, 41(8), 1099-1106. https: / / doi.org / 10.1038 / s41587-022-01618-2). Progen may contain models that are fine tuned toward specific protein families, including the oas (observed antibody space) model to generate artificial protein sequences with the ability to function as well as natural proteins.
[0186] With reference to FIG. 15, step 506 comprises combining progen and B-cell epitope scores into an overall score.
[0187] FvHallucinator, ProteinMPNN, and Progen epitope scores were obtained. These epitope scores represented the correlation of the epitope prediction of the designed antibodies with that of the stalling antibody which is known to bind an epitope. The epitope scores were a combination of the B-cell epitope score and the progen naturalness score, and represented the probability of an antibody a, given an epitope E: logP(a|E) = L((j) -I- B(a) where L(tr) is a naturalness score from progen, and B(a) is an epitope specificity score from the epitope predictor. The antibody a is represented as a pair a=( ,x) where a is the antibody’s sequence and x is the antibody’s structure. B(cr) is the cross-entropy function quantifying the difference between the antibody’s predicted epitope from the predictor and the target epitope E. Progen and Protein MPNN libraries have one CDR designed at a time, while the invfold library has a combination of CDRs varying.
[0188] A combined library of 2xlOA4 inverse fold and progen random mutagenesis designs were expressed on the surface of yeast and binding to antigen FcGRIIB was measured. Results showed antigen binding up to 4.7%, with a combination of low and high affinity binders.Antigen binding increased proportional to antigen concentration. The inverse folded libraries were scored with the B-cell epitope predictor and progen, and high scoring designs were selected.
[0189] Thus, these results demonstrate that the methods provided for herein can be used with a broad class of proteins and enzymes, including antibodies.
[0190] The present embodiments and exmaples provided herein demonstrate the surprising and unexpected results that use the robust method(s) provided for herein from generating and making proteins that are less immunogenic or have otherwise improved or desidred properties as74IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT compared to the starting sequence. These methods solve technological hurdles that were previously not able to be performed to produce protein that can be used as therapeutics with desired properties, such as reduced immunogenicity or other desired proprities as provided for herien.
[0191] The entire disclosure of each of the patent and scientific documents referred to herein is incorporated by reference for all purposes.
[0192] The invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting on the invention described herein. Scope of the invention is thus indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.75IPTS / 200129719.1
Claims
Attorney Docket No.: SES-018WO PATENTCLAIMS1. A computer-implemented method for generating a protein variant amino acid sequence of a target protein having one or more modified properties, the method comprising:(a) iteratively sampling an input amino acid sequence of the target protein, the iteratively sampling comprising:(i) mutating an amino acid residue of the input amino acid sequence to generate a single residue mutant input amino acid sequence;(ii) inputting the single residue mutant input amino acid sequence, or a portion thereof comprising the mutated amino acid residue, into at least two computational models trained to predict a relative contribution of the single residue mutant to at least one target property of the target protein, wherein the at least two computational models comprise two or more of: a post-translational modification (PTM) model; a pre-existing anti-drug antibody (pADA) model; a T-cell mediated immunogenicity (IMM) model; an activity and stability (FIT) model;(iii) generating an output of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property of the target protein from the at least two computational models that were inputted with the single residue mutant input amino acid sequence;(iv) combining the at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property to generate an individual protein score of the single residue mutant input amino acid sequence; and(b) sampling the individual protein score of at least one weighted relative contribution of the single residue mutant input amino acid sequence to the at least one target property across a plurality of other single residue mutant input amino acid sequences to generate a combined protein score, wherein the combined protein score corresponds to the protein variant comprising one or more amino acid mutations of the single residue mutant input amino acid sequences.
2. The method of claim 1, wherein sampling of step (b) comprising repeating steps (a)-(b) on the single residue mutant input amino acid sequence to generate the protein variant amino76IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT acid sequence of the target protein with one or more modified properties with a combined protein score exceeding a set threshold.
3. The method of claim 2, wherein the set threshold is an improvement in target property, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% improvement in a target property.
4. The method of claim 2, wherein the set threshold is and iteration threshold, such as at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, or more iterations.
5. The method of claims 1 or 2, the method further comprising obtaining or having obtained the input amino acid sequence of the candidate protein prior to iteratively sampling the input amino acid sequence.
6. The method of any one of claims 1-5, wherein the one or more amino acid mutations that, when included in the input amino acid sequence, improves at least one target property of the candidate protein.
7. The method of any one of claims 1-6, wherein (iv) combining the at least one weighted relative contribution of the single residue mutant to the at least one target property to generate an77IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT individual protein score comprises determining the weighted summation of the at least one weighted relative contribution of the single residue mutant to the at least one target property.
8. The method of any one of claims 1-7, wherein the (b) sampling the at least one weighted relative contribution of the single residue mutant comprises sampling at least one weighted relative contribution of at least one single residue mutant.
9. The method of any one of claims 1-8, wherein step the (b) sampling produces a combined protein score with improved scores of at least three or all of PTM, pADA, IMM, and FIT.
10. The method of claim 9, wherein the improved score for PTM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in PTMs.
11. The method of claims 9 or 10, wherein selecting for improved PTM results in at most few PTM sites in the protein variant.
12. The method of claim 11, wherein selecting for PTM results in no identified PTMs in the protein variant.
13. The method of claim 9, wherein selecting for pADA binding results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in predicted pADA binding.
14. The method of claim 13, wherein selecting for pADA results in a prediction of low pADA binding.
15. The method of claim 14, wherein selecting for pADA results in a prediction of no pADA binding.78IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT16. The method of claim 9, wherein selecting for IMM results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% reduction in immunogenicity.
17. The method of claim 16, wherein selecting for IMM results in predicted low immunogenicity.
18. The method of claim 16, wherein selecting for IMM results in a prediction of no immunogenicity.
19. The method of claim 9, wherein selecting for FIT results in at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% increase in FIT.
20. The method of claim 19, wherein selecting for FIT improves predicted FIT.
21. The method of claim 19, wherein selecting for FIT results in high FIT.
22. The method of claim 1, wherein the PTM model comprises instructions for searching the single residue mutant input amino acid sequence for regions that match PTM motifs and exceed solvent accessibility thresholds in a corresponding 3D structure of the protein.
23. The method of claim 22, wherein the PTM motifs comprise short sequence patterns that are enriched for measured PTM sites.
24. The method of claim 22, wherein the PTM motifs comprise deamidation sites (e.g., NfDNPTGSC]), DP clipping (DP), N-glycosylation sites (e.g., N[AP][ST]), integrin binding motif (RGD), oxidation sites (MW), and isomerization sites (e.g., D[DCSAG]).79IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT25. The method of any one of claims 22-24, wherein the PTM model generates a PTM score comprising a tally of all predicted PTMs throughout the single residue mutant input amino acid sequence.
26. The method of claim 1, wherein the pADA model comprises generating a negative hamming distance score between the single residue mutant input amino acid sequence and the input amino acid sequence.
27. The method of claim 1, wherein the IMM model comprises: identifying one or more sets of amino acids representing binding cores of the input amino acid sequence; determining a set of MHC alleles representative of MHC alleles of a human population; determining scores for pairs of the one or more sets of amino acids and MHC alleles; combining the determined scores to generate a combined score for each set of amino acids representing a binding core; and selecting one or more sets of amino acids representing a binding core as T-cell epitopes.
28. The method of claim 27, wherein the selected one or more sets of amino acids have higher combined scores compared to non-selected one or more sets of amino acids.
29. The method of claim 27, wherein the MHC alleles comprise class II MHC alleles or class I MHC alleles.
30. The method of claim 27, wherein the MHC alleles are selected from DRB 1*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*I5:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01, DRBl*04:02, DRBl*03:02, DRBl*03:01, DRBl*08:03, DRBl*08:07, DRB1*14:O2, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01,DRB4*01:03, DRB5*01:01, HLA-DPAl*02:02 / DPBl*26:01, HLA-DPAl*02:02 / DPB 1*13:01, HLA-DPAl*02:02 / DPBl*03:01, HLA-DPAl*01:03 / DPB 1*03:01, HLA-80IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTDPA1 *01 :03 / DPB1 *11 :01, HLA-DPAl *02:02 / DPB 1*09:01, HLA-DPA 1*01 :03 / DPB1* 17:01,HLA-DPA1*O2:O1 / DPB1*18:O1, HLA-DPAl*01:03 / DPBl* 18:01, HLA-DPAl*02:02 / DPBl*15:01, HLA-DPAl*02:02 / DPBl*19:01, HLA-DPAl*01:03 / DPBl*02:01, HLA-DPAl*01:03 / DPB 1*04:01, HLA-DPAl*02:02 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA-DPAl*01:03 / DPBl*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPA1 *01 :03 / DPB 1*19:01, HLA-DPAl*01:03 / DPB 1*05:01, HLA-DQA1 *05 :01 / DQB 1*04:02, HLA-DQA1*O1:O2 / DQB1*O5:O1, HLA-DQA 1 *05 :01 / DQB 1*05:01, HLA-DQAl*05:01 / DQBl*02:01, HLA-DQAl*01:02 / DQBl*06:04, HLA-DQAl*01:02 / DQB 1*03:01, HLA-DQAl*01:02 / DQBl*02:01, HLA-DQA1 *01 :03 / DQB 1*06:03, HLA- DQAl*01:02 / DQBl*06:02, HLA-DQA 1 *05 :01 / DQB 1*03:01, HLA- DQA 1 *05 :01 / DQB 1*03:03, or any combination thereof.
31. The method of claim 1, wherein the IMM model comprises a NetMHC model.
32. The method of claim 1, wherein the FIT model comprises an EVcouplings model.
33. The method of claim 32, wherein the FIT model generates a fitness score using log probabilities of improved fitness.
34. The method of claim 33, wherein the improved fitness comprises improvement to enzymatic activity, binding to a target protein, thermostability, melting temperature, and / or halflife.
35. The method of any one of claims 1-34, wherein the input amino acid sequence of the target protein is a wild-type sequence.
36. The method of any one of claims 1-35, wherein the target protein is a bacterial enzyme, and wherein the bacterial enzyme is a protease or a glycosidase.81IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT37. The method of claim 36, wherein the protease is an immunoglobulin-degrading protease, such as IgG, IgE, or an IgM protease.
38. The method of claim 37, wherein the protease is selected from IdeS, IdeSsuis, IdeZ, IdeE, IdeE2, IdeZ2, Ide85, or IdeC, or a variant thereof.
39. The method of any one of claims 1-38, wherein the (i) mutating an amino acid of the input amino acid sequence to generate a single residue mutant input amino acid sequence comprises introducing any of a substitution, insertion, or deletion to the input amino acid sequence.
40. A method for identifying one or more T-cell epitopes of a target protein to generate a candidate protein with fewer T-cell epitopes, the method comprising:(a) inputting a plurality of subsequences of the target protein into a computational model to determine a plurality of scores representing whether the plurality of subsequences are likely to be presented by a plurality of MHC alleles;(b) for each subsequence, transforming a corresponding score to generate a populationwide measure of antigen presentation; and(c) identifying a subset of the plurality of subsequences as candidate T-cell epitopes based on their population-wide measures of antigen presentation that indicate that the subset of the plurality of subsequences are likely bound by one or more of the plurality of MHC alleles.
41. The method of claim 40, the method further comprising:(d) introducing one or more amino acid mutations into a candidate T-cell epitope of the input amino acid sequence to generate a variant amino acid sequence; and(e) evaluating the impact of the introduced one or more amino acid mutations on T-cell IMM of a protein therapeutic comprising the variant amino acid sequence in comparison to the protein comprising the input amino acid sequence.
42. The method of claim 41, wherein step (e) further comprises:82IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENT summating population-wide measures of presentation of a plurality of subsequences of the variant amino acid sequence to generate a IMM score for the variant amino acid; and comparing the IMM score for the variant amino acid to a IMM score of the input amino acid sequence of the protein.
43. The method of claim 42, wherein a larger difference between the IMM score for the variant amino acid and the IMM score of the input amino acid sequence of the protein is indicative of a larger impact of the introduced one or more amino acid mutations.
44. The method of claim 40 or 41, wherein the computational model comprises a NetMHC pan model.
45. The method of any one of claims 40-44, wherein the input amino acid sequence is greater than 30 amino acids in length, and wherein the plurality of subsequences is less than 20 amino acids in length.
46. The method of claim 45, wherein the plurality of subsequences are about 15 amino acids in length.
47. The method of any one of claims 40-46, wherein the plurality of MHC alleles is a population-wide representation of MHC alleles, wherein the population-wide representation of MHC alleles comprises at least 50 different MHC alleles, and wherein the plurality of MHC alleles comprise class II MHC alleles or class I MHC alleles.
48. The method of any one of claims 40-47, wherein transforming the corresponding score to generate the population-wide measure of presentation comprises weighting the corresponding score by population- wide frequencies of the plurality of MHC alleles.
49. The method of claim 40, wherein the MHC alleles are selected from DRB 1*04:03, DRBl*04:01, DRBl*04:05, DRBl*09:01, DRBl*07:01, DRBl*16:01, DRB1*15:O2, DRBl*01:01, DRBl*01:03, DRBl*10:01, DRBl*01:02, DRB1*12:O1, DRBl*15:01,83IPTS / 200129719.1Attorney Docket No.: SES-018WO PATENTDRB 1 *04:02, DRB 1 *03:02, DRB 1 *03 :01 , DRB 1 *08:03, DRB 1 *08:07, DRB1 * 14:02, DRBl*l l:01, DRBl*13:03, DRB1*14:O7, DRB1*14:O1, DRBl*13:01, DRB1*11:O4, DRB1*14:O6, DRB1*13:O2, DRB3*01:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, DRB5*01:01, HLA-DPAl*02:02 / DPBl*26:01, HLA-DPAl*02:02 / DPB 1*13:01, HLA-DPAl*02:02 / DPBl*03:01, HLA-DPAl*01:03 / DPB 1*03:01, HLA-DPA 1 *01 :03 / DPB 1*11:01, HLA-DPA 1 *02:02 / DPB 1 *09:01 , HLA-DPA 1 *01 :03 / DPB 1*17:01, HLA-DPA1*O2:O1 / DPB1*18:O1, HLA-DPA l*01:03 / DPBl* 18:01, HLA- DPAl*02:02 / DPBl*15:01, HLA-DPAl*02:02 / DPBl*19:01, HLA-DPAl*01:03 / DPBl*02:01, HLA-DPA l*01:03 / DPB 1*04:01, HLA-DPAl*02:02 / DPBl*04:01, HLA-DPAl*02:02 / DPBl*01:01, HLA-DPAl*01:03 / DPBl*01:01, HLA-DPAl*02:02 / DPBl*05:01, HLA-DPA 1 *01 :03 / DPB 1*19:01, HLA-DPA l*01:03 / DPB 1*05:01, HLA-DQA1 *05 :01 / DQB 1*04:02, HLA-DQAl*01:02 / DQB 1*05:01, HLA-DQA 1 *05 :01 / DQB 1*05:01, HLA-DQAl*05:01 / DQBl*02:01, HLA- DQAl*01:02 / DQBl*06:04, HLA-DQAl*01:02 / DQB 1*03:01, HLA- DQAl*01:02 / DQB 1*02:01, HLA-DQA1 *01 :03 / DQB 1*06:03, HLA- DQAl*01:02 / DQBl*06:02, HLA-DQAl*05:01 / DQBl*03:01, HLA- DQA 1 *05 :01 / DQB 1*03:03, or any combination thereof.
50. A designed protein comprising one or more amino acid mutations relative to a protein, the designed protein exhibiting improved at least one target property in comparison to the protein, wherein the designed protein is identified by performing the method of any one of claims 1-49.84IPTS / 200129719.1
Citation Information
Cited By
A method and device for designing a short peptide of a protease substrate based on an eve model
CN122245415A
A method and device for designing a short peptide of a protease substrate based on an eve model
CN122245415B