Methods and compositions for protein expressions
Expression cassettes with optimized nucleotide sequences enhance recombinant protein yield by 1% to 1000-fold, addressing yield and stability issues in generating mussel adhesive proteins for various applications.
Patent Information
- Application Number
- PCT/CN2025/075373
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-07
AI Technical Summary
Existing methods for generating recombinant proteins, such as mussel adhesive proteins, face challenges in achieving sufficient yield and stability due to non-endogenous sequences like affinity tags, which can affect protein properties and increase costs.
The use of expression cassettes with additional nucleotides that facilitate higher expression levels of target proteins, including mussel foot proteins, by optimizing the nucleotide sequence to enhance translation efficiency and stability.
The expression cassettes enable increased yield of recombinant proteins by 1% to 1000-fold compared to conventional methods, providing a cost-effective solution for therapeutic, cosmetic, and industrial applications.
Smart Images

Figure CN2025075373_07082025_PF_FP_ABST
Abstract
Description
METHODS AND COMPOSITIONS FOR PROTEIN EXPRESSIONSCROSS-REFERENCE
[0001] This application claims the benefit of International Patent Application No. PCT / CN2024 / 075024, filed January 31, 2024, which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Natural proteins or peptides from different organisms have characteristics suitable for diverse therapeutic, cosmetic, or industrial applications.SUMMARY
[0003] Provided herein, are polynucleotides. In an aspect, a polynucleotide comprises a sequence encoding a polypeptide, wherein said sequence encoding said polypeptide comprises a start codon, a sequence encoding an affinity tag peptide, at least one additional nucleotide, and a target sequence, wherein said target sequence comprises a sequence encoding a mussel foot protein or a fragment thereof, wherein said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1 %higher than an expression level of a comparable polypeptide encoded by a comparable polynucleotide under a same condition, and wherein said comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one additional nucleotide. In some embodiments, said at least one additional nucleotide is 5’ to said sequence encoding said affinity tag peptide. In some embodiments, said at least one additional nucleotide is 3’ to said sequence encoding said affinity tag peptide. In some embodiments, said at least one additional nucleotide comprises at least two nucleotides, and wherein one of said at least two nucleotides is 5’ to said sequence encoding said affinity tag peptide, and another one of said at least two nucleotides is 3’ to said sequence encoding said affinity tag peptide. In some embodiments, said polynucleotide further comprises a sequence encoding a cleavage site. In some embodiments, said sequence encoding said cleavage site is 5’ to said target sequence. In some embodiments, said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, said at least one additional nucleotide, said sequence encoding said affinity tag peptide, said sequence encoding said cleavage site, and said target sequence. In some embodiments, said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, said sequence encoding said affinity tag peptide, said at least one additional nucleotide, said sequence encoding said cleavage site, and said target sequence. In some embodiments, said at least one additional nucleotide comprises at least two nucleotides, and wherein said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, one of said at least two nucleotides, a sequence encoding said affinity tag peptide, another one of said at least two nucleotides, said sequence encoding said cleavage site, and said target sequence. In some embodiments, said cleavage site is a protease cleavage site. In some embodiments, said protease cleavage site comprises a cleavage site of a Tobacco Etch Virus (TEV) protease, a Human rhinovirus 3C protease (HRV3C) , a thrombin, Enterokinase, Factor Xa, Chymotrypsin, Collagenase, Dispase, Endopeptidase Arg-C, Endopeptidase Asp-N, Endopeptidase Glu-C, Endopeptidase Lys-C, Ficin, Kallikrein, Papain, Pepsin, plasmin, pronase, proteinase K, Subtilisin, Thermolysin, trypsin, or a combination thereof. In some embodiments, said cleavage site is a chemical cleavage site. In some embodiments, said chemical cleavage site comprises a cleavage site of cyanogen bromide (CNBr) , hydroxylamine, formic acid, or a combination thereof. In some embodiments, said cleavage site comprises an intein cleavage site. In some embodiments, said sequence encoding said cleavage site is 3’ to said target sequence.
[0004] Provided herein, are polynucleotides. In an aspect, a polynucleotide comprises a sequence encoding a polypeptide, wherein the sequence encoding the polypeptide comprises a start codon, at least one additional nucleotide, and a target sequence, wherein said sequence encoding said polypeptide does not comprise a sequence encoding an affinity tag peptide, wherein said target sequence comprises a sequence encoding a mussel foot protein or a fragment thereof, wherein said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1 %higher than an expression level of a comparable polypeptide encoded by a comparable polynucleotide under a same condition, and wherein said comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one additional nucleotide. In some embodiments, said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, said at least one additional nucleotide, said sequence encoding said cleavage site, and said target sequence. In some embodiments, said cleavage site comprises a protease cleavage site. In some embodiments, said protease cleavage site comprises a cleavage site of a Tobacco Etch Virus (TEV) protease, a Human rhinovirus 3C protease (HRV3C) , thrombin, Enterokinase, Factor Xa, Chymotrypsin, Collagenase, Dispase, Endopeptidase Arg-C, Endopeptidase Asp-N, Endopeptidase Glu-C, Endopeptidase Lys-C, Ficin, Kallikrein, Papain, Pepsin, plasmin, pronase, proteinase K, Subtilisin, Thermolysin, trypsin, or a combination thereof. In some embodiments, said cleavage site is a chemical cleavage site. In some embodiments, said chemical cleavage site comprises a cleavage site of cyanogen bromide (CNBr) , hydroxylamine, formic acid, or a combination thereof. In some embodiments, said cleavage site comprises an intein cleavage site. In some embodiments, said sequence encoding said cleavage site is 3’ to said target sequence.
[0005] In some embodiments, said sequence encoding said cleavage site is 3’ to said target sequence. In some embodiments, said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1-fold, 5-fold, 10-fold, or 20-fold higher than said expression level of said comparable polypeptide. In some embodiments, said at least one additional nucleotide comprises at least 3 nucleotides. In some embodiments, said at least one additional nucleotide comprises multiples of 3 nucleotides. In some embodiments, said at least one additional nucleotide comprises at most about 150 nucleotides. In some embodiments, said at least one additional nucleotide comprises at most about 96 nucleotides. In some embodiments, said at least one additional nucleotide comprises at most about 84 nucleotides. In some embodiments, said at least one additional nucleotide comprises at most about 48 nucleotides. In some embodiments, said at least one additional nucleotide comprises at least about 21 nucleotides. In some embodiments, said polynucleotide is a ribonucleic acid (RNA) . In some embodiments, said RNA is a messenger RNA (mRNA) . In some embodiments, said mRNA folds to a secondary structure. In some embodiments, said at least one additional nucleotide is not base-paired with another nucleotide of said polynucleotide. In some embodiments, at least a nucleotide at a 5’ or a 3’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide. In some embodiments, at least a nucleotide at said 5’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide. In some embodiments, at least about 10 nucleotides at said 5’ end of said polynucleotide are single stranded. In some embodiments, about 12-58 nucleotides at said 5’ end of said polynucleotide are single stranded. In some embodiments, at least a nucleotide at said 3’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide. In some embodiments, at least about 10 nucleotides at said 3’ end of said polynucleotide are single stranded. In some embodiments, at least a nucleotide at said 5’ end and at least a nucleotide at said 3’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide. In some embodiments, said mRNA has a minimum free energy of at most about -10 kilocalorie per mole (kcal / mol) . In some embodiments, said mRNA has a minimum free energy of at most about -50 kcal / mol. In some embodiments, said mRNA has a minimum free energy of at least about -57 kcal / mol. In some embodiments, said mRNA has a minimum free energy of at least about -55 kcal / mol. In some embodiments, said mRNA has a minimum free energy of about -53 kcal / mol. In some embodiments, said affinity tag peptide comprises a polyhistidine, a Myc tag, a Strep tag, a V5 tag, an HA tag, a Flag tag, polyarginine, Avi Tag, SNAP-Tag, Halo Tag, small Ub-related modifier protein (SUMO) , Glutathione S-transferase (GST) , Green fluorescent protein (GFP) , maltose-binding protein (MBP) , mCheery, or a combination thereof. In some embodiments, said sequence encoding said polypeptide further comprises at least one second additional nucleotide, wherein said at least one second additional nucleotide is configured to allow said polypeptide to be expressed at a level that is at least 1%higher than a second comparable polypeptide encoded by a second comparable polynucleotide, and wherein said second comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one second additional nucleotide. In some embodiments, said at least one second additional nucleotide comprises at least one nucleotide substitution that is located within said sequence encoding said affinity tag peptide, and optionally wherein (a) said sequence encoding said affinity tag peptide comprising said at least one nucleotide substitution encodes an amino acid sequence that is more hydrophilic than an amino acid sequence encoded by a sequence encoding said affinity tag peptide not comprising said at least one nucleotide substitution; and / or (b) said sequence encoding said affinity tag peptide comprising said at least one nucleotide substitution encodes an amino acid sequence that is more positively charged than an amino acid sequence encoded by a sequence encoding said affinity tag peptide not comprising said at least one nucleotide substitution. In some embodiments, said at least one second additional nucleotide is located at 5’ or at 3’ to said sequence encoding said affinity tag peptide, and wherein said sequence encoding said affinity tag peptide comprising said at least one second additional nucleotide encodes an amino acid sequence that is more positively charged than an amino sequence encoding said affinity tag peptide not comprising said at least one second additional nucleotide. In some embodiments, said at least one second additional nucleotide comprises one nucleotide mutation that is located within said sequence encoding said mussel foot protein or said fragment thereof, optionally wherein said at least one mutation is a synonymous mutation. In some embodiments, said sequence encoding said mussel foot protein or said fragment thereof comprising said at least one second additional nucleotide has a different nucleotide sequence relative to said sequence encoding said mussel foot protein or said fragment thereof not comprising said at least one second additional nucleotide, and wherein said sequence encoding said mussel foot protein or said fragment thereof comprising said at least one second additional nucleotide encodes a same amino acid sequence relative to said sequence encoding a mussel foot protein or said fragment thereof not comprising said at least one second additional nucleotide. In some embodiments, said expression levels of said polypeptide and said comparable polypeptide are measured by levels of said polypeptide and said comparable polypeptide generated by two cells, and wherein each of said two cells comprises only one of said polynucleotide and said comparable polynucleotide. In some embodiments, said levels of said polypeptide and said comparable polypeptides generated by said two cells are measured in milligram of polypeptides per liter of a culture of cells (mg*L-1) . In some embodiments, said two cells are bacteria. In some embodiments, said bacteria is E. coli. In some embodiments, said polynucleotide further comprises a promoter, a translation initiator, a terminator, or a combination thereof.
[0006] Provided herein, are vectors. In an aspect, a vector comprises any of the polynucleotides described herein. In some embodiments, said vector comprises a bacterial vector.
[0007] Provided herein, are cells. In an aspect, a cell comprises any of the polynucleotides described herein or any of the vectors described herein. In some embodiments, said cell is a procaryotic cell. In some embodiments, said cell is a bacteria. In some embodiments, said bacteria is E. coli. In some embodiments, said cell is an eucaryotic cell. In some embodiments, said cell is a yeast cell, a fungal cell, an insect cell, or a mammalian cell.
[0008] Provided herein, are methods. In an aspect, a method comprises expressing said polypeptide encoded by any of the polynucleotides described herein or any of the vectors described herein. In some embodiments, the method further comprises purifying said polypeptide. In some embodiments, the purifying comprises purifying said polypeptide using said affinity tag peptide. In some embodiments, the method further comprises generating a polypeptide comprising said mussel foot protein or said fragment thereof but not said affinity tag peptide. In some embodiments, said generating comprises contacting said polypeptide with a protease. In some embodiments, said protease comprises a Tobacco Etch Virus (TEV) protease, a Human rhinovirus 3C protease (HRV 3C) , a thrombin, an Enterokinase, a Factor Xa, a Chymotrypsin, a Collagenase, a Dispase, an Endopeptidase Arg-C, an Endopeptidase Asp-N, an Endopeptidase Glu-C, an Endopeptidase Lys-C, a Ficin, a Kallikrein, a Papain, a Pepsin, a plasmin, a pronase, a proteinase K, a Subtilisin, a Thermolysin, a trypsin, or, or a combination thereof.
[0009] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative instances of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different instances, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. INCORPORATION BY REFERENCE
[0010] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings ( “FIG. ” or “FIGs. ” herein) , of which:
[0012] FIG. 1A shows a comparison of the physiochemical properties (molecular weight, extinction coefficient, and isoelectric point / PI) of an exemplary protein with or without a peptide affinity tag (also referred to as “purification tag, ” which is used interchangeably herein) . FIG. 1B shows a comparison of the stability of the exemplary protein with or without a peptide affinity tag at various pH. FIG. 1C depicts the design of an exemplary expression cassette for expressing an exemplary target protein encoded by the target sequence. FIG. 1D depicts 4 exemplary expression cassette designs for expressing target proteins. FIG. 1E shows that a target protein was not expressed when not fused to the first sequence depicted in FIG. 1F. FIG. 1F shows an exemplary arrangement of various sequences in an expression cassette as described herein. FIG. 1G shows the SDS-PAGE gels of an exemplary target protein expressed from various expression cassettes screened using the method described herein.
[0013] FIGs. 2A, 2B, 2C, 2D, 2E show the predicted mRNA secondary structures of a first, second, third, fourth, and fifth subset of exemplary expression cassettes, respectively.
[0014] FIG. 3A shows the predicted mRNA secondary structures of a sixth subset of exemplary expression cassettes. FIG. 3B shows the SDS-PAGE protein expression using the constructs of FIG. 3A. FIG. 3C shows the expression data of exemplary constructs with various numbers of negatively charged amino acids C-terminal to the his-tag in the first sequence as described herein and N-terminal to the TEV cleavage site. FIG. 3D shows the expression data of exemplary constructs with various numbers of hydrophilic / hydrophobic residues in the first sequence as described herein. FIG. 3E shows the expression data of exemplary constructs with various numbers of total amino acid residues (length) in the first sequence as described herein.
[0015] FIGs. 4A and 4B show the expression data of various exemplary target proteins expressed using the expression cassettes as described herein.DETAILED DESCRIPTIONOverview
[0016] Natural proteins or peptides from different organisms have characteristics suitable for diverse therapeutic, cosmetic, or industrial applications. However, various problems exist for generating a sufficient amount of the proteins / peptides. For example, extracting the proteins or peptides from the organisms may entail a high cost / time (such as for harvesting or maintaining the organisms, extracting and / or purifying the proteins / peptides from these organisms…etc. ) .
[0017] Expressing the recombinant versions of these natural proteins or peptides in model organisms can overcome the disadvantages of protein extraction from natural organisms. Using recombinant technology, the proteins or peptides can be tagged or fused with various additional sequences for processing (for example, using affinity tag protein / peptides for protein purification or boosting protein expression) . Additionally, the cost for generating recombinant proteins or peptide in model organisms can be substantially lower than the extraction methods.
[0018] However, expressing recombinant proteins / peptides can face various hurdles. For example, the presence of various non-endogenous sequences of the recombinant proteins-such as the affinity tag protein or peptide, which may facilitate or be required for processing / generation of the recombinant proteins / peptides-may have significant undesirable effects on the properties of the proteins / peptides, rendering them unsuitable for different applications. For example, 6 x histidine-tagged insulin can affect the safety of drugs used in clinical applications. Furthermore the non-endogenous sequences of the recombinant proteins may encode a relatively large amount of amino acid residues, which can decrease the yield of the expressed proteins / peptides. Due to different chemistry of various proteins / peptides, the amino acid residues encoded by the non-endogenous sequences may also negatively affect the stability or yield of the recombinant protein expressed. Thus, there exists a need for methods and compositions for expression of recombinant proteins.
[0019] The present disclosure provides compositions and methods for expressing recombinant proteins. The compositions may comprise an expression cassette. As used herein, the expression cassette comprises designs of the nucleotide (s) , amino acid (s) , and / or sequences thereof that can facilitate the production of a recombinant protein. The expression cassette may be used to express a target protein. The nucleotide (s) and / or sequence that can facilitate the production of the target protein may be located 5’ (or N-terminal when referring to the amino acid sequence subsequent to translation of the coding sequence encoded by the expression cassette) and / or 3’ (C-terminal) to the target sequence encoding the target protein; or within the target sequence encoding the target protein. The expression cassette can facilitate the expression of the target protein and thus increase the yield, relative to the existing methods. The present disclosure also provides the methods for identifying / testing / screening the expression cassettes and using the expression cassettes. Thus, the methods and reagents described herein can allow for or facilitate the expressed proteins / peptides tailored for various applications. For example, the methods and reagents described herein can facilitate the expression of mussel adhesive proteins or fragments thereof and increase yield. Mussel adhesive proteins
[0020] Mussel adhesive proteins (MAP) is the adhesive agent used by the sea mussel to attach the animal to various underwater surfaces. Mussels can attach to the reefs on the coast or to the bottom of the ship and withstand wave impacts in the offshore. Mussels can also attach extremely strongly to the substrate of any material, such as metal, wood, or glass. Such abilities derive from the MAPs that are formed and stored in the foot of the mussels. The mussels release the MAPs through the foot silk to a solid surface such as rock to form the water-resistant adhesion. Thus, MAPs have various therapeutic, tissue engineering, and / or industrial applications. However, generation of MAPs has significant challenge, whether the protein is generated by extraction from natural organisms or by recombinant protein expression, as described herein. Using the methods and compositions (such as the expression cassette) as described herein, MAPs can be generated to a sufficient amount with beneficial stability / properties (such as those described herein) . Thus, a beneficial advantage of the present disclosure is to allow for the generation of sufficient MAPs (with a relatively low cost compared to existing methods) for various therapeutic, cosmetic, or industrial applications. In some embodiments, a MAP is encoded by a sequence having at least 80%, 90%, 95%, 99%, or 100%sequence identity to any one of SEQ ID NOs: 42-44 and 70-87. In some embodiments, a MAP is encoded by a sequence having at least 80%, 90%, 95%, 99%, or 100%sequence identity to any one of SEQ ID NOs: 42. In some embodiments, a MAP comprises a sequence having at least 80%, 90%, 95%, 99%, or 100%sequence identity to any one of SEQ ID NOs: 95-112. In some embodiments, a MAP comprises a sequence having at least 80%, 90%, 95%, 99%, or 100%sequence identity to any one of SEQ ID NO: 107. In some embodiments, the MAP is a target protein encoded by the polynucleotide described herein. Expression cassette
[0021] Provided herein are expression cassettes. The expression cassette can comprise nucleotide (s) , amino acid (s) , or sequences thereof that express or facilitate expression of a target protein / peptide. The expression cassettes described herein can be a nucleic acid or a polypeptide encoded by the nucleic acid. In some cases, when referring to a particular sequence of the expression cassette and / or the target protein / peptide, the particular sequence can be a nucleotide sequence or an amino acid sequence. In some cases, when referring to a particular residue of the expression cassette and / or the target protein / peptide, the particular residue can be a nucleotide or an amino acid. In some cases, when referring to the expression cassette, a nucleotide (s) and an amino acid can be used interchangeably (when the nucleotide (s) encode the amino acid) . In some cases, the expression cassette is a polynucleotide comprising a sequence encoding a polypeptide, wherein said sequence encoding said polypeptide comprises a start codon, a sequence encoding an affinity tag peptide, at least one additional nucleotide, and a target sequence. In some cases, the expression cassette is a polynucleotide comprising a sequence encoding a polypeptide, wherein said sequence encoding said polypeptide comprises a start codon, at least one additional nucleotide, and a target sequence. In some cases, the target sequence comprises a sequence encoding a mussel foot protein or a fragment thereof. In some cases, said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1 %higher than an expression level of a comparable polypeptide encoded by a comparable polynucleotide under a same condition, wherein said comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one additional nucleotide. In some cases, the polynucleotide further comprises a sequence encoding a cleavage site.
[0022] The expression cassette can comprise a nucleic acid. In some instances, a nucleic acid may be a nucleic acid molecule. In some cases, a nucleic acid may be a species / type of nucleic acid. In some cases, a nucleic acid may comprise a polymeric form of nucleotides. In some cases, a nucleic acid may comprise a polynucleotide. In some cases, the nucleic acid may comprise a sequence of nucleotides (i.e., a nucleic acid sequence) . In some cases, a nucleic acid may comprise a modified polynucleotide. In some cases, a nucleic acid may comprise a canonical or non-canonical nucleotide. A canonical nucleotide may comprise adenosine with base types (A) , cytosine (C) , guanine (G) , thymine (T) , uracil (U) , or variants thereof. When referring to a base type of a nucleotide or a polynucleotide, T and U may be interchangeable. When referring to a sequence of a nucleic acid, the sequence may comprise the complementary form of the sequence. The complementary of the nucleic acid sequence may be based on canonical base-pairing of the nucleotides or nucleic acids. A nucleic acid may comprise one or more modified nucleotides or nucleotide analogs.
[0023] In some cases, a nucleic acid may be single-stranded, double-stranded, triple stranded, or a combination thereof. In some cases, an engineered nucleic acid may be single-stranded. In some cases, an engineered nucleic acid may be double-stranded. In some cases, an engineered nucleic acid may comprise single-stranded and double-stranded regions or portions.
[0024] In some cases, a nucleic acid may be linear. In some cases, a nucleic acid may be closed linear double-stranded (e.g., a doggybone DNA) . In some cases, a nucleic acid may be circular. In some cases, a nucleic acid may be branched. A nucleic acid may comprise a deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) . A nucleic acid may comprise a peptide engineered nucleic acid (PNA) , an unlocked engineered nucleic acid (UNA) , a locked engineered nucleic acid (LNA) , or a combination thereof. A nucleic acid may comprise at least a coding sequence, at least a non-coding sequence, or a combination thereof. A nucleic acid may comprise at least a coding sequence. A nucleic acid may comprise at least a non-coding sequence. A nucleic acid may comprise a coding or non-coding region of a gene or gene fragment, a locus defined from linkage analysis, an exons, an intron, an intein, or any combination thereof. A coding sequence may be a gene sequence, a codon-modified thereof, or a codon-optimized thereof. A coding sequence may be a variation of gene sequence. A coding sequence can comprise a target sequence (e.g., encoding MAPs) . A coding sequence can comprise a sequence encoding an affinity tag peptide (e.g., Histidine tags) . A coding sequence can comprise a start codon, a sequence encoding an affinity tag peptide, at least one additional nucleotide described herein that facilitate expression of a target sequence, and the target sequence (e.g., encoding MAPs) . A coding sequence can comprise a start codon, at least one additional nucleotide that facilitate expression of a target sequence, and the target sequence (e.g., encoding MAPs) and without a sequence encoding an affinity tag peptide. A non-coding sequence may be a gene sequence. A non-coding sequence may also be a variation of gene sequence. The variation may comprise a mutation. A mutation, in some cases, may comprise a nucleotide substitution, deletion, insertion, inversion, or a combination thereof. A mutation may alter the codon of a coding sequence or is a non-synonymous mutation. A mutation may not alter the codon of a coding sequence or is a synonymous mutation.
[0025] The expression cassette described herein can comprise the at least one nucleotide (or “an additional nucleotide, ” which is used interchangeably in the specification, or “variable region, ” which is used interchangeably in the examples and drawings) that facilitates (or is configured to facilitate) the expression of a target protein / peptide as described herein. The additional nucleotide may facilitate or allow for: (1) the ribonucleic acid form of the expression cassette to adopt a certain secondary structure as described herein; (2) the protein / polypeptide encoded by the sequence of the expression cassette to possess a certain characteristic as described herein; or (3) a combination thereof. The additional nucleotide may not overlap with the sequence (s) of the start codon, the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or a combination thereof. In other cases, the additional nucleotide may overlap with the sequence (s) of the start codon, the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or a combination thereof. The selection of whether the additional nucleotide overlaps with the sequence (s) of the start codon, the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or a combination thereof can be determined using the methods for identifying expression cassettes as described herein, depending on the yield of the protein / peptide expression, as described herein. In some cases, the at least one nucleotide forms at least one circular structure. In some cases, the at least one nucleotide forms at least one single-stranded structure at its 5’ end. In some cases, the at least one nucleotide forms at least one single-stranded structure at its 3’ end. In some cases, the at least one nucleotide forms at least one single-stranded structure at its 5’ end and / or 3’ end with multiple circular structures. In some cases, the at least one nucleotide forms a flexible RNA secondary structure that allows for efficient protein expression or translation.
[0026] The expression cassette comprising the additional nucleotide may facilitate the expression of the target protein / peptide at a level that is at least about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 100 %, 150 %, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, or 1000-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 50%higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 100%higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 2-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 5-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 10-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 15-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 20-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 50-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 100-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide can facilitate the expression of the target protein / peptide at a level that is at least about 500-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide.
[0027] The expression cassette comprising the additional nucleotide may facilitate the expression of the target protein / peptide at a level that is at most about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 100 %, 150 %, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, or 1000-fold higher than that of a comparable expression cassette having the same sequence except for the additional nucleotide. The level of the protein / peptide expression may be measured by the protein / peptide expressed by a cell (as described herein) that comprises the expression cassettes versus a cell that comprises the comparable expression cassette.
[0028] In some cases, the level of the protein / peptide expression may be measured as a yield protein / peptide expressed per volume of the cell culture. In some cases, the expression cassette described herein may facilitate the expression of the target protein / peptide at a level of at least about 1 picogram (pg) per liter, 10 pg per liter, 100 pg per liter, 1 nanogram (ng) per liter, 10 ng per liter, 100 ng per liter, 1 microgram (μg) per liter, 10 μg per liter, 100 μg per liter, 1 milligram (mg) per liter, 10 mg per liter, 100 mg per liter or more. In some cases, the expression cassette described herein may facilitate the expression of the target protein / peptide at a level of at least about 1 picogram (pg) per milliliter, 10 pg per milliliter, 100 pg per milliliter, 1 nanogram (ng) per milliliter, 10 ng per milliliter, 100 ng per milliliter, 1 microgram (μg) per milliliter, 10 μg per milliliter, 100 μg per milliliter, 1 milligram (mg) per milliliter, 3 mg per milliliter, 10 mg per milliliter, 100 mg per milliliter or more. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 500mg per L. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 1000mg per L. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 2000mg per L. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 3000mg per L. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 4000mg per L. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 4400mg per L. In some cases, the expression cassette described herein can facilitate the expression of the target protein / peptide at a level of at least about 5000mg per L. In some cases, the expression of the target protein / peptide is conducted in a 1 liter (L) batch, 5L batch, 50L batch, or 300 L batch. In some cases, the target protein is an MAP (e.g., mefp5, mgfp5, mcfp6, and mufp5) . The expression cassette described herein may facilitate the expression of the target protein / peptide at a level of at most about 1 picogram (pg) per liter, 10 pg per liter, 100 pg per liter, 1 nanogram (ng) per liter, 10 ng per liter, 100 ng per liter, 1 microgram (μg) per liter, 10 μg per liter, 100 μg per liter, 1 milligram (mg) per liter, 10 mg per liter, or 100 mg per liter. The expression cassette described herein may facilitate the expression of the target protein / peptide at a level of at most about 1 picogram (pg) per milliliter, 10 pg per milliliter, 100 pg per milliliter, 1 nanogram (ng) per milliliter, 10 ng per milliliter, 100 ng per milliliter, 1 microgram (μg) per milliliter, 10 μg per milliliter, 100 μg per milliliter, 1 milligram (mg) per milliliter, 10 mg per milliliter, or 100 mg per milliliter.
[0029] In some cases, the level of the protein / peptide expression may be measured as a yield protein / peptide expressed per numbers of host cell (such as E. coli) . For example, the expression cassette described herein may facilitate the expression of the target protein / peptide at a level of at least about 1 picogram (pg) per 1x10^10 host cells, 10 pg per 1x10^10 host cells, 100 pg per 1x10^10 host cells, 1 nanogram (ng) per 1x10^10 host cells, 10 ng per 1x10^10 host cells, 100 ng per 1x10^10 host cells, 1 microgram (μg) per 1x10^10 host cells, 10 μg per 1x10^10 host cells, 100 μg per 1x10^10 host cells, 1 milligram (mg) per 1x10^10 host cells, 10 mg per 1x10^10 host cells, 100 mg per 1x10^10 host cells or more. The expression cassette described herein may facilitate the expression of the target protein / peptide at a level of at most about 1 picogram (pg) per 1x10^10 host cells, 10 pg per 1x10^10 host cells, 100 pg per 1x10^10 host cells, 1 nanogram (ng) per 1x10^10 host cells, 10 ng per 1x10^10 host cells, 100 ng per 1x10^10 host cells, 1 microgram (μg) per 1x10^10 host cells, 10 μg per 1x10^10 host cells, 100 μg per 1x10^10 host cells, 1 milligram (mg) per 1x10^10 host cells, 10 mg per 1x10^10 host cells, or 100 mg per 1x10^10 host cells.
[0030] The expression cassette may comprise a sequence of an affinity tag protein / peptide, a sequence of a cleavage site, as described herein, or a combination thereof. Thus, the expression cassette may facilitate the expression of a fusion protein / peptide comprising an affinity tag protein / peptide, a cleavage site, a target protein / peptide, a start codon, as described herein, or a combination thereof. The fusion protein may further comprise the additional nucleotide as described herein. The expression cassette comprising the additional nucleotide may facilitate the expression of a fusion protein / peptide at a level that is at least about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 100 %, 150 %, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, or 1000-fold higher than that of a comparable expression cassette comprising the same sequence except for the additional nucleotide. The expression cassette comprising the additional nucleotide may facilitate the expression of a fusion protein / peptide at a level that is at most about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 100 %, 150 %, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, or 1000-fold higher than that of a comparable expression cassette comprising the same sequence except for the additional nucleotide. The level of a protein / peptide expression may be measured by a protein / peptide expressed by a cell (as described herein) that comprises the expression cassettes versus a cell that comprises the comparable expression cassette. In some cases, a level of a protein / peptide expression may be measured as a yield protein / peptide expressed per volume of a cell culture. For example, an expression cassette described herein may facilitate the expression of a fusion protein / peptide at a level of at least about 1 picogram (pg) per liter, 10 pg per liter, 100 pg per liter, 1 nanogram (ng) per liter, 10 ng per liter, 100 ng per liter, 1 microgram (μg) per liter, 10 μg per liter, 100 μg per liter, 1 milligram (mg) per liter, 10 mg per liter, 100 mg per liter or more. The expression cassette described herein may facilitate the expression of a fusion protein / peptide at a level of at least about 1 picogram (pg) per milliliter, 10 pg per milliliter, 100 pg per milliliter, 1 nanogram (ng) per milliliter, 10 ng per milliliter, 100 ng per milliliter, 1 microgram (μg) per milliliter, 10 μg per milliliter, 100 μg per milliliter, 1 milligram (mg) per milliliter, 10 mg per milliliter, 100 mg per milliliter or more. The expression cassette described herein may facilitate the expression of a fusion protein / peptide at a level of at most about 1 picogram (pg) per liter, 10 pg per liter, 100 pg per liter, 1 nanogram (ng) per liter, 10 ng per liter, 100 ng per liter, 1 microgram (μg) per liter, 10 μg per liter, 100 μg per liter, 1 milligram (mg) per liter, 10 mg per liter, or 100 mg per liter. The expression cassette described herein may facilitate the expression of a fusion protein / peptide at a level of at most about 1 picogram (pg) per milliliter, 10 pg per milliliter, 100 pg per milliliter, 1 nanogram (ng) per milliliter, 10 ng per milliliter, 100 ng per milliliter, 1 microgram (μg) per milliliter, 10 μg per milliliter, 100 μg per milliliter, 1 milligram (mg) per milliliter, 10 mg per milliliter, or 100 mg per milliliter. In some cases, the level of the protein / peptide expression may be measured as a yield protein / peptide expressed per numbers of host cell (such as E. coli) . For example, the expression cassette described herein may facilitate the expression of the fusion protein / peptide at a level of at least about 1 picogram (pg) per 1x10^10 host cells, 10 pg per 1x10^10 host cells, 100 pg per 1x10^10 host cells, 1 nanogram (ng) per 1x10^10 host cells, 10 ng per 1x10^10 host cells, 100 ng per 1x10^10 host cells, 1 microgram (μg) per 1x10^10 host cells, 10 μg per 1x10^10 host cells, 100 μg per 1x10^10 host cells, 1 milligram (mg) per 1x10^10 host cells, 10 mg per 1x10^10 host cells, 100 mg per 1x10^10 host cells or more. The expression cassette described herein may facilitate the expression of the fusion protein / peptide at a level of at most about 1 picogram (pg) per 1x10^10 host cells, 10 pg per 1x10^10 host cells, 100 pg per 1x10^10 host cells, 1 nanogram (ng) per 1x10^10 host cells, 10 ng per 1x10^10 host cells, 100 ng per 1x10^10 host cells, 1 microgram (μg) per 1x10^10 host cells, 10 μg per 1x10^10 host cells, 100 μg per 1x10^10 host cells, 1 milligram (mg) per 1x10^10 host cells, 10 mg per 1x10^10 host cells, or 100 mg per 1x10^10 host cells.
[0031] The additional nucleotide may have at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 200, 300, 600, 900 or more nucleotides (nt; or base pair / bp) . The additional nucleotide may have at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 nucleotides. The additional nucleotide may have at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 200, 300, 600, or 900 nt. The additional nucleotide can have at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 nucleotides. The additional nucleotide may have at most about 48 nucleotides. The additional nucleotide can have at most about 66 nucleotides. The additional nucleotide can have at most about 78 nucleotides. The additional nucleotide can have at most about 84 nucleotides. The additional nucleotide may have multiple (s) of 3 nt. For example, the additional nucleotide may have at least about 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, 63, 66, 69, 72, 75, 78, 81, 84, 87, 90, 93, 96, 99, 102, 105, 108, 111, 114, 117, 120, 123, 126, 129, 132, 135, 138, 141, 144, 147, 150, 153, 156, 159, 162, 165, 168, 171, 174, 177, 180, 183, 186, 189, 192, 195, 198, 201, 300, 600, 900 or more nt. The additional nucleotide may have at most about 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, 63, 66, 69, 72, 75, 78, 81, 84, 87, 90, 93, 96, 99, 102, 105, 108, 111, 114, 117, 120, 123, 126, 129, 132, 135, 138, 141, 144, 147, 150, 153, 156, 159, 162, 165, 168, 171, 174, 177, 180, 183, 186, 189, 192, 195, 198, 201, 300, 600, or 900 nt. In some cases, the additional nucleotide has 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, 63, 66, 69, 72, 75, 78, 81, or 84 nucleotides. In some cases, the additional nucleotide has 18 nucleotides. In some cases, the additional nucleotide has 33 nucleotides. In some cases, the additional nucleotide has 48 nucleotides. In some cases, the additional nucleotide has 63 nucleotides. In some cases, the additional nucleotide has 54 nucleotides. In some cases, the additional nucleotide has 60 nucleotides. In some cases, the additional nucleotide has 66 nucleotides. In some cases, the additional nucleotide has 69 nucleotides. In some cases, the additional nucleotide has 72 nucleotides. In some cases, the additional nucleotide has 78 nucleotides. In some cases, the additional nucleotide has 84 nucleotides. RNA structures
[0032] In some cases, the additional nucleotide may facilitate the transcribed form of the expression cassette (e.g., mRNA) to adopt certain secondary structures that facilities the expression of the target protein / peptide or fusion protein / peptide, as described herein.
[0033] Polynucleotides or RNA can fold into a various three dimensional (3D) structures. For example, transfer RNA molecule (tRNA) can fold into an L-shaped 3D structure allowing it to fit into the P and A sites of a ribosome and function as the physical link between the polypeptide coding sequence of mRNA and the amino acid sequence of the polypeptide. Since base pairing between complementary sequences of nucleobases determines the overall secondary structure of the nucleic acid molecules, sequences predicted to or known to be able to adopt a particular structure (e.g., a stem-loop) contribute significantly to the design and utility of some types of functional elements or motifs (e.g., RNA elements) .
[0034] The overall secondary structure of the RNA or mRNA can be described / measured using minimum free energy (MFE) . MFE comprises the RNA structure found by thermodynamic optimization (i.e., an implementation of the Zuker algorithm (M. Zuker and P. Stiegler., Nucleic Acids Research 9: 133-148 (1981) ) that has the lowest free energy value.
[0035] In the methods and compositions described herein, the sequence (s) and nucleotide residue (s) (or the configuration (s) of the sequences (s) or nucleotide reside (s) ) can be designed such that the ribonucleic acid or transcribed form of the expression cassette to adopt a certain RNA structure (s) , as described herein, such that the recombinant protein can be expressed to the yield as described herein.
[0036] The additional nucleotide as described herein may facilitate the transcribed form of the expression cassette to have an MFE that is at least about 1 %, at least about 2 %, at least about 3 %, at least about 4 %, at least about 5 %, at least about 6 %, at least about 7 %, at least about 8 %, at least about 9 %, at least about 10 %, at least about 20 %, at least about 30 %, at least about 40 %, at least about 50 %, at least about 60 %, at least about 70 %, at least about 80 %, at least about 90 %, at least about 100 %, at least about 150 %, at least about 2-fold, at least about 3-fold, or at least about 4-fold lower than that of a transcribed form of a comparable expression cassette that does not comprise the additional nucleotide. The additional nucleotide as described herein may facilitate the transcribed form of the expression cassette to have an MFE that is at most about 1 %, at most about 2 %, at most about 3 %, at most about 4 %, at most about 5 %, at most about 6 %, at most about 7 %, at most about 8 %, at most about 9 %, at most about 10 %, at most about 20 %, at most about 30 %, at most about 40 %, at most about 50 %, at most about 60 %, at most about 70 %, at most about 80 %, at most about 90 %, at most about 100 %, at most about 150 %, at most about 2-fold, at most about 3-fold, or at most about 4-fold lower than that of a transcribed form of a comparable expression cassette that does not comprise the additional nucleotide. MFE can be measured as Kilocalories per mole (kcal / mol) . In some cases, the MFE of the transcribed form of the expression cassette as described herein may be at least about -30 kcal / mol, -30.1 kcal / mol, -30.2 kcal / mol, -30.3 kcal / mol, -30.4 kcal / mol, -30.5 kcal / mol, -30.6 kcal / mol, -30.7 kcal / mol, -30.8 kcal / mol, -30.9 kcal / mol, -31 kcal / mol, -31.1 kcal / mol, -31.2 kcal / mol, -31.3 kcal / mol, -31.4 kcal / mol, -31.5 kcal / mol, -31.6 kcal / mol, -31.7 kcal / mol, -31.8 kcal / mol, -31.9 kcal / mol, -32 kcal / mol, -32.1 kcal / mol, -32.2 kcal / mol, -32.3 kcal / mol, -32.4 kcal / mol, -32.5 kcal / mol, -32.6 kcal / mol, -32.7 kcal / mol, -32.8 kcal / mol, -32.9 kcal / mol, -33 kcal / mol, -33.1 kcal / mol, -33.2 kcal / mol, -33.3 kcal / mol, -33.4 kcal / mol, -33.5 kcal / mol, -33.6 kcal / mol, -33.7 kcal / mol, -33.8 kcal / mol, -33.9 kcal / mol, -34 kcal / mol, -34.1 kcal / mol, -34.2 kcal / mol, -34.3 kcal / mol, -34.4 kcal / mol, -34.5 kcal / mol, -34.6 kcal / mol, -34.7 kcal / mol, -34.8 kcal / mol, -34.9 kcal / mol, -35 kcal / mol, -35.1 kcal / mol, -35.2 kcal / mol, -35.3 kcal / mol, -35.4 kcal / mol, -35.5 kcal / mol, -35.6 kcal / mol, -35.7 kcal / mol, -35.8 kcal / mol, -35.9 kcal / mol, -36 kcal / mol, -36.1 kcal / mol, -36.2 kcal / mol, -36.3 kcal / mol, -36.4 kcal / mol, -36.5 kcal / mol, -36.6 kcal / mol, -36.7 kcal / mol, -36.8 kcal / mol, -36.9 kcal / mol, -37 kcal / mol, -37.1 kcal / mol, -37.2 kcal / mol, -37.3 kcal / mol, -37.4 kcal / mol, -37.5 kcal / mol, -37.6 kcal / mol, -37.7 kcal / mol, -37.8 kcal / mol, -37.9 kcal / mol, -38 kcal / mol, -38.1 kcal / mol, -38.2 kcal / mol, -38.3 kcal / mol, -38.4 kcal / mol, -38.5 kcal / mol, -38.6 kcal / mol, -38.7 kcal / mol, -38.8 kcal / mol, -38.9 kcal / mol, -39 kcal / mol, -39.1 kcal / mol, -39.2 kcal / mol, -39.3 kcal / mol, -39.4 kcal / mol, -39.5 kcal / mol, -39.6 kcal / mol, -39.7 kcal / mol, -39.8 kcal / mol, -39.9 kcal / mol, -40 kcal / mol, -40.1 kcal / mol, -40.2 kcal / mol, -40.3 kcal / mol, -40.4 kcal / mol, -40.5 kcal / mol, -40.6 kcal / mol, -40.7 kcal / mol, -40.8 kcal / mol, -40.9 kcal / mol, -41 kcal / mol, -41.1 kcal / mol, -41.2 kcal / mol, -41.3 kcal / mol, -41.4 kcal / mol, -41.5 kcal / mol, -41.6 kcal / mol, -41.7 kcal / mol, -41.8 kcal / mol, -41.9 kcal / mol, -42 kcal / mol, -42.1 kcal / mol, -42.2 kcal / mol, -42.3 kcal / mol, -42.4 kcal / mol, -42.5 kcal / mol, -42.6 kcal / mol, -42.7 kcal / mol, -42.8 kcal / mol, -42.9 kcal / mol, -43 kcal / mol, -43.1 kcal / mol, -43.2 kcal / mol, -43.3 kcal / mol, -43.4 kcal / mol, -43.5 kcal / mol, -43.6 kcal / mol, -43.7 kcal / mol, -43.8 kcal / mol, -43.9 kcal / mol, -44 kcal / mol, -44.1 kcal / mol, -44.2 kcal / mol, -44.3 kcal / mol, -44.4 kcal / mol, -44.5 kcal / mol, -44.6 kcal / mol, -44.7 kcal / mol, -44.8 kcal / mol, -44.9 kcal / mol, -45 kcal / mol, -45.1 kcal / mol, -45.2 kcal / mol, -45.3 kcal / mol, -45.4 kcal / mol, -45.5 kcal / mol, -45.6 kcal / mol, -45.7 kcal / mol, -45.8 kcal / mol, -45.9 kcal / mol, -46 kcal / mol, -46.1 kcal / mol, -46.2 kcal / mol, -46.3 kcal / mol, -46.4 kcal / mol, -46.5 kcal / mol, -46.6 kcal / mol, -46.7 kcal / mol, -46.8 kcal / mol, -46.9 kcal / mol, -47 kcal / mol, -47.1 kcal / mol, -47.2 kcal / mol, -47.3 kcal / mol, -47.4 kcal / mol, -47.5 kcal / mol, -47.6 kcal / mol, -47.7 kcal / mol, -47.8 kcal / mol, -47.9 kcal / mol, -48 kcal / mol, -48.1 kcal / mol, -48.2 kcal / mol, -48.3 kcal / mol, -48.4 kcal / mol, -48.5 kcal / mol, -48.6 kcal / mol, -48.7 kcal / mol, -48.8 kcal / mol, -48.9 kcal / mol, -49 kcal / mol, -49.1 kcal / mol, -49.2 kcal / mol, -49.3 kcal / mol, -49.4 kcal / mol, -49.5 kcal / mol, -49.6 kcal / mol, -49.7 kcal / mol, -49.8 kcal / mol, -49.9 kcal / mol, -50 kcal / mol, -50.1 kcal / mol, -50.2 kcal / mol, -50.3 kcal / mol, -50.4 kcal / mol, -50.5 kcal / mol, -50.6 kcal / mol, -50.7 kcal / mol, -50.8 kcal / mol, -50.9 kcal / mol, -51 kcal / mol, -51.1 kcal / mol, -51.2 kcal / mol, -51.3 kcal / mol, -51.4 kcal / mol, -51.5 kcal / mol, -51.6 kcal / mol, -51.7 kcal / mol, -51.8 kcal / mol, -51.9 kcal / mol, -52 kcal / mol, -52.1 kcal / mol, -52.2 kcal / mol, -52.3 kcal / mol, -52.4 kcal / mol, -52.5 kcal / mol, -52.6 kcal / mol, -52.7 kcal / mol, -52.8 kcal / mol, -52.9 kcal / mol, -53 kcal / mol, -53.1 kcal / mol, -53.2 kcal / mol, -53.3 kcal / mol, -53.4 kcal / mol, -53.5 kcal / mol, -53.6 kcal / mol, -53.7 kcal / mol, -53.8 kcal / mol, -53.9 kcal / mol, -54 kcal / mol, -54.1 kcal / mol, -54.2 kcal / mol, -54.3 kcal / mol, -54.4 kcal / mol, -54.5 kcal / mol, -54.6 kcal / mol, -54.7 kcal / mol, -54.8 kcal / mol, -54.9 kcal / mol, -55 kcal / mol, -55.1 kcal / mol, -55.2 kcal / mol, -55.3 kcal / mol, -55.4 kcal / mol, -55.5 kcal / mol, -55.6 kcal / mol, -55.7 kcal / mol, -55.8 kcal / mol, -55.9 kcal / mol, -56 kcal / mol, -56.1 kcal / mol, -56.2 kcal / mol, -56.3 kcal / mol, -56.4 kcal / mol, -56.5 kcal / mol, -56.6 kcal / mol, -56.7 kcal / mol, -56.8 kcal / mol, -56.9 kcal / mol, -57 kcal / mol, -57.1 kcal / mol, -57.2 kcal / mol, -57.3 kcal / mol, -57.4 kcal / mol, -57.5 kcal / mol, -57.6 kcal / mol, -57.7 kcal / mol, -57.8 kcal / mol, -57.9 kcal / mol, -58 kcal / mol, -58.1 kcal / mol, -58.2 kcal / mol, -58.3 kcal / mol, -58.4 kcal / mol, -58.5 kcal / mol, -58.6 kcal / mol, -58.7 kcal / mol, -58.8 kcal / mol, -58.9 kcal / mol, -59 kcal / mol, -59.1 kcal / mol, -59.2 kcal / mol, -59.3 kcal / mol, -59.4 kcal / mol, -59.5 kcal / mol, -59.6 kcal / mol, -59.7 kcal / mol, -59.8 kcal / mol, -59.9 kcal / mol, -60 kcal / mol, -60.1 kcal / mol, -60.2 kcal / mol, -60.3 kcal / mol, -60.4 kcal / mol, -60.5 kcal / mol, -60.6 kcal / mol, -60.7 kcal / mol, -60.8 kcal / mol, -60.9 kcal / mol, -61 kcal / mol, -61.1 kcal / mol, -61.2 kcal / mol, -61.3 kcal / mol, -61.4 kcal / mol, -61.5 kcal / mol, -61.6 kcal / mol, -61.7 kcal / mol, -61.8 kcal / mol, -61.9 kcal / mol, -62 kcal / mol, -62.1 kcal / mol, -62.2 kcal / mol, -62.3 kcal / mol, -62.4 kcal / mol, -62.5 kcal / mol, -62.6 kcal / mol, -62.7 kcal / mol, -62.8 kcal / mol, -62.9 kcal / mol, -63 kcal / mol, -63.1 kcal / mol, -63.2 kcal / mol, -63.3 kcal / mol, -63.4 kcal / mol, -63.5 kcal / mol, -63.6 kcal / mol, -63.7 kcal / mol, -63.8 kcal / mol, -63.9 kcal / mol, -64 kcal / mol, -64.1 kcal / mol, -64.2 kcal / mol, -64.3 kcal / mol, -64.4 kcal / mol, -64.5 kcal / mol, -64.6 kcal / mol, -64.7 kcal / mol, -64.8 kcal / mol, -64.9 kcal / mol, -65 kcal / mol, -65.1 kcal / mol, -65.2 kcal / mol, -65.3 kcal / mol, -65.4 kcal / mol, -65.5 kcal / mol, -65.6 kcal / mol, -65.7 kcal / mol, -65.8 kcal / mol, -65.9 kcal / mol, -66 kcal / mol, -66.1 kcal / mol, -66.2 kcal / mol, -66.3 kcal / mol, -66.4 kcal / mol, -66.5 kcal / mol, -66.6 kcal / mol, -66.7 kcal / mol, -66.8 kcal / mol, -66.9 kcal / mol, -67 kcal / mol, -67.1 kcal / mol, -67.2 kcal / mol, -67.3 kcal / mol, -67.4 kcal / mol, -67.5 kcal / mol, -67.6 kcal / mol, -67.7 kcal / mol, -67.8 kcal / mol, -67.9 kcal / mol, -68 kcal / mol, -68.1 kcal / mol, -68.2 kcal / mol, -68.3 kcal / mol, -68.4 kcal / mol, -68.5 kcal / mol, -68.6 kcal / mol, -68.7 kcal / mol, -68.8 kcal / mol, -68.9 kcal / mol, -69 kcal / mol, -69.1 kcal / mol, -69.2 kcal / mol, -69.3 kcal / mol, -69.4 kcal / mol, -69.5 kcal / mol, -69.6 kcal / mol, -69.7 kcal / mol, -69.8 kcal / mol, -69.9 kcal / mol, or -70 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein may be at most about -30 kcal / mol, -30.1 kcal / mol, -30.2 kcal / mol, -30.3 kcal / mol, -30.4 kcal / mol, -30.5 kcal / mol, -30.6 kcal / mol, -30.7 kcal / mol, -30.8 kcal / mol, -30.9 kcal / mol, -31 kcal / mol, -31.1 kcal / mol, -31.2 kcal / mol, -31.3 kcal / mol, -31.4 kcal / mol, -31.5 kcal / mol, -31.6 kcal / mol, -31.7 kcal / mol, -31.8 kcal / mol, -31.9 kcal / mol, -32 kcal / mol, -32.1 kcal / mol, -32.2 kcal / mol, -32.3 kcal / mol, -32.4 kcal / mol, -32.5 kcal / mol, -32.6 kcal / mol, -32.7 kcal / mol, -32.8 kcal / mol, -32.9 kcal / mol, -33 kcal / mol, -33.1 kcal / mol, -33.2 kcal / mol, -33.3 kcal / mol, -33.4 kcal / mol, -33.5 kcal / mol, -33.6 kcal / mol, -33.7 kcal / mol, -33.8 kcal / mol, -33.9 kcal / mol, -34 kcal / mol, -34.1 kcal / mol, -34.2 kcal / mol, -34.3 kcal / mol, -34.4 kcal / mol, -34.5 kcal / mol, -34.6 kcal / mol, -34.7 kcal / mol, -34.8 kcal / mol, -34.9 kcal / mol, -35 kcal / mol, -35.1 kcal / mol, -35.2 kcal / mol, -35.3 kcal / mol, -35.4 kcal / mol, -35.5 kcal / mol, -35.6 kcal / mol, -35.7 kcal / mol, -35.8 kcal / mol, -35.9 kcal / mol, -36 kcal / mol, -36.1 kcal / mol, -36.2 kcal / mol, -36.3 kcal / mol, -36.4 kcal / mol, -36.5 kcal / mol, -36.6 kcal / mol, -36.7 kcal / mol, -36.8 kcal / mol, -36.9 kcal / mol, -37 kcal / mol, -37.1 kcal / mol, -37.2 kcal / mol, -37.3 kcal / mol, -37.4 kcal / mol, -37.5 kcal / mol, -37.6 kcal / mol, -37.7 kcal / mol, -37.8 kcal / mol, -37.9 kcal / mol, -38 kcal / mol, -38.1 kcal / mol, -38.2 kcal / mol, -38.3 kcal / mol, -38.4 kcal / mol, -38.5 kcal / mol, -38.6 kcal / mol, -38.7 kcal / mol, -38.8 kcal / mol, -38.9 kcal / mol, -39 kcal / mol, -39.1 kcal / mol, -39.2 kcal / mol, -39.3 kcal / mol, -39.4 kcal / mol, -39.5 kcal / mol, -39.6 kcal / mol, -39.7 kcal / mol, -39.8 kcal / mol, -39.9 kcal / mol, -40 kcal / mol, -40.1 kcal / mol, -40.2 kcal / mol, -40.3 kcal / mol, -40.4 kcal / mol, -40.5 kcal / mol, -40.6 kcal / mol, -40.7 kcal / mol, -40.8 kcal / mol, -40.9 kcal / mol, -41 kcal / mol, -41.1 kcal / mol, -41.2 kcal / mol, -41.3 kcal / mol, -41.4 kcal / mol, -41.5 kcal / mol, -41.6 kcal / mol, -41.7 kcal / mol, -41.8 kcal / mol, -41.9 kcal / mol, -42 kcal / mol, -42.1 kcal / mol, -42.2 kcal / mol, -42.3 kcal / mol, -42.4 kcal / mol, -42.5 kcal / mol, -42.6 kcal / mol, -42.7 kcal / mol, -42.8 kcal / mol, -42.9 kcal / mol, -43 kcal / mol, -43.1 kcal / mol, -43.2 kcal / mol, -43.3 kcal / mol, -43.4 kcal / mol, -43.5 kcal / mol, -43.6 kcal / mol, -43.7 kcal / mol, -43.8 kcal / mol, -43.9 kcal / mol, -44 kcal / mol, -44.1 kcal / mol, -44.2 kcal / mol, -44.3 kcal / mol, -44.4 kcal / mol, -44.5 kcal / mol, -44.6 kcal / mol, -44.7 kcal / mol, -44.8 kcal / mol, -44.9 kcal / mol, -45 kcal / mol, -45.1 kcal / mol, -45.2 kcal / mol, -45.3 kcal / mol, -45.4 kcal / mol, -45.5 kcal / mol, -45.6 kcal / mol, -45.7 kcal / mol, -45.8 kcal / mol, -45.9 kcal / mol, -46 kcal / mol, -46.1 kcal / mol, -46.2 kcal / mol, -46.3 kcal / mol, -46.4 kcal / mol, -46.5 kcal / mol, -46.6 kcal / mol, -46.7 kcal / mol, -46.8 kcal / mol, -46.9 kcal / mol, -47 kcal / mol, -47.1 kcal / mol, -47.2 kcal / mol, -47.3 kcal / mol, -47.4 kcal / mol, -47.5 kcal / mol, -47.6 kcal / mol, -47.7 kcal / mol, -47.8 kcal / mol, -47.9 kcal / mol, -48 kcal / mol, -48.1 kcal / mol, -48.2 kcal / mol, -48.3 kcal / mol, -48.4 kcal / mol, -48.5 kcal / mol, -48.6 kcal / mol, -48.7 kcal / mol, -48.8 kcal / mol, -48.9 kcal / mol, -49 kcal / mol, -49.1 kcal / mol, -49.2 kcal / mol, -49.3 kcal / mol, -49.4 kcal / mol, -49.5 kcal / mol, -49.6 kcal / mol, -49.7 kcal / mol, -49.8 kcal / mol, -49.9 kcal / mol, -50 kcal / mol, -50.1 kcal / mol, -50.2 kcal / mol, -50.3 kcal / mol, -50.4 kcal / mol, -50.5 kcal / mol, -50.6 kcal / mol, -50.7 kcal / mol, -50.8 kcal / mol, -50.9 kcal / mol, -51 kcal / mol, -51.1 kcal / mol, -51.2 kcal / mol, -51.3 kcal / mol, -51.4 kcal / mol, -51.5 kcal / mol, -51.6 kcal / mol, -51.7 kcal / mol, -51.8 kcal / mol, -51.9 kcal / mol, -52 kcal / mol, -52.1 kcal / mol, -52.2 kcal / mol, -52.3 kcal / mol, -52.4 kcal / mol, -52.5 kcal / mol, -52.6 kcal / mol, -52.7 kcal / mol, -52.8 kcal / mol, -52.9 kcal / mol, -53 kcal / mol, -53.1 kcal / mol, -53.2 kcal / mol, -53.3 kcal / mol, -53.4 kcal / mol, -53.5 kcal / mol, -53.6 kcal / mol, -53.7 kcal / mol, -53.8 kcal / mol, -53.9 kcal / mol, -54 kcal / mol, -54.1 kcal / mol, -54.2 kcal / mol, -54.3 kcal / mol, -54.4 kcal / mol, -54.5 kcal / mol, -54.6 kcal / mol, -54.7 kcal / mol, -54.8 kcal / mol, -54.9 kcal / mol, -55 kcal / mol, -55.1 kcal / mol, -55.2 kcal / mol, -55.3 kcal / mol, -55.4 kcal / mol, -55.5 kcal / mol, -55.6 kcal / mol, -55.7 kcal / mol, -55.8 kcal / mol, -55.9 kcal / mol, -56 kcal / mol, -56.1 kcal / mol, -56.2 kcal / mol, -56.3 kcal / mol, -56.4 kcal / mol, -56.5 kcal / mol, -56.6 kcal / mol, -56.7 kcal / mol, -56.8 kcal / mol, -56.9 kcal / mol, -57 kcal / mol, -57.1 kcal / mol, -57.2 kcal / mol, -57.3 kcal / mol, -57.4 kcal / mol, -57.5 kcal / mol, -57.6 kcal / mol, -57.7 kcal / mol, -57.8 kcal / mol, -57.9 kcal / mol, -58 kcal / mol, -58.1 kcal / mol, -58.2 kcal / mol, -58.3 kcal / mol, -58.4 kcal / mol, -58.5 kcal / mol, -58.6 kcal / mol, -58.7 kcal / mol, -58.8 kcal / mol, -58.9 kcal / mol, -59 kcal / mol, -59.1 kcal / mol, -59.2 kcal / mol, -59.3 kcal / mol, -59.4 kcal / mol, -59.5 kcal / mol, -59.6 kcal / mol, -59.7 kcal / mol, -59.8 kcal / mol, -59.9 kcal / mol, -60 kcal / mol, -60.1 kcal / mol, -60.2 kcal / mol, -60.3 kcal / mol, -60.4 kcal / mol, -60.5 kcal / mol, -60.6 kcal / mol, -60.7 kcal / mol, -60.8 kcal / mol, -60.9 kcal / mol, -61 kcal / mol, -61.1 kcal / mol, -61.2 kcal / mol, -61.3 kcal / mol, -61.4 kcal / mol, -61.5 kcal / mol, -61.6 kcal / mol, -61.7 kcal / mol, -61.8 kcal / mol, -61.9 kcal / mol, -62 kcal / mol, -62.1 kcal / mol, -62.2 kcal / mol, -62.3 kcal / mol, -62.4 kcal / mol, -62.5 kcal / mol, -62.6 kcal / mol, -62.7 kcal / mol, -62.8 kcal / mol, -62.9 kcal / mol, -63 kcal / mol, -63.1 kcal / mol, -63.2 kcal / mol, -63.3 kcal / mol, -63.4 kcal / mol, -63.5 kcal / mol, -63.6 kcal / mol, -63.7 kcal / mol, -63.8 kcal / mol, -63.9 kcal / mol, -64 kcal / mol, -64.1 kcal / mol, -64.2 kcal / mol, -64.3 kcal / mol, -64.4 kcal / mol, -64.5 kcal / mol, -64.6 kcal / mol, -64.7 kcal / mol, -64.8 kcal / mol, -64.9 kcal / mol, -65 kcal / mol, -65.1 kcal / mol, -65.2 kcal / mol, -65.3 kcal / mol, -65.4 kcal / mol, -65.5 kcal / mol, -65.6 kcal / mol, -65.7 kcal / mol, -65.8 kcal / mol, -65.9 kcal / mol, -66 kcal / mol, -66.1 kcal / mol, -66.2 kcal / mol, -66.3 kcal / mol, -66.4 kcal / mol, -66.5 kcal / mol, -66.6 kcal / mol, -66.7 kcal / mol, -66.8 kcal / mol, -66.9 kcal / mol, -67 kcal / mol, -67.1 kcal / mol, -67.2 kcal / mol, -67.3 kcal / mol, -67.4 kcal / mol, -67.5 kcal / mol, -67.6 kcal / mol, -67.7 kcal / mol, -67.8 kcal / mol, -67.9 kcal / mol, -68 kcal / mol, -68.1 kcal / mol, -68.2 kcal / mol, -68.3 kcal / mol, -68.4 kcal / mol, -68.5 kcal / mol, -68.6 kcal / mol, -68.7 kcal / mol, -68.8 kcal / mol, -68.9 kcal / mol, -69 kcal / mol, -69.1 kcal / mol, -69.2 kcal / mol, -69.3 kcal / mol, -69.4 kcal / mol, -69.5 kcal / mol, -69.6 kcal / mol, -69.7 kcal / mol, -69.8 kcal / mol, -69.9 kcal / mol, or -70 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein is at most about -40 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein is at most about -50 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein is at most about -55 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein is at most about -60 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein is about 55 kcal / mol. In some cases, the MFE of the transcribed form of the expression cassette as described herein is about -60 kcal / mol.
[0037] In some cases, the transcribed form of the expression cassette is an mRNA.
[0038] The additional nucleotide may facilitate the transcribed form of the expression cassette to adopt various secondary structures. The secondary structures can comprise stem, hairpin loop, pseudoknot, bulge, internal loop, multiloop, single-stranded region, double-stranded region, or any combination thereof. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 stems. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 stems. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 hairpin loops. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 hairpin loops. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 pseudoknots. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 pseudoknots. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 bulges. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 bulges. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 internal loops. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 internal loops. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 multiloops. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 multiloops. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 single-stranded regions. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 single-stranded regions. In some cases, the transcribed form of the expression cassette described herein may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 double-stranded regions. In some cases, the transcribed form of the expression cassette described herein may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 double-stranded regions.
[0039] In some cases, a single-stranded region of the transcribed form of the expression cassette may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 150, or 200 nucleotides. In some cases, a single-stranded region of the transcribed form of the expression cassette may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 150, or 200 nucleotides. The single-stranded region (s) may be at the 5’ or 3’ end of the expression cassette. The single-stranded region (s) may be at the 5’ and 3’ end of the expression cassette. The single-stranded region may be at the 5’ end of the expression cassette. The single-stranded region may be at the 3’ end of the expression cassette. The single-stranded region (s) may be at the 5’ end of the coding sequence of the expression cassette, 3’ end of the coding sequence of the expression cassette, within the coding sequence of the expression cassette, or a combination thereof. The single-stranded region (s) may be at the 5’ end of the coding sequence of the expression cassette. The single-stranded region (s) may be at the 3’ end of the coding sequence of the expression cassette. The single-stranded region (s) may be within the coding sequence of the expression cassette.
[0040] In some cases, a double-stranded region of the transcribed form of the expression cassette may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 200, 300, 600, or 900 nucleotides. In some cases, a double-stranded region of the transcribed form of the expression cassette may comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 200, 300, 600, or 900 nucleotides. Characteristics of target or fusion proteins / peptides
[0041] The expression cassette can comprise a sequence of a target protein / peptide and additional nucleotide. In some cases, the expression construct described herein can generate or express the target protein / peptide or the fusion protein / peptide as described herein. The additional nucleotide may facilitate the target protein / peptide or the fusion protein / peptide to possess the characteristics as described herein.
[0042] The additional nucleotide may encode at least one amino acid. The amino acid may be a canonical amino acid. “Canonical amino acids” refer to those 22 proteinogenic amino acids (20 in the standard genetic code and an additional 2 [selenocysteine and pyrrolysine] that can be incorporated by special translation mechanisms) . Table 6 below lists the 20 proteinogenic amino acids in the standard genetic code with each of their three letter abbreviations, one letter abbreviations, structures, and corresponding codons.Table 6: A list of 20 proteinogenic amino acids in the standard genetic code
[0043] The amino acid may be a non-canonical amino acid. The amino acid may be a naturally-occurring amino acid. The amino acid may be a non-naturally-occurring amino acid. The amino acid may be a proteinogenic amino acid. The amino acid may be a non-proteinogenic amino acid. The additional nucleotide may encode at least one hydrophilic amino acid.
[0044] The additional nucleotide may encode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more hydrophilic amino acids. The additional nucleotide can encode 1, 2, 3, 4, 5, or 6 hydrophilic amino acids. Hydrophilic amino acid may comprise asparagine, glutamine, serine, threonine, lysine, arginine, histidine, aspartic acid, glutamic acid, or any combination thereof. The additional nucleotide may encode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more of any one or combination of asparagine, glutamine, serine, threonine, lysine, arginine, histidine, aspartic acid, or glutamic acid. The additional nucleotide may encode at least one positively charged amino acid. The positively charged amino acid may comprise lysine, arginine, histidine, or any combination thereof. The additional nucleotide may encode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more of a positively charged amino acid. The additional nucleotide may encode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more of any one or combination of lysine, arginine, or histidine. The additional nucleotide can encode 1, 2, 3, 4, 5, or 6 positively charged amino acid, each of which is independently lysine, arginine or histidine. In some cases, the additional nucleotide may not encode any positively charged amino acid.
[0045] The at least one additional nucleotide may encode an amino acid sequence having a net charge of at least about -20, -19, -18, -17, -16, -15, -14, -13, -12, -11, -10, -9, -8, -7, -6, -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. The at least one nucleotide may encode an amino acid sequence having a net charge of at most about -20, -19, -18, -17, -16, -15, -14, -13, -12, -11, -10, -9, -8, -7, -6, -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. The fusion protein / peptide or target protein / peptide generated by the expression cassette as described herein may have a net charge of at least about -20, -19, -18, -17, -16, -15, -14, -13, -12, -11, -10, -9, -8, -7, -6, -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. The fusion protein / peptide or target protein / peptide generated by the expression cassette as described herein may have a net charge a net charge of at most about -20, -19, -18, -17, -16, -15, -14, -13, -12, -11, -10, -9, -8, -7, -6, -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.
[0046] The additional nucleotide may not encode a negatively charged amino acid. The negatively charged amino acid may comprise aspartic acid or glutamic acid. In some cases, the additional nucleotide may encode at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 negatively charged amino acid. In some cases, the additional nucleotide may encode at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid comprising any one or combination of aspartic acid or glutamic acid. The additional nucleotide may encode a negatively charged amino acid.
[0047] In some cases, the additional nucleotide may not encode a hydrophobic amino acid as described herein. Hydrophilic amino acid may comprise alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, or a tyrosine. In some cases, the additional nucleotide may not encode a glycine, proline, or cysteine. In some cases, the additional nucleotide may encode at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 hydrophobic amino acid or any one or combination of alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, glycine, proline, cysteine, or a tyrosine. In some cases, the additional nucleotide may not encode any of the hydrophobic amino acid or any one or combination of alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, glycine, proline, cysteine, or a tyrosine.
[0048] In some cases, the additional nucleotide may encode a hydrophobic amino acid. In some cases, the additional nucleotide may encode comprise alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, or a tyrosine. In some cases, the additional nucleotide may encode a glycine, proline, or cysteine. The additional nucleotide may not encode an amino acid.
[0049] In some cases, the target protein / peptide expressed using the expression cassette described herein may have a molecular weight that is within about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 11 %, 12 %, 13 %, 14 %, 15 %, 16 %, 17 %, 18 %, 19 %, 20 %, 21 %, 22 %, 23 %, 24 %, or 25 %of that of the endogenous or natural counterparts. In some cases, the target protein / peptide expressed using the expression cassette described herein may have a molecular weight that is the same as that of the endogenous or natural counterparts. The endogenous or natural counterpart of a protein may comprise a wildtype protein. For example, the endogenous or natural counterpart of a protein may be the protein that is extracted from a natural organism. In some cases, the endogenous or natural counterpart of a protein may be a protein that its sequence is documented in a reference database. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have a molecular weight of at least about 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more kilodaltons (kDa) . In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have a molecular weight of at most about 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 kDa.
[0050] In some cases, the target protein / peptide expressed using the expression cassette described herein may have an isoelectric point that is within about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 11 %, 12 %, 13 %, 14 %, 15 %, 16 %, 17 %, 18 %, 19 %, 20 %, 21 %, 22 %, 23 %, 24 %, or 25 %of that of the endogenous or natural counterparts. In some cases, the target protein / peptide expressed using the expression cassette described herein may have an isoelectric point that is the same as that of the endogenous or natural counterparts. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have an isoelectric point that is at least about 1, 2, 3, 4, 5, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9, 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, 10, 10.1, 10.2, 10.3, 10.4, 10.5, 10.6, 10.7, 10.8, 10.9, 11, 11.1, 11.2, 11.3, 11.4, 11.5, 11.6, 11.7, 11.8, 11.9, 12, 12.1, 12.2, 12.3, 12.4, 12.5, 12.6, 12.7, 12.8, 12.9, 13, 14 or more. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have an isoelectric point that is at most about 1, 2, 3, 4, 5, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9, 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, 10, 10.1, 10.2, 10.3, 10.4, 10.5, 10.6, 10.7, 10.8, 10.9, 11, 11.1, 11.2, 11.3, 11.4, 11.5, 11.6, 11.7, 11.8, 11.9, 12, 12.1, 12.2, 12.3, 12.4, 12.5, 12.6, 12.7, 12.8, 12.9, 13, or 14.
[0051] In some cases, the target protein / peptide expressed using the expression cassette described herein may have an extinction coefficient that is within about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 11 %, 12 %, 13 %, 14 %, 15 %, 16 %, 17 %, 18 %, 19 %, 20 %, 21 %, 22 %, 23 %, 24 %, or 25 %of that of the endogenous or natural counterparts. In some cases, the target protein / peptide expressed using the expression cassette described herein may have an extinction coefficient that is the same as that of the endogenous or natural counterparts. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have an extinction coefficient that is at least about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, 5, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, or 7. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have an extinction coefficient that is at most about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, 5, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, or 7.
[0052] In some cases, the target protein / peptide expressed using the expression cassette described herein may have a pH stability that is within about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 11 %, 12 %, 13 %, 14 %, 15 %, 16 %, 17 %, 18 %, 19 %, 20 %, 21 %, 22 %, 23 %, 24 %, or 25 %of that of the endogenous or natural counterparts. In some cases, the target protein / peptide expressed using the expression cassette described herein may have a pH stability that is the same as that of the endogenous or natural counterparts. The pH stability of a protein or peptide may comprise the pH at which the protein or peptide precipitates. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have a pH stability that is at least about 1, 2, 3, 4, 5, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9, 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, 10, 10.1, 10.2, 10.3, 10.4, 10.5, 10.6, 10.7, 10.8, 10.9, 11, 11.1, 11.2, 11.3, 11.4, 11.5, 11.6, 11.7, 11.8, 11.9, 12, 13, or 14. In some cases, the target protein / peptide or fusion protein / peptide expressed using the expression cassette described herein may have a pH stability that is at most about 1, 2, 3, 4, 5, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9, 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, 10, 10.1, 10.2, 10.3, 10.4, 10.5, 10.6, 10.7, 10.8, 10.9, 11, 11.1, 11.2, 11.3, 11.4, 11.5, 11.6, 11.7, 11.8, 11.9, 12, 13, or 14.
[0053] In some cases, using the expression cassette described herein, a target protein / peptide having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%, or 100 %sequence identity to the sequence of the endogenous or natural counterpart may be generated. In some cases, using the expression cassette described herein, a target protein / peptide having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%, or 100 %sequence identity to the sequence of the endogenous or natural counterpart may be generated. In some cases, using the expression cassette described herein, a target protein / peptide having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, or at most about 10 amino acid difference relative to the sequence of the endogenous or natural counterpart may be generated. In some cases, using the expression cassette described herein, a target protein / peptide having no amino acid difference as the sequence of the endogenous or natural counterpart may be generated.
[0054] The term “protein, ” “peptide, ” and “polypeptide” are used interchangeably and in their broadest sense to refer to a compound of two or more subunit amino acids, amino acid analogs or peptidomimetics. The terms also encompass an amino acid polymer that has been modified; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein the term “amino acid” refers to either natural and / or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, and amino acid analogs and peptidomimetics. The subunits may be linked by peptide bonds. In another embodiment, the subunit may be linked by other bonds, e.g., ester, ether, etc. A protein or peptide must contain at least two amino acids and no limitation is placed on the maximum number of amino acids which may comprise a protein’s or peptide's sequence. As used herein the term “amino acid” refers to either natural and / or unnatural or synthetic amino acids, including glycine and both the D and L optical isomers, amino acid analogs and peptidomimetics. As used herein, the term “fusion protein” refers to a protein comprised of domains from more than one naturally occurring or recombinantly produced protein, where generally each domain serves a different function. In this regard, the term “linker” refers to a protein fragment that is used to link these domains together –optionally to preserve the conformation of the fused protein domains and / or prevent unfavorable interactions between the fused protein domains which may compromise their respective functions.
[0055] “Homology” or “identity” or “similarity” can refer to sequence similarity between two peptides or between two nucleic acid molecules. Homology can be determined by comparing a position in each sequence which can be aligned for purposes of comparison. When a position in the compared sequence can be occupied by the same base or amino acid, then the molecules can be homologous at that position. A degree of homology between sequences can be a function of the number of matching or homologous positions shared by the sequences. An “unrelated” or “non-homologous” sequence shares less than 40%identity, or alternatively less than 25%identity, with one of the sequences of the disclosure. Sequence homology can refer to a %identity of a sequence to a reference sequence. As a practical matter, whether any particular sequence can be at least 50%, 60%, 70%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98%or 99%identical to any sequence described herein (which can correspond with a particular nucleic acid sequence described herein) , such particular polypeptide sequence can be determined conventionally using known computer programs such the Bestfit program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, 575 Science Drive, Madison, Wis. 53711) . When using Bestfit or any other sequence alignment program to determine whether a particular sequence is, for instance, 95%identical to a reference sequence, the parameters can be set such that the percentage of identity can be calculated over the full length of the reference sequence and that gaps in sequence homology of up to 5%of the total reference sequence can be facilitate.
[0056] In some cases, the identity between a reference sequence (query sequence, i.e., a sequence of the disclosure) and a subject sequence, also referred to as a global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6: 237-245 (1990) ) . In some embodiments, parameters for a particular embodiment in which identity can be narrowly construed, used in a FASTDB amino acid alignment, can include: Scoring Scheme=PAM (Percent Accepted Mutations) 0, k-tuple=2, Mismatch Penalty=1, Joining Penalty=20, Randomization Group Length=0, Cutoff Score=1, Window Size=sequence length, Gap Penalty=5, Gap Size Penalty=0.05, Window Size=500 or the length of the subject sequence, whichever can be shorter. According to this embodiment, if the subject sequence can be shorter than the query sequence due to N-or C-terminal deletions, not because of internal deletions, a manual correction can be made to the results to take into consideration the fact that the FASTDB program does not account for N-and C-terminal truncations of the subject sequence when calculating global percent identity. For subject sequences truncated at the N-and C-termini, relative to the query sequence, the percent identity can be corrected by calculating the number of residues of the query sequence that can be lateral to the N-and C-terminal of the subject sequence, which can be not matched / aligned with a corresponding subject residue, as a percent of the total bases of the query sequence. A determination of whether a residue can be matched / aligned can be determined by results of the FASTDB sequence alignment. This percentage can be then subtracted from the percent identity, calculated by the FASTDB program using the specified parameters, to arrive at a final percent identity score. This final percent identity score can be used for the purposes of this embodiment. In some cases, only residues to the N-and C-termini of the subject sequence, which can be not matched / aligned with the query sequence, can be considered for the purposes of manually adjusting the percent identity score. That is, only query residue positions outside the farthest N-and C-terminal residues of the subject sequence can be considered for this manual correction. For example, a 90-residue subject sequence can be aligned with a 100-residue query sequence to determine percent identity. The deletion occurs at the N-terminus of the subject sequence, and therefore, the FASTDB alignment does not show a matching / alignment of the first 10 residues at the N-terminus. The 10 unpaired residues represent 10%of the sequence (number of residues at the N-and C-termini not matched / total number of residues in the query sequence) so 10%can be subtracted from the percent identity score calculated by the FASTDB program. If the remaining 90 residues were perfectly matched, the final percent identity can be 90%. In another example, a 90-residue subject sequence can be compared with a 100-residue query sequence. This time the deletions can be internal deletions, so there can be no residues at the N-or C-termini of the subject sequence which can be not matched / aligned with the query. In this case, the percent identity calculated by FASTDB can be not manually corrected. Once again, only residue positions outside the N-and C-terminal ends of the subject sequence, as displayed in the FASTDB alignment, which can be not matched / aligned with the query sequence can be manually corrected for. Target protein / peptide
[0057] A target protein can be expressed using the expression cassettes described herein (e.g., comprising at least one additional nucleotide, a cleavage site, and optionally a purification tag) . The additional nucleotide and the target sequence encoding the target protein / peptide may not be the same or overlap. The additional nucleotide and the target sequence encoding the target protein / peptide may be the same or overlap.
[0058] A target sequence encoding a target protein or peptide can comprise a MAP. A target protein or peptide can comprise an antimicrobial peptide (AMP) . A target protein or peptide can comprise a MAP or an AMP. The AMP may comprise a mussel-derived AMP. A MAP can comprise Mfp-1 (Mytilus edulis (Blue mussel) ) , Mfp-1 (Mytilus galloprovincialis) , Mfp-1 (Mytlilus coruscus) , Mfp-1 (Perna viridis) , Mfp-1 (Mytlilus californianus) , Mfp-2 (Mytilus edulis (Blue mussel) ) , Mfp-2 (Mytilus galloprovincialis) , Mfp-2 (Limnoperna fortune) , Mfp-3 (Mytilus edulis (Blue mussel) ) , Mfp-3 (Mytilus galloprovincialis) , Mfp-3 (Mytlilus coruscus) , Mfp-4 (Mytlilus californianus) , Mfp-5 (Mytilus edulis (Blue mussel) ) , Mfp-5 (Mytilus galloprovincialis) , Mfp-5 (Mytlilus coruscus) , Mfp-5 (Mytlilus californianus) , Mfp-6 (Mytilus galloprovincialis) , Mfp-6 (Mytlilus coruscus) , or a combination thereof. A MAP can comprise Mfp-1 (Mytilus edulis (Blue mussel) ) . A MAP can comprise Mfp-1 (Mytilus galloprovincialis) . A MAP can comprise Mfp-1 (Mytlilus coruscus) . A MAP can comprise Mfp-1 (Perna viridis) . A MAP can comprise Mfp-1 (Mytlilus californianus) . A MAP can comprise Mfp-2 (Mytilus edulis (Blue mussel) ) . A MAP can comprise Mfp-2 (Mytilus galloprovincialis) . A MAP can comprise Mfp-2 (Limnoperna fortune) . A MAP can comprise Mfp-3 (Mytilus edulis (Blue mussel) ) . A MAP can comprise Mfp-3 (Mytilus galloprovincialis) . A MAP can comprise Mfp-3 (Mytlilus coruscus) . A MAP can comprise Mfp-4 (Mytlilus californianus) . A MAP can comprise Mfp-5 (Mytilus edulis (Blue mussel) ) . A MAP can comprise Mfp-5 (Mytilus galloprovincialis) . A MAP can comprise Mfp-5 (Mytlilus coruscus) . A MAP can comprise Mfp-5 (Mytlilus californianus) . A MAP can comprise Mfp-6 (Mytilus galloprovincialis) . A MAP can comprise Mfp-6 (Mytlilus coruscus) . An AMP can comprise AMP-1, AMP-2, AMP-3, AMP-4, AMP-5, AMP-6, AMP-7, or a combination thereof. An AMP can comprise AMP-1. An AMP can comprise AMP-2. An AMP can comprise AMP-3. An AMP can comprise AMP-4. An AMP can comprise AMP-5. An AMP can comprise AMP-6. An AMP can comprise AMP-7. The sequence encoding the target protein / peptide may have at least about 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 60, 90, 120, 150, 300, 600, 900, 1200, 1500, 3000, 6000, 9000 or more nt. The sequence encoding the target protein / peptide may have at most about 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 60, 90, 120, 150, 300, 600, 900, 1200, 1500, 3000, 6000, or 9000 nt.
[0059] A nucleic acid sequence encoding the target protein may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%sequence identity to any one of SEQ ID NOs: 42-44 and 70-94. A nucleic acid sequence encoding the target protein may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%sequence identity to any one of SEQ ID NOs: 42-44 and 70-94. A nucleic acid sequence encoding the target protein may comprise a sequence having 100 %to any one of SEQ ID NOs: 42-44 and 70-94. A nucleic acid sequence encoding the target protein may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 42-44 and 70-94. A nucleic acid sequence encoding the target protein may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 42-44 and 70-94. The target protein / peptide may thus comprise the polypeptide sequence encoded by the nucleic acid sequence as described herein. For example, the target protein may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%sequence identity to any one of SEQ ID NOs: 95-119. The target protein may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%sequence identity to any one of SEQ ID NOs: 95-119. The target protein may comprise a sequence having 100 %to any one of SEQ ID NOs: 95-119. The target protein may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 95-119. The target protein may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 95-119.
[0060] In some cases, the target protein / peptide may not comprise a starting methionine (such as those described herein) . For example, the endogenous counterpart of the target protein / peptide may comprise a signal peptide that is cleaved subsequent to protein translation. Thus, the target protein / peptide may lack a staring methionine. The N-terminal amino acid of the target protein / peptide may be matched with the cleavage site (see elsewhere in this disclosure) . Start Codon
[0061] In some cases, the expression cassette can comprise a start codon, the sequence of the target protein / peptide, and the additional nucleotide. The additional nucleotide and the starting codon may not be the same or overlap. The additional nucleotide and the starting codon may be the same or overlap. The start codon may encode a methionine or a modified methionine. Affinity tag protein / peptide
[0062] In some cases, the expression cassette can comprise a sequence encoding an affinity tag protein / peptide, the target sequence encoding the target protein / peptide, and the additional nucleotide. The additional nucleotide and the sequence of the affinity tag protein / peptide may not be the same or overlap. The additional nucleotide and the sequence of the affinity tag protein / peptide may be the same or overlap. In some cases, the expression cassette may not comprise the sequence of the affinity tag protein / peptide.
[0063] The affinity tag protein / peptide may comprise a protein or peptide that can facilitate the protein purification of the fusion protein / peptide, enhance the expression level of the fusion protein / peptide, enhance the solubility of the fusion protein / peptide, enhance the stability of the fusion protein / peptide, or a combination thereof, relative to a comparable protein / peptide without the affinity tag protein / peptide. The affinity tag protein / peptide may comprise Albumin-binding protein (ABP) , Alkaline Phosphatase (AP) , AU1 epitope, AU5 epitope, Bacteriophage T7 epitope (T7-tag) , Bacteriophage V5 epitope (V5) , Biotin-carboxy carrier protein (BCCP) , Bluetongue virus tag (B-tag) , Calmodulin binding peptide (CBP) , Cellulose binding domain (CBP) , Chitin binding domain (CBD) , Chloramphenicol Acetyl Transferase (CAT) , Choline-binding domain (CBD) , Dihydrofolate reductase (DHFR) , E2 epitope, FLAG epitope, Galactose-binding protein (GBP) , Glu-Glu (EE-tag) , Glutathione S-transferase (GST) , Green fluorescent protein (GFP) , Histidine affinity tag (HAT) , Horseradish Peroxidase (HRP) , HSV epitope, Human influenza hemagglutinin (HA) , Ketosteroid isomerase (KSI) , KT3 epitope, LacZ, Luciferase, Maltose-binding protein (MBP) , Myc epitope, NusA, PDZ domain, PDZ ligand, Polyarginine (Arg-tag) , Polyaspartate (Asp-tag) , Polycysteine (Cys-tag) , Polyhistidine (His-tag) , Polyphenylalanine (Phe-tag) , Profinity eXact, Protein C, S1-tag, Small Ubiquitin-like Modifier (SUMO) , S-tag, Staphylococcal protein A (Protein A) , Staphylococcal protein G (Protein G) , Strep-tag, Streptavadin, Streptavadin-binding peptide (SBP) , T7 epitope, Tandem Affinity Purification (TAP) , Thioredoxin (Trx) , TrpE, Ubiquitin, Universal, VSV-G, or a combination thereof. The affinity tag protein / peptide may comprise E2 epitope, Myc epitope, Polyphenylalanine (Phe-tag) , KT3 epitope, Bacteriophage T7 epitope (T7-tag) , HSV epitope, VSV-G, Bacteriophage V5 epitope (V5) , S-tag, Histidine affinity tag (HAT) , Polyhistidine (His-tag) , Polycysteine (Cys-tag) , Polyaspartate (Asp-tag) , Polyarginine (Arg-tag) , PDZ ligand, AU1 epitope, Glu-Glu (EE-tag) , Universal, Bluetongue virus tag (B-tag) , AU5 epitope, FLAG epitope, Strep-tag, S1-tag, Avi Tag, SNAP-Tag, mCherry, or a combination thereof. The affinity tag protein / peptide may comprise a Polyhistidine, a Myc tag, a Strep tag, a V5 tag, an HA tag, a Flag, tag or a combination thereof. The affinity tag protein / peptide may comprise E2 epitope. The affinity tag protein / peptide may comprise Myc epitope. The affinity tag protein / peptide may comprise Polyphenylalanine (Phe-tag) . The affinity tag protein / peptide may comprise KT3 epitope. The affinity tag protein / peptide may comprise Bacteriophage T7 epitope (T7-tag) . The affinity tag protein / peptide may comprise HSV epitope. The affinity tag protein / peptide may comprise VSV-G. The affinity tag protein / peptide may comprise Bacteriophage V5 epitope (V5-tag) . The affinity tag protein / peptide may comprise S-tag. The affinity tag protein / peptide may comprise Histidine affinity tag (HAT) . The affinity tag protein / peptide may comprise Polyhistidine (His-tag) . The affinity tag protein / peptide may comprise Polycysteine (Cys-tag) . The affinity tag protein / peptide may comprise Polyaspartate (Asp-tag) . The affinity tag protein / peptide may comprise Polyarginine (Arg-tag) . The affinity tag protein / peptide may comprise PDZ ligand. The affinity tag protein / peptide may comprise AU1 epitope. The affinity tag protein / peptide may comprise Glu-Glu (EE-tag) . The affinity tag protein / peptide may comprise Universal. The affinity tag protein / peptide may comprise Bluetongue virus tag (B-tag) . The affinity tag protein / peptide may comprise AU5 epitope. The affinity tag protein / peptide may comprise FLAG epitope. The affinity tag protein / peptide may comprise Strep-tag. The affinity tag protein / peptide may comprise S1-tag. The affinity tag protein / peptide may comprise Avi Tag. The affinity tag protein / peptide may comprise SNAP-Tag. The affinity tag protein / peptide may comprise mCherry.
[0064] The affinity tag protein / peptide may have a molecular weight of at least about 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5 kilodalton (kDA) or more. The affinity tag protein / peptide may have a molecular weight of at most about 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, or 5 kilodalton (kDA) . The affinity tag protein / peptide may have at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids. The affinity tag protein / peptide may have at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids.
[0065] An affinity tag protein / peptide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%sequence identity to any one of SEQ ID NOs: 64-69. An affinity tag protein / peptide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%sequence identity to any one of SEQ ID NOs: 64-69. An affinity tag protein / peptide may comprise a sequence having 100 %to any one of SEQ ID NOs: 64-69. An affinity tag protein / peptide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 64-69. An affinity tag protein / peptide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 64-69. The affinity tag protein / peptide may thus comprise the polypeptide sequence encoded by the nucleic acid sequence as described herein. In some cases, the affinity tag is a his-tag.
[0066] The selection of the affinity tag protein / peptide may be based on the size, physical and chemical properties and downstream applications of the protein. Using affinity tag proteins / peptides with smaller molecular weights can have minimal impact on the physical and chemical properties of the expressed protein / peptide. Additionally, once cleaved, a larger proportion of the target protein / peptide may be produced, relative to that when using an affinity tag protein / peptide with a larger molecular weight.
[0067] Wherein when the additional nucleotide is located within or overlaps with the affinity tag protein / peptide or target protein / peptide, the expression cassette comprising the additional nucleotide may facilitate expression of the fusion or target protein / peptide at a level that is at least about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 100 %, 150 %, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, or 1000-fold higher than that of a comparable expression cassette having the sequence of the fusion or target protein / peptide, respectively, but not the additional nucleotide. The expression cassette comprising the additional nucleotide may facilitate expression of the fusion or target protein / peptide at a level that is at most about 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 100 %, 150 %, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, or 1000-fold higher than that of a comparable expression cassette having the sequence of the fusion or target protein / peptide, respectively, but not the additional nucleotide. Cleavage site
[0068] In some cases, the expression cassette can comprise a sequence encoding a cleavage site. The cleavage site may be a proteolytic cleavage site. The proteolytic cleavage site may be a cleavage site of a protease as described herein. The cleavage site may be cleaved by a chemical or intein-based self-cleavage. The cleavage site may facilitate the generation of the target protein / peptide without the affinity tag or the additional nucleotide. Thus, using the expression cassette described herein, a target protein / peptide having the same or substantially the same amino acid sequence as the endogenous or natural counterpart may be generated. The cleavage site may also allow removal of any sequence (s) or amino acid reside (s) that may have a negative impact on the expression level of the target protein or peptide.
[0069] In some cases, using the expression cassette described herein, a target protein / peptide having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%, or 100 %sequence identity to the sequence of the endogenous or natural counterpart may be generated. In some cases, using the expression cassette described herein, a target protein / peptide having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%, or 100 %sequence identity to the sequence of the endogenous or natural counterpart may be generated. In some cases, using the expression cassette described herein, a target protein / peptide having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, or at most about 10 amino acid difference relative to the sequence of the endogenous or natural counterpart may be generated. In some cases, using the expression cassette described herein, a target protein / peptide having no amino acid difference as the sequence of the endogenous or natural counterpart may be generated.
[0070] The additional nucleotide and the sequence of the cleavage site may be the same or overlap. In some cases, the expression cassette may not comprise the sequence of the cleavage site.
[0071] The protease may comprise aspartic protease, glutamic protease, metalloproteases, cysteine protease, serine protease, or threonine protease. The protease may comprise aspartic protease. The protease may comprise glutamic protease. The protease may comprise metalloproteases. The protease may comprise cysteine protease. The protease may comprise serine protease. The protease may comprise threonine protease. The protease may comprise aminopeptidase, Arg-C proteinase, Asp-N endopeptidase, BNPS-Skatole, carboxypeptidase A, carboxypeptidase B, carboxypeptidase C, cathepsin C, Caspase 1, Caspase 10, Caspase 2, Caspase 3, Caspase 4, Caspase 5, Caspase 6, Caspase 7, Caspase 8, Caspase 9, Chymotrypsin, Clostripain (Clostridiopeptidase B) , CNBr, collagenase, dispase, endopeptidase Lys-C (LysC) , Enterokinase, Factor Xa, Formic acid, Glutamyl endopeptidase (Glu-C) , GranzymeB, Human rhinovirus 3C protease (HRV3C) , Hydroxylamine, Iodosobenzoic acid, Kallikrein, Neutrophil elastase, NTCB (2-nitro-5-thiocyanobenzoic acid) , Papaya Protease, Pepsin (pH>2) , Pepsin (pH1.3) , plasmin, Proline-endopeptidase, pronase, Proteinase K, Staphylococcal peptidase I, Subtilisin, SUMO Protease, TAGZyme, Thermolysin, Thrombin, tissue Protease C, Tobacco Etch Virus (TEV) protease, Trypsin, or a combination thereof. A protease may comprise Arg-C proteinase. A protease may comprise Asp-N endopeptidase. A protease may comprise BNPS-Skatole. A protease may comprise Caspase 1. A protease may comprise Caspase 10. A protease may comprise Caspase 2. A protease may comprise Caspase 3. A protease may comprise Caspase 4. A protease may comprise Caspase 5. A protease may comprise Caspase 6. A protease may comprise Caspase 7. A protease may comprise Caspase 8. A protease may comprise Caspase 9. A protease may comprise Chymotrypsin. A protease may comprise Clostripain (Clostridiopeptidase B) . A protease may comprise CNBr. A protease may comprise Enterokinase. A protease may comprise Factor Xa. A protease may comprise Formic acid. A protease may comprise Glutamyl endopeptidase. A protease may comprise GranzymeB. A protease may comprise Human rhinovirus 3C protease (HRV 3C) . A protease may comprise Hydroxylamine. A protease may comprise Iodosobenzoic acid. A protease may comprise LysC. A protease may comprise Neutrophil elastase. A protease may comprise NTCB (2-nitro-5-thiocyanobenzoic acid) . A protease may comprise Pepsin (pH>2) . A protease may comprise Pepsin (pH1.3) . A protease may comprise Proline-endopeptidase. A protease may comprise Proteinase K. A protease may comprise Staphylococcal peptidase I. A protease may comprise Thermolysin. A protease may comprise Thrombin. A protease may comprise Tobacco Etch Virus (TEV) protease. A protease may comprise Trypsin. A protease may comprise aminopeptidase. A protease may comprise carboxypeptidase A. A protease may comprise carboxypeptidase B. A protease may comprise carboxypeptidase C. A protease may comprise cathepsin C. A protease may comprise collagenase. A protease may comprise dispase. A protease may comprise Kallikrein. A protease may comprise Papaya Protease. A protease may comprise plasmin. A protease may comprise pronase. A protease may comprise subtilisin. A protease may comprise SUMO Protease. A protease may comprise TAGZYme. A protease may comprise tissue Protease C.
[0072] The sequence encoding the cleavage site may have at least about 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, 63, 66, 69, 72, 75, 78, 81, 84, 87, 90 or more nt. The sequence encoding the cleavage site may have at least about 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, 63, 66, 69, 72, 75, 78, 81, 84, 87, or 90 nt.
[0073] A cleavage site may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%sequence identity to any one of SEQ ID NOs: 59-63 and 172. A cleavage site may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%sequence identity to any one of SEQ ID NOs: 59-63 and 172. A cleavage site may comprise a sequence having 100 %to any one of SEQ ID NOs: 59-63 and 172. A cleavage site may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 59-63 and 172. A cleavage site may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 59-63 and 172. The cleavage site may thus comprise the polypeptide sequence encoded by the nucleic acid sequence as described herein. In some cases, the cleavage site is a TEV cleavage site.
[0074] In some cases, the target protein / peptide may be cleaved from other peptide sequence (s) of the expression cassette (such as the affinity protein / peptide or any additional nucleotide, as described herein) using chemical or intein-based self-cleavage methods. For example, chemical cleavage may comprise the use of cyanogen bromide (CNBr) , hydroxylamine (hydroxylamine) , formic acid, or a combination thereof. The chemical cleavage may comprise the use of CNBr. The chemical cleavage may comprise the use of hydroxylamine. The chemical cleavage may comprise the use of formic acid. The intein-based self-cleavage may comprise an insertion of intein sequence between the target protein / peptide and other peptide sequence (s) of the expression cassette. Exemplary configurations
[0075] The additional nucleotide (s) may be located 5’ and 3’ to the sequence encoding the target protein / peptide. The additional nucleotide may be located 5’ to the sequence encoding the target protein / peptide. The additional nucleotide may be located within the sequence encoding the target protein / peptide. The additional nucleotide may be located 3’ to the sequence encoding the target protein / peptide. The additional nucleotide (s) may be located 5’ and 3’ to the sequence encoding the cleavage site. The additional nucleotide may be located 5’ to the sequence encoding the cleavage site. The additional nucleotide may be located 3’ to the sequence encoding the cleavage site. The additional nucleotide (s) may be located 5’ and 3’ to the sequence encoding the affinity tag protein / peptide. The additional nucleotide may be located 5’ to the sequence encoding the affinity tag protein / peptide. The additional nucleotide may be located 3’ to the sequence encoding the affinity tag protein / peptide. The additional nucleotide may be located 5’ to the start codon. The additional nucleotide (s) may be located 5’ and 3’ to the start codon. The additional nucleotide may be located 3’ to the sequence encoding the start codon. The sequence encoding the cleavage site may be located 5’ to the sequence encoding the target protein / peptide. The sequence encoding the cleavage site may be located 3’ to the sequence encoding the target protein / peptide. The sequence encoding the cleavage site may be located 5’ to the sequence encoding the affinity tag protein / peptide. The sequence encoding the cleavage site may be located 3’ to the sequence encoding the affinity tag protein / peptide. The sequence encoding the affinity tag protein or peptide may be located 5’ to the sequence encoding the target protein / peptide. The sequence encoding the affinity tag protein or peptide may be located 3’ to the sequence encoding the target protein / peptide.
[0076] In some case, the additional nucleotide may be located within the coding sequence of the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or any combination thereof. In some cases, when located within the coding sequence of the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or any combination thereof, the additional nucleotide may not alter the amino acid sequence of the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or any combination thereof (via degenerate codon) . In some cases, when located within the coding sequence of the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or any combination thereof, the additional nucleotide may alter the amino acid sequence of the target protein / peptide, the affinity tag protein / peptide, the cleavage site, or any combination thereof
[0077] Thus, the expression cassette may have the following configurations (from 5’ to 3’ / 5’-3’) , wherein the sequence encoding the target protein / peptide is referred to as T, the sequence encoding the affinity tag protein / peptide is referred to as AP, the sequence encoding the cleavage site is referred to as CS, and the start codon is referred to as ATG: 5’-T-3’; 5’-CS-T-3’; 5’-AP-T-3’; 5’-AP-CS-T-3’; 5’-ATG-T-3’; 5’-ATG-CS-T-3’; 5’-ATG-AP-T-3’; 5’-ATG-AP-CS-T-3’; 3’-T-5’; 3’-CS-T-5’; 3’-AP-T-5’; 3’-AP-CS-T-5’; 3’-T-ATG-5’; 3’-CS-T-ATG-5’; 3’-AP-T-ATG-5’; or 3’-AP-CS-T-ATG-5’, wherein the additional nucleotide can be located at 5’, 3’, 5’ and 3’ (when the additional nucleotide contains at least two nucleotides) , or within CS, T, AP, or a combination thereof. In some cases, the additional nucleotide is located at 5’ to the CS.
[0078] The selection of the cleavage site may be based on the sequence identities of the cleavage site and the peptide placed immediately downstream (the cleavage site is 5’ or N-terminal to the peptide sequence, when referring to nucleotide and amino acid sequences, respectively) . In an illustrative example, the peptide sequence of cleavage site may have a sequence of NNNXY, and the peptide sequence of the peptide may have a sequence of YNNN, wherein X, Y, N may be any amino acids, and wherein the cleavage site is cleaved between X and Y. Subsequent to cleavage of the cleavage site, any peptides placed downstream of the cleavage site can retain amino acid Y. Thus, by matching the amino acid identities of the C-terminus of the cleavage site and the N-terminus of the peptide, the methods described herein can facilitate the expression of the peptide having the same N-terminus of the endogenous counterpart of the peptide sequence. Accordingly, in the sequences disclosed herein, wherein when the amino acid (or nucleotides or codons) identifies of the C-terminus of the cleavage site and the N-terminus of the peptide placed downstream (such as the target protein / peptide) are the same, only one of those amino acids (or nucleotides or codons) may be depicted in one of these sequences. In some cases, the cleavage site is a TEV protease cleavage site encoding serine (or glycine) in the C-terminus, and the target sequence encodes a target protein (such as mefp5) , and subsequent to the cleavage of the TEV protease cleavage site, the serine (or glycine) from C-terminus of the TEV protease cleavage site is retained at the N-terminus of the target protein. Similar arrangement can apply to applicable protease cleavage site or target protein. Exemplary sequences
[0079] In some cases, the present disclosure provides a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to any one of SEQ ID NOs: 1-187; a fragment thereof, a nucleic acid encoding such sequence when such sequence is an amino acid sequence, or an amino acid sequence encoded by such sequence when such sequence is a nucleic acid sequence. The present disclosure provides a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9%sequence identity to any one of SEQ ID NOs: 1-187; a fragment thereof, a nucleic acid encoding such sequence when such sequence is an amino acid sequence, or an amino acid sequence encoded by such sequence when such sequence is a nucleic acid sequence. The present disclosure provides a sequence having 100 %to any one of SEQ ID NOs: 1-187; a fragment thereof, a nucleic acid encoding such sequence when such sequence is an amino acid sequence, or an amino acid sequence encoded by such sequence when such sequence is a nucleic acid sequence. The present disclosure provides a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 1-187; a fragment thereof, a nucleic acid encoding such sequence when such sequence is an amino acid sequence, or an amino acid sequence encoded by such sequence when such sequence is a nucleic acid sequence. The present disclosure provides a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 1-187; a fragment thereof, a nucleic acid encoding such sequence when such sequence is an amino acid sequence, or an amino acid sequence encoded by such sequence when such sequence is a nucleic acid sequence. In the present disclosures, any nucleic acid sequence disclosed can also comprise a counterpart that comprises at least a corresponding degenerate codon (s) .
[0080] An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to any one of SEQ ID NOs: 1-41. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to any one of SEQ ID NOs: 1-41. An expression cassette may comprise a sequence having 100 %sequence identity to any one of SEQ ID NOs: 1-41. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to any one of SEQ ID NOs: 1-41. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30, or at most about 60 mutations, relative to any one of SEQ ID NOs: 1-41.
[0081] An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 1. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 1. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 1. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 1. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 1. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 2. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 2. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 2. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 2. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 2. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 3. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 3. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 3. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 3. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 3. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 4. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 4. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 4. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 4. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 4. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 5. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 5. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 5. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 5. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 5. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 6. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 6. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 6. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 6. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 6. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 7. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 7. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 7. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 7. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 7. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 8. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 8. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 8. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 8. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 8. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 9. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 9. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 9. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 9. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 9. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 10. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 10. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 10. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 10. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 10. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 11. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 11. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 11. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 11. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 11. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 12. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 12. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 12. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 12. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 12. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 13. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 13. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 13. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 13. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 13. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 14. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 14. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 14. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 14. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 14. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 15. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 15. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 15. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 15. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 15. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 16. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 16. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 16. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 16. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 16. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 17. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 17. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 17. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 17. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 17. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 18. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 18. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 18. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 18. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 18. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 19. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 19. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 19. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 19. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 19. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 20. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 20. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 20. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 20. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 20. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 21. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 21. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 21. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 21. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 21. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 22. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 22. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 22. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 22. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 22. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 23. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 23. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 23. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 23. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 23. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 24. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 24. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 24. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 24. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 24. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 25. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 25. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 25. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 25. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 25. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 26. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 26. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 26. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 26. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 26. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 27. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 27. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 27. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 27. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 27. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 28. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 28. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 28. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 28. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 28. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 29. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 29. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 29. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 29. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 29. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 30. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 30. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 30. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 30. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 30. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 31. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 31. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 31. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 31. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 31. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 32. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 32. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 32. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 32. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 32. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 33. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 33. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 33. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 33. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 33. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 34. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 34. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 34. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 34. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 34. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 35. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 35. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 35. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 35. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 35. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 36. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 36. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 36. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 36. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 36. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 37. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 37. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 37. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 37. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 37. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 38. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 38. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 38. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 38. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 38. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 39. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 39. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 39. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 39. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 39. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 40. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 40. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 40. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 40. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 40. An expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 41. An expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 41. An expression cassette may comprise a sequence having 100 %sequence identity to SEQ ID NO: 41. An expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to SEQ ID NO: 41. An expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30 mutations, or at most about 60 mutations, relative to SEQ ID NO: 41.
[0082] A polypeptide sequence may be encoded by the expression cassette. For example, the exemplary polypeptide sequence encoded by the exemplary expression cassette may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to any one of SEQ ID NOs: 48-58. The exemplary polypeptide sequence encoded by the exemplary expression cassette may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to any one of SEQ ID NOs: 48-58. The exemplary polypeptide sequence encoded by the exemplary expression cassette may comprise a sequence having 100 %sequence identity to any one of SEQ ID NOs: 48-58. The exemplary polypeptide sequence encoded by the exemplary expression cassette may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 48-58. The exemplary polypeptide sequence encoded by the exemplary expression cassette may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 48-58.
[0083] An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to any one of SEQ ID NOs: 149-169. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to any one of SEQ ID NOs: 149-169. An additional nucleotide may comprise a sequence having 100 %sequence identity to any one of SEQ ID NOs: 149-169. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 149-169. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 149-169.
[0084] An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 149. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 149. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 149. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 149. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 149. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 150. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 150. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 150. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 150. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 150. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 151. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 151. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 151. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 151. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 151. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 152. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 152. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 152. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 152. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 152. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 153. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 153. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 153. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 153. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 153. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 154. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 154. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 154. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 154. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 154. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 155. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 155. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 155. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 155. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 155. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 156. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 156. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 156. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 156. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 156. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 157. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 157. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 157. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 157. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 157. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 158. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 158. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 158. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 158. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 158. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 159. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 159. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 159. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 159. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 159. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 160. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 160. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 160. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 160. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 160. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 161. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 161. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 161. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 161. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 161. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 162. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 162. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 162. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 162. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 162. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 163. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 163. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 163. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 163. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 163. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 164. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 164. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 164. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 164. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 164. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 165. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 165. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 165. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 165. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 165. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 166. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 166. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 166. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 166. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 166. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 167. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 167. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 167. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 167. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 167. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 168. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 168. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 168. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 168. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 168. An additional nucleotide may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 169. An additional nucleotide may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 169. An additional nucleotide may comprise a sequence having 100 %sequence identity to SEQ ID NO: 169. An additional nucleotide may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 169. An additional nucleotide may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 169.
[0085] The expression cassette described herein may comprise a sequence that does not comprise the target protein / peptide or sequence (or the nucleotide sequence encoding the target protein / peptide or sequence) for facilitating the expression of the fusion or target protein / peptide, which is referred herein as the first sequence. The first sequence and target sequence can be expressed as one polypeptide sequence.
[0086] A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to any one of SEQ ID NOs: 120-148. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to any one of SEQ ID NOs: 120-148. A first sequence may comprise a sequence having 100 %sequence identity to any one of SEQ ID NOs: 120-148. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 60 or more mutations, relative to any one of SEQ ID NOs: 120-148. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, at most about 30, or at most about 60 mutations, relative to any one of SEQ ID NOs: 120-148.
[0087] A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 120. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 120. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 120. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 120. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 120. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 121. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 121. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 121. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 121. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 121. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 122. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 122. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 122. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 122. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 122. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 123. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 123. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 123. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 123. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 123. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 124. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 124. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 124. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 124. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 124. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 125. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 125. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 125. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 125. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 125. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 126. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 126. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 126. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 126. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 126. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 127. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 127. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 127. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 127. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 127. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 128. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 128. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 128. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 128. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 128. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 129. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 129. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 129. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 129. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 129. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 130. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 130. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 130. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 130. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 130. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 131. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 131. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 131. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 131. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 131. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 132. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 132. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 132. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 132. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 132. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 133. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 133. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 133. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 133. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 133. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 134. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 134. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 134. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 134. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 134. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 135. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 135. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 135. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 135. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 135. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 136. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 136. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 136. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 136. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 136. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 137. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 137. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 137. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 137. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 137. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 138. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 138. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 138. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 138. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 138. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 139. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 139. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 139. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 139. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 139. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 140. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 140. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 140. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 140. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 140. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 141. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 141. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 141. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 141. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 141. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 142. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 142. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 142. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 142. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 142. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 143. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 143. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 143. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 143. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 143. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 144. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 144. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 144. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 144. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 144. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 145. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 145. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 145. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 145. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 145. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 146. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 146. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 146. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 146. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 146. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 147. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 147. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 147. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 147. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 147. A first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to SEQ ID NO: 148. A first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to SEQ ID NO: 148. A first sequence may comprise a sequence having 100 %sequence identity to SEQ ID NO: 148. A first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 148. A first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 148.
[0088] A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 173. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 173. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 173. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 174. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 174. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 174. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 175. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 175. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 175. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 176. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 176. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 176. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 177. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 177. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 177. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 178. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 178. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 178. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 179. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 179. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 179. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 180. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 180. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 180. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 181. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 181. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 181. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 182. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 182. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 182. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 183. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 183. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 183. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 184. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 184. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 184. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having 100 %sequence identity to SEQ ID NO: 185. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to SEQ ID NO: 185. A first sequence can comprise a nucleic acid sequence encoding an amino acid sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to SEQ ID NO: 185.
[0089] A polypeptide sequence may be encoded by the first sequence. For example, the exemplary polypeptide sequence encoded by the exemplary first sequence may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.9%sequence identity to any one of SEQ ID NOs: 45-47 or 173-187. The exemplary polypeptide sequence encoded by the exemplary first sequence may comprise a sequence having at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, or at most about 99.9%sequence identity to any one of SEQ ID NOs: 45-47 or 173-187. The exemplary polypeptide sequence encoded by the exemplary first sequence may comprise a sequence having 100 %sequence identity to any one of SEQ ID NOs: 45-47 or 173-187. The exemplary polypeptide sequence encoded by the exemplary first sequence may comprise a sequence having at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30 or more mutations, relative to any one of SEQ ID NOs: 45-47 or 173-187. The exemplary polypeptide sequence encoded by the exemplary first sequence may comprise a sequence having at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 20, or at most about 30 mutations, relative to any one of SEQ ID NOs: 45-47 or 173-187.
[0090] Table 5 below describes the non-limiting exemplary sequences of this disclosure Table 5: Exemplary Sequences
[0091] Further expression elements
[0092] The expression cassette described can further comprise further expression elements. The further expression elements may allow for or facilitate the transcription and / or translation of the fusion protein / peptide or target protein / peptide. For example, the further expression elements may comprise a promoter, transcription initiation site, translation initiation site, ribosome binding site, transcription termination site, or a combination thereof. A further expression element of the expression cassette described herein may comprise a promoter. A further expression element of the expression cassette described herein may comprise a transcription initiation site. A further expression element of the expression cassette described herein may comprise a translation initiation site. A further expression element of the expression cassette described herein may comprise a ribosome binding site. A further expression element of the expression cassette described herein may comprise a transcription termination site.
[0093] Additional characteristics of the fusion or target protein / peptide
[0094] In some cases, the fusion or target protein / peptide expressed by the expression cassette as described herein may not be PEGylated. In some cases, the fusion or target protein / peptide expressed by the expression cassette as described herein may be PEGylated. In some cases, the fusion or target protein / peptide expressed by the expression cassette as described herein may not comprise a non-naturally occurring amino acid, a non-proteinogenic amino acid, or a non-canonical amino acid. In some cases, the affinity tag protein / peptide of the expression cassette as described herein may not comprise MBP. In some cases, the fusion or target protein / peptide expressed by the expression cassette as described herein may not comprise a chaperone. In some cases, the fusion or target protein / peptide expressed by the expression cassette as described herein may not comprise a small ATP-dependent chaperone. In some cases, the fusion or target protein / peptide expressed by the expression cassette as described herein may not comprise a small ATP-dependent chaperone Spy. Methods for identifying expression cassettes
[0095] Provided herein are methods for identifying the expression cassettes. The methods may comprise (a) identifying the nucleotide or sequence of the additional nucleotide. The methods may comprise randomly mutating a sequence in an expression cassette for screening. The expression cassette for screening may comprise any expression cassette as described herein. The expression cassette for screening may comprise the target protein / peptide as described herein. The expression cassette for screening may further comprise the sequence (s) of the cleavage site, the start codon, the affinity tag protein / peptide, or a combination thereof, as described herein. The nucleotide being mutated may be located 5’, 3’, 5’ and 3’, or within the sequence (s) of the cleavage site, the start codon, the affinity tag protein / peptide, or a combination thereof, as described herein. Thus, the expression cassette for screening may comprise any expression cassette configuration as described herein. The method may further comprise (b) expressing and measuring the expression level of the target protein / peptide or the fusion protein / peptide as described herein. The method may further comprise (c) identifying or selecting the expression cassette (s) having a desirable expression level of the target protein / peptide or the fusion protein / peptide (such as those described herein) . The desirable expression level may comprise the yield of the protein expression of the target protein / peptide or fusion protein / peptide as described herein. The identified / selected cassette (s) can be further mutated. Step (c) may comprise sequencing the selected expression cassettes (or the unselected for comparison) . The resultant expression cassette (s) can be used to express the target protein / peptide or fusion protein / peptide, and the expression level can be measured. The sequence or mutated sequence of the selected expression cassettes (s) may be the additional nucleotide as described herein. The method may thus comprise repeating the steps (a) - (c) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times using further mutated / identified expression cassettes. The procedures of the method for identifying the expression cassettes is also described elsewhere in this disclosure (for examples, see Examples 1-4) .
[0096] In some cases, the methods for identifying expression cassettes may be modified: The methods may comprise (a) identifying the nucleotide or sequence for expressing a fusion protein / peptide or target protein / peptide. The fusion protein / peptide may comprise the affinity tag protein / peptide, the cleavage site, the target protein / peptide, or a combination thereof, as described herein. The methods may comprise (b) expressing and measuring the expression level of the target protein / peptide or the fusion protein / peptide as described herein. For example, various combinations / configurations of the affinity tag protein / peptide, the cleavage site, the target protein / peptide in the expression cassette may be screened for the expression of the target protein / peptide or the fusion protein / peptide. The method may further comprise (c) identifying or selecting the expression cassette (s) having a desirable expression level of the target protein / peptide or the fusion protein / peptide (such as those described herein) . The desirable expression level may comprise the yield of the protein expression of the target protein / peptide or fusion protein / peptide as described herein. Step (c) may comprise sequencing the selected expression cassettes (or the unselected for comparison) . The resultant expression cassette (s) can be used to express the target protein / peptide or fusion protein / peptide, and the expression level can be measured. The method may thus comprise repeating the steps (a) - (c) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times using further mutated / identified expression cassettes. Methods for generating target protein / peptides
[0097] Provided herein are methods for generating protein / peptide or target protein / peptides using the expression cassettes as described herein. The method may comprise (a) expressing the protein / peptide encoded by the expression cassette. Step (a) may comprise cloning the expression cassette into a vector, as described herein. Step (a) may comprise introducing the expression cassette into a host cell, as described herein. The method may further comprise (b) processing the protein / peptide. Step (b) may comprise purifying the protein / peptide. The purifying may comprise using the affinity tag protein / peptide to purify the expressed protein / polypeptide. Step (b) may further comprise cleaving the expressed protein / peptide. The cleaving may comprise subjecting the purified protein / peptide or the non-purified protein / peptide with a protease specific to the cleavage site as described herein. The cleaving can separate the affinity tag protein / peptide and / or any peptide / amino acid (s) from the target protein / peptide. The cleaving may be carried out subsequent to the purifying. The cleaving may be carried out during or prior to the purifying. The method may comprise (d) storing the purified / cleaved target protein / peptide (or the non-purified or non-cleaved expressed protein / peptide) . Step (d) may comprise suspending the protein / peptide within a buffer solution as described herein. The methods for generating protein / peptide or target protein / peptides may thus comprise steps (a) - (d) . The methods for generating protein / peptide or target protein / peptides may thus comprise, from first to last, steps (a) - (d) . Vectors and host cells
[0098] Provided herein are expression vectors for carrying the expression constructs (e.g., the expression cassette described herein) or expressing the protein or polypeptide encoded by the sequence of the expression constructs as defined herein and / or comprising an isolated nucleic acid molecules encoding any of the protein or polypeptide defined herein. The present invention encompassed any suitable expression vectors including, but not limited to, plasmids, viruses, cosmids, bacmids, shuttle vectors, artificial chromosomes, etc. In one particular embodiment, the expression vector is a plasmid. Examples of potentially useful plasmid include, but are not limited to, pET24a, pLMAR, pALTER-Ex1 , pALTER-Ex2, pBAD / His, pBAD / Myc-His, pBAD / glll, pBacP, pBac, pBacPTandem, pBacTandem, pBacPTandem, pBacTandemRev, pCal-n, pCal-n-EK, pCal-c, pCal-Kc, pcDNA 2.1 , pDUAL, pET-3a-c, pET-9a-d, pET-11a-d, pET-12a-c, pET-14b, pET-15b, pET-16b, pET-17b, pET-19b, pET-20b (+) , pET-21 a-d (+) , pET-22b (+) , pET-23a-d (+) , pET-24b-d (+) , pET-25b (+) , pET-26b (+) , pET-27b (+) , pET-28a-c (+) , pET-29a-c (+) , pET-30a-c (+) , pET-31 b (+) , pET-32a-c (+) , pET-33b (+) , pET-34b (+) , pET-35b (+) , pET-36b (+) , pET-37b (+) , pET-38b (+) , pET-39b (+) , pET-40b (+) , pET-41a-c (+) , pET-42a-c (+) , pET-43a-c (+) , pETBIue-1 , pETBIue-2, pETBIue-3, pGEMEX-1 , pGEMEX-2, pGEX-I IT, pGEX-2T, pGEX-2TK, pGEX-3X, pGEX-4T, pGEX-5X, pG EX-6, P, pHAT10 / 11 / 12, pHAT20, pHAT-GF, Puv, pKK223-3, pLEX, pMAL-c2X, pMAL-c2E, pMAL-c2G, pMAL-, p2X, pMAL-, p2E, pMAL-, p2G, p, ProEX HT, p, PROLar. A, p, PROTet. E, pQE-9, pQE-16, pQE-30 / 31 / 32, pQE-40, pQE-60, pQE-70, PQE-80 / 81 / 82L, pQE-100, pRSET, pSE280, pSE380, pSE420, pThioHis, pTrc99A, pTrcHis, pTrcHis2, pTriEx-1 , pTriEx-2, pTrxFus, pBP26, pBP27, pBQ200, pGP380, pGP382, pGP3273, pGM1202.
[0099] Also provided herein are host cells for carrying the expression constructs or the vectors carrying thereof. The host cells may be used for expressing the protein or polypeptide encoded by the sequence of the expression constructs as defined herein and / or comprising an isolated nucleic acid molecules encoding any of the protein or polypeptide defined herein. The host cell may comprise bacterial cells, yeast cells, fungal cells, insect cells, mammalian cells, etc. In embodiments the cells are E. coli bacterial cells such as E. coli BL21 (DE3) , E. coli BL21 , and in E. coli ArhaB (B0002) . In embodiments the cells are E. coli BL21 (DE3) -pLysS, E. coli BL21 Star-pLysS, E. coli BL21-SI, E. coli BL21-AI, E. coli Tuner, E. coli Tuner pLysS, E. coliO rigami, E. coliO rigami B, E. coli Origami B pLysS, E. coli Rosetta, E. coli Rosetta pLysS, E. coli Rosetta-gami-pLysS, E. coli Rosetta2, E. coli Rosetta2 pLysS, E. coli BL21 CodonPlus, E. coli AD494, E. coli BL21trxB, E. coli HMS174, E. coli NovaBlue (DE3) , E. coli BLR, E. coli C41 (DE3) , E. coli C43 (DE3) , E. coli Lemo21 (DE3) , E. coli SHuffle T7, E. coli ArcticExpress, E. coli ArcticExpress (DE3) . In embodiments the cells are Streptomyces lividans, Lactoccocus lactis and Bacillus subtilis. In some cases, the host cell may comprise a cell of Saccharomyces Cerevisiae, Spodoptera frugiperda (e.g., Sf9 or Sf21) , Human Embryonic Kidney 293, Chinese hamster ovary cells, A549, Baby hamster kidney (BHK) cells, CAD, DUKX-X11, HeLa, Hep G2, HT1080, J558L, L929, MCF-7, N2a, NIH 3T3, P19, SO-Rb50, U2OS, Y79, or a combination thereof. In some cases, the host cell may comprise a cell of Saccharomyces Cerevisiae. In some cases, the host cell may comprise a cell of Spodoptera frugiperda (e.g., Sf9 or Sf21) . In some cases, the host cell may comprise a cell of Human Embryonic Kidney 293. In some cases, the host cell may comprise a cell of Chinese hamster ovary cells. In some cases, the host cell may comprise a cell of A549. In some cases, the host cell may comprise a cell of Baby hamster kidney (BHK) cells. In some cases, the host cell may comprise a cell of CAD. In some cases, the host cell may comprise a cell of DUKX-X11. In some cases, the host cell may comprise a cell of HeLa. In some cases, the host cell may comprise a cell of Hep G2. In some cases, the host cell may comprise a cell of HT1080. In some cases, the host cell may comprise a cell of J558L. In some cases, the host cell may comprise a cell of L929. In some cases, the host cell may comprise a cell of MCF-7. In some cases, the host cell may comprise a cell of N2a. In some cases, the host cell may comprise a cell of NIH 3T3. In some cases, the host cell may comprise a cell of P19. In some cases, the host cell may comprise a cell of SO-Rb50. In some cases, the host cell may comprise a cell of U2OS. In some cases, the host cell may comprise a cell of Y79. When using various host cells, suitable vector systems may be selected. For example, Baculovirus vector may be used with insect cells.
[0100] As used herein and thereof, the singular forms “a, ” “an, ” and “the” include plural references unless the context clearly dictates otherwise.
[0101] The term “about” or “approximately” as used herein when referring to a measurable value such as an amount or concentration and the like, is meant to encompass variations of 20 %, 10 %, 5 %, 1 %, 0.5 %, or even 0.1 %of the specified amount. For example, “about” can mean plus or minus 10 %, per the practice in the art. Alternatively, “about” can mean a range of plus or minus 20 %, plus or minus 10 %, plus or minus 5 %, or plus or minus 1 %of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, up to 5-fold, or up to 2-fold, of a value. Where particular values can be described in the application and claims, unless otherwise stated the term “about” may be assumed to encompass the acceptable error range for the particular value. Also, where ranges, subranges, or both, of values can be provided, the ranges or subranges can include the endpoints of the ranges or subranges. The terms “substantially, ” “substantially no, ” “substantially free, ” and “approximately” can be used when describing a magnitude, a position or both to indicate that the value described can be up to a reasonable expected range of values. For example, a numeric value can have a value that can be + / -0.1 %of the stated value (or range of values) , + / -1 %of the stated value (or range of values) , + / -2 %of the stated value (or range of values) , + / -5 %of the stated value (or range of values) , + / -10 %of the stated value (or range of values) , etc. Any numerical range recited herein can be intended to include all sub-ranges subsumed therein.
[0102] Where values are described as ranges, it may be understood that such disclosure includes the disclosure of all possible sub-ranges within such ranges, as well as specific numerical values that fall within such ranges irrespective of whether a specific numerical value or specific sub-range is expressly stated.
[0103] The terms “comprise, ” “have, ” and “include” are open-ended linking verbs. Any forms or tenses of one or more of these verbs, such as “comprises, ” “comprising, ” “has, ” “having, ” “includes, ” and “including, ” are also open-ended. For example, any method that “comprises, ” “has, ” or “includes” one or more steps is not limited to possessing only those one or more steps and also covers other unlisted steps.
[0104] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed. EXAMPLES
[0105] These examples are provided for illustrative purposes only and not to limit the scope of the claims provided herein. Example 1: Methods for identifying expression cassette having increased protein / polypeptide expression
[0106] Provided herein are methods for identifying expression cassette having increased protein or polypeptide expression.
[0107] The stability of proteins solution at different pH values is important for their application in different areas. As shown in FIGs. 1A-1B, using mussel foot proteins (Mfps) Mefp5 containing 6 histidine residues (His-tag) as an example, the fusion protein exhibited a much narrower range of pH stability (soluble protein could only be achieved at pH of 3.75 or above and below 6.5) , compared with mefp5 proteins without the 6 histidine residues (soluble protein was achieved at pH 3.75-8.60) .
[0108] The stability of Mefp5 without any histidine tags in relatively higher pH values (particularly at neutral and slightly basic conditions) is beneficial for expanding the application range of mefp5 proteins and others. In addition to decreasing the stability, His-tag in fusion proteins can bring about unexpected side effects in various application. For example, his-tagged insulin potentially can affect the safety of drugs used in clinical applications.
[0109] To enhance solution stability of mussel foot proteins (Mfps) across a wider pH range and thus facilitate application, a method for generating an expression cassette for expressing Mfps are provided. Such expression cassette can provide various beneficial advantages: (1) The application of Mfps often involves precise combination or mixture with a variety of materials and reagents with different properties, the solution stability of Mfps thus is one parameter for consideration; and / or (2) a much wider pH range stability can also facilitate long-term storage. Additionally, the expression cassette can facilitate enhanced protein production, facilitating large-scale production, and / or accelerating the industrial and therapeutic application of Mfps.
[0110] FIG. 1C depicts an exemplary expression cassette for expressing target proteins (such as Mfps) . The expression cassette comprises, from 5’ to 3’, a starting methionine, a first sequence as described herein, and the target sequence (such as nucleic acid sequence encoding Mfps) . The first sequence and target sequence can be expressed as one polypeptide sequence. A cleavage site as described herein can be inserted between the first sequence and the target sequence, such that the target protein expressed can be rid of the polypeptide encoded by the first sequence. Such configuration has various beneficial advantages. For example, while the first sequence may facilitate processing of (such as purification) of the polypeptide expressed, the sequence may affect other characteristics of the polypeptide when compared to a natural polypeptide of the target sequence. The cleavage site may allow for or facilitate the generation of a polypeptide comprising the target sequence without the first sequence subsequent to the processing.
[0111] While traditional soluble or functional fusion tag approach for enhancing protein expression is efficient, the fusion tag can contain a large amount of amino acid residues (the number of amino acids in fusion tag is much larger than that of target protein) , yielding less amount of target proteins.
[0112] To increase both the expression level and final yield of target protein, the cassette design strategy outlined in FIG. 1C comprises the first sequence and target sequence in the coding sequence region inserted into the plasmid and then transferred into the protein expressing bacterial strain.
[0113] FIG. 1D depicts 4 exemplary expression cassette designs. To ensure the removal of extra residues from the first sequence, a cleavable tag approach was used by inserting a cleavable site as described herein neighboring to the N-terminus of target sequence. To ensure higher yield, the number of amino acids in the first sequence (comprising the variable region) were minimized. For example, the expression cassette comprises only 0-32 amino acids (or less than 96 nucleotides) in the variable region. “X” denotes nucleotides that would be subjected to the methods for identifying the sequence of the expression cassette as described herein. FIG. 1E shows that when not fused to the first sequence as depicted in FIG. 1F, no obvious expression band of the mefp5 could be observed. To identify the nucleotide sequence in the first sequence for enhancing the expression of the mefp5, degenerate forward and reverse primer sequences (ATGNNNNNNGAAAACCTGTATTTTTCAG [SEQ ID NO: 170] and CTGAAAATACAGGTTTTCNNNNNNCAT [SEQ ID NO: 171] ) were used for the high-throughput screening as described in this example. N denotes random nucleotides (A, C, G, or T) to be screened. FIG. 1G shows the SDS-PAGE gels of mefp5 expressed from various expression cassettes having various nucleotide sequences (or the encoded amino acids) in the first sequence.
[0114] The effects of varying the nucleotide sequences on the protein expression is described in Examples 2 and 3. Example 2: Determining the effects of the nucleotide sequence in the first sequence and the resultant mRNA sequence secondary structures on fusion protein expression
[0115] Provided herein are methods for determining the effects of the nucleotide sequence in the first sequence and the resultant mRNA sequence secondary structures on fusion protein expression.
[0116] Various exemplary expression constructs generated using the methods described in Example 1 was listed in Table 1. Table 1: Various exemplary expression constructs “ / ” denotes N / A (not available)
[0117] The designable region of the first sequence can contain any additional purification tag as described herein. In some cases, it can contain purification tag at the 5’ end, middle or 3’ end of the variable region. Additionally, the mRNA secondary structure of the coding sequence (including the first sequence and the target sequence) can be determined. Furthermore, the effects of charges of the amino acids in the first sequence on protein expression was examined in Example 3.
[0118] FIG. 2A shows the mRNA secondary structures determined using RNAfold WebServer (which can be found at http: / / rna. tbi. univie. ac. at / cgi-bin / RNAWebSuite / RNAfold. cgi) of a first subset of expression cassettes listed in Tables 1 and 5. The mRNA hairpin structure of the first sequence in Cases 1-6 was more substantial compared to that of Cases 7-9. Consequently, the protein expression level of Cases 7-9 were substantially higher than those of Cases 1-4. Additionally, the lengths of the variable region of the first sequence of Cases 4-9 were the same. It is found that under the same length, the more relaxed the mRNA structure of the first sequence, the higher the expression level of the target protein could be achieved. The expression data also shows that when the expression constructs only carried a His-tag but not the cleavable site (Case 3) , the protein expression level achieved was low, which is consistent with literature reports.
[0119] FIG. 2B shows the mRNA secondary structures determined using RNAfold WebServer of a second subset of expression cassettes listed in Tables 1 and 5. Based on Case 7, without changing the mRNA structure of the coding sequence, the effect of the length of the variable region single-stranded structure on the protein expression level was evaluated (Cases 7 and 10-13) . Changing the overall coding sequence of the mRNA structure by changing the target sequence of Case 11, but not the first sequence, formed Case 14. The impact of the single-strand length of the variable region structure on the expression level based on Case14 was also examined (Cases 14-18) . The protein expression yield data of Cases 7 and 10-13 show that the longer the single-stranded mRNA structure of the first sequence (up to 48 bp) , the higher protein expression level could be achieved. The expression level decreased when the single-strand length is longer than 48bp. The protein expression yield data of Case 14-18 was consistent with the protein expression yield data of Case 7 and 10-13.
[0120] FIG. 2C shows the mRNA secondary structures determined using RNAfold WebServer of a third subset of expression cassettes listed in Tables 1 and 5. Considering that a longer single-stranded structure at the 5’ end of the variable region cannot always lead to more protein expression, structures other than the single-stranded structure are also evaluated for their ability of facilitating expressing proteins. Thus, a hairp...
Claims
1.A polynucleotide comprising a sequence encoding a polypeptide, wherein said sequence encoding said polypeptide comprises a start codon, a sequence encoding an affinity tag peptide, at least one additional nucleotide, and a target sequence,wherein said target sequence comprises a sequence encoding a mussel foot protein or a fragment thereof,wherein said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1 %higher than an expression level of a comparable polypeptide encoded by a comparable polynucleotide under a same condition, and wherein said comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one additional nucleotide.2.The polynucleotide of claim 1, wherein said at least one additional nucleotide is 5’ to said sequence encoding said affinity tag peptide.3.The polynucleotide of claim 1, wherein said at least one additional nucleotide is 3’ to said sequence encoding said affinity tag peptide.4.The polynucleotide of claim 1, wherein said at least one additional nucleotide comprises at least two nucleotides, and wherein one of said at least two nucleotides is 5’ to said sequence encoding said affinity tag peptide, and the other one of said at least two nucleotides is 3’ to said sequence encoding said affinity tag peptide.5.The polynucleotide of claim 1, wherein said polynucleotide further comprises a sequence encoding a cleavage site.6.The polynucleotide of claim 5, wherein said sequence encoding said cleavage site is 5’ to said target sequence.7.The polynucleotide of claim 5, wherein said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, said at least one additional nucleotide, said sequence encoding said affinity tag peptide, said sequence encoding said cleavage site, and said target sequence.8.The polynucleotide of claim 5, wherein said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, said sequence encoding said affinity tag peptide, said at least one additional nucleotide, said sequence encoding said cleavage site, and said target sequence.9.The polynucleotide of claim 4, wherein said at least one nucleotide comprises at least two nucleotides, and wherein said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, one of said at least two nucleotides, a sequence encoding said affinity tag peptide, the other one of said at least two nucleotides, said sequence encoding said cleavage site, and said target sequence.10.The polynucleotide of claim 5, wherein said cleavage site is a protease cleavage site.11.The polynucleotide of claim 10, wherein said protease cleavage site comprises a cleavage site of a Tobacco Etch Virus (TEV) protease, a Human rhinovirus 3C protease (HRV3C) , a thrombin, Enterokinase, Factor Xa, Chymotrypsin, Collagenase, Dispase, Endopeptidase Arg-C, Endopeptidase Asp-N, Endopeptidase Glu-C, Endopeptidase Lys-C, Ficin, Kallikrein, Papain, Pepsin, plasmin, pronase, proteinase K, Subtilisin, Thermolysin, trypsin, or a combination thereof.12.The polynucleotide of claim 5, wherein said cleavage site is a chemical cleavage site.13.The polynucleotide of claim 12, wherein said chemical cleavage site comprises a cleavage site of cyanogen bromide (CNBr) , hydroxylamine, formic acid, or a combination thereof.14.The polynucleotide of claim 5, wherein said cleavage site comprises an intein cleavage site.15.The polynucleotide of claim 3, wherein said sequence encoding said cleavage site is 3’ to said target sequence.16.A polynucleotide comprising a sequence encoding a polypeptide, wherein said sequence encoding said polypeptide comprises a start codon, at least one additional nucleotide, and a target sequence,wherein said sequence encoding said polypeptide does not comprise a sequence encoding an affinity tag peptide,wherein said target sequence comprises a sequence encoding a mussel foot protein or a fragment thereof,wherein said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1 %higher than an expression level of a comparable polypeptide encoded by a comparable polynucleotide under a same condition, and wherein said comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one additional nucleotide.17.The polynucleotide of claim 16, wherein said polynucleotide further comprises a sequence encoding a cleavage site.18.The polynucleotide of claim 17, wherein said sequence encoding said cleavage site is 5’ to said target sequence.19.The polynucleotide of claim 18, wherein said sequence encoding said polypeptide comprises, from 5’ to 3’, said starting codon, said at least one additional nucleotide, said sequence encoding said cleavage site, and said target sequence.20.The polynucleotide of claim 17, wherein said cleavage site comprises a protease cleavage site.21.The polynucleotide of claim 20, wherein said protease cleavage site comprises a cleavage site of a Tobacco Etch Virus (TEV) protease, a Human rhinovirus 3C protease (HRV3C) , thrombin, Enterokinase, Factor Xa, Chymotrypsin, Collagenase, Dispase, Endopeptidase Arg-C, Endopeptidase Asp-N, Endopeptidase Glu-C, Endopeptidase Lys-C, Ficin, Kallikrein, Papain, Pepsin, plasmin, pronase, proteinase K, Subtilisin, Thermolysin, trypsin, or a combination thereof.22.The polynucleotide of claim 17, wherein said cleavage site is a chemical cleavage site.23.The polynucleotide of claim 22, wherein said chemical cleavage site comprises a cleavage site of cyanogen bromide (CNBr) , hydroxylamine, formic acid, or a combination thereof.24.The polynucleotide of claim 17, wherein said cleavage site comprises an intein cleavage site.25.The polynucleotide of claim 17, wherein said sequence encoding said cleavage site is 3’ to said target sequence.26.The polynucleotide of claim 1 or 16, wherein said at least one additional nucleotide is configured to allow said polypeptide to be expressed at an expression level that is at least 1-fold, 5-fold, 10-fold, or 20-fold higher than said expression level of said comparable polypeptide.27.The polynucleotide of claim 1 or 16, wherein said at least one additional nucleotide comprises at least 3 nucleotides.28.The polynucleotide of claim 27, wherein said at least one additional nucleotide comprises multiples of 3 nucleotides.29.The polynucleotide of claim 27, wherein said at least one additional nucleotide comprises at most about 150 nucleotides.30.The polynucleotide of claim 27, wherein said at least one additional nucleotide comprises at most about 96 nucleotides.31.The polynucleotide of claim 27, wherein said at least one additional nucleotide comprises at most about 84 nucleotides.32.The polynucleotide of claim 27, wherein said at least one additional nucleotide comprises at most about 48 nucleotides.33.The polynucleotide of claim 27, wherein said at least one additional nucleotide comprises at least about 21 nucleotides.34.The polynucleotide of claim 1 or 16, wherein said polynucleotide is a ribonucleic acid (RNA) .35.The polynucleotide of claim 34, wherein said RNA is a messenger RNA (mRNA) .36.The polynucleotide of claim 35, wherein said mRNA folds to a secondary structure.37.The polynucleotide of claim 36, wherein said at least one additional nucleotide is not base-paired with another nucleotide of said polynucleotide.38.The polynucleotide of claim 36, wherein at least a nucleotide at a 5’ or a 3’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide.39.The polynucleotide of claim 38, wherein at least a nucleotide at said 5’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide.40.The polynucleotide of claim 39, wherein at least about 10 nucleotides at said 5’ end of said polynucleotide are single stranded.41.The polynucleotide of claim 40, wherein about 12-58 nucleotides at said 5’ end of said polynucleotide are single stranded.42.The polynucleotide of claim 38, wherein at least a nucleotide at said 3’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide.43.The polynucleotide of claim 42, wherein at least about 10 nucleotides at said 3’ end of said polynucleotide are single stranded.44.The polynucleotide of claim 38, wherein at least a nucleotide at said 5’ end and at least a nucleotide at said 3’ end of said polynucleotide is not base-paired with another nucleotide of said polynucleotide.45.The polynucleotide of claim 35, wherein said mRNA has a minimum free energy of at most about -10 kilocalorie per mole (kcal / mol) .46.The polynucleotide of claim 45, wherein said mRNA has a minimum free energy of at most about -50 kcal / mol.47.The polynucleotide of claim 46, wherein said mRNA has a minimum free energy of at least about -57 kcal / mol.48.The polynucleotide of claim 47, wherein said mRNA has a minimum free energy of at least about -55 kcal / mol.49.The polynucleotide of claim 48, wherein said mRNA has a minimum free energy of about -53 kcal / mol.50.The polynucleotide of claim 1 or 16, wherein said affinity tag peptide comprises a polyhistidine, a Myc tag, a Strep tag, a V5 tag, an HA tag, a Flag tag, polyarginine, Avi Tag, SNAP-Tag, Halo Tag, small Ub-related modifier protein (SUMO) , Glutathione S-transferase (GST) , Green fluorescent protein (GFP) , maltose-binding protein (MBP) , mCheery, or a combination thereof.51.The polynucleotide of claim 1 or 16, wherein said sequence encoding said polypeptide further comprises at least one second additional nucleotide, wherein said at least one second additional nucleotide is configured to allow said polypeptide to be expressed at a level that is at least 1%higher than a second comparable polypeptide encoded by a second comparable polynucleotide, and wherein said second comparable polynucleotide is identical to said polynucleotide but does not comprise said at least one second additional nucleotide.52.The polynucleotide of claim 51, wherein said at least one second additional nucleotide comprises at least one nucleotide substitution that is located within said sequence encoding said affinity tag peptide, and optionally wherein (a) said sequence encoding said affinity tag peptide comprising said at least one nucleotide substitution encodes an amino acid sequence that is more hydrophilic than an amino acid sequence encoded by a sequence encoding said affinity tag peptide not comprising said at least one nucleotide substitution; and / or (b) said sequence encoding said affinity tag peptide comprising said at least one nucleotide substitution encodes an amino acid sequence that is more positively charged than an amino acid sequence encoded by a sequence encoding said affinity tag peptide not comprising said at least one nucleotide substitution.53.The polynucleotide of claim 51, wherein said at least one second additional nucleotide is located at 5’ or at 3’ to said sequence encoding said affinity tag peptide, and wherein said sequence encoding said affinity tag peptide comprising said at least one second additional nucleotide encodes an amino acid sequence that is more positively charged than an amino sequence encoding said affinity tag peptide not comprising said at least one second additional nucleotide.54.The polynucleotide of claim 51, wherein said at least one second additional nucleotide comprises at least one nucleotide mutation that is located within said sequence encoding said mussel foot protein or said fragment thereof, optionally wherein said at least one mutation is a synonymous mutation.55.The polynucleotide of claim 54, wherein said sequence encoding said mussel foot protein or said fragment thereof comprising said at least one nucleotide mutation has a different nucleotide sequence relative to said sequence encoding said mussel foot protein or said fragment thereof not comprising said at least one nucleotide mutation, and wherein said sequence encoding said mussel foot protein or said fragment thereof comprising said at least one nucleotide mutation encodes the same amino acid sequence relative to said sequence encoding a mussel foot protein or said fragment thereof not comprising said at least one nucleotide mutation.56.The polynucleotide of claim 1 or 16, wherein said expression levels of said polypeptide and said comparable polypeptide are measured by levels of said polypeptide and said comparable polypeptide generated by two cells, and wherein each of said two cells comprises only one of said polynucleotide and said comparable polynucleotide.57.The polynucleotide of claim 56, wherein said levels of said polypeptide and said comparable polypeptides generated by said two cells are measured in milligram of polypeptides per liter of a culture of cells (mg*L-1) .58.The polynucleotide of claim 56, wherein said two cells are bacteria.59.The polynucleotide of claim 58. wherein said bacteria is E. coli.60.The polynucleotide of claim 1 or 16, wherein said polynucleotide further comprises a promoter, a translation initiator, a terminator, or a combination thereof.61.A vector comprising the polynucleotide of claim 1 or 16.62.The vector of claim 61, wherein said vector comprises a bacterial vector.63.A cell comprising the polynucleotide of claim 1 or 16or the vector of claim 61.64.The cell of claim 63, wherein said cell is a procaryotic cell.65.The cell of claim 64, wherein said cell is a bacteria.66.The cell of claim 65, wherein said bacteria is E. coli.67.The cell of claim 63, wherein said cell is an eucaryotic cell.68.The cell of claim 67, wherein said cell is a yeast cell, a fungal cell, an insect cell, or a mammalian cell.69.A method, comprising expressing said polypeptide encoded by the polynucleotide of any one of claim 1 or 16or the vector of claim 61.70.The method of claim 69, further comprising purifying said polypeptide.71.The method of claim 70, wherein said purifying comprises purifying said polypeptide using said affinity tag peptide.72.The method of claim 69, further comprising generating a polypeptide comprising said mussel foot protein or said fragment thereof but not said affinity tag peptide.73.The method of claim 72, wherein said generating comprises contacting said polypeptide with a protease.74.The method of claim 73, wherein said protease comprises a Tobacco Etch Virus (TEV) protease , a Human rhinovirus 3C protease (HRV 3C) , a thrombin , an Enterokinase , a Factor Xa , a Chymotrypsin , a Collagenase , a Dispase , an Endopeptidase Arg-C , an Endopeptidase Asp-N , an Endopeptidase Glu-C , an Endopeptidase Lys-C , a Ficin , a Kallikrein , a Papain , a Pepsin , a plasmin , a pronase , a proteinase K , a Subtilisin , a Thermolysin , a trypsin , or a combination thereof.
Citation Information
Patent Citations
Injectable self-repairing underwater protein and applications thereof
CN108948208A
Bicistronic translation coupling expression vector and application thereof
CN115786377A
Genetically engineered bacterium for exocytosis of mussel protein as well as construction method and application of genetically engineered bacterium
CN116042502A
Method for secretory expression of mussel adhesion protein in bacillus subtilis
CN116286932A
Mussel bioadhesive
CN1989152A