Engineered SELB and methods of use thereof
A SECIS-independent SelB variant with specific mutations and truncations allows for the incorporation of selenocysteine and other non-canonical amino acids into peptides, addressing the limitations of existing methods and expanding the genetic code through orthogonal elongation factor functionality.
Patent Information
- Application Number
- PCT/US2025/030815
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-28
- Filing Date
- 2025-05-23
- Publication Date
- 2025-12-04
AI Technical Summary
Existing methods struggle to incorporate selenocysteine across from stop codons without relying on the SECIS element and to have an orthogonal elongation factor that can function in parallel with EF-Tu, limiting the expansion of the genetic code.
Development of a SECIS-independent selenocysteine-specific elongation factor (SelB) variant with specific mutations and truncations, capable of incorporating non-canonical amino acids into peptides, and a system that includes EF-Tu for parallel functionality.
Enables the efficient incorporation of selenocysteine and other non-canonical amino acids into peptides without the need for SECIS, expanding the genetic code and providing an orthogonal elongation factor for broader amino acid utilization.
Smart Images

Figure IMGF000036_0001 
Figure IMGF000036_0002 
Figure IMGF000036_0003
Abstract
Description
[0001] ENGINEERED SELB AND METHODS OF USE THEREOF
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0003] This invention was made with government support under Grant numbers W91 INF-22-2-0246, W911NF-16-1-0372 and W911NF-23-P-0004 awarded by the Army Research Office. The Government has certain rights in the invention.
[0004] RELATED APPLICATION
[0005] This PCT application claims priority to, and the benefit of, U.S. Provisional Patent Application No. 63 / 652,399, filed May 28, 2024, entitled “ENGINEERED SELB AND METHODS OF USE THEREOF,” and U.S. Provisional Patent Application No. 63 / 742,654, filed January 7, 2025, entitled “ENGINEERED SELB AND METHODS OF USE THEREOF,” which are incorporated by reference herein in its entirety.
[0006] REFERENCE TO SEQUENCE LISTING
[0007] The sequence listing submitted on May 23, 2025, as an .XML file entitled “10046- 608WOl_ST26” created on May 22, 2025, and having a file size of 112,576 bytes is hereby incorporated by reference pursuant to 37 C.F.R. § 1.52(e)(5).
[0008] FIELD
[0009] The present disclosure relates to engineered selenocysteine-specific elongation factor (SelB) variant and methods of generating and using the variants thereof.
[0010] BACKGROUND
[0011] In E. coli, the selenocysteine (Sec) incorporation machinery is coded in the genome as SelABC (Bock 1991; Chung 2021). The tRNASec(SelC) is charged by seryl tRNA synthetase, and selenocysteine synthase (SelA) converts the charged serine into selenocysteine (Serrao 2018, Yoshizawa 2009). SelB is particularly important as it is an elongation factor that specifically recruits Sec-tRNASecto the ribosome, loading it adjacent to a specific RNA structure, the SECIS [selenocysteine insertion sequence] element, that is found only near a few opal (UGA) codons (Baron 1993).
[0012] In addition, the unique properties of Sec (higher nucleophilicity compared to cysteine; higher stability of diselenide bond compared to disulfide bond) makes this ‘21’ st amino acid’ a potentially useful addition to an expanded genetic code (Amer 2010; Hondal 2013). Rather than site-specifically incorporating Sec adjacent to the SECIS, previous efforts have attempted to broaden incorporation by engineering a Sec-tRNASccvariant that is capable of being recognized by the more general endogenous elongation factor, EF-Tu. Soil and colleagues achieved this by transplanting regions of tRNASecinto tRNASer8and further engineering through rational design yielding tRNAUTuX9, while Thyer and colleagues achieved this by directed evolution of anti-determinant sequences within tRNASec, making it recognizable by EF-Tu (yielding tRNASecux) (Thyer 2015).
[0013] The engineering of SelB to be SECIS-independent is also an interesting prospect since it could then serve as an ‘orthogonal’ analogue of EF-Tu. EF-Tu loads tRNAs charged with their correct canonical amino acids in part by balancing the affinity to the individual tRNA (especially the T-stem) in one set of interactions with the affinity to the amino acid in another set of interactions (LaRiviere 2004; Dale 2004; Asahara 2005; Schrader 2011). An expanded genetic code is likely to either not fit into or upset this balance, or both, given that new amino acids would be unlikely to have balanced affinities with their tRNA partners. Researchers have taken on this problem by optimizing the EF- Tu amino acid binding pocket for new amino acids such as phosphoserine or phosphotyrosine (Park 2011; Fan 2016). However, given the inherent difficulty of EF-Tu’s balancing, having an orthogonal EF-Tu functioning in parallel would allow a more ready expansion of the code than constantly changing and rebalancing the endogenous EF-Tu.
[0014] What is needed in the art is a version of SelB that is SECIS-independent, wherein this SelB variant can generally incorporate selenocysteine across from stop codons. Further, what is needed is an orthogonal EF-Tu that can function in parallel to normal EF-Tu, which provides for the ability to engineer an elongation factor that is specific for unnatural amino acid-charged tRNA.
[0015] SUMMARY
[0016] The present disclosure provides a selenocysteine-specific elongation factor (SelB) variant. The present disclosure also provides vectors, cells, and systems comprising the selenocysteinespecific elongation factor (SelB) variant. The present disclosure provides methods generating a selenocysteine-specific elongation factor (SelB) variant. The present disclosure also provides methods of using the selenocysteine- specific elongation factor (SelB) variant.
[0017] In some aspects, disclosed herein is an engineered selenocysteine-specific elongation factor (SelB) variant comprising at least 80% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6, wherein the SelB variant is capable of incorporating at least one non-canonical amino acid into a peptide. In some embodiments, the SelB variant comprises at least 90% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the SelB variant comprises SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6.
[0018] In some embodiments, the SelB variant is a truncation of SEQ ID NO: 3. In some embodiments, the truncation comprises a deletion of a carboxy terminal end. In some embodiments, the deletion comprises removal of at least 100 of amino acids from the truncation. In some embodiments, the deletion comprises removal of 127 amino acids from the truncation.
[0019] In some embodiments, the SelB variant comprises at least one amino acid mutation within the truncation. In some embodiments, the at least one amino acid mutation is selected from the group consisting of I98V, E166D, S206P, P277S A413T, S473N, W477R, and F487L. In some embodiments, the SelB variant further comprises an amino acid mutation selected from E246K and E463K within the truncation. In other embodiments, the SelB variant further comprises an amino acid mutation selected from Y44F, T193C, and / or R236S.
[0020] In one aspect, disclosed herein is an expression vector comprising a nucleic acid sequence comprising at least 80% sequence identity to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7, wherein the nucleic acid sequence encodes a Selenocysteine-specific elongation factor (SelB) variant.
[0021] In some embodiments, the nucleic acid sequence comprises at least 90% sequence identity to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7. In some embodiments, the nucleic acid sequence comprises SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7.
[0022] In some embodiments, the nucleic acid sequence encodes at least one amino acid mutation.
[0023] In some aspects, disclosed herein is a method of incorporating at least one non-canonical amino acid into a peptide during translation, the method comprising contacting an mRNA sequence with an engineered selenocysteine-specific elongation factor (SelB) variant, wherein the engineered SelB variant incorporates a non-canonical amino acid into the peptide without need of a selenocysteine insertion sequence (SECIS).
[0024] In some embodiments, translation occurs in presence of one or more translation regulatory proteins. In some embodiments, the one or more translation regulatory proteins comprise an initiation factor, an elongation factor, a release factor, and / or a ribosome. In some embodiments, the elongation factor comprises elongation Factor Tu (EF-Tu).
[0025] In some embodiments, the method comprises the SelB variant comprising at least 80% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the method comprises the SelB variant comprising at least 90% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the method comprises the SelB variant comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some aspects, disclosed herein is a peptide produced by the method of any preceding aspect, the peptide comprising at least one selenocysteine residue and / or at least one serine residue.
[0026] In some aspects, disclosed herein is a protein translation system comprising an engineered selenocysteine-specific elongation factor (SelB) variant, wherein the SelB variant comprises at 80% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6, wherein the SelB variant is capable of incorporating at least one non-canonical amino acid into a peptide, and wherein the SelB variant translates an mRNA sequence without need of a selenocysteine insertion sequence (SECIS).
[0027] In some embodiments, the system comprises the SelB variant comprising at least 80% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the system comprises the SelB variant comprising at least 90% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the system comprises the SelB variant comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6.
[0028] In some embodiments, the system further comprises an Elongation Factor Tu (EF-Tu) protein. In some embodiments, the system further comprises a plurality of nucleotides selected from the group consisting of adenine (A), uracil (U), guanine (G), and cytosine (C). In some embodiments, the system further comprises one or more translation regulatory proteins selected from an initiation factor, an elongation factor a release factor, or a ribosome.
[0029] In some aspects, disclosed herein is a cell comprising the engineered SelB variant, the vector, or the system of any preceding aspect.
[0030] BRIEF DESCRIPTION OF FIGURES
[0031] The accompanying figures, which are incorporated in and constitute a part of this specification, illustrate several aspects described below.
[0032] Figure 1 shows the crystal structure of SelB in complex with tRNASecand SECIS. The bottom region of the structure is the SECIS element, while the top and right regions of the structure are portions of the full length SelB (the right region is the C-terminal that is truncated in the present disclosure). The left region is the tRNASec. PDB ID: 51zd.
[0033] Figures 2A, 2B, and 2C show the initial activity of wild-type SelB versus truncated SelB.
[0034] Figure 2A shows the map of the plasmid used for SelB expression. The actual sequence of plasmid is shown in SEQ ID NO: 9. SelA and selC are constitutively expressed while SelB is under the control of araBAD promoter. Figure 2B shows the crystal structure of NMC-A showing the essential disulfide bond between C69 and C238 (PDB: 1BUE). Figure 2C shows the growth in the presence of carbenicillin following induction of either WT-SelB (blue) or truncated-SelB (red). Each data point represents the mean of three independent experiments ± s.d (standard deviation). Figures 3A and 3B show the directed evolution of SelB. Figure 3A shows the schematic of the directed evolution of SelB. A plasmid-based SelB library was transformed into C321 dSelABC and grown in the presence of both carbenicillin and selenium. Grown cells were collected, and plasmid was recovered, then SelB variants were amplified by PCR and cloned back into the same backbone. A detailed summary of the rounds of selection is shown in Tables la and lb. Figure 3B shows the growth assay in the presence of cabenicillin with WT-SelB (blue), truncated SelB (red), SelB-vl (green) and SelB-v2 (purple). Each data point represents the mean of three independent experiments ± s.d (standard deviation).
[0035] Figures 4A and 4B show the monitoring Sec incorporation with fluorescent proteins. Figure 4A shows the smURFP expression in the presence or absence of expression of SelB WT or SelB-v2, and in the presence or absence of the addition of selenium to the media. Fluorescence was measured at 633 nm / 695 nm excitation / emission. Each data point represents the mean of three independent experiments + s.d (standard deviation). Figure 4B shows the sfGFP expression in the presence or absence of SelB WT or SelB-v2, and in the presence or absence of selenium in the media. Fluorescence was measured at 485 nm / 528 nm excitation / emission. Each data point represents the mean of three independent experiments + s.d (standard deviation).
[0036] Figures 5A, 5B, and 5C show the mass spectrometric and functional characterization of SelB- v2 insertion of selenocysteine. Figure 5A shows that SelB-v2 was expressed in C321 dSelABC along with either DHFR P39S (top) or DHFR P39TAG (bottom), and the DHFR proteins were purified in the presence of selenium and MS analyses carried out. The top shows the deconvoluted mass spectrum of DHFR, with serine at position 39 (calculated mass is 19029.48 and observed mass is 19028.15). The bottom is the deconvoluted mass spectrum of DHFR with TAG at position 39 (calculated mass for selenocysteine incorporated DHFR is 19090.42 and observed mass is 19089.95). Figure 5B shows the EThcD (Electron-Transfer / higher-energy collision Dissociation) fragmentation pattern of the purified DHFR P39S (top) and DHFR P39TAG (bottom) which shows the selenium-sulfur bond. Slash marks represent the cleavage sites for different N-terminal and C-terminal ions; regions that are ‘protected’ by covalent crosslinking drop out of the mass spectra which was also reported in previously reported selenocysteine containing DHFR. Figure 5C shows the benzyl viologen colorimetric assay. The NMC-A reporter was replaced with a formate dehydrogenase H (fdhF) reporter. The reduction of benzyl viologen (monitored by color change) by expressed fdhF confirms the successful incorporation of selenocysteine into the protein. The functional fdHF expression assay was carried out with SelB-WT, SelB-v2 and SelB-truncated.
[0037] Figures 6A, 6B, 6C, and 6D show the characterization of SelB mutants for serine incorporation. Figure 6A shows the growth assay in the presence of carbenicillin with expression of SelB-v2 and NMC-A S71TAG, which replaces a key serine with a stop codon. The blue bar indicates growth without induction while the red bar is with induction of SelB-v2 by IPTG. Each data point represents the mean of three independent experiments ± s.d (standard deviation). Figure 6B shows the growth assay in the presence of carbenicillin with expression of NMC-A S71TAG and either induced SelB-v2 (red) or the evolved SelB-v2-Ser (green). Each data point represents the mean of three independent experiments ± s.d (standard deviation). Figure 6C shows the growth assay in the presence of carbenicillin with expression of NMC-A C70TAG (which is sensitive to selenocysteine insertion) and either induced SelB-v2 (red) or SelB-v2-Ser (green). Each data point represents the mean of three independent experiments ± s.d (standard deviation). Figure 6D shows the intact mass spectra of DHFR P39TAG expressed and purified in the presence of selenium and SelB-v2 (top) or SelB-v2-Ser (bottom). The calculated mass of DHFR P39S is 19029.48 and DHFR P39Sec with wild-type SelB is 19090.42, whereas when SelB-v2 is co-expressed, the observed mass for DHFR P39S is 19028.08 and for DHFR P39Sec is 19089.03. When SelB-v2-Ser is co-expressed, the observed mass for P39S is 19029.06 and for DHFR P39Sec is 19091.01.
[0038] Figure 7 shows the plasmid map of pSelC-tac-SelB. Compared to pSelAC-araBAD-SelB, SelA gene was removed as well as the promoter of SelB was changed from araBAD to tac-lac promoter. araC was replaced with LacI as well.
[0039] Figure 8 shows the plasmid map of SelB -vl -containing plasmid recovered after round 4. After round 4, the plasmid backbone had duplications in NMC-A and SelC (highlighted with red letters), while SelB was in- frame with the duplicated SelA. The sequence is shown in SEQ ID NO: 28.
[0040] Figures 9A and 9B show the NMC-A assay of SelB-vl. Figure 9A shows the NMC-A assay of SelB-vl having one TAG codon in the NMC-A gene, (b) NMC-A assay of SelB-vl using two TAG codons in the NMC-A gene.
[0041] Figure 10 shows the LC-MS result of trypsin-digested sfGFP expressed in the presence of selenium and SelB-v2. Ion fragments corresponding to dehydroalanine incorporation were observed in LC-MS / MS. Shown is an example of the spectrum of a trypsin-digested peptide from sfGFP whose sequence is LEYNFNSHXVYITADK (m / z 1881.89; SEQ ID NO: 8), where X corresponds to dehydroalanine. Both b and y ions from this peptide fragments were observed, indicating dehydroalanine incorporation.
[0042] Figure 11 show mass spectrometry of DHFR P61TAG in a C321dSelABC strain background. The deconvoluted mass spectrum of purified DHRF P39TAG protein which was expressed in the presence of SelB-v2 in C321 dSelABC. Mass spectrometry shows that both serine-containing DHFR and selenocysteine-containing DHFR were present in the sample mixture: the calculated mass for DHRF P39S (serine-containing) is 19029.48 and the observed mass is 19028.08, while the calculated mass for DHFR P61U (selenocysteine-containing) is 19090.42 and the observed mass is 19089.03. Calc., calculated; Obs., observed.
[0043] Figures 12A and 12B show the NMC-A assay with S71 mutants. Figure 12A shows the NMC- A assay using wild-type NMC-A. The blue bar refers to growth without induction of SelB-v2, while the red bar indicates growth with induction with SelB-v2. Each data point represents the mean of three independent experiments ± s.d (standard deviation). Figure 12B shows the NMC-A assay of the NMC-A S71C mutant. The blue bar refers to growth without induction of SelB-v2, while the red bar indicates growth with induction with SelB-v2. Each data point represents the mean of three independent experiments ± s.d (standard deviation).
[0044] Figure 13 shows the SDS-PAGE of purified DHFR expressed in different strains. SDS-PAGE analysis of purified DHFR expressed in either the DH10B dSelABC strain or the C321 dSelABC strain. Lane 1 corresponds to purified DHFR P39S while lanes 2-4 are purified DHFR P61TAG variants. For lanes 1-3 SelB-v2 was induced, while for lane 4, SelB-v2-Ser was induced, instead of SelB-v2. For the DHFR P61TAG protein that was expressed in DH10B dSelABC with SelB-v2, only DHFR P39U (selenocysteine) was observed, while no DHFR P61S (serine) (lane 2). In the context of C321 dSelABC, however, some DHFR P61S (serine) was observed (lane 3). When DHFR P61TAG was expressed in the presence of SelB-v2-Ser, however, primarily DHFR P61S (serine) was observed (lane 4).
[0045] Figure 14 shows the non-truncated SelB RNA-binding domain undergoes conformational change upon S206P and P277S mutations. Structures generated with AlphaFold3, pTM scores of 0.81, 0.82, 0.82, 0.83 for WT, S206P, P277S, and [S206P, P277S] double mutant, respectively. Residues 206 and 277 are highlighted in red. The RNA-binding domain (right side) undergoes a large conformational change upon the introduction of these mutations.
[0046] Figure 15 shows a total of four rounds in LB medium and 1 round of plate-based screening.
[0047] Figure 16 shows the map of the plasmid used for SelB expression.
[0048] Figure 17 shows the NMC-A assay of SelB full, SelB-trunc, and SelB-vl. Figure 17 also shows the assay of SelB-vl having one TAG codon in the NMC-A gene. Figure 17 also shows the NMC-A assay of SelB-vl using two TAG codons in the NMC-A gene.
[0049] Figure 18 shows a total of seven rounds in LB medium using NMC-A with two stop codons.
[0050] Figure 19 shows the ten hand-picked clones after seven rounds of screening. All clones had an E463K mutation, nine clones had an E246K mutation. Both mutations were introduced into SelB vl.
[0051] Figure 20 shows the beta-lactamase assay with two TAG codons. The presence of the E246K and E463K mutations increased Sec incorporation. Figure 21 shows beta-lactamase assay with a single TAG codon. The effect of E246K and E463K mutations was confirmed with a single TAG codon bearing beta-lactamase assay.
[0052] Figure 22 shows the use of sfGFP to check Sec incorporation (sfGFP with TAG is located at position 149). When SelB induction and selenium presence, sfGFP expression was observed (use of DH10B dSelABC).
[0053] Figure 23 shows the comparison of corresponding EF-Tu residues.
[0054] Figure 24 shows the mutated positions within the SelB protein structure.
[0055] Figures 25 A and 25B show directed evolution of SelB. Figure 25 A shows the schematic of the directed evolution of SelB. A plasmid-based SelB library was transformed into C321 dSelABC and grown in the presence of both carbenicillin and selenium. Grown cells were collected, and plasmid was recovered, then SelB variants were amplified by PCR and cloned back into the same backbone. A detailed summary of the rounds of selection is shown in Table 1. Figure 25B shows the growth assay in the presence of cabenicillin with WT-SelB (blue), truncated SelB (red), SelB-vl (green) and SelB-v2 (purple). Each data point represents the mean of three independent experiments ± s.d (standard deviation).
[0056] Figure 26 shows the beta-lactamase assay of SelB constructs without addition of selenium to the media.
[0057] Figure 27 shows a diagram depicting the directed evolution of a SECIS -independent SelB protein that can incorporate selenocysteine and other amino acids, such as for example serine.
[0058] DETAILED DESCRIPTION
[0059] The following description of the disclosure is provided as an enabling teaching of the disclosure in its best, currently known embodiment(s). To this end, those skilled in the relevant art will recognize and appreciate that many changes can be made to the various embodiments of the invention described herein, while still obtaining the beneficial results of the present disclosure. It will also be apparent that some of the desired benefits of the present disclosure can be obtained by selecting some of the features of the present disclosure without utilizing other features. Accordingly, those who work in the art will recognize that many modifications and adaptations to the present disclosure are possible and can even be desirable in certain circumstances and are a part of the present disclosure. Thus, the following description is provided as illustrative of the principles of the present disclosure and not in limitation thereof.
[0060] Reference will now be made in detail to the embodiments of the invention, examples of which are illustrated in the drawings and the examples. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Terminology
[0061] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. The term “comprising” and variations thereof as used herein is used synonymously with the term “including” and variations thereof and are open, non-limiting terms. Although the terms “comprising” and “including” have been used herein to describe various embodiments, the terms “consisting essentially of’ and “consisting of’ can be used in place of “comprising” and “including” to provide for more specific embodiments and are also disclosed. As used in this disclosure and in the appended claims, the singular forms “a”, “an”, “the”, include plural referents unless the context clearly dictates otherwise.
[0062] The following definitions are provided for the full understanding of terms used in this specification.
[0063] Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. It is also understood that when a value is disclosed that “less than or equal to” the value, “greater than or equal to the value” and possible ranges between values are also disclosed, as appropriately understood by the skilled artisan. For example, if the value “10” is disclosed the “less than or equal to 10” as well as “greater than or equal to 10” is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and starting points, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point 15 are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed. “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0064] “Composition” refers to any agent that has a beneficial biological effect. Beneficial biological effects include both therapeutic effects, e.g., treatment of a disorder or other undesirable physiological condition, and prophylactic effects, e.g., prevention of a disorder or other undesirable physiological condition. The terms also encompass pharmaceutically acceptable, pharmacologically active derivatives of beneficial agents specifically mentioned herein, including, but not limited to, a vector, polynucleotide, cells, salts, esters, amides, proagents, active metabolites, isomers, fragments, analogs, and the like. When the term “composition” is used, then, or when a particular composition is specifically identified, it is to be understood that the term includes the composition per se as well as pharmaceutically acceptable, pharmacologically active vector, polynucleotide, salts, esters, amides, proagents, conjugates, active metabolites, isomers, fragments, analogs, etc.
[0065] "Comprising" is intended to mean that the compositions, methods, etc. include the recited elements, but do not exclude others. "Consisting essentially of' when used to define compositions and methods, shall mean including the recited elements, but excluding other elements of any essential significance to the combination. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives, and the like. "Consisting of’ shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions provided and / or claimed in this disclosure. Embodiments defined by each of these transition terms are within the scope of this disclosure.
[0066] An "increase" can refer to any change that results in a greater amount of a symptom, disease, composition, condition, or activity. An increase can be any individual, median, or average increase in a condition, symptom, activity, composition in a statistically significant amount. Thus, the increase can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100% or more increase so long as the increase is statistically significant.
[0067] A "decrease" can refer to any change that results in a smaller amount of a symptom, disease, composition, condition, or activity. A substance is also understood to decrease the genetic output of a gene when the genetic output of the gene product with the substance is less relative to the output of the gene product without the substance. Also, for example, a decrease can be a change in the symptoms of a disorder such that the symptoms are less than previously observed. A decrease can be any individual, median, or average decrease in a condition, symptom, activity, composition in a statistically significant amount. Thus, the decrease can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100%, or more decrease so long as the decrease is statistically significant.
[0068] Reference also is made herein to peptides, polypeptides, proteins, and compositions comprising peptides, polypeptides, and proteins. As used herein, a polypeptide and / or protein is defined as a polymer of amino acids, typically of length>100 amino acids (Garrett & Grisham, Biochemistry, 2nd edition, 1999, Brooks / Cole, 110). A peptide is defined as a short polymer of amino acids, of a length typically of 20 or less amino acids, and more typically of a length of 12 or less amino acids (Garrett & Grisham, Biochemistry, 2nd edition, 1999, Brooks / Cole, 110).
[0069] The peptides, polypeptides, and proteins disclosed herein may be modified to include nonamino acid moieties. Modifications may include but are not limited to carboxylation (e.g., N-terminal carboxylation via addition of a di-carboxylic acid having 4-7 straight-chain or branched carbon atoms, such as glutaric acid, succinic acid, adipic acid, and 4,4-dimethylglutaric acid), amidation (e.g., C- terminal amidation via addition of an amide or substituted amide such as alkylamide or dialkylamide), PEGylation (e.g., N-terminal or C-terminal PEGylation via additional of polyethylene glycol), acylation (e.g., O-acylation (esters), N-acylation (amides), S-acylation (thioesters)), acetylation (e.g., the addition of an acetyl group, either at the N-terminus of the protein or at lysine residues), formylation lipoylation (e.g., attachment of a lipoate, a C8 functional group), myristoylation (e.g., attachment of myristate, a C14 saturated acid), palmitoylation (e.g., attachment of palmitate, a C16 saturated acid), alkylation (e.g., the addition of an alkyl group, such as an methyl at alysine or arginine residue), isoprenylation or prenylation (e.g., the addition of an isoprenoid group such as farnesol or geranylgeraniol), amidation at C-terminus, glycosylation (e.g., the addition of a glycosyl group to either asparagine, hydroxylysine, serine, or threonine, resulting in a glycoprotein). Distinct from glycation, which is regarded as a nonenzymatic attachment of sugars, polysialylation (e.g., the addition of polysialic acid), glypiation (e.g., glycosylphosphatidylinositol (GPI) anchor formation, hydroxylation, iodination (e.g., of thyroid hormones), and phosphorylation (e.g., the addition of a phosphate group, usually to serine, tyrosine, threonine, or histidine).
[0070] The phrases “percent identity” and “% identity,” as applied to polypeptide sequences, refer to the percentage of residue matches between at least two polypeptide sequences aligned using a standardized algorithm. Methods of polypeptide sequence alignment are well-known. Some alignment methods consider conservative amino acid substitutions. Such conservative substitutions, explained in more detail above, generally preserve the charge and hydrophobicity at the site of substitution, thus preserving the structure (and therefore function) of the polypeptide. Percent identity for amino acid sequences may be determined as understood in the art. (See, e.g., U.S. Pat. No. 7,396,664, which is incorporated herein by reference in its entirety). A suite of commonly used and freely available sequence comparison algorithms is provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST) (Altschul, S. F. et al. (1990) J. Mol. Biol. 215:403 410), which is available from several sources, including the NCBI, Bethesda, Md., at its website. The BLAST software suite includes various sequence analysis programs including “blastp,” that is used to align a known amino acid sequence with other amino acids sequences from a variety of databases.
[0071] Percent identity may be measured over the length of an entire defined polypeptide sequence or may be measured over a shorter length, for example, over the length of a fragment taken from a larger, defined polypeptide sequence, for instance, a fragment of at least 15, at least 20, at least 30, at least 40, at least 50, at least 70 or at least 150 contiguous residues. Such lengths are exemplary only, and it is understood that any fragment length may be used to describe a length over which percentage identity may be measured.
[0072] The term “amino acid,” includes but is not limited to amino acids contained in the group consisting of alanine (Ala or A), cysteine (Cys or C), aspartic acid (Asp or D), glutamic acid (Glu or E), phenylalanine (Phe or F), glycine (Gly or G), histidine (His or H), isoleucine (He or I), lysine (Lys or K), leucine (Leu or L), methionine (Met or M), asparagine (Asn or N), proline (Pro or P), glutamine (Gin or Q), arginine (Arg or R), serine (Ser or S), threonine (Thr or T), valine (Vai or V), tryptophan (Trp or W), and tyrosine (Tyr or Y) residues. The term “amino acid residue” also may include amino acid residues contained in the group consisting of homocysteine, 2-Aminoadipic acid, N- Ethylasparagine, 3-Aminoadipic acid, Hydroxylysine, P-alanine. p-Amino-propionic acid, allo- Hydroxylysine acid, 2-Aminobutyric acid, 3-Hydroxyproline, 4-Aminobutyric acid, 4- Hydroxyproline, piperidinic acid, 6-Aminocaproic acid, Isodesmosine, 2-Aminoheptanoic acid, allo- Isoleucine, 2-Aminoisobutyric acid, N-Methylglycine, sarcosine, 3 -Aminoisobutyric acid, N- Methylisoleucine, 2-Aminopimelic acid, 6-N-Methyllysine, 2,4-Diaminobutyric acid, N- Methylvaline, Desmosine, Norvaline, 2,2'-Diaminopimelic acid, Norleucine, 2,3-Diaminopropionic acid, Ornithine, and N-Ethylglycine. Typically, the amide linkages of the peptides are formed from an amino group of the backbone of one amino acid and a carboxyl group of the backbone of another amino acid.
[0073] Hie term “variant” means a polypeptide derived from a parent polypeptide by one or more (several) alteralion(s), i.e.. a substitution, insertion, and / or deletion, at one or more (several) positions. A substitution means a replacement of an amino acid occupying a position with a different amino acid: a deletion means removal of an amino acid occupying a position: and an insertion means adding 1 or more, such as 1, 2, 3, 4, 5, 6, 7 , 8, 9 or 10, preferably 1-3 amino acids immediately adjacent an amino acid occupying a position. In relation io substitutions, ‘immediately adjacent’ may be to the N-side (‘upstream’) or C-side (‘downstream’) of the amino acid occupying a position (‘the named amino acid’). Therefore, for an amino acid named / n umbered ‘X,’ the insertion may be at position ‘X+l’ ('downstream’) or at position ‘X-l ’ (‘upstream’).
[0074] A “variant” of a particular polypeptide sequence may be defined as a polypeptide sequence having at least 50% sequence identity to the particular polypeptide sequence over a certain length of one of the polypeptide sequences using blastp with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information's website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), “Blast 2 sequences — a new tool for comparing protein and nucleotide sequences”, FEMS Microbiol Lett. 174:247-250). In some embodiments a variant polypeptide may show, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length relative to a reference polypeptide.
[0075] A variant polypeptide may have substantially the same functional activity as a reference polypeptide. For example, a variant polypeptide may exhibit or more biological activities associated with binding a ligand and / or binding DNA at a specific binding site.
[0076] Variants comprising deletions relative to a reference amino acid sequence contemplated herein. A “deletion” refers to a change in the amino acid sequence that results in the absence of one or more amino acid residues relative to a reference sequence. A deletion removes at least 1, 2, 3, 4, 5, 10, 20, 50, 100, or 200 amino acids residues. A deletion may include an internal deletion or a terminal deletion (e.g., an N-terminal truncation or a C-terminal truncation or both of a reference polypeptide).
[0077] Variants comprising a fragment of a reference amino acid sequence are contemplated herein. A “fragment” is a portion of an amino acid sequence which is identical in sequence to but shorter in length than the reference sequence. A fragment may comprise up to the entire length of the reference sequence, minus at least one amino acid residue. For example, a fragment may comprise from 5 to 1000 contiguous amino acid residues of a reference polypeptide, respectively. In some embodiments, a fragment may comprise at least 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 150, 250, or 500 contiguous amino acid residues of a reference polypeptide, respectively. Fragments may be preferentially selected from certain regions of a molecule, for example the N- terminal region and / or the C-terminal region of a polypeptide. The term “at least a fragment” encompasses the full length polypeptide.
[0078] The word “vector” or “expression vector” refers to any vehicle that carries a polynucleotide into a cell for the expression of the polynucleotide in the cell. The vector may be, for example, a plasmid, a virus, a phage particle, or a nanoparticle. Once transformed into a suitable host, the vector may replicate and function independently of the host genome, or may in some instances, integrate into the genome itself. In some embodiments, the vector is a DNA construct containing a DNA sequence which is operably linked to a suitable control sequence capable of affecting the expression of the DNA in a suitable host cell. Such control sequences can include a promoter to effect transcription, an optional operator sequence to control such transcription, a sequence encoding suitable mRNA ribosome binding sites, and sequences which control the termination of transcription and translation.
[0079] A “nucleotide” is a compound consisting of a nucleoside, which consists of a nitrogenous base and a 5-carbon sugar, linked to a phosphate group forming the basic structural unit of nucleic acids, such as DNA or RNA. The four types of nucleotides are adenine (A), cytosine (C), guanine (G), and thymine (T), each of which are bound together by a phosphodiester bond to form a nucleic acid molecule.
[0080] A “nucleic acid” is a chemical compound that serves as the primary information-carrying molecules in cells and make up the cellular genetic material. Nucleic acids comprise nucleotides, which are the monomers made of a 5-carbon sugar (usually ribose or deoxyribose), a phosphate group, and a nitrogenous base. A nucleic acid can also be a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA). A chimeric nucleic acid comprises two or more of the same kind of nucleic acid fused together to form one compound comprising genetic material.
[0081] The terms “percent identity” and “% identity,” as applied to polynucleotide sequences, refer to the percentage of residue matches between at least two polynucleotide sequences aligned using a standardized algorithm. Such an algorithm may insert, in a standardized and reproducible way, gaps in the sequences being compared in order to optimize alignment between two sequences and therefore achieve a more meaningful comparison of the two sequences. Percent identity for a nucleic acid sequence may be determined as understood in the art. (See, e.g., U.S. Pat. No. 7,396,664, which is incorporated herein by reference in its entirety). A suite of commonly used and freely available sequence comparison algorithms is provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST) (Altschul, S. F. et al. (1990) J. Mol. Biol. 215:403 410), which is available from several sources, including the NCBI, Bethesda, Md., at its website. The BLAST software suite includes various sequence analysis programs including “blastn,” that is used to align a known polynucleotide sequence with other polynucleotide sequences from a variety of databases. Also available is a tool called “BLAST 2 Sequences” that is used for direct pairwise comparison of two nucleotide sequences. “BLAST 2 Sequences” can be accessed and used interactively at the NCBI website. The “BLAST 2 Sequences” tool can be used for both blastn and blastp (discussed above).
[0082] Percent identity may be measured over the length of an entire defined polynucleotide sequence or may be measured over a shorter length, for example, over the length of a fragment taken from a larger, defined sequence, for instance, a fragment of at least 20, at least 30, at least 40, at least 50, at least 70, at least 100, or at least 200 contiguous nucleotides. Such lengths are exemplary only, and it is understood that any fragment length may be used to describe a length over which percentage identity may be measured.
[0083] A “full length” polynucleotide sequence is one containing at least a translation initiation codon (e.g., methionine) followed by an open reading frame and a translation termination codon. A “full length” polynucleotide sequence encodes a “full length” polypeptide sequence.
[0084] A “variant,” “mutant,” or “derivative” of a particular nucleic acid sequence may be defined as a nucleic acid sequence having at least 50% sequence identity to the particular nucleic acid sequence over a certain length of one of the nucleic acid sequences using blastn with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information's website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), “Blast 2 sequences — a new tool for comparing protein and nucleotide sequences”, FEMS Microbiol Lett. 174:247-250). In some embodiments a variant polynucleotide may show, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length relative to a reference polynucleotide.
[0085] As used herein, a “mutation” refers to changing the structure of a gene, resulting in a variant form that may be transmitted to later generations. A mutation is caused by the alteration of single nucleotides in DNA, or the deletion, insertion, or rearrangement of larger sections of genes. A mutation can lead to the expression of a protein that has been changed physically or functionally leading to lethality, non-lethal dysfunction effects, or no effects.
[0086] SelB Variants and Translation Systems
[0087] The present disclosure provides a selenocysteine-specific elongation factor (SelB) variant. The present disclosure also provides vectors, cells, and systems comprising the selenocysteinespecific elongation factor (SelB) variant. The present disclosure provides methods generating a selenocysteine-specific elongation factor (SelB) variant. The present disclosure also provides methods of using the selenocysteine-specific elongation factor (SelB) variant.
[0088] The present disclosure is advantageous in a sense that the optimal Sec-tRNA formation is be affected by using the variants disclosed herein. Furthermore, the achievement of SECIS-independent SelB formation also renders an orthogonal EF-Tu that can function in parallel to normal EF-Tu, which provides for the ability to engineer an elongation factor that is specific for unnatural amino acid- charged tRNA.
[0089] Disclosed herein are variants of SelB. These can be found in SEQ ID NO: 1 (Variant 1, also referred to herein as SelB-vl), SEQ ID NO: 2 (Variant 2, also referred to herein as SelB-v2), and SEQ ID NO: 6 (Variant 3, which is also referred to herein as SelB v2-Ser or SelvB v3). These variants are all based on the wild type sequence SEQ ID NO: 3. It is noted that these sequences can comprise additional modifications, or can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more combinations of the mutations found in SEQ ID NO: 1, 2, and 6 when compared to wild type SEQ ID NO: 3. These variants are encoded by SEQ ID NOS: 4, 5, and 7, respectively.
[0090] In some aspects, disclosed herein is an engineered selenocysteine-specific elongation factor (SelB) variant comprising at least 70% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6, wherein the SelB variant is capable of incorporating at least one non-canonical amino acid into a peptide. In some embodiments, the SelB variant comprises at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the SelB variant comprises SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6.
[0091] As used herein, a “truncation” refers to the shortening of an amino acid, polypeptide, or protein sequence, typically due to premature termination of translation or other genetic modifications. Truncation of an amino acid, polypeptide, or protein sequence can be caused by various factors, including, but not limited to nonsense mutations, frameshift mutations, and / or deletion mutations.
[0092] In some embodiments, the SelB variant is a truncation of SEQ ID NO: 3. In some embodiments, the truncation comprises a deletion of a carboxy terminal end. In some embodiments, the deletion comprises removal of at least 100 of amino acids from the truncation. In some embodiments the deletion comprises removal of 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129,
[0093] 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149,
[0094] 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169,
[0095] 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189,
[0096] 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, or more amino acids from the carboxy terminal end.
[0097] In some embodiments, the SelB variant comprises at least one amino acid mutation within the truncation. In some embodiments, the at least one amino acid mutation is selected from the group consisting of I98V, E166D, S206P, P277S A413T, S473N, W477R, and F487L. In some embodiments, the SelB variant further comprises an amino acid mutation selected from E246K and E463K within the truncation. In some embodiments, the SelB variant comprises both the E246K and E463K within the truncation. In some embodiments, the SelB variant comprises at least one of Y44F, T193C, and / or R236S within the truncation. When all three are present, this is represented in SEQ ID NO: 6.
[0098] In some aspects, disclosed herein is a peptide produced by any method disclosed herein, the peptide comprising at least one selenocysteine residue and / or at least one serine residue.
[0099] In some embodiments, the peptide further comprises at least one natural amino acid residue including, but not limited to alanine (Ala or A), cysteine (Cys or C), aspartic acid (Asp or D), glutamic acid (Glu or E), phenylalanine (Phe or F), glycine (Gly or G), histidine (His or H), isoleucine (He or I), lysine (Lys or K), leucine (Leu or L), methionine (Met or M), asparagine (Asn or N), proline (Pro or P), glutamine (Gin or Q), arginine (Arg or R), serine (Ser or S), threonine (Thr or T), valine (Vai or V), tryptophan (Trp or W), tyrosine (Tyr or Y), homocysteine, selenocysteine, 2- Aminoadipic acid, N-Ethylasparagine, 3-Aminoadipic acid, Hydroxylysine, P-alanine, P-Amino-propionic acid, allo- Hydroxylysine acid, 2-Aminobutyric acid, 3-Hydroxyproline, 4-Aminobutyric acid, 4- Hydroxyproline, piperidinic acid, 6-Aminocaproic acid, Isodesmosine, 2-Aminoheptanoic acid, allo- Isoleucine, 2-Aminoisobutyric acid, N-Methylglycine, sarcosine, 3 -Aminoisobutyric acid, N- Methylisoleucine, 2-Aminopimelic acid, 6-N-Methyllysine, 2,4-Diaminobutyric acid, N- Methylvaline, Desmosine, Norvaline, 2,2'-Diaminopimelic acid, Norleucine, 2,3-Diaminopropionic acid, Ornithine, and N-Ethylglycine residues. In some embodiments, the peptide further comprises at least one amino acid residue including, but not limited to acetylated, hydroxylated, nitrated, and methylated analogs of alanine, cysteine, aspartic acid, glutamic acid, phenylalanine, glycine, histidine , isoleucine, lysine, leucine, methionine, asparagine, proline, glutamine, arginine, serine, threonine, valine, tryptophan, or tyrosine.
[0100] In some embodiments, the peptide includes, but is not limited to a polypeptide, an oligopeptide, a protein, an antibody, a hormone, a polymer, a coenzyme, a cofactor, an antibiotic, and an anti-microbial peptide. In some embodiments, the peptide of any preceding aspect is generated with a non-canonical amino acid (such as, for example selenocysteine) replacing a canonical amino acid (such as, for example cysteine) to make said peptide more stable.
[0101] In some aspects, disclosed herein is a protein translation system comprising an engineered selenocysteine-specific elongation factor (SelB) variant, as described above. The SelB variant can comprise 70% or more sequence identity to SEQ ID NO: 1 or SEQ ID NO: 2, wherein the SelB variant is capable of incorporating at least one non-canonical amino acid into a peptide, and wherein the SelB variant translates an mRNA sequence without need of a selenocysteine insertion sequence (SECIS). In some embodiments, the SelB variant comprises at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% to SEQ ID NO: 1 or SEQ ID NO: 2. In some embodiments, the system comprises the SelB variant comprising SEQ ID NO: 1 or SEQ ID NO: 2.
[0102] In some embodiments, the system further comprises an Elongation Factor Tu (EF-Tu) protein. EF-Tu is a prokaryotic, G-protein elongation factor responsible for catalyzing the binding of one or more aminoacyl-tRNA (aa-tRNA) to the ribosome. EF-Tu is a crucial factor for generation of proteins containing the standard amino acids of any preceding aspect. EF-Tu is also found as the most abundant and highly conserved proteins in prokaryotes. The eukaryotic analog of EF-Tu is elongation factor Tu, mitochondria (TUFM). It should be understood that the EF-Tu can be interchanged for TUFM in the system of any preceding aspect. In some embodiments, the system further comprises a plurality of nucleotides selected from the group consisting of adenine (A), uracil (U), guanine (G), and cytosine (C). In some embodiments, the system further comprises one or more translation regulatory proteins selected from an initiation factor, an elongation factor a release factor, or a ribosome.
[0103] In some embodiments, the system or SelB variant of any preceding aspect can be combined with or incorporated into an in vitro transcription system. In some embodiments, the system or SelB variant of any preceding aspect can be combined with or incorporated into an in vitro translation system.
[0104] In some aspects, disclosed herein is a cell comprising the engineered SelB variant, the vector, or the system of any preceding aspect.
[0105] Expression Vectors
[0106] In one aspect, disclosed herein is an expression vector comprising a nucleic acid sequence comprising at least 80% sequence identity to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7, wherein the nucleic acid sequence encodes a selenocysteine-specific elongation factor (SelB) variant.
[0107] In some embodiments, the nucleic acid sequence comprises at least 70% sequence identity to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7. In some embodiments, the nucleic acid sequence comprises at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7.
[0108] In some embodiments, the nucleic acid sequence encodes at least one amino acid mutation. In some embodiments, the nucleic acid sequence encodes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid mutations. In some embodiments, the nucleic acid sequence encodes at least one amino acid mutation selected from the group consisting of I98V, E166D, S206P, P277S A413T, S473N, W477R, and F487L. In some embodiments, the nucleic acid sequence further encodes an amino acid mutation selected from E246K and E463K.
[0109] In some embodiments, the expression vector comprises a plasmid or a virus or viral vector. A plasmid or a viral vector can be capable of extrachromosomal replication or, optionally, can integrate into the host genome. As used herein, the term "integrated" used in reference to an expression vector (e.g., a plasmid or viral vector) means the expression vector, or a portion thereof, is incorporated (physically inserted or ligated) into the chromosomal DNA of a host cell. As used herein, a “viral vector” refers to a virus-like particle containing genetic material which can be introduced into a eukaryotic cell without causing substantial pathogenic effects to the eukaryotic cell. A wide range of viruses or viral vectors can be used for transduction but should be compatible with the cell type the virus or viral vector are transduced into (e.g., low toxicity, capability to enter cells). Suitable viruses and viral vectors include adenovirus, lentivirus, retrovirus, among others.
[0110] A) Retroviral Vectors
[0111] A retrovirus is an animal virus belonging to the virus family of Retroviridae, including any types, subfamilies, genus, or tropisms. Retroviral vectors, in general, are described by Verma, I.M., Retroviral vectors for gene transfer.
[0112] A retrovirus is essentially a package which has packed into its nucleic acid cargo. The nucleic acid cargo carries with it a packaging signal, which ensures that the replicated daughter molecules will be efficiently packaged within the package coat. In addition to the package signal, there are a number of molecules which are needed in cis, for the replication, and packaging of the replicated virus. Typically a retroviral genome, contains the gag, pol, and env genes which are involved in the making of the protein coat. It is the gag, pol, and env genes which are typically replaced by the foreign DNA that it is to be transferred to the target cell. Retrovirus vectors typically contain a packaging signal for incorporation into the package coat, a sequence which signals the start of the gag transcription unit, elements necessary for reverse transcription, including a primer binding site to bind the tRNA primer of reverse transcription, terminal repeat sequences that guide the switch of RNA strands during DNA synthesis, a purine rich sequence 5' to the 3' LTR that serve as the priming site for the synthesis of the second strand of DNA synthesis, and specific sequences near the ends of the LTRs that enable the insertion of the DNA state of the retrovirus to insert into the host genome. The removal of the gag, pol, and env genes allows for about 8 kb of foreign sequence to be inserted into the viral genome, become reverse transcribed, and upon replication be packaged into a new retroviral particle. This amount of nucleic acid is sufficient for the delivery of a one to many genes depending on the size of each transcript. It is preferable to include either positive or negative selectable markers along with other genes in the insert. Since the replication machinery and packaging proteins in most retroviral vectors have been removed (gag, pol, and env), the vectors are typically generated by placing them into a packaging cell line. A packaging cell line is a cell line which has been transfected or transformed with a retrovirus that contains the replication and packaging machinery but lacks any packaging signal. When the vector carrying the DNA of choice is transfected into these cell lines, the vector containing the gene of interest is replicated and packaged into new retroviral particles, by the machinery provided in cis by the helper cell. The genomes for the machinery are not packaged because they lack the necessary signals.
[0113] B) Adenoviral Vectors
[0114] The construction of replication-defective adenoviruses has been described (Berkner et al., J. Virology 61: 1213-1220 (1987); Massie et al., Mol. Cell. Biol. 6:2872-2883 (1986); Haj-Ahmad et al., J. Virology 57:267-274 (1986); Davidson et al., J. Virology 61: 1226-1239 (1987); Zhang "Generation and identification of recombinant adenovirus by liposome-mediated transfection and PCR analysis" BioTechniques 15:868-872 (1993)). The benefit of the use of these viruses as vectors is that they are limited in the extent to which they can spread to other cell types, since they can replicate within an initial infected cell, but are unable to form new infectious viral particles. Recombinant adenoviruses have been shown to achieve high efficiency gene transfer after direct, in vivo delivery to airway epithelium, hepatocytes, vascular endothelium, CNS parenchyma and a number of other tissue sites (Morsy, J. Clin. Invest. 92:1580-1586 (1993); Kirshenbaum, J. Clin. Invest. 92:381-387 (1993); Roessler, J. Clin. Invest. 92: 1085-1092 (1993); Moullier, Nature Genetics 4: 154-159 (1993); La Salle, Science 259:988-990 (1993); Gomez-Foix, J. Biol. Chem. 267:25129-25134 (1992); Rich, Human Gene Therapy 4:461-476 (1993); Zabner, Nature Genetics 6:75-83 (1994); Guzman, Circulation Research 73:1201-1207 (1993); Bout, Human Gene Therapy 5:3-10 (1994); Zabner, Cell 75:207-216 (1993); Caillaud, Eur. J. Neuroscience 5:1287-1291 (1993); and Ragot, J. Gen. Virology 74:501-507 (1993)). Recombinant adenoviruses achieve gene transduction by binding to specific cell surface receptors, after which the virus is internalized by receptor-mediated endocytosis, in the same manner as wild type or replication-defective adenovirus (Chardonnet and Dales, Virology 40:462-477 (1970); Brown and Burlingham, J. Virology 12:386- 396 (1973); Svensson and Persson, J. Virology 55:442-449 (1985); Seth, et al., J. Virol. 51:650-655 (1984); Seth, et al., Mol. Cell. Biol. 4:1528-1533 (1984); Varga et al., J. Virology 65:6061-6070 (1991); Wickham et al., Cell 73:309-319 (1993)).
[0115] A viral vector can be one based on an adenovirus which has had the El gene removed and these virions are generated in a cell line such as the human 293 cell line. In another preferred embodiment both the El and E3 genes are removed from the adenovirus genome. C) Adeno -associated viral vectors
[0116] Another type of viral vector is based on an adeno-associated virus (AAV). This defective parvovirus is a preferred vector because it can infect many cell types and is nonpathogenic to humans. AAV type vectors can transport about 4 to 5 kb and wild type AAV is known to stably insert into chromosome 19. Vectors which contain this site-specific integration property are preferred. An especially preferred embodiment of this type of vector is the P4.1 C vector produced by Avigen, San Francisco, CA, which can contain the herpes simplex virus thymidine kinase gene, HSV-tk, and / or a marker gene, such as the gene encoding the green fluorescent protein, GFP.
[0117] In another type of AAV virus, the AAV contains a pair of inverted terminal repeats (ITRs) which flank at least one cassette containing a promoter which directs cell-specific expression operably linked to a heterologous gene. Heterologous in this context refers to any nucleotide sequence or gene which is not native to the AAV or B19 parvovirus.
[0118] Typically, the AAV and B19 coding regions have been deleted, resulting in a safe, noncytotoxic vector. The AAV ITRs, or modifications thereof, confer infectivity and site- specific integration, but not cytotoxicity, and the promoter directs cell-specific expression. United states Patent No. 6,261,834 is herein incorporated by reference for material related to the AAV vector.
[0119] D) Large pay load viral vectors
[0120] Molecular genetic experiments with large human herpesviruses have provided a means whereby large heterologous DNA fragments can be cloned, propagated and established in cells permissive for infection with herpesviruses (Sun et al., Nature genetics 8: 33-41, 1994; Cotter and Robertson,. Curr Opin Mol Ther 5: 633-644, 1999). These large DNA viruses (herpes simplex virus (HSV) and Epstein-Barr virus (EBV), have the potential to deliver fragments of human heterologous DNA > 150 kb to specific cells. EBV recombinants can maintain large pieces of DNA in the infected B-cells as episomal DNA. Individual clones carried human genomic inserts up to 330 kb appeared genetically stable. The maintenance of these episomes requires a specific EBV nuclear protein, EBNA1, constitutively expressed during infection with EBV. Additionally, these vectors can be used for transfection, where large amounts of protein can be generated transiently in vitro. Herpesvirus amplicon systems are also being used to package pieces of DNA > 220 kb and to infect cells that can stably maintain DNA as episomes.
[0121] Other useful systems include, for example, replicating and host-restricted non-replicating vaccinia virus vectors.
[0122] In some embodiments, the expression vector encoding a chimeric polypeptide is a naked DNA or is comprised in a nanoparticle (e.g., liposomal vesicle, porous silicon nanoparticle, gold-DNA conjugate particle, polyethyleneimine polymer particle, cationic peptides, etc.). Methods
[0123] In some aspects, disclosed herein is a method of incorporating at least one non-canonical amino acid into a peptide during translation, the method comprising contacting an mRNA sequence with an engineered selenocysteine-specific elongation factor (SelB) variant, wherein the engineered SelB variant incorporates a non-canonical amino acid into the peptide without need of a selenocysteine insertion sequence (SECIS).
[0124] In some embodiments, method incorporates 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16,
[0125] 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43,
[0126] 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70,
[0127] 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,
[0128] 98, 99, 100, or more non-canonical amino acids into a peptide during translation. In some embodiments, the method uses the engineered SelB factor, the peptide, the translation system, or expression vector of any preceding aspect.
[0129] In some embodiments, translation occurs in presence of one or more translation regulatory proteins. In some embodiments, the one or more translation regulatory proteins comprise an initiation factor, an elongation factor, a release factor, and / or a ribosome. In some embodiments, the elongation factor comprises elongation Factor Tu (EF-Tu).
[0130] In some embodiments, the method comprises the SelB variant comprising at least 70% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the method comprises the SelB variant comprising at 70%, 75%, 80%, 85%, 90%, 95%, or 99%, 100% to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6. In some embodiments, the method comprises the SelB variant comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6.
[0131] In some aspects, disclosed herein is a method of expanding canonical translation mechanisms, the method comprising initiating translation of a messenger RNA (mRNA) into a peptide chain; elongating the peptide chain, wherein an engineered SelB peptide and an EF-Tu factor are simultaneously utilized to incorporate the canonical and non-canonical amino acids of any preceding aspect into the peptide chain; and terminating the peptide chain once it reaches a pre-determined length.
[0132] In some aspects, disclosed herein is a method of generating an engineered peptide comprising a non-canonical amino acid (such as, for example selenocysteine) in replacement of a canonical amino acid (such as, for example cysteine), wherein the method increases the stability of the engineered peptide relative to a control peptide. In some embodiments, the engineered peptide includes, but is not limited to a polypeptide, an oligopeptide, a protein, an antibody, a hormone, a polymer, a coenzyme, a cofactor, an antibiotic, and an anti-microbial peptide.
[0133] It should be understood that a non-canonical amino acid refers to an amino acid that is not part of the standard genetic code or are amino acids that are not encoded by standard mRNA codons. In some embodiments, the non-canonical amino acid includes, but is not limited to selenocysteine, hydroxyproline, hydroxylysine, citrulline, pyrrolysine, azidohomoalanine, azidonorleucine, and azidophenylalanine. In some embodiments, the canonical amino acid includes, but is not limited to alanine (Ala or A), cysteine (Cys or C), aspartic acid (Asp or D), glutamic acid (Glu or E), phenylalanine (Phe or F), glycine (Gly or G), histidine (His or H), isoleucine (He or I), lysine (Lys or K), leucine (Leu or L), methionine (Met or M), asparagine (Asn or N), proline (Pro or P), glutamine (Gin or Q), arginine (Arg or R), serine (Ser or S), threonine (Thr or T), valine (Vai or V), tryptophan (Trp or W), tyrosine (Tyr or Y).
[0134] A number of embodiments of the disclosure have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention. Accordingly, other embodiments are within the scope of the following claims.
[0135] By way of non-limiting illustration, examples of certain embodiments of the present disclosure are given below.
[0136] EXAMPLES
[0137] The following examples are set forth below to illustrate the compositions, devices, methods, and results according to the disclosed subject matter. These examples are not intended to be inclusive of all aspects of the subject-matter disclosed herein, but rather to illustrate representative methods and results. These examples are not intended to exclude equivalents and variations of the present invention which are apparent to one skilled in the art.
[0138] Example 1: Directed Evolution of a SelB Variant that does not Require a SECIS Element for Function
[0139] The introduction of non-canonical amino acids (ncAA) into proteins has enabled researchers to modify physicochemical and functional properties of proteins. However, some type of ncAAs such as beta-amino acid or D-amino acids still struggle for their low incorporation efficiency. This is partially due to the weakened binding affinity of tRNA charged with ncAAs to EF-Tu; an elongation factor that recruits all-natural aminoacyl-tRNA to the ribosome. Researchers have therefore attempted to engineer the amino acid binding pocket of EF-Tu to restore its weak affinity to ncAAs and increase its incorporation efficiency. However, mutating EF-Tu has the ability to alter its recruitment ability of tRNA charged with natural amino acids as well. Thus, a separate elongation factor that functions in parallel to EF-Tu is a promising engineering platform for ncAA incorporation. To achieve this, a homologue of EF-Tu called SelB is the focus herein. SelB is an elongation factor that is specific for selenocysteineyl-tRNASec (Sec-tRNASec). SelB specifically interact with selenocysteineyl-tRNASec and not with other natural tRNAs which makes SelB a perfect candidate of ncAA specific elongation factor. However, in order for SelB to recruit Sec- tRNA to ribosome, it needs a specific RNA sequence, the SECIS (selenocysteine insertion sequence) element, on the mRNA which is an undesired function for engineering purpose. Therefore, the present disclosure provides development of SECIS independent SelB with the use of directed evolution.
[0140] Example 2: Directed evolution of a SelB variant that does not require a SECIS element for function
[0141] In bacteria the incorporation of selenocysteine is achieved through the interaction of the selenocysteine specific elongation factor (SelB) with selenocysteine-charged tRNASecand a selenocysteine insertion sequence (SECIS) element adjacent to an opal stop codon in a mRNA. The more generalized, SECIS-independent incorporation of selenocysteine is of interest because of the high nucleophilicity of selenium and the greater durability of diselenide bonds. It is likely that during the course of evolution selenocysteine insertion originally arose without the presence of a SECIS element, relying only on SelB. Herein, the experiments were undertaken to evolve an ancestral version of SelB that is SECIS-independent and show that not only can this protein (SelB-v2) generally incorporate selenocysteine across from stop codons, but also that the new, orthogonal translation factor can be repurposed to other amino acids, such as serine. Given the delicate energetic balancing act already performed by EF-Tu, this achievement raises the greatly expanded genetic codes that relied in part on SelB-based loading can now be contrived.
[0142] Designing SelB to function without its SECIS
[0143] The crystal structure of full length SelB in complex with a ribosome loaded with a mRNA- bearing SECIS element and Sec-tRNASechas been solved. This structure reveals that binding of SelB to the SECIS essentially tethers Sec-tRNASecto the ribosome, with few direct interactions with the ribosome except between the C-terminal domain of SelB, helix hl6 of 16S ribosomal RNA (rRNA), and protein S4. In contrast, EF-Tu is recruited by interacting directly with the L7 / 12 stalk of the 50S ribosomal subunit. Previously mutating interactions between SelB and the SECIS resulted in a 600- fold reduction of GTPase activity, showing that the SECIS :SelB interaction was functionally important.
[0144] This structure proved invaluable for designing SelB variants that operates independently of the SECIS. Given that the N-terminal region (Domains I to III) of SelB is a homologue of EF-Tu, while the C-terminal domain (Domain IV) consists of four winged-helix (WH) motif, wherein the last two WHs directly interact with the SECIS element, it was first contemplated that SECIS independence is obtained through the simple expedient of deleting WH3 and WH4. This was achieved by removing WH3 / WH4 (S488 to K614). This join cleanly deletes the C-terminal 381 residues and leaves a SelB variant that is only 487 residues long: this truncation was used as the starting point for further studies (Figure 1).
[0145] Assaying truncated SelB for selenocysteine incorporation
[0146] Previously, the machinery for selenocysteine incorporation in C321dA, an ‘Amber-less’ E. coli strain, was manipulated. The SelABC gene cluster was further deleted from the genome of this strain, and the necessary components for selenocysteine incorporation were instead re-introduced on a plasmid (Figure 2A). In particular, SelA, SelC, were constitutively expressed, while SelB was placed under the control of an araBAD promoter.
[0147] In order to assay and select for selenocysteine incorporation, a key cysteine residue in a disulfide-dependent beta-lactamase, NMC-A, was previously replaced with selenocysteine (Figure 2B). Formation of a chimeric cysteine-selenocysteine bond promotes growth in the presence of carbenicillin. Surprisingly, the truncated SelB variant actually showed improved Sec incorporation compared to the wild-type enzyme, at least in the NMC-A assay (Figure 2C). These results indicate a different role for the SECIS than is typically reported wherein the SECIS does not enable incorporation but rather limits mis-incorporation.
[0148] Directed evolution of the SECIS-independent SelB variant
[0149] The demonstrated dependence of selenocysteine incorporation for NMC-A beta-lactamase activity enables the directed evolution of Sec incorporation machinery (Figure 3A). Since it wasn’t clear exactly what portions of truncated SelB needed to be optimized, a library of SelB variants was generated by error-prone PCR with a mutation rate of up to 12 substitutions per gene. The transformation efficiency in the first and subsequent rounds of selection always exceeded 10A6. The first round of selection was conducted at 5 pg / mL carbenicillin, followed by three rounds of selection at 10 pg / mL carbenicillin (Table 1). Error prone PCR was conducted after Round 2, with a mutation frequency of an additional 10 mutations per gene.
[0150] After four rounds of selection, six individual colonies were picked (at 100 pg / mL carbenicillin), and sequencing revealed a predominant (5 / 6) SelB variant that shared a total of twelve substitutions, including four silent mutations (Table 2). At the same time, the plasmid backbone had duplicated the NMC-A gene, and that a recombination event had yielded a SelA-SelB fusion protein. These additional modifications to the plasmid should both have led to the ability to survive the carbenicillin challenge: additional copies of the NMC-A gene would yield greater beta-lactamase activity, while the fusion led to constitutive production of SelB, as opposed to the inducible araBAD promoter (Figure 8 and SEQ ID NO: 28).
[0151] The enriched SelB variant was re-cloned to the original plasmid configuration, yielding SelB- vl. This plasmid was then tested for beta-lactamase activity, sans gene duplications and fusions. Sec incorporation was indeed enhanced at 100 pg / mL carbenicillin compared to truncated SelB (Figure 3B).
[0152] SelB-vl was subjected to additional rounds of directed evolution. The library for selection was prepared by error-prone PCR of SelB-vl and had a mutation frequency of up to 4 mutations per SelB gene. In order to increase the stringency for selenocysteine incorporation, two TAG codons were introduced into the NMC-A gene, which reduced NMC-A expression (as observed by carbenicillin resistance; Figures 9A and 9B). Six additional rounds of selection were carried out at 10 pg / mL carbenicillin, and a final round at 100 pg / mL (Table 1). Error prone PCR was also conducted after Rounds 3 and 5.
[0153] Some 10 colonies were picked for sequencing: all contained E463K, and 9 out of 10 contained E246K. The addition of these two predominant mutations to SelB-vl yields SelB-v2. SelB-v2 showed enhanced incorporation efficiency compared to SelB-vl via the NMC-A assay with increasing carbenicillin concentrations (Figure 3B).
[0154] Validating selenocysteine incorporation with SelB-v2
[0155] In order to further validate selenocysteine incorporation by SelB-v2, there were attempts to introduce a fluorescent protein that depends on the incorporation of Sec for its fluorescence (smURFP29) into the C321dA strain. The smURFP protein covalently binds the chromophore biliverdin (BV), whose production requires the expression of the pbsAl heme oxidase for chromophore maturation. However, transformation and expression of this pathway proved difficult, partially due to the growth defects associated with C321dA. Therefore, it was switched to the more commonly used E. coli cloning and expression strain DH10B; this switch should also provide additional evidence of the generality of the utility of SelB- v2 for general selenocysteine suppression. The genomic SelABC gene in DH10B was removed to avoid competition with the native selenocysteine insertion machinery.
[0156] The fluorescence of the smURFP reporter was rendered dependent upon selenocysteine incorporation machinery by converting the cysteine codon at amino acid position 62 to TAG). For this selenocysteine-dependent smURFP construct in the modified DH10B strain, fluorescence was only observed when (i) SelB-v2 (not wild-type SelB) was induced, and (ii) selenium was added to the media (Figure 4A). In addition, for a sfGFP reporter bearing a TAG codon at position 150, fluorescence could only be observed when SelB-v2 was induced and selenium was added to the media (Figure 4B).
[0157] Sec incorporation in the modified DH10B strain was further confirmed via LC-MS. The sfGFP protein expressed in the presence of selenium and SelB-v2 via a pendant strep-tag, digested with trypsin, and the digested peptides were analyzed for Sec incorporation. Interestingly, while a mass shift was noted, it corresponded to dehydroalanine, rather than selenocysteine (Figure 10). It is well-known that Sec can be converted to dehydroalanine through beta-elimination, and thus these results are consistent with selenocysteine incorporation. In order to more directly monitor Sec incorporation, a E. coli dihydrofolate reductase (DHFR) gene engineered to contain a selenyl- sulfhydryl bond was introduced into the modified DH10B strain. Top-down mass spectrometry showed Sec incorporation only upon SelB-v2 induction and selenium addition, and no detectable Ser incorporation (Figure 5A). EThcD fragmentation, which uses both electron transfer dissociation (ETD) and higher-energy collisional dissociation (HCD) for fragmentation, confirmed the formation of a selenide- sulfur bond upon incorporation of Sec at position 39 (Figure 5B). In contrast, when identical experiments were performed in C321dA strain, 84 % of selenocysteine incorporation and 16 % of Ser incorporation was observed (Figure 11), again indicating that this strain may in some cases be problematic for monitoring protein expression and suppression. Previous work has also hinted at this: in the context of the C321dA ‘Amber-less’ strain, translation of either glutathione peroxidases(GPx) or mammalian TrxRl lacking a SECIS element yielded Sec incorporation at the TAG codon but also misincorporation of glutamine and lysine. It is also important to note that when the DH10B strain and SelB-v2 were used for DHFR expression, the yield was 5 mg / L. This is comparable to extant SECIS independent selenoprotein synthesis methods, such as using an Amberless strain and canonical SelB mediated Sec incorporation (1-3 mg / L) or using allo-tRNA for EF-Tu mediated Sec incorporation (up to 10 mg / L). Having used engineered reporters (NMC-A) to show that SelBv-2 mediated selenocysteine incorporation, the present disclosure attempted to determine whether SelB-v2 could also incorporate selenocysteine in a natural context, via formate dehydrogenase (fdHF), which contains SECIS element in its coding sequence. This native selenoprotein oxidizes formate into carbon dioxide, which is important for maintaining cellular redox balance. FdhF activity is strictly dependent on selenocysteine incorporation. The genomic fdHF was removed from C321 dSelABC and instead expressed the formate dehydrogenase gene from the same plasmid as SelB-v2. fdHF still relied on its endogenous promoter, while SelB-v2 was expressed from a tac promoter and induced via IPTG. Plate-based assays with benzyl viologen assay showed functional expression of formate dehydrogenase with SelB-v2, although overall enzymatic activity was slightly lower than with the SelB wild-type (Figure 5C), indicating that further evolutionary optimization is still possible.
[0158] Directed evolution of the SelB-v2 into Ser-tRNASecspecific elongation factor
[0159] Since SelB-v2 showed specific recruitment of Sec-tRNASecto the ribosome in the absence of a SECIS element, it was explored whether it could be used as an orthogonal equivalent of EF-Tu in an even more general manner. Given that the biosynthetic pathway for Sec-tRNASccinvolves the formation of Ser-tRNASecas an early intermediate, it was contemplated whether serine loading might be handled by SelB-v2.
[0160] To determine whether specific serine loading could be accomplished, the catalytic serine (position 71) in the beta-lactamase gene was changed to the TAG stop codon, all in the context of the C321dA strain. NMC-A S71C mutant was confirmed to not show growth in all the carbenicillin condition tested indicating NMC-A S71TAG to be a Ser-dependent reporter (Figures 12A and 12B). Upon induction of SelB-v2, the NMC-A S71TAG gene yielded increase in cell growth at low- to midconcentrations of carbenicillin tested, indicating the serine incorporation was likely taking place (Figure 6A). In order to change or broaden the specificity of SelBv-2 for other amino acids, the residues within the amino acid binding pocket (within 4 Angstroms of the amino acid: Y42, Y44, R181, T193, R236) were randomized, and the SelB-v2 library was plated on LB-agar containing 50 pg / mL Carb. Plate-based assays were used because the TAG codon can mutate to Ser via just one mutation (as opposed to two for Cys), and thus we anticipated (and initially found) that liquid culturing could easily enrich such single mutants. Roughly 150 colonies were observed on the 50 pg / mL Carb plate and 96 colonies were picked and sequenced.
[0161] Sequencing of over 96 colonies picked from the plate revealed convergence of mutations with the top mutant constituting (SelB-v2-Ser; Y42Y, Y44F, R181R, T193C, R236S) representing a quarter of the clones. SelB-v2-Ser showed improved incorporation of Ser compared to the original SelB-v2 (Figure 6B), as active as the wild-type enzyme.
[0162] In order to see if the mutated binding pocket of SelB-v2-Ser also affects the incorporation of Sec, the original NMC-A C70TAG gene, which was used to engineer SelB-v2, was used to assess the Sec incorporation efficiency. It should be noted that the C321 strain retains SelD in the genome, which means that Ser-tRNASeccan be converted to Sec-tRNASecby SelA using selenophosphate as the substrate when selenium is present. While SelB-v2 previously allowed growth at higher concentrations of carbenicillin (25-100 pg / mL), this growth advantage was lost in SelB-v2-Ser (Figure 6C). Ser incorporation was further measured in DHFR P61TAG in the presence of selenium and induction of SelB-v2-Ser. Intact mass analysis revealed that despite selenocysteine being fully available for incorporation, when the C321 strain was used for expression roughly 74% of the protein contained serine, as opposed to only 17% with the parental SelB-v2 (Figure 6D) indicating that tRNA selectivity can be modulated by mutating the amino acid binding pocket. The change in aminoacyl- tRNA selectivity was also confirmed via SDS-PAGE gel analysis (Figure 13).
[0163] The traditional view of the role of SECIS element is to enhance selenocysteine incorporation by recruiting SelB to the ribosome, enabling tRNASccloading. This view has been validated by multiple results in which mutation of the SECIS element decreased selenoprotein translation. The notion that the SECIS enhances selenoprotein production is also seemingly inherent in its mechanism, wherein tRNASecis first recognized by SelA and then transferred to SelB:GTP complex, which in turn is recognized by the SECIS element, leading to correct delivery of Sec-tRNASecto ribosome.
[0164] However, a consideration of the possible origins of selenocysteine incorporation tell a slightly different story. If selenoprotein production proved important for the evolution of metabolism, the simplest innovation would have been to have EF-Tu adopt a selenocysteine-charged tRNA; the introduction of the complex system that currently exists for selenocysteine incorporation would have emerged much later. The ostensible reason for its emergence have been to not enhance selenocysteine incorporation, but rather to restrict it, limiting it to only certain stop codons. This explanation has also been embraced in the literature.
[0165] It is believed that Sec decoding machinery arose before the division to the three domains of life and has been lost in some lineages such as plants and some fungi. Over time, SelB could have emerged by duplication and divergence of EF-Tu, or possibly other translation factors such as IF-2. This would have been the key step that allowed further elaboration of the remaining selenocysteine incorporation machinery and would have occurred prior to the Last Universal Common Ancestor (LUCA). What is less clear is why this would have been necessary. Other suppressor systems arise with regularity, and do not greatly diminish organismal fitness. Previous work from our lab on a different system for selenocysteine incorporation, one that involved altering Sec-tRNA to be used by EF-Tu for loading, showed that large scale selenocysteine incorporation, even in an Amber-less background, was highly toxic to cells. Thus, there have been significant selective pressure following the initial, EF-Tu mediated incorporation of selenocysteine into proteins to generate a system that would limit selenocysteine introduction to only specified codons.
[0166] Overall, these previous results show that the role of SelB is not to promote efficiency, but rather to greatly limit suppression which was hinted by previous research as well. In this model, coupling SelB to the SECIS is not functionally required, but rather is an augmented function, following the original divergence of SelB. The largely structurally separate RNA-binding domain on SelB is unrelated to other translation factors, and would have most likely been added following divergence, further supporting this hypothesis; the more ancestral version of SelB is more akin to our own SelB-v2.
[0167] If this idea is valid, then not only is the RNA-binding domain of SelB a late addition, but it should also be functionally separable, without compromising SelB’s tRNA loading activity. This idea was the basis for the experiments described herein, where it was in fact demonstrated that a SelB variant that was completely separable from the SECIS could readily be generated by only a few rounds of directed evolution. Interestingly, when the SECIS-binding domain is left in place, the S206P and P277S mutations (singly or in tandem) are predicted to lead to a complete reorientation of the SECIS-binding domain relative to the rest of the protein (Figure 14), perturbing function. This further shows the likely ancestral independence of the ancestral-like SelB-v2.
[0168] This idea also impacts further alteration of the genetic code. While the SECIS may have once been required to stabilize an expanded code, it obviously now limits expansion, and the more ancestral-like SelB-v2 provides an intriguing opportunity for genetic code expansion. EF-Tu is of course known to balance the binding of various tRNA-amino acid conjugates, in order to help ensure the overall accurate loading across from a wide variety of codons. This ‘balancing act’ is largely mediated through a complex series of interactions between tRNAs and the surface of SelB and likely has its limits. The introduction of a new, orthogonal loading factor, SelB-v2, can open the way to expansions of the genetic code much larger than just 21 amino acids, as it takes over the loading of an entirely new expansion set via its own balancing act. The fact that the present disclosure is able to not only introduce selenocysteine, but also serine via SelB-v2 directed evolution provides an existence proof for further expansion of its role in loading non-canonical amino acids in further expanded codes. Splitting the duties between EF-Tu and SelB would require multi-step engineering wherein a set of tRNAs would initially be engineered to exclusively be recognized by SelB and not EF-Tu, and then new AARS:tRNA orthogonal pairs could be generated, with the charged tRNAs ultimately being loaded by SelB and not EF-Tu.
[0169] Material and methods
[0170] Strain Construction
[0171] Strains used in this study were kindly gifted from Prof. Ross Thyer and were used in explorations of selenocysteine incorporation. One exception was the strain used in the benzyl viologen assay, wherein the fdHF gene was further deleted from the genome of C321 lacking SelAB and SelC using the lamda Red system from Datsenko and Wanner and replaced with a chloramphenicol resistance marker.
[0172] Library Construction and Selection
[0173] The overall design of the plasmid (pSelAC-araBAD-SelB) for selecting SECIS-independent SelB is shown in SEQ ID NO: 9. Libraries were created by using error-prone PCR to amplify the coding region of SelB (GeneMorph II Random Mutagenesis Kit, Agilent, #200550). The library (pSelAC-araBAD-SelB) was transformed into an electrocompetent C321 dSelABC strain which was then grown for 1 hour at 37°C in 2 mL of SOC, and later transferred to 200 mL of LB medium in the presence of 10 pM Na2SeO3, 0.01 % (v / v) arabinose, 50 pg / mL kanamycin and various concentrations of carbenicillin (Table 1) for 16 hours at 37°C. Cells were recovered, plasmids are isolated, and the coding region of SelB was amplified by PCR (either using Platinum SuperFi II DNA polymerase (Thermo Fisher, #12361010) or the GeneMorph II Random Mutagenesis Kit). The PCR products (Oligol-4) were introduced into the same plasmid via Gibson assembly. The same procedure was used to derive SelB-v2, except cloning was carried out via Golden Gate assembly (Oligo5-8).
[0174] The overall design of the plasmid (pSelC-tat-SelB) for identifying serine-incorporating variants of SelB-v2 is shown in Figure 7 and SEQ ID NO: 27. The plasmid pSelC-tat-SelB was initially prepared by amplifying four fragments using four sets of primers that had NNS primers targeting 5 positions (Y42, Y44, R181, T193, R236) (Oligo 9-11, 12-13, 14-15, 6-16; Table 3). These four fragments were combined using assembly PCR and then cloned into the vector (pSelC-tat-SelB) using Golden Gate assembly (Oligo 6,8,10,17; Table 3). The library was transformed into electrocompetent C321 dSelABC, incubated for 1 hour at 37°C in 2.4 ml of SOC and then plated onto four 15 cm LB-plates containing 50 pM IPTG, 50 pg / mL carbenicillin and 50 pg / mL kanamycin. Roughly 150 colonies were picked after 16 hours at 37°C and then individually checked for sequencing. NMC-A C70TAG assay
[0175] Plasmids (pSelAC-araBAD-SelB) encoding SelB variants were freshly transformed into C321dSelABC and grown on LB-agar plates containing 50 pg / mL kanamycin. Three colonies were picked from a plate and grown overnight in LB media containing 50 pg / mL kanamycin at 37 °C. The overnight culture was diluted 1:100 in fresh LB supplemented with 50 pg / mL kanamycin, 10 pM Na2SeO3 and 0.01 % (v / v) arabinose, and 0, 5, 10, 25, 50, 100 pg / mL carbenicillin. Cultures were incubated for 16 h at 37 °C with shaking. OD600 values were measured using a TEC AN plate reader.
[0176] NMC-A S71TAG assay
[0177] Some 10 ng of plasmids (pSelC-tac-SelB) containing corresponding SelB variants were transformed into electrocompetent C321dSelABC in triplicate and grown for 2 hours. Cultures were then diluted 1:100 in fresh LB supplemented with 50 pg / mL kanamycin, 50 pM 1PTG and 0, 5, 10, 25, 50, 100 pg / mL carbenicillin, and incubated for 16 h at 37 °C with shaking. OD600 values were measured using a TECAN plate reader. sfGFP and smURFP assay
[0178] The sfGPF reporter plasmid was prepared by swapping the NMC-A gene in pSelAC-araBAD- SelB with sfGFP, and the promoter was changed to the Ipp5 promoter. The smURFP reporter plasmid was prepared by swapping the NMC-A gene in pSelAC-araBAD-SelB with the smURFP-HO-1 gene (retaining the NMC-A promoter). Plasmids encoding the WT-SelB or SelB-v2 were freshly transformed into DH10B dSelABC and grown on LB-agar plates containing 50 pg / mL kanamycin. Three colonies were picked from a plate and grown overnight in LB media containing 50 pg / mL kanamycin at 37 °C. Overnight cultures were diluted 1: 100 in fresh LB supplemented with 50 pg / mL kanamycin, 10 pM Na2SeO3 and 0.01 % (v / v) arabinose, and incubated for 16 h at 37 °C with shaking. OD600 as well as fluorescence (ex 485 nm Zem 528 nm for sfGFP, and ex 633 nm / em 695 nm for smURFP) was measured using a TECAN plate reader.
[0179] Protein purification
[0180] To express and purify protein for MS analyses, the NMC-A gene in pSelAC-araBAD-SelB was replaced with the folA gene containing the mutation P61S or P61TAG followed by a strep-tag II (WSHPQFEK). Plasmids encoding SelB-v2 or SelB-v2-Ser were freshly transformed into C321 dSelABC or DH10B dSelABC and grown on LB-agar plates containing 50 pg / mL kanamycin. A single colony was picked from a plate and grown overnight in LB media containing 50 pg / mL kanamycin at 37 °C. Overnight cultures were diluted 1: 100 in fresh LB (50 ml) supplemented with 50 g / mL kanamycin, 1 pM Na2SeO3 and 0.002 % (v / v) arabinose for the DH10B dSelABC strain, and 10 pM Na2SeO3 and 0.01 % (v / v) arabinose for the C321 dSelABC strain, and then grown overnight. Cells were harvested by centrifugation at 10,000 x g for 10 min and resuspended in 3 mL of BugBuster (Novagen #70584-3). Following 20 min incubation on a rotating mixer at 20 rpm at room temperature, cells were clarified by centrifugation at 16,000 x g for 20 min at 4 C. Lysates were passed through a 0.2 pm filter and DHFR or sfGFP samples were purified using MagStrept Strep- Tactin XT beads following the manufacturer's instructions (IBA Lifesciences #2-4090-002). After purification, the buffer was exchanged to 0.1 % formic acid in water by acetone precipitation.
[0181] Benzyl Viologen Assay
[0182] To carry out the benzyl viologen assay, a reporter plasmid was prepared by replacing the promoter and NMC-A gene with the endogenous promoter of the fdhF gene and a fdhF gene with a TGA140TAG mutation. E. coli C321 lacking the SelAB, SelC, and fdhF loci was transformed with the reporter plasmid that contained either wild-type SelB, a SelB truncation, or SelB-v2. Transformants were grown overnight at 37 °C in LB medium containing 50 pg / mL kanamycin. Overnight cultures were diluted 1:20 and incubated for three hours in 1.5 mL of LB medium and cultures were then normalized to OD600 = 0.5 by the addition of LB medium. A number of 5 pl aliquots were dotted on LB agar plates containing 50 pg / mL kanamycin, 5 mM sodium formate, 1 pM Na2MoO4, 1 pM Na2SeO3. Plates were incubated at 37 °C for 3 h under aerobic conditions and then grown for 60 h at 37 °C under anaerobic conditions using the pouch system (BD B260683). Upon removal from the pouch, plates were immediately overlaid with 0.75 % agar containing 1 mg / mL benzyl viologen, 250 mM sodium formate and 25 mM KH2PO4 at pH 7.0. Plates were then photographed within 1 h of overlaying.
[0183] AlphaFold3 Structure Prediction
[0184] AlphaFold3 (https: / / alphafoldserver.com / ) was used to generate the following structures for non-truncated SelB: wild type, S206P, P277S, and [S206P, P277S]. The following seeds were used for each of these structures: 865694845, 601147261, 1701205798, and 1572407154, respectively, and the top ranked structures “_model_0” were used.
[0185] It will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the scope or spirit of the invention. Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the methods disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.
[0186] TABLES
[0187] Table 1. Selection conditions of each round. Table la shows that to obtain SelB-vl, 4 rounds of selection were carried out. Table lb shows that to obtain SelB-v2, 7 rounds of selection were carried out, in liquid culture. N / A refers to not applicable.
[0188] Table 2. List of mutations found in clones picked after round 5. A total of six clones were picked and compared with wild-type SelB.
[0189]
[0190] Table 3. Oligonucleotide sequences.
[0191] *4 refers to any nucleotides; S refers to either cytosine (c) or guanine (G). SEQUENCES
[0192] 1. SEQ ID NO: 1 - SelB Variant 1
[0193] MIIATAGHVDHGKTTLLQAITGVNADRLPEEKKRGMTIDLGYAYWPQPDGRVPGFIDVPG HEKFLSNMLAGVGGIDHALLVVACDDGVMAQTREHLAVLQLTGNPMLTVALTKADRVD
[0194] EARVDEVERQVKEVLREYGFAEAKLFITAATEGRGMDALREHLLQLPDREHASQHSFRLAI DRAFTVKGAGLVVTGTALSGEVKVGDPLWLTGVNKPMRVRALHAQNQPTETANAGQRIA LNIAGDAEKEQINRGDWLLADVPPEPFTRVIVELQTHTSLTQWQPLHIHHAASHVTGRVSL LEDNLAELVFDTPLWLADNDRL VERDIS ARNTLAGARVVMLNPPRRGKRKPEYLQWLASL
[0195] ARAQSDADALSVHLERGAVNLADFAWARQLNGEGMRELLQQPGYIQAGYSLLNTPVAAR WQRKILDTLATYHEQHRDEPGPGRERLRRMALPMEDEALVLLLIEKMRESGDIHNHHGRL HLPDHKAGL
[0196] 2. SEQ ID NO: 2 - SelB Variant 2
[0197] MIIATAGHVDHGKTTLLQAITGVNADRLPEEKKRGMTIDLGYAYWPQPDGRVPGFIDVPG HEKFLSNMLAGVGGIDHALLVVACDDGVMAQTREHLAVLQLTGNPMLTVALTKADRVD
[0198] EARVDEVERQVKEVLREYGFAEAKLFITAATEGRGMDALREHLLQLPDREHASQHSFRLAI DRAFTVKGAGLVVTGTALSGEVKVGDPLWLTGVNKPMRVRALHAQNQPTETANAGQRIA LNIAGDAKKEQINRGDWLLADVPPEPFTRVIVELQTHTSLTQWQPLHIHHAASHVTGRVSL
[0199] LEDNLAELVFDTPLWLADNDRLVLRDISARNTLAGARVVMLNPPRRGKRKPEYLQWLASL ARAQSDADALSVHLERGAVNLADFAWARQLNGEGMRELLQQPGYIQAGYSLLNTPVAAR WQRKILDTLATYHEQHRDEPGPGRERLRRMALPMEDEALVLLLIKKMRESGDIHNHHGRL HLPDHKAGL
[0200] 3. SEQ ID NO: 3 - Full-length SelB
[0201] MIIATAGHVDHGKTTLLQAITGVNADRLPEEKKRGMTIDLGYAYWPQPDGRVPGFIDVPG
[0202] HEKFLSNMLAGVGGIDHALLVVACDDGVMAQTREHLAILQLTGNPMLTVALTKADRVDE ARVDEVERQVKEVLREYGFAEAKLFITAATEGRGMDALREHLLQLPEREHASQHSFRLAID
[0203] RAFTVKGAGLVVTGTALSGEVKVGDSLWLTGVNKPMRVRALHAQNQPTETANAGQRIAL NIAGDAEKEQINRGDWLLADVPPEPFTRVIVELQTHTPLTQWQPLHIHHAASHVTGRVSLL
[0204] EDNLAELVFDTPLWLADNDRL VERDIS ARNTLAGARVVMLNPPRRGKRKPEYLQWLASL ARAQSDADALSVHLERGAVNLADFAWARQLNGEGMRELLQQPGYIQAGYSLLNAPVAAR WQRKILDTLATYHEQHRDEPGPGRERLRRMALPMEDEALVLLLIEKMRESGDIHSHHGWL HLPDHKAGFSEEQQAIWQKAEPLFGDEPWWVRDLAKETGTDEQAMRLTLRQAAQQGIITA
[0205] 31 IVKDRYYRNDRIVEFANMIRDLDQECGSTCAADFRDRLGVGRKLAIQILEYFDRIGFTRRRG
[0206] NDHLLRDALLFPEK
[0207] 4. SEQ ID NO: 4 - SelB Variant 1 Nucleic Acid
[0208] ATGATTATCGCGACTGCCGGACATGTGGATCATGGAAAGACAACATTGTTGCAGGCGA
[0209] TTACTGGCGTAAATGCTGACCGTCTGCCGGAAGAAAAAAAGCGCGGCATGACCATAGA
[0210] TCTCGGCTATGCCTACTGGCCGCAGCCGGATGGTCGCGTGCCTGGTTTTATCGACGTTC
[0211] CCGGTCATGAAAAGTTTCTTTCCAACATGCTGGCGGGCGTTGGTGGTATCGATCACGCG
[0212] CTGTTGGTGGTGGCATGCGATGACGGCGTGATGGCACAGACCCGTGAGCATCTGGCGG
[0213] TTTTGCAGCTGACCGGTAACCCGATGCTGACAGTGGCGCTGACCAAAGCCGATCGCGT
[0214] GGACGAAGCGCGTGTTGATGAGGTTGAACGCCAGGTAAAGGAGGTTCTGCGGGAATAC
[0215] GGTTTTGCTGAGGCAAAACTGTTTATCACCGCAGCAACCGAAGGTCGGGGAATGGATG
[0216] CCCTGCGCGAGCATCTGCTTCAGTTGCCGGATCGCGAGCACGCCAGCCAACATAGTTTC
[0217] CGCCTCGCGATTGACCGCGCATTTACCGTAAAAGGTGCCGGCCTGGTCGTCACCGGTAC
[0218] GGCGTTAAGCGGGGAAGTGAAGGTAGGCGATCCACTCTGGCTGACTGGTGTAAATAAA
[0219] CCGATGCGTGTACGTGCGCTGCATGCGCAAAACCAGCCAACAGAAACCGCCAATGCCG
[0220] GGCAGCGTATCGCGCTTAACATCGCGGGTGATGCGGAAAAAGAGCAGATTAACCGTGG
[0221] CGACTGGCTGCTTGCCGATGTGCCGCCAGAGCCGTTCACACGGGTGATTGTCGAGCTTC
[0222] AAACCCATACATCGCTGACCCAGTGGCAGCCGCTGCATATTCACCACGCCGCCAGCCA
[0223] CGTCACGGGACGCGTTTCACTGCTGGAAGATAACCTTGCTGAACTGGTCTTCGACACCC
[0224] CGTTATGGCTGGCAGATAACGACCGCCTGGTATTGCGCGATATCTCTGCCCGCAACACG
[0225] CTGGCCGGAGCGCGCGTCGTGATGCTTAACCCGCCGCGTCGCGGTAAACGTAAGCCGG
[0226] AATATCTGCAATGGCTGGCGTCACTTGCACGGGCACAGAGCGATGCCGATGCGTTATC
[0227] TGTTCATCTGGAACGCGGCGCGGTTAACCTTGCGGATTTCGCCTGGGCGCGCCAGCTCA
[0228] ACGGCGAAGGGATGCGCGAATTGCTGCAACAGCCTGGTTATATTCAGGCTGGTTATAG
[0229] CTTGTTGAATACGCCGGTTGCCGCCCGCTGGCAGCGGAAAATTCTCGACACATTAGCG
[0230] ACTTATCATGAGCAACATCGCGATGAACCTGGCCCTGGGCGCGAACGTCTGCGACGTA
[0231] TGGCGTTGCCAATGGAAGATGAAGCGCTGGTACTGTTGCTGATTGAAAAGATGCGCGA
[0232] AAGCGGCGACATCCACAACCATCACGGCCGGCTGCATCTGCCAGATCACAAAGCGGGC CTC
[0233] 5. SEQ ID NO: 5 - SelB Variant 2 Nucleic Acid
[0234] ATGATTATCGCGACTGCCGGACATGTGGATCATGGAAAGACAACATTGTTGCAGGCGA
[0235] TTACTGGCGTAAATGCTGACCGTCTGCCGGAAGAAAAAAAGCGCGGCATGACCATAGA TCTCGGCTATGCCTACTGGCCGCAGCCGGATGGTCGCGTGCCTGGTTTTATCGACGTTC
[0236] CCGGTCATGAAAAGTTTCTTTCCAACATGCTGGCGGGCGTTGGTGGTATCGATCACGCG
[0237] CTGTTGGTGGTGGCATGCGATGACGGCGTGATGGCACAGACCCGTGAGCATCTGGCGG
[0238] TTTTGCAGCTGACCGGTAACCCGATGCTGACAGTGGCGCTGACCAAAGCCGATCGCGT
[0239] GGACGAAGCGCGTGTTGATGAGGTTGAACGCCAGGTAAAGGAGGTTCTGCGGGAATAC
[0240] GGTTTTGCTGAGGCAAAACTGTTTATCACCGCAGCAACCGAAGGTCGGGGAATGGATG
[0241] CCCTGCGCGAGCATCTGCTTCAGTTGCCGGATCGCGAGCACGCCAGCCAACATAGTTTC
[0242] CGCCTCGCGATTGACCGCGCATTTACCGTAAAAGGTGCCGGCCTGGTCGTCACCGGTAC
[0243] GGCGTTAAGCGGGGAAGTGAAGGTAGGCGATCCACTCTGGCTGACTGGTGTAAATAAA
[0244] CCGATGCGTGTACGTGCGCTGCATGCGCAAAACCAGCCAACAGAAACCGCCAATGCCG
[0245] GGCAGCGTATCGCGCTTAACATCGCGGGTGATGCGAAAAAAGAGCAGATTAACCGTGG
[0246] CGACTGGCTGCTTGCCGATGTGCCGCCAGAGCCGTTCACACGGGTGATTGTCGAGCTTC
[0247] AAACCCATACATCGCTGACCCAGTGGCAGCCGCTGCATATTCACCACGCCGCCAGCCA
[0248] CGTCACGGGACGCGTTTCACTGCTGGAAGATAACCTTGCTGAACTGGTCTTCGACACCC
[0249] CGTTATGGCTGGCAGATAACGACCGCCTGGTATTGCGCGATATCTCTGCCCGCAACACG
[0250] CTGGCCGGAGCGCGCGTCGTGATGCTTAACCCGCCGCGTCGCGGTAAACGTAAGCCGG
[0251] AATATCTGCAATGGCTGGCGTCACTTGCACGGGCACAGAGCGATGCCGATGCGTTATC
[0252] TGTTCATCTGGAACGCGGCGCGGTTAACCTTGCGGATTTCGCCTGGGCGCGCCAGCTCA
[0253] ACGGCGAAGGGATGCGCGAATTGCTGCAACAGCCTGGTTATATTCAGGCTGGTTATAG
[0254] CTTGTTGAATACGCCGGTTGCCGCCCGCTGGCAGCGGAAAATTCTCGACACATTAGCG
[0255] ACTTATCATGAGCAACATCGCGATGAACCTGGCCCTGGGCGCGAACGTCTGCGACGTA
[0256] TGGCGTTGCCAATGGAAGATGAAGCGCTGGTACTGTTGCTGATTAAAAAGATGCGCGA
[0257] AAGCGGCGACATCCACAACCATCACGGCCGGCTGCATCTGCCAGATCACAAAGCGGGC CTC
[0258] 6. SEQ ID NO: 6- SelB Variant 2-selenium (SelB-v2 Ser) (also referred to herein as SelB Variant
[0259] 3)
[0260] MIIATAGHVDHGKTTLLQAITGVNADRLPEEKKRGMTIDLGYAFWPQPDGRVPGFIDVPGH
[0261] EKFLSNMLAGVGGIDHALLVVACDDGVMAQTREHLAVLQLTGNPMLTVALTKADRVDE
[0262] ARVDEVERQVKEVLREYGFAEAKLFITAATEGRGMDALREHLLQLPDREHASQHSFRLAID
[0263] RAFTVKGAGLVVCGTALSGEVKVGDPLWLTGVNKPMRVRALHAQNQPTETANAGQSIAL
[0264] NIAGDAKKEQINRGDWLLADVPPEPFTRVIVELQTHTSLTQWQPLHIHHAASHVTGRVSLL
[0265] EDNLAELVFDTPLWLADNDRLVLRDISARNTLAGARVVMLNPPRRGKRKPEYLQWLASL
[0266] ARAQSDADALSVHLERGAVNLADFAWARQLNGEGMRELLQQPGYIQAGYSLLNTPVAAR WQRKILDTLATYHEQHRDEPGPGRERLRRMALPMEDEALVLLLIKKMRESGDIHNHHGRL
[0267] HLPDHKAG
[0268] 7. SEQ ID NO: 7: SelB Variant 2- selenium (aka SelB Variant 3) Nucleic Acid
[0269] ATGATCATCGCCACGGCGGGTCATGTGGATCATGGAAAGACAACATTGTTGCAGGCGA
[0270] TTACTGGCGTAAATGCTGACCGTCTGCCGGAAGAAAAAAAGCGCGGCATGACCATAGA
[0271] TCTCGGCTACGCCTTCTGGCCGCAGCCGGATGGTCGCGTGCCTGGTTTTATCGACGTTC
[0272] CCGGTCATGAAAAGTTTCTTTCCAACATGCTGGCGGGCGTTGGTGGTATCGATCACGCG
[0273] CTGTTGGTGGTGGCATGCGATGACGGCGTGATGGCACAGACCCGTGAGCATCTGGCGG
[0274] TTTTGCAGCTGACCGGTAACCCGATGCTGACAGTGGCGCTGACCAAAGCCGATCGCGT
[0275] GGACGAAGCGCGTGTTGATGAGGTTGAACGCCAGGTAAAGGAGGTTCTGCGGGAATAC
[0276] GGTTTTGCTGAGGCAAAACTGTTTATCACCGCAGCAACCGAAGGTCGGGGAATGGATG
[0277] CCCTGCGCGAGCATCTGCTTCAGTTGCCGGATCGCGAGCACGCCAGCCAACATAGTTTC
[0278] CGCCTCGCGATTGACCGCGCATTTACCGTAAAAGGTGCCGGCCTGGTCGTCTGCGGTAC
[0279] GGCGTTAAGCGGGGAAGTGAAGGTAGGCGATCCACTCTGGCTGACTGGTGTAAATAAA
[0280] CCGATGCGTGTACGTGCGCTGCATGCGCAAAACCAGCCAACAGAAACCGCCAATGCCG
[0281] GGCAGAGCATCGCGCTTAACATCGCGGGTGATGCGAAAAAAGAGCAGATTAACCGTG
[0282] GCGACTGGCTGCTTGCCGATGTGCCGCCAGAGCCGTTCACACGGGTGATTGTCGAGCTT
[0283] CAAACCCATACATCGCTGACCCAGTGGCAGCCGCTGCATATTCACCACGCCGCCAGCC
[0284] ACGTCACGGGACGCGTTTCACTGCTGGAAGATAACCTTGCTGAACTGGTCTTCGACACC
[0285] CCGTTATGGCTGGCAGATAACGACCGCCTGGTATTGCGCGATATCTCTGCCCGCAACAC
[0286] GCTGGCCGGAGCGCGCGTCGTGATGCTTAACCCGCCGCGTCGCGGTAAACGTAAGCCG
[0287] GAATATCTGCAATGGCTGGCGTCACTTGCACGGGCACAGAGCGATGCCGATGCGTTAT
[0288] CTGTTCATCTGGAACGCGGCGCGGTTAACCTTGCGGATTTCGCCTGGGCGCGCCAGCTC
[0289] AACGGCGAAGGGATGCGCGAATTGCTGCAACAGCCTGGTTATATTCAGGCTGGTTATA
[0290] GCTTGTTGAATACGCCGGTTGCCGCCCGCTGGCAGCGGAAAATTCTCGACACATTAGC
[0291] GACTTATCATGAGCAACATCGCGATGAACCTGGCCCTGGGCGCGAACGTCTGCGACGT
[0292] ATGGCGTTGCCAATGGAAGATGAAGCGCTGGTACTGTTGCTGATTAAAAAGATGCGCG
[0293] AAAGCGGCGACATCCACAACCATCACGGCCGGCTGCATCTGCCAGATCACAAAGCGGG CCTC
[0294] 8. SEQ ID NO: 8: trypsin-digested peptide from sfGFP
[0295] LEYNFNSHXVYITADK 9. SEQ ID NO: 9 - Plasmid of pSelAC-araB AD-SelB
[0296] TCGACCGATGCCCTTGAGACGTTGATCGGCACGTATGGCAGATAGTAAATTTTATAGAT
[0297] TTAAATCAACATCTATAACTGTATAATTCGTTTTCTCAACTCATTACAACACTCCGTTAG
[0298] TAATGAAGCTCATTCTATACAATGACAGTTAATAGGTAAAGTTATGTCACTTAATGTAA
[0299] AGCAAAGTAGAATAGCCATCTTGTTTAGCTCTTGTTTAATTTCAATATCATTTTTCTCAC
[0300] AGGCCAATACGAAGGGCATTGATGAGATTAAAAACCTTGAAACAGATTTCAATGGCAG
[0301] GATTGGTGTCTACGCTTTAGACACTGGCTCGGGTAAATCATTTTCGTACAGAGCAAATG
[0302] AACGATTTCCATTATAGAGTTCTTTTAAAGGTTTTTTAGCTGCTGCTGTATTAAAAGGCT
[0303] CTCAAGATAATCGACTTAATCTTAATCAGATTGTGAATTATAATACAAGAAGTTTAGAG
[0304] TTCCATTCACCCATCACAACTAAATATAAAGATAATGGAATGTCATTAGGTGATATGGC
[0305] TGCTGCTGCTTTACAATATAGCGACAATGGTGCTACTAATATTATTCTTGAACGTTATA
[0306] TCGGTGGTCCAGAGGGTATGACTAAATTCATGCGGTCGATTGGAGATGAAGATTTTAG
[0307] ACTCGATCGTTGGGAGTTAGATCTAAACACAGCTATTCCAGGCGATGAGCGTGACACA
[0308] TCTACACCTGCAGCAGTAGCCAAGAGTCTGAAAACCCTTGCTCTGGGTAACATACTTAG
[0309] TGAACATGAAAAGGAAACCTATCAGACATGGTTAAAGGGTAACACAACCGGTGCAGC
[0310] GCGTATTCGTGCTAGCGTACCAAGCGATTGGGTAGTTGGCGATAAAACTGGTAGTTAG
[0311] GGAGCATACGGTACGGCAAATGATTATGCGGTAGTCTGGCCAAAGAACCGGGCTCCTC
[0312] TTATAATTTCTGTATACACAACAAAAAACGAAAAAGAAGCCAAGCATGAGGATAAAGT
[0313] AATCGCAGAAGCTTCAAGAATTGCAATTGATAACCTTAAATAACAAAGCCCGCCGAAA
[0314] GGCGGGCTTTTTTTTGGATCCTGTAGAAACGCAAAAAGGCCATCCGCATGACTAACAT
[0315] GAGAATTACAACTATCACCGCAAGGGATAAATATCTAACACCGTTATGACAACTTGAC
[0316] GGCTACATCATTCACTTTTTCTTCACAACCGGCACGGAACTCGCTCGGGCTGGCCCCGG
[0317] TGCATTTTTTAAATACCCGCGAGAAGTAGAGTTGATCGTCAAAACCAACATTGCGACC
[0318] GACGGTGGCGATAGGCATCCGGGTGGTGCTCAAAAGCAGCTTCGCCTGGCTGATACGT
[0319] TGGTCCTCGCGCCAGCTTAAGACGCTAATCCCTAACTGCTGGCGGAAAAGATGTGACA
[0320] GACGCGACGGCGACAAGCAAACATGCTGTGCGACGCTGGCGATATCAAAATTGCTGTC
[0321] TGCCAGGTGATCGCTGATGTACTGACAAGCCTCGCGTACCCGATTATCCATCGGTGGAT
[0322] GGAGCGACTCGTTAATCGCTTCCATGTGCCGCAGTAACAATTGCTCAAGCAGATTTATC
[0323] GCCAGCAGCTCCGAATAGCGCCCTTCCCCTTGCCCGGCGTTAATGATTTGCCCAAACAG
[0324] GTCGCTGAAATGCGGCTGGTGCGCTTCATCCGGGCGAAAGAACCCCGTATTGGCAAAT
[0325] ATTGACGGCCAGTTAAGCCATTCATGCCAGTAGGCGCGCGGACGAAAGTAAACCCACT
[0326] GGTGATACCATTCGCGAGCCTCCGGATGACGACCGTAGTGATGAATCTCTCCTGGCGG
[0327] GAACAGCAAAATATCACCCGGTCGGCAAACAAATTCTCGTCCCTGATTTTTCACCACCC
[0328] CCTGACCGCGAATGGTGAGATTGAGAATATAACCTTTCATTCCCAGCGGTCGGTCGATA AAAAAATCGAGATAACCGTTGGCCTCAATCGGCGTTAAACCCGCCACCAGATGGGCAT
[0329] TAAACGAGTATCCCGGCAGCAGGGGATCATTTTGCGCTTCAGCCATACTTTTCATACTC
[0330] CCGCCATTCAGAGAAGAAACCAATTGTCCATATTGCATCAGACATTGCCGTCACTGCGT
[0331] CTTTTACTGGCTCTTCTCGCTAACCAAACCGGTAACCCCGCTTATTAAAAGCATTCTGT
[0332] AACAAAGCGGGACCAAAGCCATGACAAAAACGCGTAACAAAAGTGTCTATAATCACG
[0333] GCAGAAAAGTCCACATTGATTATTTGCACGGCGTCACACTTTGCTATGCCATAGCATTT
[0334] TTATCCATAAGATTAGCGGATCCTACCTGACGCTTTTTATCGCAACTCTCTACTGTTTCT
[0335] CCATACCCGTTAGGAGGATTGTATGATTATCGCGACTGCCGGACATGTGGATCATGGA
[0336] AAGACAACATTGTTGCAGGCGATTACTGGCGTAAATGCTGACCGTCTGCCGGAAGAAA
[0337] AAAAGCGCGGCATGACCATCGATCTCGGCTATGCCTACTGGCCGCAGCCGGATGGTCG
[0338] CGTGCCTGGTTTTATCGACGTTCCCGGTCATGAAAAGTTTCTTTCCAACATGCTGGCGG
[0339] GCGTTGGTGGTATCGATCACGCGCTGTTGGTGGTGGCGTGCGATGACGGCGTGATGGC
[0340] ACAGACCCGTGAGCATCTGGCGATTTTGCAGCTGACCGGTAACCCGATGCTGACAGTG
[0341] GCGCTGACCAAAGCCGATCGCGTGGACGAAGCGCGTGTTGATGAGGTTGAACGCCAGG
[0342] TAAAGGAGGTTCTGCGGGAATACGGTTTTGCTGAGGCAAAACTGTTTATCACCGCAGC
[0343] AACCGAAGGTCGGGGAATGGATGCCCTGCGCGAGCATCTGCTTCAGTTGCCGGAACGC
[0344] GAGCACGCCAGCCAACATAGTTTCCGCCTCGCGATTGACCGCGCATTTACCGTAAAAG
[0345] GTGCCGGGCTGGTCGTCACCGGTACGGCGTTAAGCGGGGAAGTGAAGGTAGGCGATTC
[0346] ACTCTGGCTGACTGGTGTAAATAAACCGATGCGTGTACGTGCGCTGCATGCGCAAAAC
[0347] CAGCCAACAGAAACCGCCAATGCCGGGCAGCGTATCGCGCTTAACATCGCGGGTGATG
[0348] CGGAAAAAGAGCAGATTAACCGTGGCGACTGGCTGCTTGCCGATGTGCCGCCAGAGCC
[0349] GTTCACACGGGTGATTGTCGAGCTTCAAACCCATACACCGCTGACCCAGTGGCAGCCG
[0350] CTGCATATTCACCACGCCGCCAGCCACGTCACGGGACGCGTTTCACTGCTGGAAGATA
[0351] ACCTTGCTGAACTGGTCTTCGACACCCCGTTATGGCTGGCAGATAACGACCGCCTGGTA
[0352] TTGCGCGATATCTCTGCCCGCAACACGCTGGCCGGAGCGCGCGTCGTGATGCTTAACCC
[0353] GCCGCGTCGCGGTAAACGTAAGCCGGAATATCTGCAATGGCTGGCGTCACTTGCACGG
[0354] GCGCAGAGCGATGCCGATGCGTTATCTGTTCATCTGGAACGCGGCGCGGTTAACCTTGC
[0355] GGATTTCGCCTGGGCGCGCCAGCTCAACGGCGAAGGGATGCGCGAATTGCTGCAACAG
[0356] CCTGGTTATATTCAGGCTGGTTATAGCTTGTTGAATGCGCCGGTTGCCGCCCGCTGGCA
[0357] GCGGAAAATTCTCGACACATTAGCGACTTATCATGAGCAACATCGCGATGAACCTGGC
[0358] CCTGGGCGCGAACGTCTGCGACGTATGGCGTTGCCAATGGAAGATGAAGCGCTGGTAC
[0359] TGTTGCTGATTGAAAAGATGCGCGAAAGCGGCGACATCCACAGCCATCACGGCTGGCT
[0360] GCATCTGCCAGATCACAAAGCGGGCTTCTAACATTGTCTTATAAATACCAAAAATTTCA
[0361] CGTCTCTTGTTTAAAGAGGCGTGCTTTAATATTTTCTTAGAGTCCAATACATGAGAGTTT TAATTTCGTTTTTTTTCCTTATTTTTTCTATTAGCGTTAGAGCTGACTGTGGCCTTGATAG
[0362] ATCTTCGAAGCTTGGGCCCGAACAAAAACTCATCTCAGAAGAGGAGAGATAAATTGCA
[0363] CTGAAATCTAGAGTAACGGAATAGCTGTTCGTTGACTTGATAGACCGATTGATTCATCA
[0364] TCTCATAAATAAAGAAAAACCACCGCTACCAACGGTGGTTTTCTCAAGGTTCGCTGAG
[0365] CTACCAACTCTTTGAACCAAGGTAAGTGGGTTGGAGGACCGCACTCACCAAAATCTGT
[0366] TCTTTCAGTTTAGCCTTAACAGGTGCATAACTTCAAGACAAAGTCCTCTAAATCAGTTA
[0367] CCAATGGCTGCTGCCAGTGGCGATAAGTCGTGTCTTACCGGGTTGGACTCAAGACGAT
[0368] AGTTACCGGATAAGGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAG
[0369] CTTGGAGCGAACGACCTACACCGAACTGAGATACCAACAGCGTGAGCTATGAGAAAGC
[0370] GCCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCGGA
[0371] ACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTG
[0372] TCGGGTTTCGCCACCTCTGGCTTGAGCGTCGATTTTTGTGATGCTCGTCAGGGGGGCGG
[0373] AGCCTATGGAAAAACGCCTGCGGCGTTGGCTTCTTCCGGTGCTTTGCTTTTTGCTCACA
[0374] TGTTCTTTCCGGCTTTATCCCCTGATTCTGTGGATAACCGTATTACCGCTTTTGAGTGAG
[0375] CTGACACCGCTCGCCGCAGTCGAACGACCGAGCGTAGCGAGTCAGTGAGCGAGGAAG
[0376] CGGAAGAGCGCTGCATGCCTATTTGTTTATTTTTCTAAATACATTCAAATATGTATCCG
[0377] CTCATGAGACAATAACCCTGATAAATGCTTCAATAATATTGAAAAAGGAAGAGTATGA
[0378] GCCATATTCAACGGGAAACGTCTTGCTCTAGGCCGCGATTAAATTCCAACATGGATGCT
[0379] GATTTATATGGGTATAAATGGGCTCGCGATAATGTCGGGCAATCAGGTGCGACAATCT
[0380] ATCGATTGTATGGGAAGCCCGATGCGCCAGAGTTGTTTCTGAAACATGGCAAAGGTAG
[0381] CGTTGCCAATGATGTTACAGATGAGATGGTCAGACTAAACTGGCTGACGGAATTTATG
[0382] CCTCTTCCGACCATCAAGCATTTTATCCGTACTCCTGATGATGCATGGTTACTCACCACT
[0383] GCGATCCCCGGGAAAACAGCATTCCAGGTATTAGAAGAATATCCTGATTCAGGTGAAA
[0384] ATATTGTTGATGCGCTGGCAGTGTTCCTGCGCCGGTTGCATTCGATTCCTGTTTGTAATT
[0385] GTCCTTTTAACAGCGACCGCGTATTTCGTCTCGCTCAGGCGCAATCACGAATGAATAAC
[0386] GGTTTGGTTGATGCGAGTGATTTTGATGACGAGCGTAATGGCTGGCCTGTTGAACAAGT
[0387] CTGGAAAGAAATGCATAAACTTTTGCCATTCTCACCGGATTCAGTCGTCACTCATGGTG
[0388] ATTTCTCACTTGATAACCTTATTTTTGACGAGGGGAAATTAATAGGTTGTATTGATGTT
[0389] GGACGAGTCGGAATCGCAGACCGATACCAGGATCTTGCCATCCTATGGAACTGCCTCG
[0390] GTGAGTTTTCTCCTTCATTACAGAAACGGCTTTTTCAAAAATATGGTATTGATAATCCT
[0391] GATATGAATAAATTGCAGTTTCATTTGATGCTCGATGAGTTTTTCTAACTGCAATAAGG
[0392] TTGTTTTGCCGTGGTCAACGTGTCCGGCAGTCGCAATAATCATTTCAACAACATCTCCA
[0393] AAAACCGTTGCTCATCTTCAAGGCAGCGTAAATCCAGCCACAATCGTCCGTCATAAAT
[0394] ACGACCAATCACCGGCACTGGCAATTCACGCCAGCGGGCGGCTAATGACTCAAGGTGG CTACCGCGTCCATCATGGGGTGTAAACGTTAATGCCGCGCTCGGCAGGCGATCAACCG
[0395] GCAGCGAACCACTGCCAATCTGCGAAAGACATGGCATAACCTGTACCGCAAACTCCGC
[0396] GCCGTAATGTGCGGCAAGGGGGGCCTGTAAACGTTGTGCCTGGATTTGAATGACCTCT
[0397] GCGCTGCGGGTAAGCAGGCGCAGGGTCGGTAATTTTTCACTCAGAGCTTCAGGGTGTA
[0398] AATAAAGACGCAACGTGGCTTCCAGCGCCGCGAGGGTCATTTTATCCGCGCGTAATGC
[0399] ACGCTTCAGCGGGTGGCTTTGCAGGCGGGCGATCATCTCTTTTTTACCAACAATAATTC
[0400] CTGCCTGCGGCCCGCCTAACAACTTGTCGCCGGAGAAACTCACCAGACTGACGCCCGC
[0401] CGCAATCAACTCCTGCGGCATTGGCTCTTTCGGCAAACCGTACTGGCTAAGATCGACCA
[0402] GCGAGCCACTGCCTAAATCAGTCACTACGGGAACATCCAGCTCTTTGCCGAGCGCCAC
[0403] CAGTTCCGCTTCATCTATCGCTTTGGTGAACCCCTGAATGCTGTAGTTACTGGTATGTAC
[0404] TTTCATCAACAGTGCGGTATTTTCATTCACCGCCTGACGATAATCATTCGCGTGCGTGC
[0405] GGTTGGTGGTCCCTACTTCGTGTAGGGTGCAGCCTGCCTGACGCATAACATCGGGAATA
[0406] CGAAACGCGCCGCCAATCTCCACCAGTTCGCCGCGAGATACCACCACCTCTTTTCCGCT
[0407] GGCAGTGGCCGCCAACATCAATAACACCGCCGCCGCATTGTTATTGACGATACAGGCA
[0408] TCTTCCGCCCCCGTAATACGGCACAGCAGCTGCGCCAGCGCCCGATCGCGATGTCCGC
[0409] GTCCGGCGTCGTCCAGATCATACTCGAGGGTCACTGGCGAACGCATAGCCTGCGCAAC
[0410] GGCTTCCACCGCGGCTTCCGCCTGTAAAGCTCGCCCAAGGTTGGTATGCAGCACGGTTC
[0411] CCGTCAGGTTGATCACCGGACGCAGCGCGCTCTGCGCTTCTTTCGTCAACCGGGCATCG
[0412] ACTTCTTGCGCCCAGTTTTCACACCACGCAGGCAGCGTCTGGCTGCCACGAATCACTTC
[0413] TCGCGCTTCGTCGAGCATCTGACGCAACAATTCCACCACGCGGGTGTGACCATAAGTAT
[0414] CACGCAAAGAAAGGAAGGAGCTATCGCGCAATAAGCGATCAATAGCCGGAAGTTGAC
[0415] TATAGAGGAAACGCGTTTCGGTTGTCATGGTTTAGTTCCTCACCTTGTCGTATTATACTA
[0416] TGCCGATATACTATGCCGATGATTAGGAAACCTGGCTGATCAAGGCCCTCTCACACGG
[0417] AGAAGGGCGTTTAACATAACCACGGATTGTAACGTGAGATGGGTCAGGAGGACATATC
[0418] GCGCATCAAGCCTTTGGCGGTTCGGTGCATCCCATGGCGATTGAATATTGAGCAGTAG
[0419] AAAATATCTGGATTGACAGTATTACAAGGGGTTGGAGAACGTCAGGAGAGGCATTTTG
[0420] GCGGAAGATCACAGGAGTCGAACCTGCCCGGGACCGCTGGCGGCCCCAACTGGATTTA
[0421] GAGTCCAGCCGCCTCACCGGAGACGACGATCTTCCGCGCCTCGATTGCTACATGGAGG
[0422] CGGGGCGCATTATAGCTACTTCCTTGAGTTTCTACATCCCCCAGATCAGTTTTTTCCGCC
[0423] TAATACACGCAGCGTATTATCACGAAAGCATTATAGTTCATATAAAACCCATATTTTAC
[0424] AATGATATATGAGCGGAATTAACGCTACCACGATCACATAAAACAAGAATCATAAATT
[0425] AATAACCAGATATCGGAATATTCGCTCTCACAGGGATGGCGGCCGC 10. SEQ ID NO: 10 - Oligo 1
[0426] CATACAATCCTCCTAACGGGT
[0427] 11. SEQ ID NO: 11 - Oligo 2
[0428] TAACATTGTCTTATAAATACCAAAAATTTCAC
[0429] 12. SEQ ID NO: 12 - Oligo 3
[0430] ACCCGTTAGGAGGATTGTATG
[0431] 13. SEQ ID NO: 13 - Oligo 4
[0432] GTGAAATTTTTGGTATTTATAAGACAATGTTA
[0433] 14. SEQ ID NO: 14 - Oligo 5
[0434] GCATGGTCTCTTACCCGTTAGGAGGATTGTATG
[0435] 15. SEQ ID NO: 15 - Oligo 6
[0436] GCATGGTCTCTGTGAAATTTTTGGTATTTATAAGACAATGTTA
[0437] 16. SEQ ID NO: 16 - Oligo 7
[0438] GCATGGTCTCTGGTATGGAGAAACAGTAGAGAG
[0439] 17. SEQ ID NO: 17 - Oligo 8
[0440] GCATGGTCTCTTCACGTCTCTTGTTTAAAGAGGC
[0441] 18. SEQ ID NO: 18 - Oligo 9
[0442] GCATGGTCTCTCTAGAGAAAGAGGGGAAATACTAGATGATCATCGCCACGGCGGGTCA
[0443] TGTGGATCATGGAAAGACAACA
[0444] 19. SEQ ID NO: 19 - Oligo 10
[0445] GCATGGTCTCTCTAGTATTAAACAAAATTATTTGTAGAGG
[0446] 20. SEQ ID NO: 20 - Oligo 11
[0447] GCCGAGATCTATGGTCATGC 21. SEQ ID NO: 21 - Oligo 12
[0448] GCATGACCATAGATCTCGGCNNSGCCNNSTGGCCGCAGCCGGATGG
[0449] N refers to any nucleotides; S refers to either cytosine (c) or guanine (G).
[0450] 22. SEQ ID NO: 22 - Oligo 13
[0451] GGCCGGCACCTTTTACGGTAAATGCSNNGTCAATCGCGAGGCGGAAA
[0452] N refers to any nucleotides; S refers to either cytosine (c) or guanine (G).
[0453] 23. SEQ ID NO: 23 - Oligo 14
[0454] CCGTAAAAGGTGCCGGCCTGGTCGTCNNSGGTACGGCGTTAAGCGGG
[0455] N refers to any nucleotides; S refers to either cytosine (c) or guanine (G).
[0456] 24. SEQ ID NO: 24 - Oligo 15
[0457] CTGCCCGGCATTGGCGGT
[0458] 25. SEQ ID NO: 25 - Oligo 16
[0459] AAACCGCCAATGCCGGGCAGNNSATCGCGCTTAACATCGCGG
[0460] N refers to any nucleotides; S refers to either cytosine (c) or guanine (G).
[0461] 26. SEQ ID NO: 26 - Oligo 17
[0462] GCATGGTCTCTCTAGAGAAAGAGGGGAAATACTAG
[0463] 27. SEQ ID NO: 27 - Plasmid of pSelC-tac-SelB
[0464] TCGACCGATGCCCTTGAGACGTTGATCGGCACGTATGGCAGATAGTAAATTTTATAGAT
[0465] TTAAATCAACATCTATAACTGTATAATTCGTTTTCTCAACTCATTACAACACTCCGTTAG
[0466] TAATGAAGCTCATTCTATACAATGACAGTTAATAGGTAAAGTTATGTCACTTAATGTAA
[0467] AGCAAAGTAGAATAGCCATCTTGTTTAGCTCTTGTTTAATTTCAATATCATTTTTCTCAC
[0468] AGGCCAATACGAAGGGCATTGATGAGATTAAAAACCTTGAAACAGATTTCAATGGCAG
[0469] GATTGGTGTCTACGCTTTAGACACTGGCTCGGGTAAATCATTTTCGTACAGAGCAAATG
[0470] AACGATTTCCATTATGTTAGTCTTTTAAAGGTTTTTTAGCTGCTGCTGTATTAAAAGGCT
[0471] CTCAAGATAATCGACTTAATCTTAATCAGATTGTGAATTATAATACAAGAAGTTTAGAG
[0472] TTCCATTCACCCATCACAACTAAATATAAAGATAATGGAATGTCATTAGGTGATATGGC
[0473] TGCTGCTGCTTTACAATATAGCGACAATGGTGCTACTAATATTATTCTTGAACGTTATA
[0474] TCGGTGGTCCAGAGGGTATGACTAAATTCATGCGGTCGATTGGAGATGAAGATTTTAG ACTCGATCGTTGGGAGTTAGATCTAAACACAGCTATTCCAGGCGATGAGCGTGACACA
[0475] TCTACACCTGCAGCAGTAGCCAAGAGTCTGAAAACCCTTGCTCTGGGTAACATACTTAG
[0476] TGAACATGAAAAGGAAACCTATCAGACATGGTTAAAGGGTAACACAACCGGTGCAGC
[0477] GCGTATTCGTGCTAGCGTACCAAGCGATTGGGTAGTTGGCGATAAAACTGGTAGTTGC
[0478] GGAGCATACGGTACGGCAAATGATTATGCGGTAGTCTGGCCAAAGAACCGGGCTCCTC
[0479] TTATAATTTCTGTATACACAACAAAAAACGAAAAAGAAGCCAAGCATGAGGATAAAGT
[0480] AATCGCAGAAGCTTCAAGAATTGCAATTGATAACCTTAAATAACAAAGCCCGCCGAAA
[0481] GGCGGGCTTTTTTTTGGATCCTGTAGAAACGCAAAAAGGCCATCCGCATGACTAACAT
[0482] GAGAATTACAACTGCGGCGCGCCATCGAATGGCGCAAAACCTTTCGCGGTATGGCATG
[0483] ATAGCGCCCGGAAGAGAGTCAATTCAGGGTGGTGAATATGAAACCAGTAACGTTATAC
[0484] GATGTCGCAGAGTATGCCGGTGTCTCTTATATGACCGTTTCCCGCGTGGTGAACCAGGC
[0485] CAGCCACGTTTCTGCGAAAACGCGGGAAAAAGTGGAAGCGGCGATGGTGGAGCTGAA
[0486] TTACATTCCCAACCGCGTGGCACAACAACTGGCGGGCAAACAGTCGTTGCTGATTGGC
[0487] GTTGCCACCTCCAGTCTGGCCCTGCACGCGCCGTCGCAAATTGTCGCGGCGATTAAATC
[0488] TCGCGCCGATCAACTGGGTGCCAGCGTGGTGGTGTCGATGGTAGAACGAAGCGGCGTC
[0489] GAAGCCTGTAAAGCGGCGGTGCACAATCTTCTCGCGCAACGCGTCAGTGGGCTGATCA
[0490] TTAACTATCCGCTGGATGACCAGGATGCCATTGCTGTGGAAGCTGCCTGCACTAATGTT
[0491] CCGGCGTTATTTCTTGATGTCTCTGACCAGACACCCATCAACAGTATTATTTACTCCCAT
[0492] GAGGACGGTACGCGACTGGGCGTGGAGCATCTGGTCGCATTGGGTCACCAGCAAATCG
[0493] CGCTGTTAGCGGGCCCATTAAGTTCTGTCTCGGCGCGTCTGCGTCTGGCTGGCTGGCAT
[0494] AAATATCTCACTCGCAATCAAATTCAGCCGATAGCGGAACGGGAAGGCGACTGGAGTG
[0495] CCATGTCCGGTTTTCAACAAACCATGCAAATGCTGAATGAGGGCATCGTTCCCACTGCG
[0496] ATGCTGGTTGCCAACGATCAGATGGCGCTGGGCGCAATGCGCGCCATTACCGAGTCCG
[0497] GGCTGCGCGTTGGTGCGGATATCTCGGTAGTGGGATACGACGATACCGAAGATAGCTC
[0498] ATGTTATATCCCGCCGTTAACCACCATCAAACAGGATTTTCGCCTGCTGGGGCAAACCA
[0499] GCGTGGACCGCTTGCTGCAACTCTCTCAGGGCCAGGCGGTGAAGGGCAATCAGCTGTT
[0500] GCCAGTCTCACTGGTGAAAAGAAAAACCACCCTGGCGCCCAATACGCAAACCGCCTCT
[0501] CCCCGCGCGTTGGCCGATTCATTAATGCAGCTGGCACGACAGGTTTCCCGACTGGAAA
[0502] GCGGGCAGTGATAAGGATCCTAATTGGTAACGAATCAGACAATTGACGGCTCGAGGGA
[0503] GTAGCATAGGGTTTGCAGAATCCCTGCTTCGTCCATTTGACAGGCACATTATGCATCGA
[0504] TGATAAGCTGTCAAACATGAGCAGATCCTCTACGCCGGACGCATCGTGGCCGGCATCA
[0505] CCGGCGCCACAGGTGCGGTTGCTGGCGCCTATATCGCCGACATCACCGATGGGGAAGA
[0506] TCGGGCTCGCCACTTCGGGCTCATGAGCAAATATTTTATCTGGCTTAACGATCGTTGGC
[0507] TGTGTTGACAATTAATCATCGGCTCGTATAATGTGTGGAATTGTGAGCGCTCACAATTA GCTGTCACCGGATGTGCTTTCCGGTCTGATGAGTCCGTGAGGACGAAACAGCCTCTACA
[0508] AATAATTTTGTTTAATACTAGAGAAAGAGGGGAAATACTAGATGATTATCGCGACTGC
[0509] CGGACATGTGGATCATGGAAAGACAACATTGTTGCAGGCGATTACTGGCGTAAATGCT
[0510] GACCGTCTGCCGGAAGAAAAAAAGCGCGGCATGACCATAGATCTCGGCTATGCCTACT
[0511] GGCCGCAGCCGGATGGTCGCGTGCCTGGTTTTATCGACGTTCCCGGTCATGAAAAGTTT
[0512] CTTTCCAACATGCTGGCGGGCGTTGGTGGTATCGATCACGCGCTGTTGGTGGTGGCATG
[0513] CGATGACGGCGTGATGGCACAGACCCGTGAGCATCTGGCGGTTTTGCAGCTGACCGGT
[0514] AACCCGATGCTGACAGTGGCGCTGACCAAAGCCGATCGCGTGGACGAAGCGCGTGTTG
[0515] ATGAGGTTGAACGCCAGGTAAAGGAGGTTCTGCGGGAATACGGTTTTGCTGAGGCAAA
[0516] ACTGTTTATCACCGCAGCAACCGAAGGTCGGGGAATGGATGCCCTGCGCGAGCATCTG
[0517] CTTCAGTTGCCGGATCGCGAGCACGCCAGCCAACATAGTTTCCGCCTCGCGATTGACCG
[0518] CGCATTTACCGTAAAAGGTGCCGGCCTGGTCGTCACCGGTACGGCGTTAAGCGGGGAA
[0519] GTGAAGGTAGGCGATCCACTCTGGCTGACTGGTGTAAATAAACCGATGCGTGTACGTG
[0520] CGCTGCATGCGCAAAACCAGCCAACAGAAACCGCCAATGCCGGGCAGCGTATCGCGCT
[0521] TAACATCGCGGGTGATGCGAAAAAAGAGCAGATTAACCGTGGCGACTGGCTGCTTGCC
[0522] GATGTGCCGCCAGAGCCGTTCACACGGGTGATTGTCGAGCTTCAAACCCATACATCGCT
[0523] GACCCAGTGGCAGCCGCTGCATATTCACCACGCCGCCAGCCACGTCACGGGACGCGTT
[0524] TCACTGCTGGAAGATAACCTTGCTGAACTGGTCTTCGACACCCCGTTATGGCTGGCAGA
[0525] TAACGACCGCCTGGTATTGCGCGATATCTCTGCCCGCAACACGCTGGCCGGAGCGCGC
[0526] GTCGTGATGCTTAACCCGCCGCGTCGCGGTAAACGTAAGCCGGAATATCTGCAATGGC
[0527] TGGCGTCACTTGCACGGGCACAGAGCGATGCCGATGCGTTATCTGTTCATCTGGAACGC
[0528] GGCGCGGTTAACCTTGCGGATTTCGCCTGGGCGCGCCAGCTCAACGGCGAAGGGATGC
[0529] GCGAATTGCTGCAACAGCCTGGTTATATTCAGGCTGGTTATAGCTTGTTGAATACGCCG
[0530] GTTGCCGCCCGCTGGCAGCGGAAAATTCTCGACACATTAGCGACTTATCATGAGCAAC
[0531] ATCGCGATGAACCTGGCCCTGGGCGCGAACGTCTGCGACGTATGGCGTTGCCAATGGA
[0532] AGATGAAGCGCTGGTACTGTTGCTGATTAAAAAGATGCGCGAAAGCGGCGACATCCAC
[0533] AACCATCACGGCCGGCTGCATCTGCCAGATCACAAAGCGGGCCTCTAACATTGTCTTAT
[0534] AAATACCAAAAATTTCACGTCTCTTGTTTAAAGAGGCGTGCTTTAATATTTTCTTAGAG
[0535] TCCAATACATGAGAGTTTTAATTTCGTTTTTTTTCCTTATTTTTTCTATTAGCGTTAGAGC
[0536] TGACTGTGGCCTTGATAGATCTTCGAAGCTTGGGCCCGAACAAAAACTCATCTCAGAA
[0537] GAGGAGAGATAAATTGCACTGAAATCTAGAGTAACGGAATAGCTGTTCGTTGACTTGA
[0538] TAGACCGATTGATTCATCATCTCATAAATAAAGAAAAACCACCGCTACCAACGGTGGT
[0539] TTTCTCAAGGTTCGCTGAGCTACCAACTCTTTGAACCAAGGTAAGTGGGTTGGAGGACC
[0540] GCACTCACCAAAATCTGTTCTTTCAGTTTAGCCTTAACAGGTGCATAACTTCAAGACAA AGTCCTCTAAATCAGTTACCAATGGCTGCTGCCAGTGGCGATAAGTCGTGTCTTACCGG
[0541] GTTGGACTCAAGACGATAGTTACCGGATAAGGCGCAGCGGTCGGGCTGAACGGGGGGT
[0542] TCGTGCACACAGCCCAGCTTGGAGCGAACGACCTACACCGAACTGAGATACCAACAGC
[0543] GTGAGCTATGAGAAAGCGCCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCGG
[0544] TAAGCGGCAGGGTCGGAACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCT
[0545] GGTATCTTTATAGTCCTGTCGGGTTTCGCCACCTCTGGCTTGAGCGTCGATTTTTGTGAT
[0546] GCTCGTCAGGGGGGCGGAGCCTATGGAAAAACGCCTGCGGCGTTGGCTTCTTCCGGTG
[0547] CTTTGCTTTTTGCTCACATGTTCTTTCCGGCTTTATCCCCTGATTCTGTGGATAACCGTAT
[0548] TACCGCTTTTGAGTGAGCTGACACCGCTCGCCGCAGTCGAACGACCGAGCGTAGCGAG
[0549] TCAGTGAGCGAGGAAGCGGAAGAGCGCTGCATGCCTATTTGTTTATTTTTCTAAATACA
[0550] TTCAAATATGTATCCGCTCATGAGACAATAACCCTGATAAATGCTTCAATAATATTGAA
[0551] AAAGGAAGAGTATGAGCCATATTCAACGGGAAACGTCTTGCTCTAGGCCGCGATTAAA
[0552] TTCCAACATGGATGCTGATTTATATGGGTATAAATGGGCTCGCGATAATGTCGGGCAAT
[0553] CAGGTGCGACAATCTATCGATTGTATGGGAAGCCCGATGCGCCAGAGTTGTTTCTGAA
[0554] ACATGGCAAAGGTAGCGTTGCCAATGATGTTACAGATGAGATGGTCAGACTAAACTGG
[0555] CTGACGGAATTTATGCCTCTTCCGACCATCAAGCATTTTATCCGTACTCCTGATGATGC
[0556] ATGGTTACTCACCACTGCGATCCCCGGGAAAACAGCATTCCAGGTATTAGAAGAATAT
[0557] CCTGATTCAGGTGAAAATATTGTTGATGCGCTGGCAGTGTTCCTGCGCCGGTTGCATTC
[0558] GATTCCTGTTTGTAATTGTCCTTTTAACAGCGACCGCGTATTTCGTCTCGCTCAGGCGCA
[0559] ATCACGAATGAATAACGGTTTGGTTGATGCGAGTGATTTTGATGACGAGCGTAATGGC
[0560] TGGCCTGTTGAACAAGTCTGGAAAGAAATGCATAAACTTTTGCCATTCTCACCGGATTC
[0561] AGTCGTCACTCATGGTGATTTCTCACTTGATAACCTTATTTTTGACGAGGGGAAATTAA
[0562] TAGGTTGTATTGATGTTGGACGAGTCGGAATCGCAGACCGATACCAGGATCTTGCCATC
[0563] CTATGGAACTGCCTCGGTGAGTTTTCTCCTTCATTACAGAAACGGCTTTTTCAAAAATA
[0564] TGGTATTGATAATCCTGATATGAATAAATTGCAGTTTCATTTGATGCTCGATGAGTTTTT
[0565] CTAACTGCAATAAGGTTGTTTTGCCGTGGTCAACGTGTCCGGCAGTCGCAATAATCATG
[0566] GAGAAGGGCGTTTAACATAACCACGGATTGTAACGTGAGATGGGTCAGGAGGACATAT
[0567] CGCGCATCAAGCCTTTGGCGGTTCGGTGCATCCCATGGCGATTGAATATTGAGCAGTAG
[0568] AAAATATCTGGATTGACAGTATTACAAGGGGTTGGAGAACGTCAGGAGAGGCATTTTG
[0569] GCGGAAGATCACAGGAGTCGAACCTGCCCGGGACCGCTGGCGGCCCCAACTGGATTTA
[0570] GAGTCCAGCCGCCTCACCGGAGACGACGATCTTCCGCGCCTCGATTGCTACATGGAGG
[0571] CGGGGCGCATTATAGCTACTTCCTTGAGTTTCTACATCCCCCAGATCAGTTTTTTCCGCC
[0572] TAATACACGCAGCGTATTATCACGAAAGCATTATAGTTCATATAAAACCCATATTTTAC AATGATATATGAGCGGAATTAACGCTACCACGATCACATAAAACAAGAATCATAAATT
[0573] AATAACCAGATATCGGAATATTCGCTCTCACAGGGATGGCGGCCGC
[0574] 28. SEQ ID NO: 28 - Plasmid of the plasmid recovered after 4 rounds of initial selection.
[0575] AAAAAAAAACCACCGCTACCAACAGTGGTTTTCTCAAGGTTCGCTGAGCTACCAACTC
[0576] TTTGAACCAAGGTAAGTGGGTTGGAGGACCGCACTCACCAAAATCTGTTCTTTCAGTTT
[0577] AGCCTTAACAGGTGCATAACTTCAAGACAAAGTCCTCTAAATCAGTTACCAATGGCTG
[0578] CTGCCAGTGGCGATAAGTCGTGTCTTACCGGGTTGGACTCAAGACGATAGTTACCGGA
[0579] TAAGGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGGAGCGA
[0580] ACGACCTACACCGAACTGAGATACCAACAGCGTGAGCTATGAGAAAGCGCCACGCTTC
[0581] CCGAAGGGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCGGAACAGGAGAGC
[0582] GCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCGGGTTTCGC
[0583] CACCTCTGGCTTGAGCGTCGATTTTTGTGATGCTCGTCAGGGGGGCGGAGCCTATGGAA
[0584] AAACGCCTGCGGCGTTGGCTTCTTCCGGTGCTTTGCTTTTTGCTCACATGTTCTTTCCGG
[0585] CTTTATCCCCTGATTCTGTGGATAACCGTATTACCGCTTTTGAGTGAGCTGACACCGCTC
[0586] GCCGCAGTCGAACGACCGAGCGTAGCGAGTCAGTGAGCGAGGAAGCGGAAGAGCGCT
[0587] GCATGCCTATTTGTTTATTTTTCTAAATACATTCAAATATGTATCCGCTCATGAGACAAT
[0588] AACCCTGATAAATGCTTCAATAATATTGAAAAAGGAAGAGTATGAGCCATATTCAACG
[0589] GGAAACGTCTTGCTCTAGGCCGCGATTAAATTCCAACATGGATGCTGATTTATATGGGT
[0590] ATAAATGGGCTCGCGATAATGTCGGGCAATCAGGTGCGACAATCTATCGATTGTATGG
[0591] GAAGCCCGATGCGCCAGAGTTGTTTCTGAAACATGGCAAAGGTAGCGTTGCCAATGAT
[0592] GTTACAGATGAGATGGTCAGACTAAACTGGCTGACGGAATTTATGCCTCTTCCGACCAT
[0593] CAAGCATTTTATCCGTACTCCTGATGATGCATGGTTACTCACCACTGCGATCCCCGGGA
[0594] AAACAGCATTCCAGGTATTAGAAGAATATCCTGATTCAGGTGAAAATATTGTTGATGC
[0595] GCTGGCAGTGTTCCTGCGCCGGTTGCATTCGATTCCTGTTTGTAATTGTCCTTTTAACAG
[0596] CGACCGCGTATTTCGTCTCGCTCAGGCGCAATCACGAATGAATAACGGTTTGGTTGATG
[0597] CGAGTGATTTTGATGACGAGCGTAATGGCTGGCCTGTTGAACAAGTCTGGAAAGAAAT
[0598] GCATAAACTTTTGCCATTCTCACCGGATTCAGTCGTCACTCATGGTGATTTCTCACTTGA
[0599] TAACCTTATTTTTGACGAGGGGAAATTAATAGGTTGTATTGATGTTGGACGAGTCGGAA
[0600] TCGCAGACCGATACCAGGATCTTGCCATCCTATGGAACTGCCTCGGTGAGTTTTCTCCT
[0601] TCATTACAGAAACGGCTTTTTCAAAAATATGGTATTGATAATCCTGATATGAATAAATT
[0602] GCAGTTTCATTTGATGCTCGATGAGTTTTTCTAACTGCAATAAGGTTGTTTTGCCGTGGT
[0603] CAACGTGTCCGGCAGTCGCAATAATCATTTCAACAACATCTCCAAAAACCGTTGCTCAT
[0604] CTTCAAGGCAGCGTAAATCCAGCCACAATCGTCCGTCATAAATACGACCAATCACCGG CACTGGCAATTCACGCCAGCGAGCGGCTAATGACTCAAGGTGGCTACCGCGTCCATCA
[0605] TGGGGTGTAAACGTTAATGCCGCGCTCGGCAGGCGATCAACCGGCAGCGAACCACTGC
[0606] CAATCTGCGAAAGACATGGCATAACCTGTACCGCAAACTCCGCGCCGTAATGTGCGGC
[0607] AAGGGGGGCCTGTAAACGTTGTGCCTGGATTTGAATGACCTCTGCGCTGCGGGTAAGC
[0608] AGGCGCAGGGTCGGTAATTTTTCACTCAGAGCTTCAGGGTGTAAATAAAGACGCAACG
[0609] TGGCTTCCAGCGCCGCGAGGGTCATTTTATCCGCGCGTAATGCACGCTTCAGCGGGTGG
[0610] CTTTGCAGGCGGGCGATCATCTCTTTTTTACCAACAATAATTCCTGCCTGCGGCCCGCC
[0611] TAACAACTTGTCGCCGGAGAAACTCACCAGACTGACGCCCGCCGCAATCAACTCCTGC
[0612] GGCATTGGCTCTTTCGGCAAACCGTACTGGCTAAGATCGACCAGCGAGCCACTGCCTA
[0613] AATCAGTCACTACGGGAACATCCAGCTCTTTGCCGAGCGCCACCAGTTCCGCTTCATCT
[0614] ATCGCTTTGGTGAACCCCTGAATGCTGTAGTTACTGGTATGTACTTTCATCAACAGTGC
[0615] GGTATTTTCATTCACCGCCTGACGATAATCATTCGCGTGCGTGCGGTTGGTGGTCCCTA
[0616] CTTCGTGTAGGGTGCAGCCTGCCTGACGCATAACATCGGGAATACGAAACGCGCCGCC
[0617] AATCTCCACCAGTTCGCCGCGAGATACCACCACCTCTTTTCCGCTGGCAGTGGCCGCCA
[0618] ACATCAATAACACCGCCGCCGCATTGTTATTGACGATACAGGCATCTTCCGCCCCCGTA
[0619] ATACGGCACAGCAGCTGCGCCAGCGCCCGATCGCGATGTCCGCGTCCGGCGTCGTCCA
[0620] GATCATACTCGAGGGTCACTGGCGAACGCATAGCCTGCGCAACGGCTTCCACCGCGGC
[0621] TTCCGCCTGTAAAGCTCGCCCAAGGTTGGTATGCAGCACGGTTCCCGTCAGGTTGATCA
[0622] CCGGACGCAGCGCGCTCTGCGCTTCTTTCGTCAACCGGGCATCGACTTCTTGCGCCCAG
[0623] TTTTCACACCACGCAGGCAGCGTCTGGCTGCCACGAATCACTTCTCGCGCTTCGTCGAG
[0624] CATCTGACGCAACAATTCCACCACGCGGGTGTGACCATAAGTATCACGCAAAGAAAGG
[0625] AAGGAGCTATCGCGCAATAAGCGATCAATAGCCGGAAGTTGACTATAGAGGAAACGC
[0626] GTTTCGGTTGTCATGGTTTAGTTCCTCACCTTGTCGTATTATACTATGCCGATATACTAT
[0627] GCCGATGATTAGGAAACCTGGCTGATCAAGGCCCTCTCACACGGAGAAGGGCGTTTAA
[0628] CATAACCACGGATTGTAACGTGAGATGGGTCAGGAGGACATATCGCGCATCAAGCCTT
[0629] TGGCGGTTCGGTGCATCCCATGGCGATTGAATATTGAGCAGTAGAAAATATCTGGATT
[0630] GACAGTATTACAAGGGGTTGGAGAACGTCAGGAGAGGCATTTTGGCGGAAGATCACA
[0631] GGAGTCGAACCTGCCCGGGACCGCTGGCGGCCCCAACTGGATTTAGAGTCCAGCCGCC
[0632] TCACCGGAGACGACGATCTTCCGCGCCTCGATTGCTACATGGAGGCGGGGCGCATTAT
[0633] AGCTACTTCCTTGAGTTTCTACATCCCCCAGATCAGTTTTTTCCGCCTAATACACGCAGC
[0634] GTATTATCACGAAAGCATTATAGTTCATATAAAACCCATATTTTACAATGATATATGAG
[0635] CGGAATTAACGCTACCACGATCACATAAAACAAGAATCATAAATTAATAACCAGATAT
[0636] CGGAATATTCGCTCTCACAGGGATGGCGGCCGCTCGACCGATGCCCTTGAGACGTTGA
[0637] TCGGCACGTATGGCAGATAGTAAATTTTATAGATTTAAATCAACATCTATAACTGTATA ATTCGTTTTCTCAACTCATTACAACACTCCGTTAGTAATGAAGCTCATTCTATACAATG
[0638] ACAATTAATAGGTAAAGTTATGTCACTTAATGTAAAGCAAAGTAGAATAGCCATCTTG
[0639] TTTAGCTCTTGTTTAATTTCAATATCATTTTTCTCACAGGCCAATACGAAGGGCATTGAT
[0640] GAGATTAAAAACCTTGAAACAGATTTCAATGGCAGGATTGGTGTCTACGCTTTAGACA
[0641] CTGGCTCGGGTAAATCATTTTCGTACAGAGCAAATGAACGATTTCCATTATAGAGTTCT
[0642] TTTAAAGGTTTTTTAGCTGCTGCTGTATTAAAAGGCTCTCAAGATAATCGACTTAATCTT
[0643] AATCAGATTGTGAATTATAATACAAGAAGTTTAGAGTTCCATTCACCCATCACAACTAA
[0644] ATATAAAGATAATGGAATGTCATTAGGTGATATGGCTGCTGCTGCTTTACAATATAGCG
[0645] ACAATGGTGCTACTAATATTATTCTTGAACGTTATATCGGTGGTCCAGAGGGTATGACT
[0646] AAATTCATGCGGTCGATTGGAGATGAAGATTTTAGACTCGATCGTTGGGAGTTAGATCT
[0647] AAACACAGCTATTCCAGGCGATGAGCGTGACACATCTACACCTGCAGCAGTAGCCAAG
[0648] AGTCTGAAAACCCTTGCTCTGGGTAACATACTTAGTGAACATGAAAAGGAAACCTATC
[0649] AGACATGGTTAAAGGGTAACACAACCGGTGCAGCGCGTATTCGTGCTAGCGTACCAAG
[0650] CGATTGGGTAGTTGGCGATAAAACTGGTAGTTGCGGAGCATACGGTACGGCAAATGAT
[0651] TATGCGGTAGTCTGGCCAAAGAACCGGGCTCCTCTTATAATTTCTGTATACACAACAAA
[0652] AAACGAAAAAGAAGCCAAGCATGAGGATAAAGTAATCGCAGAAGCTTCAAGAATTGC
[0653] AATTGATAACCTTAAATAACAAAGCCCGCCGAAAGGCGGGCTTTTTTTTTGGATCCTGT
[0654] AGAAACGCAAAAAGGCCATCCGCATGACTAACATGAGAATTACAACTATCACCGCAAG
[0655] GGATAAATATCTAACACCGTTATGACAACTTGACGGCTACATCATTCACTTTTTCTTCA
[0656] CAACCGGCACGGAACTCGCTCGGGCTGGCCCCGGTGCATTTTTTAAATACCCGCGAGA
[0657] AGTAGAGTTGATCGTCAAAACCAACATTGCGACCGACGGTGGCGATAGGCATCCGGGT
[0658] GGTGCTCAAAAGCAGCTTCGCCTGGCTGATACGTTGGTCCTCGCGCCAGCTTAAGACGC
[0659] TAATCCCTAACTGCTGGCGGAAAAGATGTGACAGACGCGACGGCGACAAGCAAACAT
[0660] GCTGTGCGACGCTGGCGATATCAAAATTGCTGTCTGCCAGGTGATCGCTGATGTACTGA
[0661] CAAGCCTCGCGTACCCGATTATCCATCGGTGGATGGAGCGACTCGTTAATCGCTTCCAT
[0662] GTGCCGCAGTAACAATTGCTCAAGCAGATTTATCGCCAGCAGCTCCGAATAGCGCCCTT
[0663] CCCCTTGCCCGGCGTTAATGATTTGCCCAAACAGGTCGCTGAAATGCGGCTGGTGCGCT
[0664] TCATCCGGGCGAAAGAACCCCGTATTGGCAAATATTGACGGCCAGTTAAGCCATTCAT
[0665] GCCAGTAGGCGCGCGGACGAAAGTAAACCCACTGGTGATACCATTCGCGAGCCTCCGG
[0666] ATGACGACCGTAGTGATGAATCTCTCCTGGCGGGTTTAACGCCGATTGAGGCCAACGG
[0667] TTATCTCGATTTTTTTATCGACCGACCGCTGGGAATGAAAGGTTATATTCTCAATCTCAC
[0668] CATTCGCGGTCAGGGGGTGGTGAAAAATCAGGGACGAGAATTTGTTTGCCGACCGGGT
[0669] GATATTTTGCTGTTCCCGCCAGGAGAGATTCATCACTACGGTCGTCATCCGGAGGCTCG
[0670] CGAATGGTATCACCAGTGGGTTTACTTTCGTCCGCGCGCCTACTGGCATGAATGGCTTA ACTGGCCGTCAATATTTGCCAATACGGGGTTCTTTCGCCCGGATGAAGCGCACCAGCCG
[0671] CATTTCAGCGACCTGTTTGGGCAAATCATTAACGCCGGGCAAGGGGAAGGGCGCTATT
[0672] CGGAGCTGCTGGCGATAAATCTGCTTGAGCAATTGTTACTGCGGCACATGGAAGCGAT
[0673] TAACGAGTCGCTCCATCCACCGATGGATAATCGGGTACGCGAGGCTTGTCAGTACATC
[0674] AGCGATCACCTGGCAGACAGCAATTTTGATATCGCCAGCGTCGCACAGCATGTTTGCTT
[0675] GTCGCCGTCGCGTCTGTCACATCTTTTCCGCCAGCAGTTAGGGATTAGCGTCTTAAGCT
[0676] GGCGCGAGGACCAACGTATCAGCCAGGCGAAGCTGCTTTTGAGCACCACCCGGATGCC
[0677] TATCGCCACCGTCGGTCGCAATGTTGGTTTTGACGATCAACTCTACTTCTCGCGGGTAT
[0678] TTAAAAAATGCACCGGGGCCAGCCCGAGCGAGTTCCGTGCCGGTTGTGAAGAAAAAGT
[0679] GAATGATGTAGCCGTCAAGTTGTCATAACGGTGTTAGATATTTATCCCTTGCGGTGATA
[0680] GTTGTAATTCTCATGTTAGTCATGCGGATGGCCTTTTTGCGTTTCTACAGGATCCAAAA
[0681] AAAGCCCGCCTTTCGGCGGGCTTTGTTATTTAAGGTTATCAATTGCAATTCTTGAAGCT
[0682] TCTGCGATTACTTTATCCTCATGCTTGGCTTCTTTTTCGTTTTTTGTTGTGTATACAGAAA
[0683] TTATAAGAGGAGCCCGGTTCTTTGGCCAGACTACCGCATAATCATTTGCCGTACCGTAT
[0684] GCTCCGCAACTACCAGTTTTATCGCCAACTACCCAATCGCTTGGTACGCTAGCACGAAT
[0685] ACGCGCTGCACCGGTTGTGTTACCCTTTAACCATGTCTGATAGGTTTCCTTTTCATGTTC
[0686] ACTAAGTATGTTACCCAGAGCAAGGGTTTTCAGACTCTTGGCTACTGCTGCAGGTGTAG
[0687] ATGTGTCACGCTCATCGCCTGGAATAGCTGTGTTTAGATCTAACTCCCAACGATCGAGT
[0688] CTAAAATCTTCATCTCCAATCGACCGCATGAATTTAGTCATACCCTCTGGACCACCGAT
[0689] ATAACGTTCAAGAATAATATTAGTAGCACCATTGTCGCTATATTGTAAAGCAGCAGCA
[0690] GCCATATCACCTAATGACATTCCATTATCTTTATATTTAGTTGTGATGGGTGAATGGAA
[0691] CTCTAAACTTCTTGTATTATAATTCACAATCTGATTAAGATTAAGTCGATTATCTTGAGA
[0692] GCCTTTTAATACAGCAGCAGCTAAAAAACCTTTAAAAGAACTCTATAATGGAAATCGT
[0693] TCATTTGCTCTGTACGAAAATGATTTACCCGAGCCAGTGTCTAAAGCGTAGACACCAAT
[0694] CCTGCCATTGAAATCTGTTTCAAGGTTTTTAATCTCATCAATGCCCTTCGTATTGGCCTG
[0695] TGAGAAAAATGATATTGAAATTAAACAAGAGCTAAACAAGATGGCTATTCTACTTTGC
[0696] TTTACATTAAGTGACATAACTTTACCTATTAATTGTCATTGTATAGAATGAGCTTCATTA
[0697] CTAACGGAGTGTTGTAATGAGTTGAGAAAACGAATTATACAGTTATAGATGTTGATTTA
[0698] AATCTATAAAATTTACTATCTGCCATACGTGCCGATCAACGTCTCAAGGGCATCGGTCG
[0699] AGCGGCCGCCATCCCTGTGAGAGCGAATATTCCGATATCTGGTTATTAATTTATGATTC
[0700] TTGTTTTATGTGATCGTGGTAGCGTTAATTCCGCTCATATATCATTGTAAAATATGGGTT
[0701] TTATATGAACTATAATGCTTTCGTGATAATACGCTGCGTGTATTAGGCGGAAAAAACTG
[0702] ATCTGGGGGATGTAGAAACTCAAGGAAGTAGCTATAATGCGCCCCGCCTCCATGTAGC
[0703] AATCGAGGCGCGGAAGATCGTCGTCTCCGGTGAGGCGGCTGGACTCTAAATCCAGTTG GGGCCGCCAGCGGTCCCGGGCAGGTTCGACTCCTGTGATCTTCCGCCAAAATGCCTCTC
[0704] CTGACGTTCTCCAACCCCTTGTAATACTGTCAATCCAGATATTTTCTACTGCTCAATATT
[0705] CAATCGCCATGGGATGCACCGAACCGCCAAAGGCTTGATGCGCGATATGTCCTCCTGA
[0706] CCCATCTCACGTTACAATCCGTGGTTATGTTAAACGCCCTTCTCCGTGTGAGAGGGCCT
[0707] TGATCAGCCAGGTTTCCTAATCATCGGCATAGTATATCGGCATAGTATAATACGACAAG
[0708] GTGAGGAACTAAACCATGACAACCGAAACGCGTTTCCTCTATAGTCAACTTCCGGCTAT
[0709] TGATCGCTTATTGCGCGATAGCTCCTTCCTTTCTTTGCGTGATACTTATGGTCACACCCG
[0710] CGTGGTGGAATTGTTGCGTCAGATGCTCGACGAAGCGCGAGAAGTGATTCGTGGCAGC
[0711] CAGACGCTGCCTGCGTGGTGTGAAAACTGGGCGCAAGAAGTCGATGCCCGGTTGACGA
[0712] AAGAAGCGCAGAGCGCGCTGCGTCCGGTGATCAACCTGACGGGAACCGTGCTGCATAC
[0713] CAACCTTGGGCGAGCTTTACAGGCGGAAGCCGCGGTGGAAGCCGTTGCGCAGGCTATG
[0714] CGTTCGCCAGTGACCCTCGAGTATGATCTGGACGACGCCGGACGCGGACATCGCGATC
[0715] GGGCGCTGGCGCAGCTGCTGTGCCGTATTACGGGGGCGGAAGATGCCTGTATCGTCAA
[0716] TAACAATGCGGCGGCGGTGTTATTGATGTTGGCGGCCACTGCCAGCGGAAAAGAGGTG
[0717] GTGGTATCTCGCGGCGAACTGGTGGAGATTGGCGGCGCGTTTCGTATTCCCGATGTTAT
[0718] GCGTCAGGCAGGCTGCACCCTACACGAAGTAGGGACCACCAACCGCACGCACGCGAAT
[0719] GATTATCGTCAGGCGGTGAATGAAAATACCGCACTGTTGATGAAAGTACATACCAGTA
[0720] ACTACAGCATTCAGGGGTTCACCAAAGCGATAGATGAAGCGGAACTGGTGGCGCTCGG
[0721] CAAAGAGCTGGATGTTCCCGTAGTGACTGATTTAGGCAGTGGCTCGCTGGTCGATCTTA
[0722] GCCAGTACGGTTTGCCGAAAGAGCCAATGCCGCAGGAGTTGATTGCGGCGGGCGTCAG
[0723] TCTGGTGAGTTTCTCCGGCGACAAGTTGTTAGGCGGGCCGCAGGCAGGAATTATTGTTG
[0724] GTAAAAAAGAGATGATCGCCCGCCTGCAAAGCCACCCGCTGAAGCGTGCATTACGCGC
[0725] GGATAAAATGACCCTCGCGGCGCTGGAAGCCACGTTGCGTCTTTATTTACACCCTGAAG
[0726] CTCTGAGTGAAAAATTACCGACCCTGCGCCTGCTTACCCGCAGCGCAGAGGTCATTCA
[0727] AATCCAGGCACAACGTTTACAGGCCCCCCTTGCCGCACATTACGGCGCGGAGTTTGCG
[0728] GTACAGGTTATGCCATGTCTTTCGCAGATTGGCAGTGGTTCGCTGCCGGTTGATCGCCT
[0729] GCCGAGCGCGGCATTAACGTTTACACCCCATGATGGACGCGGTAGCCACCTTGAGTCA
[0730] TTAGCCGCTCGCTGGCGTGAATTGCCAGTGCCGGTGATTGGTCGTATTTATGACGGACG
[0731] ATTGTGGCTGGATTTACGCTGCCTTGAAGATGAGCAACGGTTTTTGGAGGATTGTATGA
[0732] TTATCGCGACTGCCGGACATGTGGATCATGGAAAGACAACATTGTTGCAGGCGATTAC
[0733] TGGCGTAAATGCTGACCGTCTGCCGGAAGAAAAAAAGCGCGGCATGACCATAGATCTC
[0734] GGCTATGCCTACTGGCCGCAGCCGGATGGTCGCGTGCCTGGTTTTATCGACGTTCCCGG
[0735] TCATGAAAAGTTTCTTTCCAACATGCTGGCGGGCGTTGGTGGTATCGATCACGCGCTGT
[0736] TGGTGGTGGCATGCGATGACGGCGTGATGGCACAGACCCGTGAGCATCTGGCGGTTTT GCAGCTGACCGGTAACCCGATGCTGACAGTGGCGCTGACCAAAGCCGATCGCGTGGAC
[0737] GAAGCGCGTGTTGATGAGGTTGAACGCCAGGTAAAGGAGGTTCTGCGGGAATACGGTT
[0738] TTGCTGAGGCAAAACTGTTTATCACCGCAGCAACCGAAGGTCGGGGAATGGATGCCCT
[0739] GCGCGAGCATCTGCTTCAGTTGCCGGATCGCGAGCACGCCAGCCAACATAGTTTCCGC
[0740] CTCGCGATTGACCGCGCATTTACCGTAAAAGGTGCCGGCCTGGTCGTCACCGGTACGG
[0741] CGTTAAGCGGGGAAGTGAAGGTAGGCGATCCACTCTGGCTGACTGGTGTAAATAAACC
[0742] GATGCGTGTACGTGCGCTGCATGCGCAAAACCAGCCAACAGAAACCGCCAATGCCGGG
[0743] CAGCGTATCGCGCTTAACATCGCGGGTGATGCGGAAAAAGAGCAGATTAACCGTGGCG
[0744] ACTGGCTGCTTGCCGATGTGCCGCCAGAGCCGTTCACACGGGTGATTGTCGAGCTTCAA
[0745] ACCCATACATCGCTGACCCAGTGGCAGCCGCTGCATATTCACCACGCCGCCAGCCACG
[0746] TCACGGGACGCGTTTCACTGCTGGAAGATAACCTTGCTGAACTGGTCTTCGACACCCCG
[0747] TTATGGCTGGCAGATAACGACCGCCTGGTATTGCGCGATATCTCTGCCCGCAACACGCT
[0748] GGCCGGAGCGCGCGTCGTGATGCTTAACCCGCCGCGTCGCGGTAAACGTAAGCCGGAA
[0749] TATCTGCAATGGCTGGCGTCACTTGCACGGGCACAGAGCGATGCCGATGCGTTATCTGT
[0750] TCATCTGGAACGCGGCGCGGTTAACCTTGCGGATTTCGCCTGGGCGCGCCAGCTCAAC
[0751] GGCGAAGGGATGCGCGAATTGCTGCAACAGCCTGGTTATATTCAGGCTGGTTATAGCT
[0752] TGTTGAATGCGCCGGTTGCCGCCCGCTGGCAGCGGAAAATTCTCGACACATTAGCGAC
[0753] TTATCATGAGCAACATCGCGATGAACCTGGCCCTGGGCGCGAACGTCTGCGACGTATG
[0754] GCGTTGCCAATGGAAGATGAAGCGCTGGTACTGTTGCTGATTGAAAAGATGCGCGAAA
[0755] GCGGCGACATCCACAACCATCACGGCCGGCTGCATCTGCCAGATCACAAAGCGGGCCT
[0756] CTAACATTGTCTTATAAATACCAAAAATTTCACATCTCTTGTTTAAAGAGGCGTGCTTT
[0757] AATATTTTCTTAGAGTCCAATACATGAGAGTTTTAATTTCGTTTTTTTTCCTTATTTTTTC
[0758] TATTAGCGTTAGAGCTGACTGTGGCCTTGATAGATCTTCGAAGCTTGGGCCCGAACAAA
[0759] AACTCATCTCAGAAGAGGATCTGAGAGATAAATTGCACTGAAATCTAGAGTAACGGAA
[0760] TAGCTGTTCGTTGACTTGATAGACCGATTGATTCATCATCTCATAAAT
[0761] 29. SEQ ID NO: 29 - DHFR P39S
[0762] MISLIAALAVDRVIGMENAMPWNLPADLAWFKRNTLNKSVIMGRHTWESIGRPLPGRKNII
[0763] LSSQPGTDDRVTWVKSVDEAIAACGDVPEIMVIGGGRVYEQFLPKAQKLYLTHIDAEVEG
[0764] DTHFPDYEPDDWESVFSEFHDADAQNSHSYCFEILERRWSHPQFEK
[0765] 30. SEQ ID NO: 30 - DHFR P39TAG MISLIAALAVDRVIGMENAMPWNLPADLAWFKRNTLNKUVIMGRHTWESIGRPLPGRKNII
[0766] LSSQPGTDDRVTWVKSVDEAIAACGDVPEIMVIGGGRVYEQFLPKAQKLYLTHIDAEVEG
[0767] DTHFPDYEPDDWESVFSEFHDADAQNSHSYCFEILERRWSHPQFEK
[0768] 31. SEQ ID NO: 31 - Trypsin-digested peptide
[0769] LEYNFNSH
[0770] 32. SEQ ID NO: 32 - Trypsin-digested peptide
[0771] VYITADK
[0772] 33. SEQ ID NO: 33 - Trypsin-digested peptide
[0773] KDAT1YV
[0774] 34. SEQ ID NO: 34 - Trypsin-digested peptide
[0775] HSNFNYEL
[0776] 35. SEQ ID NO: 35 - Ec SelB
[0777] MIIATAGHVDHGKTTLLQAITGVNADRLPEEKKRGMTIDLGYAYWPQPDGRVPGFIDVPG
[0778] HEKFLSNMLAGVGGIDHALLVVACDDGVMAQTREHLAILQLTGNPMLTVALTKADRVDE
[0779] ARVDEVERQVKEVLLREYGFAEAKLFITAATEGRGMDALREHLLQLPEREHASQHSFRLAI
[0780] DRAFTVKGAGLVVTGTALSGEVKVGDSLWLTGVNKPMRVRALHAQNQPTETANAGQRIA
[0781] LNIAGDAEKEQINRGDWLLADVPPEPFTRVIVELQTHTPLTQWQPLHIHHAASHVTGRVSL
[0782] LEDNLAELVFDTPLWLADNDRLVLRDISARNTLAGARVVMLNPP
[0783] 36. SEQ ID NO: 36 - Ec EF-Tu
[0784] VNVGTIGHVDHGKTTLTAAITTVLAKTYGGAARAFDQIDNAPEEKARGITINTSHVEYDTP
[0785] TRHYAHVDCPGHADYVKNMITGAAQMDGAILVVAATDGPMPQTREHILLGRQVGVPYIIV
[0786] FLNKCDMVDDEELLELVEMEVRELLSQYDFPGDDTPIVRGSALKALEGDAEWEAKILELA
[0787] GFLDSYIPEPERAIDKPFLLPIEDVFSISGRGTVVTGRVERGIIKVGEEVEIVGIKETQKSTCTG
[0788] VEMFRKLLDEGRAGENVGVLLRGIKREEIERGQVLAKPGTIKPHTKFESEVYILSKDEGGRH
[0789] TPFFKGYRPQFYFRTTDVTGTIELPEGVEMVMPGDNIKMVVTLIHPIAMDDGLRFAIREGGR
[0790] TVGAGVVAKVLS
[0791] 37. SEQ ID NO: 37 - Taq EF-Tu VNVGTIGHVDHGKTTLTAALTYVAAAENPNVEVKDYGDIDKAPEERARGITINTAHVEYE
[0792] TAKRHYSHVDCPGHADYIKNMITGAAQMDGAILVVSAADGPMPQTREHILLARQVGVPYI
[0793] VVFMNKVDMVDDPELLDLVEMEVRDLLNQYEFPGDEVPVIRGSALLALEEMHKNPKTKR
[0794] GENEWVDKIWELLDAIDEYIPTPVRDVDKPFLMPVEDVFTITGRGTVATGRERGKVKVGD
[0795] EVEIVGLAPETRKTVVTGVEMHRKTLQEGIAGDNVGLLLRGVSREEVERGQVLAKPGSITP
[0796] HTKFEASVYILKKEEGGRHTGFFTGYRPQFYFRTTDVTGVVRLPQGVEMVMPGDNVTFTV
[0797] ELIKPVALEEGLRFAIREGGRTVGAGVVTKILE
[0798] 38. SEQ ID NO: 38 - Consensus
[0799] VNVGTIGHVDHGKTTLTAAITXVXAXXXXXXXXXXXXIDXAPEEKARGITINTXHVEYXT
[0800] PXRHYXHVDCPGHADYXKNMITGAAQMDGAILVVAAXDGPMPQTREHILLXRQVGVPYI
[0801] XVFLNKXDMVDDEELLDLVEMEVRELLXQYXFPGDEXPXXRGSALXALEXXXEWXXK1X
[0802] ELXXXLDXYIPEPERAXDKPFLLPIEDVFTIXGRGTVVTGRXERGXVKVGDEVEIVGXXET
[0803] XKXXVTGVEMXRKXLXEGXAGXNVGLLLRGXXREEIERGQVLAKPGXIXPHTKFEXXVYI
[0804] LXKXEGGRHTXFFXGYRPQFYFRTTDVTGXVXLPXGVEMVMPGDNXXXXVXLIXPXALX
[0805] DGLRFAIREGGRTVGAGVVXKXLX
[0806] X indicates either any amino acid or any three amino acids compared between E. Coli SelB, E. Coli
[0807] EF-Tu, and Taq EF-Tu.
Claims
CLAIMSWhat is claimed is:
1. An engineered selenocysteine-specific elongation factor (SelB) comprising at least 90% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6, wherein the SelB variant is capable of incorporating at least one non-canonical amino acid into a peptide.
2. The engineered elongation factor of claim 1 , wherein the SelB variant is a truncation of SEQ ID NO: 3.
3. The engineered elongation factor of claim 2, wherein the truncation comprises a deletion of a carboxy terminal end.
4. The engineered elongation factor of claim 2, wherein the deletion comprises removal of at least 100 of amino acids from the truncation.
5. The polypeptide of claim 2, wherein the deletion comprises removal of about 127 amino acids from the truncation.
6. The polypeptide of claim 2, wherein the SelB variant comprises at least one amino acid mutation within the truncation.
7. The polypeptide of claim 6, wherein the at least one amino acid mutation is selected from the group consisting of I98V, E166D, S206P, P277S A413T, S473N, W477R, and F487L.
8. The polypeptide of claim 7, wherein the SelB variant further comprises an amino acid mutation selected from E246K and / or E463K within the truncation.
9. The polypeptide of claim 7, wherein the SelB variant further comprises an amino acid mutation selected from Y44F, T193C, and / or R236S.
10. An expression vector comprising a nucleic acid sequence comprising at least 90% sequence identity to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 7 wherein the nucleic acid sequence encodes a Selenocysteine-specific elongation factor (SelB) variant.
11. The vector of claim 10, wherein the nucleic acid sequence encodes at least one amino acid mutation.
12. A cell comprising the vector of claim 10.
13. A method of incorporating at least one non-canonical amino acid into a peptide during translation, the method comprising contacting an mRNA sequence with an engineered selenocysteinespecific elongation factor (SelB) variant, wherein the engineered SelB variant incorporates a non- canonical amino acid into the peptide without need of a selenocysteine insertion sequence (SECIS).
14. The method of claim 13, wherein translation occurs in presence of one or more translation regulatory proteins.
15. The method of claim 14, wherein the one or more translation regulatory proteins comprise an initiation factor, an elongation factor, a release factor, and / or a ribosome.
16. The method of claim 15, wherein the elongation factor comprises elongation Factor Tu (EF- Tu).
17. The method of claim 13, wherein the SelB variant comprises at least 90% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6.
18. A peptide produced by the method of claim 13, wherein the peptide comprises at least one selenocysteine residue or at least one serine residue.
19. A protein translation system comprising an engineered selenocysteine-specific elongation factor (SelB) variant, wherein the SelB variant comprises at 90% sequence identity to SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 6, wherein the SelB variant is capable of incorporating at least one non-canonical amino acid into a peptide, and wherein the SelB variant translates an mRNA sequence without need of a selenocysteine insertion sequence (SECIS).
20. A cell comprising the system of claim 19.
Citation Information
Patent Citations
Method for synthesizing selenoprotein by inserting selenocysteine into lactic acid bacteria at fixed point and application of selenoprotein
CN117887748A
Selenocysteine fixed-point insertion method and application thereof
CN118147181A
Compositions and methods for making selenocysteine containing polypeptides
US20200332336A1