Genetic code
By re-encoding the genetic code of the cell and making it orthogonal to the general genetic code, the problem of horizontal gene transfer between mobile genetic elements and cells is solved, and effective defense and cell resistance are achieved for horizontal gene transfer.
Patent Information
- Application Number
- CN202380063771.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-18
- Filing Date
- 2023-07-19
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively prevent horizontal gene transfer between mobile genetic elements and cells, resulting in host organisms being invaded by selfish genetic elements.
By re-encoding the cell's genetic code, it exhibits semantic and functional orthogonality relative to the general genetic code, creating an orthogonal horizontal gene transfer system. Specific methods include reassigning the sense codon, decoding it into amino acids that are not associated with the normative genetic code, and modifying the tRNA to ensure that it can only decode a specific type of codon.
Effective defense against horizontal gene transfer is achieved, reducing the invasiveness of mobile genetic elements, and improving cell resistance to mutations, preventing harmful genes from escaping from synthetic organisms to natural organisms.
Smart Images

Figure CN120202291A_ABST
Abstract
Description
Technical Field
[0001] This article provides cells that are resistant to mobile genetic elements or horizontal gene transfer, as well as methods for obtaining such cells. Also provided are methods for preventing the horizontal transfer of genetic information between mobile genetic elements and cells, cells utilizing new genetic codon schemes and related topics, kits comprising mutually orthogonal cells, and mobile genetic elements. Also provided are methods for altering the susceptibility of genes to mutations that alter the encoded amino acid sequence, methods for evolving or improving proteins, and methods for making target genes more resistant to mutations. Additionally, provided are uses of cells for the manufacture of polymers and methods comprising using cells for the manufacture of polymers. Background Art
[0002] The nearly universal genetic code defines the correspondence between codons in genes and amino acids in proteins (1, 2). Because all forms of life use essentially the same genetic code, evolutionary innovations can be shared between organisms via horizontal gene transfer (HGT) (3, 4). The sharing of genetic information between organisms is a major driving force in the evolution of prokaryotes and some eukaryotes (5).
[0003] However, the nearly universal genetic code is also a burden to organisms; mobile genetic elements (or selfish genetic elements), including transposons, viruses, and plasmids, exploit the universality of the code and coerce the host cell's machinery to read their genes and reproduce themselves at the expense of the host organism. There is a clear tension between maintaining a common genetic code to allow beneficial innovations through HGT and excluding selfish genetic elements that exploit the common code for their own purposes (3, 6).
[0004] Several deviations from the standard genetic code have been documented in mitochondria and chloroplasts, and the vast majority of characteristic codon reassignments involve stop codons (7-9). Known sense codon reassignments in nuclear genomes are rare. "CTG yeast" decodes CUG codons (which encode leucine in the standard code) predominantly as serine (97%, with the remaining 3% still assigned to leucine) (10). Viruses are essentially unknown in CTG yeast, suggesting that sense codon reassignment may provide protection from viruses (11). There are currently no experimentally verified examples of sense codon reassignment in bacteria, although recent work provides computational evidence for the reassignment of arginine codons in Bacillus (12).
[0005] Genome synthesis (13-15) and editing provide opportunities to rewrite an organism's genetic code (15-17). We synthesized a 4-Mb Escherichia coli genome in which we compressed the genetic code by removing all annotated occurrences of the TCG and TCA sense codons encoding serine and the TAG stop codon; this created a new strain, Syn61 (15). We then further evolved the strain and deleted the tRNA genes that decode TCG and TCA codons (serU, tRNA CGA Ser and serT, tRNA UGA Ser ) and an RF-1 gene (prfA) that terminates protein synthesis at the TAG stop codon. The resulting organism, Syn61Δ3, is unable to read all codons in the nearly universal genetic code and is therefore unable to read horizontally transferred genes containing codons deleted from its genome, as exemplified by resistance to a range of bacteriophages (18).
[0006] It has been widely hypothesized that reorganizing the genetic code by reassigning sense codons to different canonical amino acids would create organisms with novel properties and would create a genetic firewall to limit the escape of genetic information from synthetic organisms to natural organisms (4, 6, 19, 20). However, these hypotheses remain untested. Summary of the Invention
[0007] In experiments disclosed herein, the genetic code of a synthetic E. coli strain was reconfigured to exhibit semantic and functional orthogonality relative to the universal genetic code, allowing for the creation of an orthogonal horizontal gene transfer system.
[0008] In one aspect, a cell is provided that: comprises a genome in which at least a first type of sense codon has been recoded such that a first endogenous tRNA is dispensable; does not express the first endogenous tRNA; expresses a first modified tRNA capable of decoding the first type of sense codon, wherein the first modified tRNA carries a first amino acid that is not a naturally cognate amino acid for the first type of sense codon; and comprises a gene required for viability, wherein the gene comprises at least one occurrence of the first type of sense codon, and the cell is viable when the first type of sense codon in the gene decodes to the first amino acid.
[0009] In another aspect, a cell is provided that: comprises a genome in which a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable; does not express the first endogenous tRNA and the second endogenous tRNA; expresses a first anticodon-exchanged tRNA derived from a naturally occurring first parent tRNA, wherein the first anticodon-exchanged tRNA carries a first amino acid and the first parent tRNA is an isoacceptor for the first amino acid, and wherein the first amino acid is not a naturally cognate amino acid for the first type of sense codon; and expresses a second anticodon-exchanged tRNA derived from a naturally occurring second parent tRNA, wherein the second anticodon-exchanged tRNA carries a second amino acid and the second parent tRNA is an isoacceptor for the second amino acid, and wherein the second amino acid is not a naturally cognate amino acid for the second type of sense codon; wherein the first and / or second modified tRNA is incapable of decoding any type of codon other than the first type of sense codon and / or the second type of sense codon.
[0010] In another aspect, a cell is provided, comprising: a genome wherein a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable; the first endogenous tRNA and the second endogenous tRNA are not expressed; a first modified tRNA capable of decoding the first type of sense codon is expressed, wherein the first modified tRNA carries a first amino acid that is not a naturally cognate amino acid for the first type of sense codon; and a second modified tRNA capable of decoding the second type of sense codon is expressed, wherein the second modified tRNA carries a second amino acid that is not a naturally cognate amino acid for the second type of sense codon; wherein: i) the first amino acid is alanine and the second amino acid is alanine; ii) the first amino acid is alanine and the second amino acid is histidine; iii) the first amino acid is tRNA. The first amino acid is alanine and the second amino acid is leucine; iv) the first amino acid is alanine and the second amino acid is proline; v) the first amino acid is histidine and the second amino acid is alanine; vi) the first amino acid is histidine and the second amino acid is histidine; vii) the first amino acid is histidine and the second amino acid is leucine; viiii) the first amino acid is histidine and the second amino acid is proline; ix) the first amino acid is leucine and the second amino acid is alanine; x) the first amino acid is leucine and the second amino acid is histidine; xi) the first amino acid is leucine and the second amino acid is proline; xii) the first amino acid is proline and the second amino acid is alanine; xiii) the first amino acid is proline and the second amino acid is histidine; xiv) the first amino acid is proline and the second amino acid is leucine; or xv) the first amino acid is proline and the second amino acid is proline.
[0011] In another aspect, a cell with increased resistance to horizontal gene transfer or mobile genetic elements is provided, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid that is not associated with a sense codon in the canonical genetic code, and the cell comprises a gene required for viability that is functional when decoded according to the reassigned genetic code and is not functional when decoded according to the canonical genetic code.
[0012] In another aspect, a method is provided for increasing resistance of a cell to mobile genetic elements or horizontal gene transfer, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid that is not associated with a sense codon in the canonical genetic code, the method comprising: modifying a gene required for viability to include at least one occurrence of the reassigned sense codon, wherein the cell is viable if the reassigned sense codon in the gene decodes to the reassigned amino acid, and the cell is non-viable if the reassigned sense codon in the gene decodes according to the canonical genetic code, or wherein the reassigned sense codon in the gene, if decoded according to the canonical genetic code, contributes at least in part to loss of viability.
[0013] In another aspect, a kit is provided comprising a first cell recoded according to a first orthogonal encoding scheme and a second cell recoded according to a second orthogonal encoding scheme, wherein the first and second encoding schemes are orthogonal to each other.
[0014] In another aspect, mobile genetic elements are provided that are recoded according to an orthogonal coding scheme.
[0015] In another aspect, a method for preventing horizontal transfer of genetic information between a mobile genetic element and a first cell is provided, the method comprising incubating a mobile genetic element and a first cell, wherein the mobile genetic element is a mobile genetic element as disclosed herein, and the first cell comprises tRNA that decodes codons according to the canonical genetic code or according to a coding scheme orthogonal to the mobile genetic element.
[0016] In another aspect, a method of altering the susceptibility of a gene to a mutation that changes the encoded amino acid sequence is provided, the method comprising: i) identifying a target gene; and ii) incubating a cell comprising the target gene, wherein the cell comprises a tRNA capable of decoding at least one sense codon into a reassigned amino acid.
[0017] In another aspect, a use of a cell disclosed herein for producing a polymer is provided. In one embodiment, a method for producing a polymer is provided, the method comprising: culturing a cell disclosed herein, providing the cell with a nucleic acid sequence encoding the polymer, and obtaining the polymer. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 .The compressed genetic code is non-orthogonal. (A) Relationship between TCG and TCA codons in a gene, decoders for these codons in cells with wild-type (WT) decoding and Syn61Δ3 decoding (Δ3), and the corresponding protein sequences synthesized. The anticodon (decoder) indicates the tRNA that reads the TCG or TCA codon. The amino acid (aa) used by the tRNA is indicated. A decoder in gray indicates that the tRNA is charged with serine. A codon in gray indicates that the codon is located in a non-codon compressed gene and that decoding it to serine will produce the correct protein sequence. A decoder / amino acid pair in pink indicates that the tRNA is missing. A codon in pink indicates that the codon is not present in the gene because the gene has been designed with codon compression. (B) Functional evaluation of wild-type (SpecR WT) and codon-compressed spectinomycin-resistant (recSpecR (ΔTCG, TCA)) genes in cells that decode all codons in the reading frame using the full tRNA complement (Syn61 WT, left panel) and cells in which tRNAs that decode TCG and TCA codons have been deleted (Syn61Δ3, right panel). Cells were spotted on agar plates in the presence or absence of spectinomycin and incubated overnight. Growth of cells in the presence of spectinomycin indicates that the indicated SpecR gene is functional in the indicated strain. (C, D) Predicted protein synthesis and horizontal gene transfer outcomes from mobile genetic elements and recipient cells with the indicated decoders, and codons in essential genes. (C) A mobile genetic element that encodes its gene according to the canonical genetic code in which TCG and TCA encode serine cannot be horizontally transferred to Syn61Δ3 cells that lack decoders for TCG and TCA codons. Translation will stall at TCG and TCA codons, and full-length protein will not be synthesized from essential genes within mobile genetic elements containing TCG and TCA codons. (D) Mobile genetic elements that encode their genes according to the canonical genetic code and also carry genes for tRNAs that decode TCG and TCA codons can be horizontally transferred into Syn61Δ3 cells. tRNAs encoded by mobile genetic elements can rescue the decoding of TCG and TCA codons within essential genes in the mobile genetic element to produce the correct protein. (E) Transfer and tRNA encoding of mobile genetic elements by conjugation; colony counts indicate successful transfection conjugates received from ~106 cells. The WT mobile genetic element (F WT) can be transferred into cells with the WT translation machinery (Syn61 WT), but cannot be transferred into cells lacking tRNAs that decode TCG and TCA (Syn61Δ3). The WT mobile genetic element (F (WT+serT)) that encodes tRNAs that decode TCG and TCA codons into serine can be transferred into both Syn61 WT and Syn61Δ3.
[0019] Figure 2 . Sense codon reassignment generates new genetic codes. (A) Total synthesis of a codon-compacted genome, followed by tRNA and release factor deletion, produced Syn61Δ3. Discovery of tRNAs that guide the incorporation of different natural amino codons. (B) Isoacceptor tRNAs for the indicated amino acids with anticodons changed to the Watson-Crick complement of TCG or TCA codons were introduced into cells in the indicated paired combinations. We used a GFP gene with a TCG or TCA codon at position 3 and electrospray ionization mass spectrometry (ESIMS) to read out the identity of the amino acid incorporated into each codon. When paired isoacceptors for different amino acids were used, each codon resulted in the specific incorporation of the amino acid attached to the Watson-Crick paired isoacceptor. The secondary peak measured in proline incorporation is due to incomplete cleavage of the N-terminal methionine. A complete list of discoveries and expected masses is provided in Data File S1. (C) 16 new genetic codes in which TCG and TCA codons were reassigned to Ala, His, Leu, and Pro.
[0020] Figure 3 Semantic orthogonality in genetic systems. (A) Relationship between TCG and TCA codons in genes that are cleaved by tRNA in decoders in cells with wild-type (WT) decoding and in Syn61Δ3. Ala CGA tRNA His UGA Decoding, and the corresponding protein sequence synthesized. Indicates the anticodon (decoder) of the tRNA that reads the TCG or TCA codon. Indicates the amino acid (aa) used by the tRNA. The gray color of the codon indicates that it decodes to serine and will produce the correct protein sequence. The yellow color of the codon indicates that it decodes to alanine and will produce the correct protein sequence. The green color of the codon indicates that it decodes to histidine and will produce the correct protein sequence. (B) Functional evaluation of SpecR WT (which uses the natural genetic code) and O-SpecR (TCG-Ala, TCA-His), O-SpecR is a codon compressed according to the Syn61 recoding scheme, and has an Ala codon replaced with TCG and a His codon replaced with TCA. In cells with WT translation machinery (Syn61 WT) and cells in which TCG decodes to Ala and TCA decodes to His (Syn61Δ3(tRNA CGA Ala , tRNA UGA His)) read the gene. Cells were spotted on agar plates in the presence or absence of spectinomycin and incubated overnight. Growth of the cells in the presence of spectinomycin indicates that the indicated SpecR gene is functional in the indicated strain.
[0021] Figure 4 .Orthogonal and mutually orthogonal horizontal gene transfer systems.
[0022] (A) Horizontal gene transfer between two bacterial strains is prohibited (grey dashed arrows), while horizontal gene transfer between cells sharing a common genetic code is possible (solid arrows). (B) Orthogonal horizontal transfer of mobile genetic elements. Colony counts indicate the number of transconjugants received from ~106 donor cells carrying the indicated mobile genetic elements. The WT mobile genetic element (F WT) was transferred into cells with the WT translation machinery (Syn61 WT), but not into cells in which TCG was redistributed to alanine and TCA to histidine (Syn61Δ3(tRNA CGA Ala , tRNA UGA His )). The orthogonal mobile genetic element (OF1) was transferred to Syn61Δ3 (tRNA CGA Ala , tRNA UGA His ) but not transferred into Syn61WT. (C) Mutually orthogonal horizontal genetic systems. Colony counts showed that from ~10 6 The orthogonal mobile genetic element (O-F2) was transferred only to cells in which TCG was redistributed to histidine and TCA was redistributed to alanine (Syn61Δ3(tRNA CGA His , tRNA UGA Ala )); O-F2 was not transferred to Syn61Δ3(tRNA CGA Ala , tRNA UGA His ) or Syn61 WT. Both F WT and O-F1 could not transfer (Syn61Δ3(tRNA CGA His , tRNA UGA Ala (D) Expands the experiment shown in (C) and demonstrates similar specificity of O-F3 and O-F4.
[0023] Figure 5.Orthogonal codon locking prevents invading codons. (A, B) Expected protein synthesis and horizontal gene transfer outcomes from mobile genetic elements and recipient cells with the indicated decoders, and codons in essential genes. (A) A WT mobile genetic element encoding tRNA that decodes TCG and TCA codons as Ser is transferred into cells in which TCG is reassigned to Ala and TCA is reassigned to His. Essential genes in WT mobile genetic elements containing TCG and TCA codons will be missynthesized, with each TCG and TCA codon in the gene randomly decoded as Ser or His / Ala. This is expected to attenuate horizontal gene transfer. (B) A WT mobile genetic element encoding tRNA that decodes TCG and TCA codons as serine is transferred into cells in which TCG is reassigned to Ala and TCA is reassigned to His. Essential genes in WT mobile genetic elements containing TCG and TCA codons will be missynthesized, with each TCG and TCA codon in the gene randomly decoded as Ser or His / Ala. In addition, essential genes in the host cell (where TCG is used to encode Ala and TCA is used to encode His) will be synthesized incorrectly. It is expected that this can ablate horizontal gene transfer. (C) Horizontal gene transfer is ablated in cells that use orthogonal genetic codes in the mobile genetic element and the essential genes of the recipient cell. F (WT + serT) was used as a mobile genetic element in all experiments. Colony counts showed successful transformation conjugates received from ~ 106 cells. The spectinomycin resistance gene variant (SpecR gene) in the recipient cell and the recipient cell was indicated. It is crucial to correctly read the indicated SpecR gene in the recipient cell by adding spectinomycin. (D) Encoding seryl-tRNA UGA The T4-like phage infected Syn61Δ3 but not cells carrying the orthogonal genetic code. Plaque counts showed that the number of cells obtained from the use of 1.1x10 10 PFU / mL (phage 12) and 7.5x10 9 The number of successfully replicating phage obtained by infection with 10 PFU / mL (phage 6). Cells contain the homologous spectinomycin resistance gene as in C; all experiments were performed in the presence of spectinomycin.
[0024] Figure 6 (Figure S1). Compressed genetic codes are non-orthogonal. Functional evaluation of wild-type (HygR WT) and codon-compressed hygromycin resistance (recHygR (ΔTCG, TCA)) genes in cells that decode all codons in the reading frame using complete tRNA complements (Syn61 WT, left panel) and cells in which the tRNA genes that decode TCG and TCA codons have been deleted (Syn61Δ3, right panel). Cells were spotted on agar plates in the presence or absence of hygromycin and incubated overnight. Growth of cells in the presence of hygromycin indicates that a given HygR gene is functional in a given strain.
[0025] Figure 7 (Figure S2). tRNA import breaks genetic isolation. The WT mobile genetic element (FWT) encoding the chloramphenicol resistance gene and the lux operon encoding the enzyme for synthesizing luciferin were conjugated from donor cells (Syn61 WT) into recipient cells (Syn61Δ2: a strain derived from Syn61Δ3 in which prfA was reintroduced via lambda red recombination), which contained the pSC101-Hyg plasmid encoding hygromycin resistance and the pKW20 plasmid (35). After conjugation, cells were selected on agar plates containing hygromycin and chloramphenicol, so that only cells containing the pSC101 vector and FWT survived. (A) Chemiluminescent images of the selection plate from the conjugation assay. The three colonies that survived on the selection plate after conjugation glowed due to luciferin production, indicating that they had received FWT. (B) Genotyping of two colonies picked from the selection plate (the third colony did not grow in liquid culture). Controls included Syn61WT (containing genomically encoded prfA, serT, and serU; these cells do not contain the pSC101-Hyg plasmid) and Syn61Δ3 (in which prfA, serT, and serU have been deleted from the genome; these cells contain the pSC101-Hyg plasmid). Genotyping of the pSC101-Hyg plasmid confirmed that the clones were recipient cells. Genotyping of the prfA locus revealed that both clones that survived selection contained the prfA gene at the endogenous genomic site. Genotyping of the serT locus revealed that both clones carried serT at the endogenous genomic site. Genotyping of the serU locus revealed that the clones did not carry serU at the endogenous genomic site. (C) Next-generation sequencing (NGS) sequence alignment of the recipient and donor strains against a reference genome. We observed that at the serT locus, the cloned sequence matched the donor reference (no gaps or insertions were observed when aligned with the Syn61 WT sequence; insertions were observed when aligned with the Syn61Δ3 sequence). However, at the serU locus, the cloned sequence matched the Syn61Δ3 sequence (insertions were observed when aligned with the Syn61 WT sequence. No gaps or insertions were observed when aligned with the Syn61Δ3 sequence). Vertical colored lines indicate mismatches between sequencing reads and the reference. Grey shows paired-end reads where the pairs are correctly oriented and positioned relative to each other. Reads shown in green / red indicate aberrant paired-end reads where the pairs are misoriented or mispositioned relative to each other. (D) Sequencing allowed us to define the origin of the genomic sequence of each clone; Syn61Δ3 (and therefore Syn61Δ2) contains mutations resulting from its evolution from Syn61 WT, which act as watermarks, allowing us to delineate whether the DNA sequence is derived from the donor (Syn61 WT) or recipient (Syn61Δ2) genome.We found that the majority of the genome was that of the recipient. However, within a ~400 kb stretch, multiple segments of the donor DNA integrated into the recipient genome. The genomic DNA integration pattern was unique for each sequenced clone, but both clones contained the serT locus.
[0026] FIG8 (FIG. S3). Isoacceptor tRNAs with altered anticodons are active and specific. (A) sfGFP-His6 was produced from the sfGFP3TCG or sfGFP3TCA genes in Syn61Δ3 cells containing the indicated isoacceptor tRNA chimeras with CGA or UGA anticodons. serT and serU decode both TCG and TCA codons, while the chimeric isoacceptor exhibits specificity for codons with Watson-Crick complementarity at all three bases of the codon-anticodon interaction. With the exception of proM with a UGA anticodon, the levels of GFP produced by all chimeric tRNAs based on their Watson-Crick complement codons were at least comparable to those produced from the WT GFP gene without TCG or TCA codons; this indicates that the chimeric isoacceptor tRNAs are efficient and specific. Protein yield of WT sfGFP was 17 mg / L. sfGFP-3-TCG / TCA expression was comparable to WT sfGFP from the isoacceptor. (B) The chimeric isoacceptor retains specificity for aminoacylation of the amino acid specified by the parent isoacceptor. The identity of the amino acid incorporated at position 3 of sfGFP in response to TCG or TCA in cells containing a chimeric tRNA with an anticodon that is the Watson-Crick complement of the codon at position 3 of sfGFP was confirmed by electrospray ionization mass spectrometry (ESI-MS). Only masses corresponding to the correct amino acid were detected. The secondary peak measured in proline incorporation is due to incomplete cleavage of the methionine at the N-terminus. Expected and actual masses of sfGFP-3TCG expression; expected mass sfGFP-3-Ser: 27755.13 Da, actual mass: 27756.00 Da. Expected mass sfGFP-3-Ala: 27739.13 Da, actual mass: 27740.60 Da. Expected mass sfGFP-3-His: 27805.19 Da, actual mass: 27806.20 Da. Expected mass of sfGFP-3-Leu: 27781.21 Da, actual mass: 27782.00 Da. Expected mass of sfGFP-3-Pro: 27765.17 Da, actual mass: 27765.40 Da. Expected and actual mass determination of sfGFP-3TCA expression: Expected mass of sfGFP-3-Ser: 27755.13 Da, actual mass: 27756.00 Da. Expected mass of sfGFP-3-Ala: 27739.13 Da, actual mass: 27741.20 Da. Expected mass of sfGFP-3-His: 27805.19 Da, actual mass: 27805.60 Da. Expected mass of sfGFP-3-Leu: 27781.21 Da, actual mass: 27782.00 Da.Expected mass of sfGFP-3-Pro: 27765.17 Da, actual mass: 27765.20 Da.
[0027] Figure 9 (Figure S4). Semantic orthogonality in genetic systems. (A) Genes written in the canonical genetic code (SpecR WT) and codon-reassigned spectinomycin resistance (O-SpecR(His-TCG, Ala-TCA), O-SpecR(Ala-TCG, Leu-TCA), O-SpecR(Leu-TCG, Leu-TCA), O-SpecR(Pro-TCG, Leu-TCA), O-SpecR(Ala-TCG, Ala-TCA), O-SpecR(Ala-TCG, Pro-TCA)) in cells with the canonical coding (Syn61 Functional evaluation was performed in cells expressing the indicated SpecR gene (WT) and in cells in which the TCG and TCA codons were reassigned according to the appropriate codons (Syn61Δ3(tRNACGAHis, tgNAUGAAla), Syn61Δ3(tRNACGAAla, tRNAUGALeu), Syn61Δ3(tRNACGALeu, tRNAUGALeu), Syn61Δ3(tRNACGAPro, tRNAUGALeu), Syn61Δ3(tRNACGAAla, tRNAUGAAla), Syn61Δ3(tRNACGAAla, tRNAUGAPro). Cells were spotted on agar plates in the presence or absence of spectinomycin and incubated overnight. Growth of the cells in the presence of spectinomycin indicates that the indicated SpecR gene is functional in the indicated strain. (B) Functional evaluation of HygR WT (which uses the native genetic code) and O-HygR (TCG-Ala, TCA-His), which is codon-constricted according to the Syn61 recoding scheme and has Ala codons replaced by TCG and His codons replaced by TCA. CGA Ala , tRNA UGA His ))'s cells read genes.
[0028] Figure 10 (Figure S5). Mutational landscape of the WT and codon-compressed genetic codes. Depiction of the mutational landscape of various amino acids in the WT and codon-compressed genetic codes. Amino acids showing the mutational landscape are highlighted (serine in gray, alanine in yellow, histidine in green, leucine in orange, and proline in blue). Amino acids with only one point mutation removed from the highlighted amino acids are connected by lines and labeled.
[0029] Figure 11 (Figure S6). Reconstructed genetic code alters mutational landscape. Depiction of the reconstructed genetic code mutational landscape created in this work. TCG reallocations are shown on the left. TCA reallocations are shown at the top. For each codon, serine and the reallocated amino acid are highlighted. Amino acids with only one point mutation removed from the highlighted amino acids are connected by lines and labeled.
[0030] Figure 12 . Original spectrum. (A) shows the integrated mass spectrum (pre-deconvolution) of sfGFP-3-TCG measurement (as Figure 2 (B) shows the intensity of each mass as total ion counts. (B) shows the integrated mass spectrum (pre-deconvolution) of sfGFP-3-TCA measurement (as shown in Figure 2 B). The intensity of each mass is shown as total ion counts.
[0031] Figure 13 . Fidelity of TCG TCA decoded by isoacceptor tRNA at position 11 of ubiquitin. Isoacceptor tRNAs for the indicated amino acids with the anticodon changed to the Watson-Crick complement of the TCG or TCA codon were introduced into Syn61Δ3 in the indicated paired combinations. We used ubiquitin genes with TCG or TCA codons at position 11 and electrospray ionization mass spectrometry (ESI-MS) to read out the identity of the amino acid incorporated into each codon. When paired isoacceptors of different amino acids were used, each codon resulted in the specific incorporation of the amino acid attached to the Watson-Crick paired isoacceptor. We calculated the specific incorporation of the amino acid attached to the Watson-Crick paired isoacceptor in the presence of tRNA CGA XXX and tRNA UGA YYY The lowest specificity for decoding TCG or TCA codons in both cases, where XXX and YYY are different amino acids (Methods). CGA Ala and tRNA UGA Leu We estimate that ≥78.2% Leu is incorporated at this codon; this may be due in part to the tRNA UGA Leu Misacylation of tRNA and / or CGA Ala For all other profiles, the specificity of decoding codons with correct versus incorrect anticodons ranged from ≥96% to ≥99.8%.
[0032] Figure 14.Sense codon reassignment does not produce detectable off-target incorporation at the TCT codon. (A) As indicated, isoacceptor tRNAs for the indicated amino acids with anticodons changed to the Watson-Crick complement of the TCG or TCA codons were introduced into cells in paired combinations. We used electrospray ionization mass spectrometry (ESI-MS) to read out the identity of the amino acids incorporated into the TCT codon at position 3 of GFP. The mass detected in the presence of all isoacceptor pairs corresponds to the incorporation of serine at the TCT codon. The expected mass of serine is 27755.13 Da; alanine: 27739.13 Da; histidine: 27805.19 Da; leucine 27781.21 Da; proline: 27765.17 Da. All measured masses (27754.5 ± 0.5 Da) correspond to the incorporation of serine at TCT. The fidelity measurement limits (methods) for these spectra were in the range of 97.4%-99.6%, and we did not observe peaks for amino acids other than serine. (B) shows the integrated mass spectrum (pre-deconvolution) of the sfGFP-3-TCT measurement (as in A). The intensity of each mass is shown as the total ion count.
[0033] Figure 15 .Sense codon reassignment does not produce detectable off-target incorporation at TCC codons. (A) As indicated, isoacceptor tRNAs for the indicated amino acids with anticodons changed to the Watson-Crick complement of TCG or TCA codons were introduced into cells in paired combinations. We used electrospray ionization mass spectrometry (ESI-MS) to read out the identity of the amino acid incorporated into the TCC codon at position 3 of GFP. The mass detected in the presence of all isoacceptor pairs corresponds to the incorporation of serine at the TCT codon. The expected mass of serine is 27755.13 Da; alanine: 27739.13 Da; histidine: 27805.19 Da; leucine 27781.21 Da; proline: 27765.17 Da. All measured masses (27754.5 ± 0.5 Da) correspond to the incorporation of serine at TCC. The limits of measurement fidelity (methods) for these spectra were in the 98.0%–99.5% range, and no peaks were observed for amino acid incorporation other than serine. (B) Shows the pre-deconvoluted composite mass spectrum of sfGFP-3-TCC (as in A). The intensity of each mass is shown as the total ion count.
[0034] Figure 16 . Raw spectrum of ubiquitin. (A) shows the integrated mass spectrum (pre-deconvolution) of Ub-11-TCG measurement (as Figure 13 (B) shows the integrated mass spectrum (pre-deconvolution) of the Ub-11-TCA measurement (as shown). Figure 13). The intensity of each mass is shown as total ion counts.
[0035] Figure 17 .MS-MS of amino acids incorporated at position 11 of ubiquitin (A) is due to the presence of tRNA UGA Ser Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCG gene. The y-ions are labeled in red; the b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The serine at position 5 of the peptide (marked with a green asterisk) confirms that tRNA UGA Ser Correct decoding of the TCG codon in Ub-11-TCG. (B) is due to the presence of tRNA UGA Ser Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCA gene. Y-ions are labeled in red; b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The serine at position 5 of the peptide (marked with a green asterisk) confirms that tRNA UGA Ser Correct decoding of the TCA codon in Ub-11-TCA.
[0036] Figure 18 .MS-MS of amino acids incorporated at position 11 of ubiquitin (A) is due to the presence of tRNA CGA Ala Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCG gene. Y-ions are labeled in red; b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. Alanine at position 5 of the peptide (marked with a green asterisk) confirms the presence of tRNA. CGA Ala Correct decoding of the TCG codon in Ub-11-TCG. (B) is due to the presence of tRNA UGA Ala Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCA gene. Y-ions are labeled in red; b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. Alanine at position 5 of the peptide (marked with a green asterisk) confirms the presence of tRNA. UGA Ala Correct decoding of the TCA codon in Ub-11-TCA.
[0037] Figure 19 .MS-MS of amino acids incorporated at position 11 of ubiquitin (A) is due to the presence of tRNA CGA HisOverlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCG gene. The y-ions are labeled in red; the b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The histidine at position 5 of the peptide (marked with a green asterisk) confirms the presence of tRNA. CGA His Correct decoding of the TCG codon in Ub-11-TCG. (B) is due to the presence of tRNA UGA His Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCA gene. The product ions are marked in red; the b-ions are marked in blue. The peptide sequence is shown at the bottom of each spectrum. The histidine at position 5 of the peptide (marked with a green asterisk) confirms the presence of tRNA. UGA His Correct decoding of the TCA codon in Ub-11-TCA.
[0038] Figure 20 .MS-MS of amino acids incorporated at position 11 of ubiquitin (A) is due to the presence of tRNA CGA Leu Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCG gene. Y-ions are labeled in red; b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The leucine at position 5 of the peptide (marked with a green asterisk) confirms the presence of tRNA. CGA Leu Correct decoding of the TCG codon in Ub-11-TCG. (B) is due to the presence of tRNA UGA Leu Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCA gene. Y-ions are labeled in red; b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The leucine at position 5 of the peptide (marked with a green asterisk) confirms the presence of tRNA. UGA Leu Correct decoding of the TCA codon in Ub-11-TCA.
[0039] Figure 21 .MS-MS of amino acids incorporated at position 11 of ubiquitin (A) is due to the presence of tRNA CGA Pro Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCG gene. The y-ions are labeled in red; the b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The proline at position 5 of the peptide (marked with a green asterisk) confirms that tRNA CGAPro Correct decoding of the TCG codon in Ub-11-TCG. (B) is due to the presence of tRNA UGA Pro Overlay of tandem mass spectrometry spectra of a peptide at position 11 of the ubiquitin protein after digestion of the ubiquitin protein expressed from the Ub-11-TCA gene. The y-ions are labeled in red; the b-ions are labeled in blue. The peptide sequence is shown at the bottom of each spectrum. The proline at position 5 of the peptide (marked with a green asterisk) confirms that tRNA UGA Pro Correct decoding of the TCA codon in Ub-11-TCA.
[0040] Figure 22 . Screening of Anticodon-Modified tRNAs (A) We expressed and purified ubiquitin-His6 with a TCG codon at position 11 (Ub-11-TCG-His6) in the presence of the indicated tRNAs with a CGA anticodon; expression was performed in Syn61Δ3 and purified by Ni-NTA chromatography. To assess the amino acid specificity of anticodon-modified tRNAs, we performed whole protein profiling. The expected and measured masses of each tRNA are indicated. The expected mass is the mass of the parental isoacceptor amino acid; in cases where the mass differs from this, the mass of the closest canonical amino acid incorporated is also shown as an additional expected mass. The following tRNAs incorporated amino acids that differed from the expected parental isoacceptor (or gave results that could not be unambiguously assigned to the amino acid of the parental isoacceptor) and were not considered further: AsnT, CysT, GlnV, LysQ, MetV, MetY, PheU, ValV, ValW. The following tRNAs incorporated the amino acids of the parental isoacceptor: ArgU, ArgX, ArgQ, ArgW, GltU, GlyU, HisR, ProK, ProL, ProM, ThrT, TrpT, TyrV. (B) For a subset of tRNAs incorporating the parental isoacceptor amino acids, we expressed Ub-11-TCG-His6 in the presence of the indicated tRNAs with a CGA anticodon; expression was performed in Syn61Δ3 and lysates were probed with anti-His6 after SDS-PAGE (we replaced ThrU with ThrT). From these experiments, we selected ProM and HisR as good candidates for further characterization. In control experiments without tRNA expression, we did not detect ubiquitin from Ub-11-TCG. SerU is the natural TCG decoder and serves as a positive control.
[0041] Figure 23Doubling time The doubling time of Syn61Δ3 was measured in 2xYT in the presence of different pairwise combinations of anticodon-modified tRNAs. The strains were either tRNA-free (-) or in the presence of serT (tRNA UGA Ser ) served as a control. Most anticodon-modified tRNAs resulted in no or modest changes in the doubling time of Syn61Δ3. To facilitate parallel doubling time measurements, the assay was performed in a 96-well format with a 200 μL volume (Methods). For comparison, when grown in shake flasks, the doubling time of Syn61Δ3 (no tRNA(-) control) was 49.77 ± 0.8 min.
[0042] Figure 24. Mutually orthogonal horizontal gene transfer system. Colony counts indicate successful transconjugants received from ~10<6 >donor cells carrying the indicated mobile genetic elements. The mobile genetic elements - (F(WT(TCG-Ser, TCA-Ser)), O-F1(TCG-Ala, TCA-His), O-F2(TCG-His, TCA-Ala), O-F3(TCG-Ala, TCA-Ala), O-F4(TCG-Ala, TCA-Pro)) - were transferred only to cells with a homologous reassignment of TCG and TCA codons (Syn61 WT(tRNA CGA Ser , tRNA UGA Ser )、Syn61Δ3(tRNA CGA Ala , tRNA UGA His )、Syn61Δ3(tRNA CGA His , tRNA UGA Ala )、Syn61Δ3(tRNA CGA Ala , tRNA UGA Ala )、Syn61Δ3(tRNA CGA Ala , tRNA UGA Pro No transposition of non-homologous cipher / decoder systems has been observed.
[0043] Figure 25 Orthogonal codon locking prevents horizontal gene transfer of the invading codon WT mobile genetic element (F(WT+serT)) by ablation of the mobile genetic element in the cells using the reconstructed genetic code and in the recipient cells. Colony counts showed that the number of cells that had undergone the reconstructed genetic code increased from 10 to 10. 6The successful transfection of the cells is indicated. The recipient cells and the hygromycin resistance gene variant (HygR gene) in the recipient cells are indicated. It is crucial that the indicated HygR gene is correctly read in the recipient cells by adding hygromycin.
[0044] Figure 26 Orthogonal codon locking prevents horizontal gene transfer of the invading codon WT mobile genetic element (F(WT+serT)) by ablation of the mobile genetic element in the cells using the reconstructed genetic code and in the recipient cells. Colony counts showed that the number of cells that had undergone the reconstructed genetic code increased from 10 to 10. 6 The successful transformation of the cells received by the recipient cells indicates the presence of the spectinomycin-resistant gene variant (HygR gene) in the recipient cells. It is crucial that the spectinomycin-resistant gene is correctly read in the recipient cells by adding spectinomycin.
[0045] Figure 27 Whole-genome sequencing of purified phages reveals seryl-tRNA genes. (A) We evaluated 31 phage enriched from environmental samples in the presence of spectinomycin in a codon-compacted strain (Syn61Δ3) and a strain with a remodeled and locked genetic code (Syn61Δ3(tRNACGA Ala , tRNAUGA His )O-SpecR (Ala-TCG, His-TCA)). Thirteen environmental samples formed plaques on Syn61A3, while none formed plaques on a strain with a reconstructed genetic code (Methods). Two separate phages (named 6 and 12) were sequentially purified from single plaques in Syn61WT (as described in the Methods). High-coverage NGS enabled de novo genome assembly of these phages. Their genome sizes were 164,924 and 167,593, respectively. Both phages belong to the genus Tequatrovirus (T4-like phages) and show >97% identity with Citrobacter phage ZZ23 (NC_054901) and Escherichia phage U115 (MZ753803), respectively. Both genomes contain a tRNA encoding Φ UGA Ser (B) The predicted secondary structure of the bacteriophage encoding seryl-tRNA (ΦtRNA UGA Ser ) and serT (endogenous E. coli tRNA UGA Ser All E. coli seryl-aaRS identity elements are present in ΦtRNA UGA SerThe sequence shown in the figure is SEQ ID NO: 69 (endogenous E. coli tRNA UGA Ser gene) and SEQ ID NO: 70 (ΦtRNA UGA Ser (C) Electron microscopy images of isolated phages 6 and 12. Both phages displayed characteristic Myoviridae morphology, consistent with the phage genome sequences.
[0046] Figure 28 tRNA from phages 6 and 12 UGA Ser We evaluated the ability of T4 phage to form plaques in different strains. Although T4 formed plaques in Syn61 WT, Syn61Δ3 was resistant to T4 plaque formation due to the lack of tRNA to decode TCG and TCA codons. T4 plaque formation was inhibited in Syn61Δ3 by expressing serT (encoding E. coli tRNA). UGA Ser ) or ΦtRNA encoded on the genomes of both phage 12 and phage 06 UGA Ser To rescue. Cell infection has ~10 6 to 10 3 PFU / mL.
[0047] Figure 29 .Genetic code reconstruction and code locking block viral replication (A) Encoding ΦtRNA on its genome UGA Ser The cloned pools of phages 12 and 06 formed plaques in both Syn61 WT and Syn61Δ3. Various strains with different reconstructed and locked genetic codes showed complete ablation of plaque formation. (B) Titration (~10 7 to 10 3 PFU / mL) of phages 6 and 12. No plaque formation was observed in strains with reconstructed and locked genetic codes.
[0048] Figure 30Phage replication is a more complex biological function than conjugation. Plaque formation from T-4-like phage infection is a more complex biological process than colony formation after successful horizontal gene transfer from conjugation. (A) Comparative genomic analysis of T-4-like phages (phage 06 and phage 12) and RK2 F plasmid. On the x-axis, size is expressed in kilobases. The positions where TCA codons occur are marked with vertical yellow lines (top). The positions where TCG codons occur are marked with vertical red lines (bottom). (B) Codon usage in mobile genetic elements. The total frequency of target codons (TCA and TCG) is represented on the x-axis. TCA codon frequency is represented in yellow, and TCG codon frequency is represented in red. (C) In order to form plaques, phage particles first need to attach to bacterial cells and inject their DNA into the cytosol. Subsequently, genes from the phage genome are transcribed and translated by the host machinery that produces viral proteins. In addition, the phage genome replicates within the cell. Finally, these components need to mature into functional phage particles, which escape from the cell and infect neighboring cells. This complex process requires many structural and regulatory proteins to be correctly expressed from the phage genome. (D) In order to form colonies from recipient cells carrying horizontally transferred conjugative plasmids, donor cells first need to attach to the recipient. Subsequently, the F plasmid is transferred through the mating channel and recycled within the recipient cell. The process of attachment, transfer and recycling depends only on the proteins expressed in the donor cell. Proteins are then expressed from the conjugative elements that allow them to replicate and separate during cell division. The division of the recipient cell carrying the transferred plasmid results in colony formation without the need for further conjugation events.
[0049] Figure 31 .Codon locking ensures the stability of the reconstructed genetic code. (A) Cells were passaged every 12 h and evaluated for the presence of the reconstructed code. Although cells with codon locking retained the genetic code in all cases and the cells demonstrated the stability of the codon reconstructed code, in cells without codon locking, the reconstructed code was not stable. Cells contained the codon-compressed SpecR resistance gene (unlocked: -) or the homologous SpecR gene (SpecR TCG-Ala, TCA-His / SpecR TCG-Leu, TCA-Leu) (locked: +); all experiments were performed in the presence of spectinomycin. (B) Encoding seryl-tRNA UGA Phages infected cells with unstable genetic codes but not cells with stably reconstructed codes. Cells from the last point in the time course (A) (with and without code lock) were infected with a T4-like phage (phage 06 / 12). Plaque counts indicated that ~5x10 8 PFU (phage 12) and ~1x10 7The number of successfully replicating phage obtained is 6 PFU (phage 6). Cells contained either a codon-compressed SpecR resistance or a homologous SpecR gene (as in (A)); all experiments were performed in the presence of spectinomycin.
[0050] Figure 32 Structure of the major capsid protein gp23. (A) Protein structure of the major capsid protein gp23 from T4 bacteriophage. Three serine residues are present (highlighted in orange) that are encoded by TCAs and are therefore subject to ambiguous decoding in cells with a reconstructed genetic code. (B) Binding interface of the major capsid protein gp23; the three subunits of gp23 are shown, which are part of the hexameric capsid subunit. The serine residue on the central subunit (grey) encoded by TCAs is shown in orange.
[0051] Figure 33 Phage propagation assay. Encoding seryl-tRNA (tRNA UGA Ser ) successfully infected Syn61Δ3 but not cells with the reconstructed and locked genetic code. Plaque counts indicate the number of phage particles in a 7.5 μL volume after 24 h of propagation in culture containing the indicated cells.
[0052] Figure 34 The effects of code remodeling and code locking on horizontal gene transfer patterns. UGA Ser ) successfully infected Syn61Δ3 but not cells with the reconstructed genetic code. Plaque counts showed that 1.1×10 10 The number of successfully replicating phages obtained was 7.5×109 plaque forming units (PFU) / ml (phage 12) and 7.5×109 PFU / ml (phage 6). DETAILED DESCRIPTION
[0053] Password lock
[0054] It is necessary to prevent mobile genetic elements such as viruses from contaminating cells. For example, industrial-scale fermentation of bacteria for commercial product production may be contaminated by mobile genetic elements such as viruses. This can lead to economic losses and can destroy important supply requests. There are existing methods for protecting cells from such contamination (see WO2020 / 229592A1 or Robertson et al., Science, 4Jun 2021, Vol372, Issue 6546, pp.1057-1062, both of which are incorporated herein by reference), but the inventors show in this article that there is still a risk of mobile genetic elements from comprising tRNA. Attempts have been made to reduce the risk from this mobile genetic element (see Nyerges et al., " Swapped genetic code blocks viral infections and gene transfer ", https: / / doi.org / 10.1101 / 2022.07.08.499367), but technology is still needed to make cells resistant to mobile genetic elements comprising tRNA.
[0055] Provided herein are "codon-locked" cells. The genomes of these cells have been recoded to reduce or remove, for example, at least one type of sense codon, which then allows for the removal of endogenous cognate tRNAs, as they are now dispensable to the cell (see Robertson et al., Science, 4 Jun 2021, Vol 372, Issue 6546, pp. 1057-106). The inventors have found that comprising tRNAs specific for the removed sense codons but with amino acids that are not naturally associated with the sense codons reduces, but does not eliminate, the risk of contamination by mobile genetic elements containing the relevant tRNAs (see the disclosure of Figure 5 ). Others have noted this (see Nyerges et al.). The inventors have overcome this problem by generating cells containing a gene required for cell viability, wherein the gene has been recoded so that the sense codon is reallocated to an amino acid that is different from the canonical genetic code. The inventors have surprisingly found that inclusion of such a gene reduces or eliminates horizontal gene transfer from mobile genetic elements that utilize the canonical genetic code or a genetic code that does not match the target cell. The inventors have also found that this approach results in the maintenance of the exogenous tRNA that has been introduced into the codon remodeling. Therefore, this approach also results in the maintenance of resistance to horizontal gene transfer or mobile genetic elements.
[0056] Thus, in a first aspect, a cell is provided that: comprises a genome in which at least a first type of sense codon has been recoded such that a first endogenous tRNA is dispensable; does not express the first endogenous tRNA; expresses a first modified tRNA capable of decoding the first type of sense codon, wherein the first modified tRNA carries a first amino acid that is not a naturally cognate amino acid for the first type of sense codon; and comprises a gene required for viability, wherein the gene comprises at least one occurrence of the first type of sense codon, and the cell is viable when the first type of sense codon in the gene decodes to the first amino acid.
[0057] The cell may have increased resistance to horizontal gene transfer or mobile genetic elements, as discussed in the following section. Thus, a fourth aspect provided herein is a cell having increased resistance to horizontal gene transfer or mobile genetic elements, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid that is not associated with a sense codon in the canonical genetic code, and the cell comprises a gene required for viability that is functional when decoded according to the reassigned genetic code and is not functional when decoded according to the canonical genetic code. The gene may be required for viability alone or in combination with other genes.
[0058] Cells that have been modified in the manner described above exhibit improved maintenance of resistance to horizontal gene transfer or mobile genetic elements, as described in Example 7. Thus, the increased resistance can be resistance maintained over a longer period of time compared to cell cultures that do not contain code-locked bacteria.
[0059] The gene required for vigor can be an exogenous gene. For example, the gene can be a gene commonly used as a positive selectable marker. In some instances, the gene is an antibiotic resistance gene. Illustrative embodiments include a spectinomycin resistance gene or a hygromycin resistance gene.
[0060] In other examples, a gene required for viability may be an essential gene within the genome of a cell. As used herein, a gene is "essential" if its product is required for cell viability. For example, a gene is considered essential if preventing the expression of a functional form of the protein encoded by the gene would result in non-viability of the cell.
[0061] A gene required for activity may contain at least one reassigned codon, wherein a mutation of the corresponding residue in the translation product may result in loss of function. In particular, the reassigned codon may be positioned so that decoding the codon according to the canonical genetic code results in a loss of function in the product. For example, if a cell contains a tRNA capable of decoding a codon that is normally associated with serine but carries an alanine, and the gene contains said codon, wherein alanine would be present in the natural product, the product of the gene may be a product in which serine would be non-functional at said position. The examples of specific amino acids described above are purely illustrative, and any one may be used. In particular, Figure 2 Any reassignment scheme in C. Thus, a cell can comprise a gene required for viability, wherein the gene comprises at least one occurrence of a first type of sense codon, and the cell is non-viable when the first type of sense codon is decoded according to the canonical genetic code.
[0062] Genes required for vitality may include multiple reallocated codons. Multiple reallocated codons may be positioned individually or cumulatively so that products containing non-reallocated amino acids (as described in the preceding paragraphs) will be non-functional. Thus, a cell may include a gene required for vitality, wherein the gene includes multiple occurrences of a first type of sense codon, and when the occurrences are decoded according to the canonical genetic code, the cell is non-viable. If decoded according to the canonical genetic code, at least one occurrence of the first type of sense codon in a gene required for vitality may at least partially contribute to loss of vitality, and may be combined with other features (such as other reallocated codons or other types of reallocated codons) to contribute to complete loss of vitality. In some instances, multiple reallocated codons (possibly of multiple types) may be present in a gene required for vitality, or multiple genes required for vitality may be present. If decoded according to the canonical genetic code, any individual instance of a reallocated codon may at least partially contribute to loss of vitality, and complete loss of vitality may be due to the effects of translating multiple reallocated codons according to the canonical genetic code.
[0063] Cells of the present disclosure may comprise more than one gene required for viability comprising at least one reassigned codon.
[0064] The cells of the present disclosure may comprise a genome that has been recoded with respect to sense codons of the second type.
[0065] In some embodiments, the genome of the cell is recoded such that the first endogenous tRNA is dispensable and the second endogenous tRNA is dispensable. The cell may not express or contain the first or second endogenous tRNA. In an example, the cell expresses or contains a second modified tRNA that is capable of decoding a second type of sense codon. The second modified tRNA carries a second amino acid, and the second amino acid is not the natural cognate amino acid for the second type of sense codon.
[0066] A gene required for viability can comprise at least one occurrence of a second type of sense codon, wherein when the second type of sense codon in the gene decodes to a second amino acid, the cell is viable. The gene can be the same gene required for viability and comprise a first type of sense codon, or it can be a different gene.
[0067] In some instances, when the second type of sense codons in genes required for vitality are decoded according to the canonical genetic code, the cell is non-viable. A gene may comprise multiple occurrences of the second type of sense codons, and when the occurrences are decoded according to the canonical genetic code, the cell may be non-viable. If decoded according to the canonical genetic code, at least one occurrence of the second type of sense codons in genes required for vitality may contribute at least in part to loss of vitality, and may contribute to complete loss of vitality in combination with other features (such as other reassigned codons or other types of reassigned codons). Complete loss of vitality may be due to the effects of translating multiple reassigned codons according to the canonical genetic code.
[0068] The cell may comprise at least one gene required for viability comprising a first type of sense codon and at least one different gene required for viability comprising a second type of sense codon. The cell may comprise a gene required for viability comprising both the first and second types of sense codons. Combinations of genes required for viability are also possible, as well as any combination comprising reassigned codons.
[0069] Cells of the present disclosure may be viable when genes are decoded according to the reassigned genetic code, but may be non-viable when genes are decoded at least in part according to the canonical genetic code.
[0070] Modified tRNA is derived from a tRNA (which may be a naturally occurring tRNA) that has been altered so that a codon can be decoded into an amino acid that the codon is not associated with in the canonical genetic code. For example, the anticodon residue of a tRNA can be substituted so that the tRNA has a different codon assignment, and such a tRNA can be referred to as an anticodon-swapped tRNA. Alternatively, the tRNA can be made to carry an amino acid that is not naturally associated with it, thereby enabling the tRNA to decode a codon into an amino acid that the codon is not associated with in the canonical genetic code. Modified tRNAs can also be modified in other ways, such as by adding additional sequences. The modified tRNA can carry a natural amino acid that the parent tRNA naturally associates with.
[0071] The modified tRNA can be derived from a naturally occurring tRNA (which can be referred to as a parent tRNA). For example, the modified tRNA can be derived from a tRNA that is endogenous to the cell in question. The modified tRNA can be derived from an isoacceptor tRNA for a specific amino acid in the cell. For example, if the cell is E. coli, the modified tRNA can be derived from an E. coli tRNA that is an isoacceptor for the first or second amino acid. The modified tRNA can be derived from a naturally occurring tRNA found in a mobile genetic element, such as a viral tRNA. The modified tRNA can comprise an identity element that is recognized by an aminoacyl-tRNA synthetase that is endogenous to the cell. The modified tRNA can retain the identity element of the parent tRNA.
[0072] The inventors have proved in this article that the episome encoding the translation machinery composition can be necessary for cells, and the composition is required for the gene required for the translation vigor of the genetic code redistributed. Therefore, these episomes are stably maintained by the cell of the first aspect. Therefore, in one embodiment, the tRNA of the first modification can be encoded by the episome in the cell, such as a plasmid. The episome can further comprise other genes that the expectation stably maintains.
[0073] In a specific example, the cell is E. coli and comprises a genome that has been recoded with respect to first and second types of sense codons (e.g., TCA and TCG). The first modified tRNA can be an E. coli isoacceptor tRNA for a first amino acid (e.g., alanine) that has been altered to include an anticodon that is complementary to the first type of sense codon (e.g., TCA). The second modified tRNA can be an E. coli isoacceptor tRNA for a second amino acid (e.g., histidine) that has been altered to include an anticodon that is complementary to the first sense codon (e.g., TCG).
[0074] In some examples, the first modified tRNA cannot decode the second type of sense codon, and / or the second modified tRNA cannot decode the first type of sense codon. In further examples, the first modified tRNA cannot decode any type of codon other than the first type of sense codon, and / or the second modified tRNA cannot decode any type of codon other than the second type of sense codon.
[0075] The first and second amino acids can be the same amino acid, for example, they can both be alanine, histidine, leucine, or proline. In other examples, the first and second amino acids can be different. For example, one can be alanine and the other can be leucine, etc. Some exemplary reassignment schemes are shown in Figure 2 C.
[0076] Cells can be described as containing tRNA XXX Xaa (where XXX is the type of sense codon and Xaa is the charged amino acid). Thus, cells utilizing a particular reallocation scheme can be defined as containing the relevant tRNA. For example, in a particular embodiment, cells utilizing TCG reallocation to alanine and TCA reallocation to histidine contain tRNA CGA Ala and tRNA UGA His ; and cells that utilize TCG to redistribute to histidine and TCA to redistribute to alanine contain tRNA CGA His and tRNA UGA Ala .
[0077] The first and / or second amino acid can be a naturally occurring amino acid. The naturally occurring amino acid can be any natural proteinogenic amino acid. "Natural proteinogenic amino acid" is any of L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan and L-tyrosine, L-pyrrolysine and L-selenocysteine. The naturally occurring amino acid can be a canonical amino acid. The "canonical amino acids" are any one of L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan and L-tyrosine.
[0078] The cell of the first aspect may be of any species or type as disclosed herein. For example, the cell may be a bacterial cell having a genome recoded according to the codons TCA and TCG, which lacks tRNA Ser UGA and tRNA Ser CGA , and wherein TCA and TCG have been redistributed. The genome of the cell may have been recoded in any manner as discussed herein. The redistribution scheme of the cell of the first aspect may be any scheme disclosed herein, for example Figure 2 One of the options shown in C.
[0079] The cell may be Syn61, a strain derived from Syn61, or recoded in the same manner as Syn61. The cell may be Syn61Δ3, a strain derived from Syn61Δ3, or may be modified in the same manner as Syn61Δ3.
[0080] Features of the first aspect related to "code lock" can be applied to the recoding schemes of the second or third aspects. Thus, any features of the first, second, and third aspects can be combined and are not mutually exclusive. Features of the second and third aspects, such as tRNA and coding schemes, can be applied to the first aspect.
[0081] The cells of the first aspect may have increased resistance to mobile genetic elements. The cells of the first aspect may have improved maintenance of resistance to mobile genetic elements; for example, resistance may be maintained for a longer period of time in cell culture when compared to a control culture that does not contain code-locked cells.
[0082] Orthogonal coding scheme
[0083] The use of genetically modified organisms is increasing, and there is a need to limit the transfer of genetic information from these organisms to native organisms. The inventors herein provide orthogonal encoding schemes that prevent the transfer of genetic information to native organisms or to organisms that utilize alternative orthogonal encoding schemes. For example, mobile genetic elements utilizing one of the orthogonal encoding schemes cannot be transferred to or expressed by native organisms.
[0084] Others have attempted to generate synthetic genetic information to prevent horizontal gene transfer (Nyerges et al., "Swapped genetic code blocks viral infections and gene transfer", https: / / doi.org / 10.1101 / 2022.07.08.499367). However, the inventors herein provide a screening method that allows the development of active and specific tRNAs (see Figure 8 (Figure S3)). As the skilled artisan will appreciate, as disclosed in the Examples, this screening is applicable to codons other than TCA and TCG and allows the development of active and specific tRNAs that do not decode off-target codons.
[0085] Thus, in a second aspect, a cell is provided which: comprises a genome wherein a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable; does not express the first endogenous tRNA and the second endogenous tRNA; expresses a first anticodon-exchanged tRNA derived from a naturally occurring first parent tRNA, wherein the first anticodon-exchanged tRNA carries a first amino acid and the first parent tRNA is an isoacceptor for the first amino acid, and wherein the first amino acid is not a naturally cognate amino acid for the first type of sense codon; and expresses a second anticodon-exchanged tRNA derived from a naturally occurring second parent tRNA, wherein the second anticodon-exchanged tRNA carries a second amino acid and the second parent tRNA is an isoacceptor for the second amino acid, and wherein the second amino acid is not a naturally cognate amino acid for the second type of sense codon; wherein the first and / or second modified tRNA is incapable of decoding any type of codon other than the first type of sense codon and / or the second type of sense codon.
[0086] In certain embodiments, the tRNA does not decode a particular codon when the misincorporation rate is undetectable or too low to affect the fitness of the cell in the screening methods disclosed herein. Thus, the tRNA of the second aspect may have no detectable misincorporation rate of non-target codons, or may have no relevant misincorporation rate given the size of the relevant host genome. In one embodiment, the tRNA of the second aspect is as active and specific as the tRNA exemplified in Example 3, i.e., the tRNA CGA Ala tRNA UGA Ala tRNA CGA His tRNA UGA His tRNA CGALeu tRNA UGA Leu and tRNA CGA Leu tRNA UGA Leu Any of the .
[0087] An anticodon-exchanged tRNA is a tRNA in which the residue of the anticodon is substituted so that the tRNA has a different codon specificity. The anticodon-exchanged tRNA can also be modified in other ways, for example, additional sequences can be added.
[0088] The present inventors have surprisingly found that sense codons that canonically encode the same amino acid and are canonically decoded by the same tRNA or overlapping tRNAs due to wobble base pairing can be used to encode multiple alternative amino acids. It is expected that such sense codons only allow a single reallocation. The inventors herein provide a screening method that allows the development of tRNAs with the desired activity and specificity. This discovery is advantageous because in an exemplary organism with, for example, two reallocated serine codons, the inventors were able to generate many different orthogonal codons. For example, the invention has generated 16 reconstructed codons from only two reallocated sense codons and four amino acids.
[0089] Thus, in some instances, due to wobble base pairing, the first and second types of sense codons will be canonically decoded by the same tRNA or overlapping tRNAs. The first anticodon-exchanged tRNA may not be able to decode any codon type other than the first type of sense codon, and the second anticodon-exchanged tRNA may not be able to decode any codon type other than the second type of sense codon. This allows the use of the first and second types of sense codons to encode two different amino acids without misincorporation.
[0090] The first type of sense codon and the second type of sense codon can belong to the formula XXN. This means that the first and second bases are the same, while the third base is different. In an example, the tRNA of the first anticodon exchange cannot decode the second type of sense codon, and the tRNA of the second anticodon exchange cannot decode the first type of sense codon.
[0091] In some examples, the first anticodon-exchanged tRNA cannot decode any codon type except the first type of sense codons and / or the second anticodon-exchanged tRNA cannot decode any codon type except the second type of sense codons.
[0092] In a specific embodiment, the tRNA of the first anticodon exchange does not decode TCC or TCT codons, and the tRNA of the second anticodon exchange does not decode TCC or TCT codons. For example, when the first or second type of recoded sense codon is TCA or TCG, this can be advantageous because misincorporation of TCC or TCT codons can reduce the adaptability of the cell. For example, some E. coli genomes contain 9,999 TCC codons and 9,566 TCT codons, and therefore misincorporation can affect adaptability. In a specific embodiment, when the misincorporation rate is undetectable in the screening method disclosed herein or is too low to affect the adaptability of the cell, the tRNA does not decode a specific codon. Therefore, the tRNA may have no detectable misincorporation rate at TCC or TCT, or may have no relevant misincorporation rate considering the size of the relevant host genome. In one embodiment, the misincorporation rate of the tRNA at TCC or TCT is no higher than that of the tRNA as shown in Example 3. CGA Ala tRNA UGA Ala tRNA CGA His tRNA UGA His tRNA CGA Leu tRNA UGA Leu and tRNA CGA Leu tRNA UGA Leu Any of the .
[0093] The anticodon-exchanged tRNA of the second aspect carries the natural amino acid with which it is naturally associated. Thus, the anticodon-exchanged tRNA carries the same amino acid as the parent tRNA from which it is derived. The anticodon-exchanged tRNA may comprise an identity element that is recognized by an aminoacyl-tRNA synthetase endogenous to the cell.
[0094] The anticodon-swapped tRNA can be derived from a tRNA that is endogenous to the cell of interest. The anticodon-swapped tRNA can be derived from an isoacceptor tRNA for a specific amino acid within the cell. For example, if the cell is E. coli, the anticodon-swapped tRNA can be derived from an E. coli tRNA that serves as an isoacceptor for the first or second amino acid. The anticodon-swapped tRNA can be derived from naturally occurring tRNAs found in mobile genetic elements, such as viral tRNAs. The anticodon-swapped tRNA can retain the identity elements of the parental tRNA.
[0095] In an example, the first and second types of sense codons can both canonically encode serine, can both canonically encode alanine, or can both canonically encode leucine.
[0096] In certain embodiments, the first type of sense codon is TCA and the second type of sense codon is TCG.
[0097] The first and / or second type of sense codons can be reassigned to any natural proteinogenic amino acid or canonical amino acid. In an illustrative embodiment, the first type of sense codon, such as a canonical serine codon, can be reassigned to one of alanine, histidine, leucine, and proline, while the second type of sense codon, such as a canonical serine codon, can be reassigned to one of alanine, histidine, leucine, and proline.
[0098] In certain examples, TCA can be reallocated to any non-serine naturally occurring proteinogenic amino acid or canonical amino acid and / or TCG can be reallocated to any non-serine naturally occurring proteinogenic amino acid or canonical amino acid. In further illustrative embodiments, TCA can be reallocated to one of alanine, histidine, leucine, and proline, and TCG can be reallocated to one of alanine, histidine, leucine, and proline. In some examples, the reallocation scheme is as follows: Figure 2 As shown in C.
[0099] In other examples, the first amino acid and the second amino acid are different types of amino acids. The following is an exemplary reallocation scheme:
[0100] TCG is redistributed to alanine and TCA is redistributed to histidine
[0101] TCG redistributes to alanine and TCA redistributes to leucine
[0102] TCG is redistributed to alanine and TCA is redistributed to proline
[0103] TCG is redistributed to histidine and TCA is redistributed to alanine
[0104] TCG redistributes to histidine and TCA redistributes to leucine
[0105] TCG is redistributed to histidine and TCA is redistributed to proline
[0106] TCG redistributes to leucine and TCA redistributes to alanine
[0107] TCG redistributes to leucine and TCA redistributes to histidine
[0108] TCG is redistributed to leucine and TCA is redistributed to proline
[0109] TCG is redistributed to proline and TCA is redistributed to alanine
[0110] TCG is redistributed to proline and TCA is redistributed to histidine
[0111] TCG is redistributed to proline and TCA is redistributed to leucine.
[0112] The tRNA in which the first or second anticodon is exchanged can be derived from a parent tRNA encoded by ArgQ, ArgU, GltU, HisR, ProK, ProL, ProM, TrpT, ThrU, ThrT, TyrU, TyrV, AlaT or LeuQ. Thus, the tRNA in which the first or second anticodon is exchanged can be encoded by any of the genes described, wherein the anticodon has been modified so that the tRNA recognizes a type of sense codon that is not canonically associated with the amino acid carried by the parent tRNA. In a specific example, the tRNA in which the first or second anticodon is exchanged can be derived from a parent tRNA encoded by HisR, ProM, AlaT or LeuQ. The gene encoding the tRNA can be derived from Escherichia coli.
[0113] In some examples, the first and second anticodon exchanged tRNA is derived from a parent tRNA encoded by: ArgQ, ArgU, GltU, HisR, ProK, ProL, ProM, TrpT, ThrU, ThrT, TyrU, TyrV, AlaT, and LeuQ. In some examples, the first and second anticodon exchanged tRNA is derived from a parent tRNA encoded by: HisR, ProM, AlaT, and LeuQ.
[0114] In some instances, except for the anticodon, the ArgQ, ArgU, GltU, HisR, ProK, ProL, ProM, TrpT, ThrU, ThrT, TyrU, TyrV, AlaT or LeuQ gene is unmodified. In other instances, the gene may comprise additional sequences or be truncated. In specific instances, the encoded tRNA may comprise an identity element of the parent tRNA. In other instances, the gene may comprise one or more modifications, wherein the encoded tRNA is still functional. These genes may comprise 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, 4 or less, 3 or less, 2 or less, 1 or no substitutions, additions or replacements.
[0115] In some instances, the anticodon is exchanged to UGA or CGA. In some instances, the first anticodon-exchanged tRNA is exchanged to UGA and the second anticodon-exchanged tRNA is exchanged to CGA.
[0116] In a specific example, the tRNA derived from ArgQ is according to SEQ ID NO: 43 or 44, the tRNA derived from GltU is according to SEQ ID NO: 19 or 20, the tRNA derived from HisR is according to SEQ ID NO: 49 or 50, the tRNA derived from ProK is according to SEQ ID NO: 55 or 56, the tRNA derived from ProL is according to SEQ ID NO: 57 or 58, the tRNA derived from ProM is according to SEQ ID NO: 59 or 60, the tRNA derived from TrpT is according to SEQ ID NO: 61 or 62, the tRNA derived from ThrU is according to SEQ ID NO: 25 or 26, the tRNA derived from ThrT is according to SEQ ID NO: 23 or 24, the tRNA derived from TyrV is according to SEQ ID NO: 63 or 64, the tRNA derived from AlaT is according to SEQ ID NO: 65 or 66, and the tRNA derived from LeuQ is according to SEQ ID NO: ID NO: 67 or 68. Any of these sequences may contain one or more modifications, wherein the encoded tRNA is still functional. Any of these sequences may contain 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, 1, or no substitutions, additions, or replacements. Modifications may be in regions that do not encode identity elements. Any of these sequences may be modified to encode a different anticodon. The alternative anticodon may not be the naturally associated anticodon.
[0117] The cell of the second aspect may be of any species or type as disclosed herein. For example, the cell may be a bacterial cell having a genome recoded according to the codons TCA and TCG, which lacks tRNA Ser UGA and tRNA Ser CGA The genome of the cell may have been recoded in any manner as discussed herein. The cell may be Syn61, a strain derived from Syn61, or recoded in the same manner as Syn61. The cell may be Syn61Δ3, a strain derived from Syn61Δ3, or may be modified in the same manner as Syn61Δ3.
[0118] Redistribution schemes can vary in their efficiency. For example, Figure 4C compared two different genetic encoding schemes and noted differences in colony counts. Reassignment schemes can affect cell fitness. Therefore, the contribution of validated reassignment schemes is a valuable one.
[0119] Thus, in a third aspect, a cell is provided, comprising: a genome wherein a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable; the first endogenous tRNA and the second endogenous tRNA are not expressed; a first modified tRNA capable of decoding the first type of sense codon is expressed, wherein the first modified tRNA carries a first amino acid that is not a naturally cognate amino acid for the first type of sense codon; and a second modified tRNA capable of decoding the second type of sense codon is expressed, wherein the second modified tRNA carries a second amino acid that is not a naturally cognate amino acid for the second type of sense codon; wherein: i) the first amino acid is alanine and the second amino acid is alanine; ii) the first amino acid is alanine and the second amino acid is histidine; iii) the first amino acid The first amino acid is alanine and the second amino acid is leucine; iv) the first amino acid is alanine and the second amino acid is proline; v) the first amino acid is histidine and the second amino acid is alanine; vi) the first amino acid is histidine and the second amino acid is histidine; vii) the first amino acid is histidine and the second amino acid is leucine; viiii) the first amino acid is histidine and the second amino acid is proline; ix) the first amino acid is leucine and the second amino acid is alanine; x) the first amino acid is leucine and the second amino acid is histidine; xi) the first amino acid is leucine and the second amino acid is proline; xii) the first amino acid is proline and the second amino acid is alanine; xiii) the first amino acid is proline and the second amino acid is histidine; xiv) the first amino acid is proline and the second amino acid is leucine; or xv) the first amino acid is proline and the second amino acid is proline.
[0120] The modified tRNA may be as discussed for the first or second aspects. In particular, the first modified tRNA may not be able to decode the second type of sense codon and / or the second modified tRNA may not be able to decode the first type of sense codon. In some instances, the first modified tRNA cannot decode any type of codon other than the first type of sense codon, and / or the second modified tRNA cannot decode any type of codon other than the second type of sense codon. This high specificity is achieved for the first time by the screening methods disclosed herein.
[0121] Modified tRNA is derived from tRNA (which may be naturally occurring tRNA), which has been changed so that a codon can be decoded into an amino acid that the codon is not associated with in the canonical genetic code. For example, the anticodon residue of a tRNA can be substituted so that the tRNA has a different codon assignment, and such tRNA can be referred to as an anticodon-exchanged tRNA. Modified tRNA can also be modified in other ways, such as by adding another sequence. The modified tRNA can carry the natural amino acid that the parent tRNA naturally associates with. The modified tRNA can be derived from a naturally occurring tRNA (which can be referred to as a parent tRNA). For example, the modified tRNA can be derived from a tRNA that is endogenous to the relevant cell. The modified tRNA can be derived from an isoacceptor tRNA for a specific amino acid in the cell. For example, if the cell is Escherichia coli, the modified tRNA can be derived from an Escherichia coli tRNA that is an isoacceptor for the first or second amino acid. The modified tRNA can be derived from a naturally occurring tRNA found in a mobile genetic element, such as a viral tRNA. The modified tRNA may comprise an identity element that is recognized by an aminoacyl-tRNA synthetase endogenous to the cell.The modified tRNA may retain the identity element of the parent tRNA.
[0122] The recoding scheme may be any as discussed herein. In a particular embodiment, the first type of sense codon is TCA and the second type of sense codon is TCG.
[0123] The cell of the third aspect may be of any species or type as disclosed herein. For example, the cell may be a bacterial cell having a genome recoded according to the codons TCA and TCG, which lacks tRNA Ser UGA and tRNA Ser CGA The cell may be Syn61, a strain derived from Syn61, or recoded in the same manner as Syn61. The cell may be Syn61Δ3, a strain derived from Syn61Δ3, or may be modified in the same manner as Syn61Δ3.
[0124] Kits containing cells utilizing mutually orthogonal encoding schemes
[0125] The inventors discovered that cells using a first orthogonal code can be orthogonal to cells using a second orthogonal code (see Figure 4 C) Such cells can therefore coexist with each other, and optionally with cells utilizing a canonical genetic code, and horizontal gene transfer will not be able to occur between cells that do not belong to the same coding scheme.
[0126] Thus, in a sixth aspect of the present invention, there is provided a kit comprising a first cell re-encoded according to a first orthogonal encoding scheme and a second cell re-encoded according to a second orthogonal encoding scheme, wherein the first and second encoding schemes are mutually orthogonal.
[0127] In one example, the kit comprises a first cell of the first, second or third aspect and a second cell of the first, second or third aspect, wherein the first and second cells utilize mutually orthogonal encoding schemes.
[0128] In some examples, the first and / or second orthogonal genetic encoding scheme is any of the ones disclosed herein, such as Figure 2 Any orthogonal genetic code of C.
[0129] In one example, the kit may include a first cell, which may be a cell cultured using Figure 2 The second cell of the kit, which may be used, is a bacterial cell, such as E. coli, of the redistribution scheme shown in C. Figure 2 The different redistribution schemes shown in C are bacterial cells, such as E. coli. In some examples, these cells are based on or derived from Syn61.
[0130] In certain embodiments, the first cell comprises tRNA CGA Ala and tRNA UGA His or containing tRNA CGA His and tRNA UGA Ala .
[0131] The kit may further comprise cells utilizing a canonical genetic code.
[0132] The kit may further comprise a first mobile genetic element that has been re-encoded according to a first orthogonal encoding scheme. The kit may further comprise a second mobile genetic element that has been re-encoded according to a second orthogonal encoding scheme. The first and / or second mobile genetic element may be a mobile genetic element as disclosed herein.
[0133] In one example, the kit may comprise a mobile genetic element in which at least one, multiple, or each instance of a codon that canonically encodes alanine (i.e., a GCN codon) in at least one gene required for horizontal transfer of genetic information has been replaced by TCG or TCA, and the kit may comprise a cell expressing a modified tRNA capable of decoding TCG or TCA as alanine. The tRNA may be as disclosed for the first, second, or third aspect.
[0134] In one example, the kit may comprise a mobile genetic element in which at least one, more than one, or each instance of a codon that canonically encodes histidine (i.e., a CAT / C codon) in at least one gene required for horizontal transfer of genetic information has been replaced by TCG or TCA, and may comprise cells expressing a modified tRNA capable of decoding TCG or TCA as histidine. The tRNA may be as disclosed for the first, second, or third aspect.
[0135] In a specific example, the kit may comprise a mobile genetic element wherein at least one, multiple, or each instance of a GCN codon in at least one gene required for horizontal transfer of genetic information has been replaced by a TCG codon; and at least one, multiple, or each instance of a CAT / C codon in at least one gene required for horizontal transfer of genetic information has been replaced by a TCA codon.
[0136] The kit can further comprise a third, fourth, fifth or further cell, wherein each of the third, fourth, fifth or further cells utilizes an encoding scheme that is mutually orthogonal to each of the other cells.
[0137] Mobile genetic elements
[0138] The inventors have demonstrated that mobile genetic elements using orthogonal genetic codes cannot be transferred to cells using canonical genetic codes or cells using mutually orthogonal genetic codes. Horizontal gene transfer was shown to be prevented in mobile genetic elements, such as F plasmids that are transferred via conjugation. This is an improvement over orthogonal genetic elements that would require electroporation because such elements are not truly mobile.
[0139] Thus, in a seventh aspect, there is provided a mobile genetic element recoded according to an orthogonal coding scheme.
[0140] The orthogonal coding scheme can be any of the ones discussed herein, including Figure 2 Any of C.
[0141] In one embodiment, a mobile genetic element is provided wherein at least one, multiple or each instance of a specific type of sense codon in at least one gene required for horizontal transfer of genetic information is replaced by a sense codon that canonically encodes a different amino acid.
[0142] In one embodiment, a mobile genetic element is provided in which at least one, multiple, or each instance of a codon that canonically encodes alanine, leucine, histidine, proline, or any combination thereof in at least one gene required for horizontal transfer of genetic information is replaced by a sense codon that does not encode the corresponding amino acid. In some instances, the new sense codon can canonically encode serine, such as TCA or TCG.
[0143] In one embodiment, a mobile genetic element is provided, wherein at least one, multiple or each instance of a codon that canonically encodes alanine, leucine, histidine, proline or any combination thereof in at least one gene required for horizontal transfer of genetic information is replaced by TCG or TCA.
[0144] For example, at least one, multiple, or each occurrence of a codon that canonically encodes alanine (GCN codon) in at least one gene required for horizontal transfer of genetic information can be replaced in the mobile genetic element by a codon that has been reassigned to alanine. In one embodiment, a mobile genetic element is provided, wherein at least one, multiple, or each instance of a codon that canonically encodes alanine (i.e., a GCN codon) in at least one gene required for horizontal transfer of genetic information has been replaced by TCG or TCA.
[0145] Alternatively or additionally, at least one, more or each instance of a codon that canonically encodes histidine (i.e., CAT / C codon) in at least one gene required for horizontal transfer of genetic information may be replaced by a codon that has been reassigned to histidine. In one embodiment, a mobile genetic element is provided wherein at least one, more or each instance of a codon that canonically encodes histidine (i.e., CAT / C codon) in at least one gene required for horizontal transfer of genetic information has been replaced by TCG or TCA.
[0146] In one embodiment, a mobile genetic element is provided, wherein at least one, multiple or each instance of a codon that canonically encodes alanine (GCN codon) in at least one gene required for horizontal transfer of genetic information has been replaced by TCG, and at least one, multiple or each instance of a codon that canonically encodes histidine (CAT / C codon) in at least one gene required for horizontal transfer of genetic information has been replaced by TCA.
[0147] Mobile genetic element comprises at least one gene required for the horizontal transfer of genetic information that has been recoded according to the reallocation scheme.Mobile genetic element can comprise two, three, four or more such genes.Gene that the horizontal transfer of genetic information in the mobile genetic element is not needed can be recoded as the coding scheme with compression, for example, there can not be one or more types of sense codons in the gene.Gene that the horizontal transfer of genetic information in the mobile genetic element is not needed can also comprise one or more codons (for example, replaced by another codon according to the reallocation scheme) that have been recoded.Therefore, in some embodiments, all genes in the mobile genetic element have been recoded so that there is not one or more types of sense codons, and one or more genes, and mobile genetic element comprises at least one gene required for the horizontal transfer of genetic information that has been recoded according to the reallocation scheme.
[0148] In some examples, the mobile genetic element can be a plasmid or a virus. The mobile genetic element can be a bacteriophage. The mobile genetic element can be an F plasmid.
[0149] In one aspect of the present invention, a kit is provided, comprising a first mobile genetic element as disclosed herein and a first cell of the first, second or third aspect as disclosed herein, wherein the first mobile genetic element and the first cell utilize the same genetic encoding scheme.
[0150] In one example, the kit may comprise a mobile genetic element in which at least one, multiple, or each instance of a codon (GCN codon) encoding alanine canonically in at least one gene required for horizontal transfer of genetic information has been replaced by TCG or TCA, and may comprise cells expressing a modified tRNA capable of decoding TCG or TCA as alanine. The tRNA may be as disclosed for the first, second, or third aspects. Alternatively, or in addition, the kit may comprise a mobile genetic element in which at least one, multiple, or each instance of a codon (CAT / C codon) encoding histidine canonically in at least one gene required for horizontal transfer of genetic information has been replaced by TCG or TCA, and may comprise cells expressing a modified tRNA capable of decoding TCG or TCA as histidine.
[0151] In other examples, the kits can include mobile genetic elements and cells utilizing any of the orthogonal genetic codes disclosed herein. For example, Figure 2 Any orthogonal genetic code of C.
[0152] The kit may comprise a second mobile genetic element as disclosed herein and a second cell of the first, second or third aspect, wherein the second mobile genetic element and the second cell utilize the same genetic coding scheme, and wherein the first mobile genetic element and the second mobile genetic element utilize different genetic coding schemes. In some examples, the genetic coding scheme of the second mobile genetic element and the second cell is any of the ones disclosed herein, such as Figure 2 Any orthogonal genetic code of C.
[0153] The kit may further comprise a third, fourth, fifth or further mobile genetic element and cell, wherein each pair of mobile genetic element and cell is compatible and orthogonal to every other pair.
[0154] Methods for increasing resistance of cells to mobile genetic elements or horizontal gene transfer
[0155] In a fifth aspect of the invention, a method is provided for increasing resistance of a cell to mobile genetic elements or horizontal gene transfer, wherein the cell has been modified to reallocate at least one type of sense codon to an amino acid that is not associated with a sense codon in the canonical genetic code, the method comprising: modifying a gene required for viability to include at least one occurrence of the reallocated sense codon, wherein the cell is viable if the reallocated sense codon in the gene decodes to the reallocated amino acid, and the cell is non-viable if the reallocated sense codon in the gene decodes according to the canonical genetic code, or wherein the reallocated sense codon in the gene, if decoded according to the canonical genetic code, contributes at least in part to loss of viability.
[0156] The increased resistance can be resistance that is maintained for a longer period of time. For example, resistance that is not lost during extended cell culture (see Example 7). Thus, compared to non-crypto-locked cells, the cells of the fifth aspect can exhibit resistance to horizontal gene transfer or mobile genetic elements that is maintained over a longer period of time.
[0157] The genome of the cell may have been recoded to remove at least one type of sense codon. The recoding may be any of those disclosed herein, such as recoding TCA or TCG. The cell may not express at least one endogenous tRNA, such as tRNA Ser UGA or tRNA Ser CGA The reassignment may be due to the insertion of one or more modified tRNAs. The modified tRNAs may be any of those disclosed herein, for example, an isoacceptor tRNA with an anticodon exchanged for any of alanine, leucine, histidine, or proline.
[0158] The gene required for viability can be any one, including any one disclosed herein. For example, the gene can be an essential gene or a positive selectable marker.
[0159] The resulting cell may be a cell according to the first, second or third aspect of the invention and the method may be modified accordingly.
[0160] Resistance to horizontal gene transfer
[0161] The cells of the present disclosure, including the cells of the first, second, and third aspects, can be resistant to horizontal gene transfer. For example, the cells can be resistant to the transfer of genetic information from mobile genetic elements, including plasmids (such as F plasmids), viruses (including bacteriophages), and the like.
[0162] In particular, the cells of the present disclosure can be resistant to the transfer of genetic information from a mobile genetic element comprising a cognate tRNA.The cognate tRNA can be a tRNA that can decode one or more reassigned codons according to the canonical genetic code.
[0163] In some embodiments, cells can reduce or completely eliminate the transfer of F plasmids containing relevant tRNAs. This property can be used to Figure 5 Method test of C.
[0164] The cell may be a bacterium and may be resistant to a bacteriophage disclosed in Nyerges et al., "Swapped genetic code blocks viral infections and gene transfer," https: / / doi.org / 10.1101 / 2022.07.08.499367. The cell may be Escherichia coli and may be resistant to the bacteriophage.
[0165] The cells may also be resistant to horizontal gene transfer from the cell to other types of cells. For example, the cells of the first, second, or third aspects, or as created by the method of the fifth aspect, may not be able to transfer synthetic genes into wild-type bacteria or bacteria. The cells of the present disclosure may not be able to transfer synthetic genes into wild-type bacteria of the same species. The synthetic genes may be according to any reassigned coding scheme as disclosed herein.
[0166] In addition, the cells of the present disclosure are also resistant to horizontal gene transfer from the cells to other cells that do not use the same reassigned coding scheme. Thus, the cells of the present disclosure are unable to transfer synthetic genes to bacteria that are unable to decode the synthetic genes according to the specific reassigned coding scheme. Other bacteria may also utilize reassigned coding schemes, but if the schemes are orthogonal to the cells of the present disclosure, horizontal gene transfer will be prevented.
[0167] Methods for altering the susceptibility of a gene to mutations that alter the encoded amino acid sequence
[0168] The cells, codes, and techniques disclosed herein allow for methods to alter the susceptibility of a gene to mutations that alter the encoded amino acid sequence. Thus, the remodeling secrets disclosed herein can be used to accelerate or slow the rate of protein evolution.
[0169] The standard genetic code is conservative to some extent, because point mutations may not change the encoded amino acid. In addition, point mutations may change the encoded amino acid into another amino acid with similar properties (e.g., conservative substitutions) or different properties (e.g., non-conservative substitutions). The number of differences between codon types may vary, and this may affect the probability that point mutations will result in no change, conservative changes, or non-conservative changes in the encoded amino acid. The inventors have provided techniques for changing the mutation landscape and for implementing such codes (see, e.g., Figure 10 (Figure S5) and Figure 11 (Figure S6)).
[0170] Thus, in one aspect, a method is provided for altering the susceptibility of a gene to a mutation that alters the encoded amino acid sequence, the method comprising:
[0171] i) identifying target genes; and
[0172] ii) incubating a cell comprising the target gene, wherein the cell comprises a tRNA capable of decoding at least one sense codon into the reassigned amino acid.
[0173] The target gene can be one or more target genes. The target gene can be a synthetic or natural gene. Suitably, the synthetic gene can change the use of codons to support the evolutionary trajectory. In some aspects, the target gene can be based on a compressed genetic code.
[0174] The cell may be any of those disclosed herein, for example, any of the first, second, or third aspects. The cell may be a bacterial cell, such as E. coli, that has been recoded with respect to the first and second types of sense codons. The cell may be Syn61, derived from Syn61, or recoded in the same manner as Syn61. The cell may be Syn61Δ3, derived from Syn61Δ3, or recoded in the same manner as Syn61Δ3.
[0175] In an example, the reallocation scheme may be as follows Figure 2 Any of the options shown in C. These strategies can be used to alter the mutational landscape as shown in Figure 11 (Figure S6).
[0176] Cells can be incubated under conditions that are likely or expected to result in mutations. This method can be used to evolve or improve proteins. This method can be used to make a target gene more resistant to mutations, for example, to protect cells from harmful mutations.
[0177] Recoding of sense codons
[0178] This section further describes exemplary embodiments of re-encoding and is applicable to all aspects disclosed herein.
[0179] If the endogenous tRNA is not present in a form that would allow it to decode its (one or more) cognate codons, then it is considered that the endogenous tRNA is not expressed. Therefore, any method that will prevent the production of a functional form of the endogenous tRNA in the cell can be used to remove the endogenous tRNA. For example, an endogenous gene or a portion of a gene can be deleted to prevent expression. Regulatory sequences can be deleted or altered to prevent expression. Alternatively, nonsense, frameshift or missense mutations can prevent the expression of tRNA in a functional form.
[0180] As used herein, "recoding" is the replacement of one type of codon with a different codon so that the codon is removed from the genome. The recoded sense codon can be replaced with a synonymous codon to result in the use of a different codon without changing the encoded polypeptide. Alternatively, the sense codon can be replaced with a non-synonymous codon, for example, if the change in the encoded polypeptide sequence does not affect viability. The deleted endogenous tRNAs are those that are dispensable in view of the recoding. As used herein, "dispensable" means that it is not required for cell viability.
[0181] Viable cells are those that are metabolically active. In certain embodiments, viable cells may be capable of growth when cultured in an appropriate culture medium and under appropriate conditions for a particular species or strain. Such cells may be referred to as culturable. As an example, if the cells are bacterial cells, such as E. coli, viability may be assessed by culturing the bacteria at 37°C in a culture medium comprising LB medium or on agar containing LB agar. The culture medium or agar may be supplemented with 2% glucose. Standard methods, such as measuring OD 600 , to monitor bacterial growth. Alternative methods, or methods adapted to specific cells, bacterial strains, bacterial species or in view of the inclusion of marker genes, are known to the skilled person.
[0182] The endogenous tRNA that decodes one or more sense codons that have been replaced (or deleted) can be deleted, and if the tRNA only decodes one or more sense codons that have been replaced (or deleted); or alternatively, if the tRNA decodes one or more sense codons that have been replaced (or deleted) and one or more sense codons that have not been replaced (or deleted), if the tRNA is dispensable for the one or more sense codons that have not been replaced (or deleted) (i.e., one or the remaining sense codons decoded by the tRNA are decoded by one or more alternative tRNAs), the cell will still be viable. For example, if the genome of a prokaryotic cell lacks the TCA sense codon, then the tRNA encoding Ser UGA The serT can be deleted, and / or if the genome lacks TCG sense codon, the encoding tRNA Ser CGAThus, in one embodiment, the cell expresses neither tRNA Ser UGA , nor does it express tRNA Ser CGA .
[0183] The number of times that the first and / or second type of recoded sense codon occurs is sufficient to achieve the removal of the cognate tRNA corresponding to the sense codon while maintaining the viability of the cell. For example, this can be achieved by removing all natural occurrences of the first and second type sense codons from essential genes. In particular, if a "blank" codon in a gene (i.e., a codon that does not contain the corresponding tRNA or release factor) causes loss of cell viability, the gene is considered to be essential. Therefore, in one embodiment, all genes of a cell that cannot tolerate a blank codon without losing viability are recoded, but genes that can tolerate a blank codon may not be recoded. Therefore, a technician can evaluate whether all essential genes have been recoded by evaluating whether the cognate tRNA is dispensable. It is noteworthy that some embodiments require at least one essential gene to include the first and / or second type of sense codon; however, in such embodiments, the codon is reallocated rather than being in a naturally occurring position.
[0184] The cells of the present disclosure, including the cells of the first, second, and third aspects, can be recoded with respect to the first, second, third, fourth, fifth, or further types of sense codons. Recoding of the first and second types of sense codons is exemplified herein, and the skilled artisan will appreciate that this principle can be extended to recoding further types of sense codons within the genome of the cell and thereby reducing their occurrence. For example, further types of sense codons can be replaced by synonymous codons to eliminate specific occurrences without changing the encoded sequence, and sufficient numbers of a particular type of sense codon can be removed such that at least one further endogenous tRNA is dispensable and does not need to be expressed by the cell.
[0185] In certain embodiments, the genome comprises 100 or more, 200 or more, or 300 or more essential genes without the natural occurrence of the first and / or second type sense codons. For example, all or substantially all essential genes in the genome may not comprise the natural occurrence of the first and / or second type sense codons. In some embodiments, the essential genes may be selected from one or more of the following list: ribF, lspA, ispH, dapB, folA, imp, yabQ, ftsL, ftsI, murE, murF, mraY, murD, ftsW, murG, murC, ftsQ, ftsA, ftsZ, lpxC, secM, secA, can, folK, heml, yadR, dapD, map, rpsB, tsf, pyrH, frr, dxr, ispU, cdsA, yaeL, yaeT, lpxD, fabZ, lpxA, lpxB, dnaE, accA, tilS, proS, yafF, hemB, secD, secF, ribD, ribE, thiL, dxs, ispA, dnaX, adk, hemH, lpxH, cysS, folD, entD, mrdB, mrd A, nadD, holA, rlpB, leuS, lnt, glnS, fldA, cydA, infA, cydC, ftsK, lolA, serS, rpsA, msbA, lpxK, kdsB, mukF, mukE, mukB, asnS, fabA, mviN, rne, fabD, fabG, acpP, tmk, holB, lolC, lolD, lolE, purB, minE, minD, pth, prsA, ispE, lolB, hemA, prfA, prmC, kdsA, topA, ribA , fabI, tyrS, ribC, ydiL, pheT, pheS, rplT, infC, thrS, nadE, gapA, yeaZ, aspS, argS, pgsA, yefM, metG, folE, yejM, gyrA, nrdA, nrdB, folC, accD, fabB, gltX, ligA, zipA, dapE, dapA, der, hisS, ispG, suhB, tadA, acpS, era, rnc, lepB, rpoE, pssA, yfiO, rplS, trmD, rpsP , ffh, grpE, csrA, ispF, ispD, ftsB, eno, pyrG, chpR, lgt, fbaA, pgk, yqgD, metK, yqgF, plsC, ygiT, parE, ribB, cca, ygjD, tdcF, yraL,yhbV, infB, nusA, ftsH, obgE, rpmA, rplU, ispB, murA, yrbB, yrbK, yhbN, rpsI, rplM, degS, mreD, mreC, mreB, accB, accC, yrdC, def, fmt, rplQ, rpoA, rpsD, rpsK, rpsM, secY, rplO, rpmD, rpsE, rplR, rplF, rpsH, rpsN, rplE, rplX, rplN, rpsQ, rpmC, rplP, rpsC, rplV, rpsS, rplB, rplW, rplD, rplC, rpsJ, fusA, rpsG, rpsL, trpS, yrfF, asd, rpoH, ftsX, ftsE, ftsY, yhhQ, bcsB, glyQ, gpsA, rfaK, kdtA, coaD, rpmB, dfp, dut, gmk, spoT, gyrB, dnaN, dnaA, rpmH, rnpA, yidC, tnaB, glmS, glmU, wzyE, hemD, hemC, yigP, ubiB, ubiD, hemG, yihA, ftsN, murI, murB, birA, secE, nusG, rplJ, rplE, rpoB, rpoC, ubiA, plsB, lexA, dnaB, ssb, alsK, groS, psd, orn, yjeE, rpsR, chpS, ppa, valS, yjgP, yjgQ and dnaC.
[0186] RibF、lsp is the name of the group. A、ispH、dapB、folA、imp、yabQ、lpxC、 secM, secA, can, folK, hemL, yadR, dapD, map, rpsB, tsf, pyrH, frr, dxr, isp U、cdsA、yaeL、yaeT、lpxD、fabZ、lpxA、lpxB、dnaE、accA、tilS、proS、yafF、 hemB、secD、secF、ribD、ribE、thiL、dxs、ispA、dnaX、adk、hemH、lpxH、cysS、 folD、entD、mrdB、mrdA、nadD、holA、rlpB、leuS、lnt、glnS、fldA、cydA、inf A、cydC、ftsK、lolA、serS、rpsA、msbA、lpxK、kdsB、mukF、mukE、mukB、asnS、f abA、mviN、rne、fabD、fabG、acpP、tmk、holB、lolC、lolD、lolE、purB、minE、 minD, pth, prsA, ispE, lolB, hemA, prfA, prmC, kdsA, topA, ribA, fabI, tyrS 、ribC、ydiL、pheT、pheS、rplT、infC、thrS、nadE、gapA、yeaZ、aspS、argS、p gsA、yefM、metG、folE、yejM、gyrA、nrdA、nrdB、folC、accD、fabB、gltX、ligA 、zipA、dapE、dapA、der、hisS、ispG、suhB、tadA、acpS、era、rnc、lepB、rpoE pssA、yfiO、rplS、trmD、rpsP、ffh、grpE、csrA、ispF、ispD、ftsB、eno、pyrG 、chpR、lgt、fbaA、pgk、yqgD、metK、yqgF、plsC、ygiT、parE、ribB、cca、ygjD 、tdcF、yraL、yhbV、infB、nusA、ftsH、obgE、rpmA、rplU、ispB、murA、yrbB、yr bK、yhbN、rpsI、rplM、degS、mreD、mreC、mreB、accB、accC、yrdC、def、fmt、rp lQ、rpoA、rpsD、rpsK、rpsM、secY、rplO、rpmD、rpsE、rplR、rplF、rpsH、rpsN、rplE, rplX, rplN, rpsQ, rpmC, rplP, rpsC, rplV, rpsS, rplB, rplW, rplD, rplC, rpsJ, fusA, rpsG, rpsL, trpS, yrfF, asd , rpoH, ftsX, ftsE, ftsY, yhhQ, bcsB, glyQ, gpsA, rfaK, kdtA, coaD, rpmB, dfp, dut, gmk, spoT, gyrB, dnaN, dnaA, rpmH, rnpA, yidC, tnaB, glmS, glmU, wzyE, hemD, hemC, yigP, ubiB, ubiD, hemG, yihA, ftsN, murI, murB, birA, secE, nusG, rpl J, rplL, rpoB, rpoC, ubiA, plsB, lexA, dnaB, ssb, alsK, groS, psd, orn, yjeE, rpsR, chpS, ppa, valS, yjgP, yjgQ, and dnaC. ,
[0187] In other embodiments, the cell may comprise a genome comprising 5 or fewer natural occurrences of the first and / or second type of sense codons. The genome may be derived from a parental genome and may comprise less than 10%, 5%, 2%, 1%, 0.5%, 0.1% of the occurrences of the first and / or second type of sense codons relative to the parental genome. The genome may comprise 100 or more, 200 or more, or 1000 or more genes without natural occurrences of the first and / or second type of sense codons. In particular, all or substantially all genes in the genome may lack natural occurrences of the first and / or second type of sense codons. Thus, the genome of the cell may comprise 5, 4, 3, 2, 1, or no natural occurrences of the first type of sense codons and 5, 4, 3, 2, 1, or no natural occurrences of the second type of sense codons.
[0188] In some embodiments, the genome comprises 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more or 2000 or more recoded genes. In some embodiments, these genes are those with evidence of translation and / or predicted protein products. For example, the genome can comprise 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, or 2000 or more recoded genes for which there is evidence of translation and / or a predicted protein product.
[0189] In one embodiment, all annotated open reading frames within the genome are free of naturally occurring first and second types of sense codons. The cell may be a bacterial cell, preferably Escherichia coli, and the genome of Escherichia coli may not contain naturally occurring first and second types of sense codons as annotated by GenBank Accession No. CP040347.1.
[0190] In certain embodiments, the protein coding gene does not have the natural occurrence of the first and second types of sense codons. In certain embodiments, no protein is translated from any remaining natural occurrences of the first and / or second types of sense codons, and / or the remaining natural occurrence genes comprising the first and / or second types of sense codons are putative or non-coding genes. In some embodiments, translation of the remaining natural occurrence genes comprising the first and / or second types of sense codons is reduced and / or prevented (e.g., these genes may comprise a stop codon in the 5' sequence).
[0191] Any remaining natural occurrences of sense codons may be necessary to ensure that the genome is viable. For example, one or more, particularly all, of the remaining natural occurrences of the first and / or second type of sense codons in the genome may be present in regulatory elements of essential genes; and / or one or more, particularly all, of the remaining natural occurrences of the first and / or second type of sense codons may be located in genes for which there is no evidence of translation or predicted protein products (i.e., putative or non-coding genes).
[0192] As used herein, a "sense codon" is a nucleotide triplet that encodes an amino acid. Thus, sense codons can be identified in a genome by gene prediction, i.e., by identifying protein-encoding genomic regions (i.e., genes) and corresponding open reading frames (ORFs). In general, the genome naturally comprises 61 sense codons: GCT, GCC, GCA, GCG, CGT, CGC, CGA, CGG, AGA, AGG, AAT, AAC, GAT, GAC, TGT, TGC, CAA, CAG, GAA, GAG, GGT, GGC, GGA, GGG, CAT, CAC, ATT, ATC, ATA, TTA, TTG, CTT, CTC, CTA, CTG, AAA, AAG, ATG, TTT, TTC, CCT, CCC, CCA, CCG, TCT, TCC, TCA, TCG, AGT, AGC, ACT, ACC, ACA, ACG, TGG, TAT, TAC, GTT, GTC, GTA and GTG (read from 5' to 3' on the DNA coding chain). The standard genetic code encodes 20 kinds of canonical amino acids using 61 triplet codons. 18 of the 20 amino acids are encoded by more than one synonymous codon. The first or second type of sense codon may be a natural sense codon, ie, a sense codon present in the parental genome.
[0193] The 61 sense codons in the DNA are transcribed into corresponding mRNAs and then decoded by one or more tRNAs. tRNAs carry amino acids to the ribosomes as directed by the sense codons in the mRNA. tRNAs can recognize one or more sense codons via complementary anticodons. The sequence of sense codons is then translated into a polypeptide (i.e., an amino acid sequence). The codons and anticodons in the E. coli genome interact as described in WO2020 / 229592. Figure 17 (incorporated herein by reference).
[0194] Genome-wide removal of the first and / or second type of sense codons, but not other sense codons, results in the depletion of cognate tRNAs corresponding to the first or second type of sense codons without removing the ability to decode the remaining sense codons in the genome.
[0195] The recoded sense codon may be selected from the group consisting of: TCG, TCA, TCT, TCC, AGT, or AGC. In a specific embodiment, the first and second types of sense codons are TCA and TCG.
[0196] In order to realize the removal of sense codon, they can be replaced by synonymous sense codon.This is preferred for guaranteeing that the protein sequence of coding does not change.For example, cell can have a kind of genome, wherein the first or second type sense codon in parent's genome 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more or 100% appearance is replaced by synonymous sense codon.Those skilled in the art can infer that suitable synonymous sense codon replaces.For example, in intestinal bacteria, generally TCG, TCA, TCT, TCC, AGT and AGC all encode serine.
[0197] In some embodiments, the replacement is a defined replacement, i.e., one sense codon is replaced by a single synonymous sense codon. Preferably, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the natural occurrences of the first or second type of sense codon in the parental genome are replaced by a defined (i.e., single) synonymous sense codon.
[0198] For example, a substitution may be defined as: TCG is replaced by any one of TCT, TCC, AGT, or AGC; or TCA is replaced by any one of TCT, TCC, AGT, or AGC. In particular, the substitution is one or more selected from the group consisting of: TCG is replaced by AGT or AGC; or TCA is replaced by AGT or AGC. In a specific embodiment, TCG is replaced by AGC and TCA is replaced by AGT.
[0199] Preferably, none of these codon substitutions will affect the ribosome binding site (AGGAGG), a highly conserved regulatory sequence in E. coli. Selected codon substitutions can be tested on a small test region (e.g., a 20 kb region of the genome rich in essential target genes and target codons) to evaluate activity. If codon substitutions are not feasible on a small test region, they can be ignored.
[0200] When replacing the sense codon in the parental genome with a defined replacement synonymous sense codon does not result in a viable cell, an alternative replacement synonymous sense codon can be used. For example, 99.9% of the occurrences of the first and / or second type sense codons in the parental genome can be replaced by a defined (i.e., single) synonymous sense codon, while the remaining 0.1% are replaced by an alternative synonymous sense codon. For example, 99.9% of the natural occurrences of TCG can be replaced by AGC, while 0.1% are replaced by TCT, TCC, AGT, or AGC; and / or 99.9% of the occurrences of TCA can be replaced by AGT, while 0.1% are replaced by TCT, TCC, AGT, or AGC.
[0201] In some cases, a particular occurrence of a sense codon may not be replaced by any potential synonymous sense codon without affecting viability. In order to maintain viability, a sense codon may be replaced by a non-synonymous sense codon that does not affect viability. For example, 99.9% of the occurrences of the first and / or second type of sense codon in the parental genome may be replaced by a defined (i.e., single) synonymous sense codon, and the remaining 0.1% may be replaced by an alternative non-synonymous sense codon.
[0202] Recoding of stop codons
[0203] This section further describes exemplary embodiments involving recoding stop codons and applies to all aspects disclosed herein.
[0204] In some examples, the first type of stop codon has been recoded within the genome of the cell such that the first endogenous release factor is dispensable, and the cell does not express the first endogenous release factor.
[0205] The removal of the first endogenous release factor can be performed in a cell in which the genome has been recoded to remove the occurrence of the first type of stop codon. Optionally, the removed stop codon is replaced by a synonymous codon. The deleted endogenous release factor is a factor that is dispensable in view of the recoding.
[0206] In a specific example, the cell does not express the first endogenous tRNA, the second endogenous tRNA, and the first endogenous release factor; and the genome has been recoded to remove a plurality of sense codons to which the first and second endogenous tRNAs are homologous, and to remove a plurality of stop codons to which the first endogenous release factor is homologous.
[0207] If the endogenous release factor is not present in a form that would allow it to decode its (one or more) cognate codons, then it is considered that the endogenous release factor is not expressed. Therefore, any method that will prevent the production of a functional form of the endogenous release factor in the cell can be used to remove the endogenous release factor. For example, the endogenous gene or a portion of the gene can be deleted to prevent expression. Regulatory sequences can be deleted or altered to prevent expression. Alternatively, nonsense, frameshift or missense mutations can prevent the expression of the release factor in a functional form.
[0208] As used herein, a "stop codon" is a nucleotide triplet that codes for the termination of translation into a protein. Typically, genomes naturally contain three stop codons: TAA ("ochre"), TGA ("opal" or "umber"), and TAG ("amber").
[0209] The natural occurrence number of the first type of termination codon removed is enough to make the homology release factor corresponding to the termination codon be removed while keeping the vigor of the cell.Therefore, in some instances, the essential gene of the cell does not contain the occurrence of the first type of termination codon.Essential gene can be any as discussed herein, particularly about removing those of the first or second type sense codon discussion.In a specific instance, the genome comprises 100 or more, 200 or more or 300 or more essential genes, and does not have the natural occurrence of the first type of termination codon.For example, all or substantially all essential genes in the genome may not comprise the occurrence of the first type of termination codon.
[0210] For example, the genome may comprise 100 or more, 200 or more, or 300 or more essential genes without naturally occurring occurrences of the first type of sense codons, the second type of sense codons, and the first type of stop codons. In particular, all or substantially all of the essential genes in the genome may not comprise naturally occurring occurrences of the first type of sense codons, the second type of sense codons, and the first type of stop codons.
[0211] In some embodiments, the genome comprises 10 or fewer, 5 or fewer, or no natural occurrences of the first type of stop codon, such as 5, 4, 3, 2, 1, or no natural occurrences of the first type of stop codon.
[0212] In certain embodiments, the first type of stop codon is TAG and the first endogenous release factor is RF-1. In this embodiment, there are 10 or less times, 5 or less times or no natural occurrence of amber stop codon (TAG). In other embodiments, 90% or more, 95% or more, 98% or more, 99% or more or all occurrences of TAG in the parental genome are replaced by TAA (ochre stop codon). In certain embodiments, the genome does not comprise the occurrence of amber stop codon (TAG), and optionally wherein the occurrences of all TAGs in the parental genome are replaced by TAA (ochre stop codon).
[0213] In one embodiment, all annotated open reading frames in the genome are free of the occurrence of the first type of stop codon. The cell of the present disclosure can be a bacterial cell, such as E. coli, and the genome of E. coli can be free of the occurrence of the first type of stop codon as annotated with GenBank Accession No. CP040347.1.
[0214] In some instances, the protein-coding gene has no natural occurrence of the first type of stop codon. In particular instances, no protein is translated from any remaining occurrence of the first type of stop codon, and / or the gene comprising the remaining occurrences of the first type of stop codon is putative or is a non-coding gene. In some instances, translation of the gene comprising the remaining occurrences of the first type of stop codon is reduced and / or prevented (e.g., the gene may comprise a stop codon in the 5' sequence).
[0215] Any remaining occurrences of the first type of stop codons may be necessary to ensure that the genome is viable. For example, one or more, and particularly all, of the remaining natural occurrences of the first type of stop codons in the genome may be present in regulatory elements of essential genes; and / or one or more, and particularly all, of the remaining occurrences of the first type of stop codons may be located in genes for which there is no evidence of translation or predicted protein products (i.e., putative or non-coding genes).
[0216] Genome recodes sense and stop codons
[0217] This section further describes exemplary embodiments of re-encoding and applies to all aspects disclosed herein.
[0218] Thus, in some examples, the genome comprises no occurrences of the first and second types of sense codons, and no occurrences of a stop codon, preferably an amber stop codon (TAG). In specific examples, the genome comprises no occurrences of the sense codons TCG and TCA, and no occurrences of an amber stop codon (TAG), optionally wherein TCG, TCA, and TAG in the parental genome are replaced by synonymous codons, e.g., 99.9% or more occurrences of TCG in the parental genome are replaced by AGC, 99.9% or more occurrences of TCA in the parental genome are replaced by AGT, and all occurrences of TAG in the parental genome are replaced by TAA.
[0219] In a particular example, the genome of the cell has been recoded such that the sense codon TCG has been replaced by AGC, the sense codon TCA has been replaced by AGT, and the stop codon TAG has been replaced by TAA, and wherein a sufficient number of said codons have been recoded such that two cognate tRNAs and one cognate release factor are dispensable.
[0220] In a specific example, the cell of the present disclosure is an E. coli cell, which does not express tRNA Ser UGA tRNA Ser CGA or RF-1, the presence of the sense codon TCA has been recoded so that tRNA Ser UGA is dispensable (e.g. the presence of TCA in an essential gene of the parent strain has been replaced by AGT), the presence of the sense codon TCG has been recoded so that tRNA Ser CGA is dispensable (e.g., occurrences of TCG in essential genes of the parental strain have been replaced by AGC), and the occurrence of the stop codon TAG has been recoded, making RF-1 dispensable (e.g., occurrences of TAG in essential genes of the parental strain have been replaced by TAA).
[0221] In some embodiments, the genome of the disclosed cells comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to the sequence provided as GenBank Accession No. CP040347.1, and wherein the genome has been further altered such that tRNA Ser UGA and tRNA Ser CGAThe RF-1 gene is not functionally expressed (e.g., by deletion of serT and serU). The genome can be altered even further such that RF-1 is not functionally expressed (e.g., by deletion of prfA). The E. coli strain comprising the genome according to GenBank Accession No. CP040347.1 is referred to as Syn61 WT in the examples disclosed herein.
[0222] In some embodiments, the genome of the disclosed cells comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to SEQ ID NO: 1, and wherein the genome has been further altered such that tRNA Ser UGA and tRNA Ser CGA is not functionally expressed (e.g. by deletion of serT and serU). The genome may be altered even further such that RF-1 is not functionally expressed (e.g. by deletion of prfA). An E. coli strain comprising the genome according to SEQ ID NO: 1 may be designated Syn61 (ev1).
[0223] In some embodiments, the genome of the disclosed cells comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to SEQ ID NO: 2, and wherein the genome has been further altered such that tRNA Ser UGA and tRNA Ser CGA is not functionally expressed (e.g. by deletion of serT and serU). The genome may be altered even further such that RF-1 is not functionally expressed (e.g. by deletion of prfA). The E. coli strain comprising the genome according to SEQ ID NO: 2 may be designated Syn61 (ev2).
[0224] In some embodiments, the genome of the cell comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to SEQ ID NO: 3. The E. coli strain comprising the genome according to SEQ ID NO: 3 can be referred to as Syn61Δ3.
[0225] In some embodiments, the genome of the cell comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to SEQ ID NO: 4. The E. coli strain comprising the genome according to SEQ ID NO: 4 can be designated Syn61Δ3 (ev3).
[0226] In some embodiments, the genome of the cell comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to SEQ ID NO: 5. The E. coli strain comprising the genome according to SEQ ID NO: 5 can be designated Syn61Δ3(ev4).
[0227] In some embodiments, the genome of the cell comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to SEQ ID NO: 6 (also provided as GenBank Accession No. CP071799.1). The E. coli strain comprising the genome according to SEQ ID NO: 6 can be designated Syn61Δ3(ev5).
[0228] Provided herein are prokaryotic cells comprising a genome that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9%, or 100% identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, and SEQ ID NO: 6. The prokaryotic cell can be a bacterium, such as E. coli. In some embodiments, the calculation of the percent sequence identity excludes any sequences that have been inserted to further modify the cell. The calculation of the percent sequence identity can further exclude any exogenous sequences that have been further introduced into the genome. For example, any additional tRNAs, selectable markers, changes in genes required for viability according to the present disclosure, constructs for industrial expression of peptide or protein products, and the like.
[0229] Cell species
[0230] The cells of the present disclosure, including those of all aspects disclosed herein, may be prokaryotic cells. The bacterial cells may be of any species suitable for heterologous protein production, particularly polypeptide production. Suitable bacterial host cells include: Escherichia (e.g., Escherichia coli), Caulobacteria (e.g., Caulobacter crescentus), phototrophic bacteria (e.g., Rodhobacter sphaeroides), cold-adapted bacteria (e.g., Pseudoalteromonas haloplanktis), Shewanella sp.) strain Ac10), Pseudomonas (e.g., Pseudomonas fluorescens, Pseudomonas putida, Pseudomonas aeruginosa), halophilic bacteria (e.g., Halomonas elongate, Chromohalobacter salexigens), Streptomyces (e.g., Streptomyces lividans, Streptomyces griseus), Nocardia (e.g., Nocardia lactamdurans), Mycobacteria (e.g., Mycobacterium smegmatis), Corynebacterium (e.g., Corynebacterium glutamicum, Corynebacterium ammoniagenes), ammoniagenes), Brevibacterium lactofermentum), Bacillus (e.g., Bacillus subtilis, Bacillus brevis, Bacillus megaterium, Bacillus licheniformis, Bacillus amyloliquefaciens), Vibrio bacteria (e.g., Vibrio cholera, Vibrio natriegens), and lactic acid bacteria (e.g., Lactococcus lactis, Lactobacillus plantarum, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri). In some examples, the bacteria are Gram-negative bacteria.
[0231] In a specific example, the bacterium is Escherichia coli, Salmonella enterica or Shigella dysenteriae. More preferably, the cell is Escherichia coli. Suitable Escherichia coli cells include K-12, MG1655, BL21, BL21 (DE3), AD494, Origami, HMS174, BLR (DE3), HMS174 (DE3), Tuner (DE3), Origami2 (DE3), Rosetta2 (DE3), Lemo21 (DE3), NiCo21 (DE3), T7 Express, SHuffle Express, C41 (DE3), C43 (DE3) and m15pREP4 or its derivatives (Rosano, GL and Ceccarelli, EA, 2014. Frontiers in microbiology, 5, p.172). In particular, the cell can be MG1655 or BL21 or its derivatives. MG1655 is considered a wild-type strain of E. coli. The GenBank ID for the genomic sequence of this strain is U00096. BL21 is widely commercially available. For example, it can be purchased from New England BioLabs under the catalog number C2530H.
[0232] The cell may contain a genome derived from the same species or strain, or may be derived from a different species. For example, if the cell is E. coli, the genome may be an E. coli genome.
[0233] The cells of the present disclosure, including those of all aspects disclosed herein, can be biocontained cells. Thus, the cells of the present disclosure can be viable or capable of proliferation only under conditions that do not exist in nature. Such cells can be considered to comprise a biocontainment system.
[0234] For example, cells may be viable or capable of proliferation only in the presence of an agent not present in their natural environment. Such cells can be cultured in the presence of the agent, but will not remain as a cell population if the cells are placed in an environment without the agent. Examples of such agents include unnatural amino acids, which may be required for the functional translation of one or more essential genes. Other examples include ligands required for the expression or activity of essential genes / proteins.
[0235] In another example, a cell may contain a gene that prevents viability or proliferation, wherein the gene is inactive in the presence of an agent not found in the natural environment. The gene may be referred to as a "kill switch" and may, for example, encode a toxin.
[0236] Polymer production
[0237] As disclosed herein, the cells of the present disclosure may be suitable for polymer production. Thus, in a seventh aspect of the present disclosure, there is provided the use of any cell disclosed herein for producing a polymer.
[0238] In one embodiment, a method for producing a polymer is provided, the method comprising: culturing a cell as disclosed herein, providing the cell with a nucleic acid sequence encoding the polymer, and obtaining the polymer.
[0239] The polymer can be a polypeptide. The polymer can be a heterologous protein. The polymer can include monomers that can be incorporated by charged tRNA, such as canonical amino acids, natural amino acids, unnatural amino acids, beta amino acids, hydroxy acids, alpha hydroxy acids, etc.
[0240] Further information
[0241] Sequence comparison can be performed with the aid of readily available sequence comparison programs. These publicly available and commercially available computer programs can calculate the sequence identity between two or more sequences.
[0242] The skilled person will appreciate how to calculate the percent identity between two nucleotide sequences. In order to calculate the percent identity between two nucleotide sequences, the comparison of the two sequences must first be prepared, and then the sequence identity value is calculated. The percent identity of two sequences can depend on the following and adopt different values: (i) method for aligning sequences, such as Needleman-Wunsch algorithm (such as applied by Needle (EMBOSS) or Stretcher (EMBOSS)), Smith-Waterman algorithm (such as applied by Water (EMBOSS)) or LALIGN application (such as applied by Matcher (EMBOSS)); and (ii) parameters used by the alignment method, such as local relative to global alignment, the matrix used, and the parameters applied to spaces.
[0243] After alignment, there are many different ways to calculate the percent identity between two sequences. For example, the number of identities can be divided by: (i) the length of the shortest sequence; (ii) the length of the alignment; (iii) the average length of the sequences; (iv) the number of non-gap positions; or (iv) the number of equivalent positions excluding overhangs. Furthermore, it will be appreciated that percent identity is also strongly dependent on length. Thus, the shorter a pair of sequences is, the higher the chance of sequence identity being expected.
[0244] Calculation of percent identity between two nucleic acid sequences can then be calculated from this alignment as (N / T)*100, where N is the number of positions at which the sequences share the same residue, and T is the total number of positions compared, including gaps but excluding overhangs.
[0245] The sequence alignment can be a paired sequence alignment. Suitable services include Needle (EMBOSS), Stretcher (EMBOSS), Water (EMBOSS), Matcher (EMBOSS), LALIGN or GeneWise. In one example, the identity between two amino acid sequences can be calculated using the service Needle (EMBOSS) set to default parameters, such as matrix (BLOSUM62), gap open (10), gap extension (0.5), end gap penalty (false), end gap open (10) and end gap extension (0.5). In another example, the identity between two amino acid sequences can be calculated using the service Matcher (EMBOSS) set to default parameters, such as matrix (BLOSUM62), gap open (14), gap extension (4), alternative match (1). In one example, the identity between two nucleic acid sequences can be calculated using the service Needle (EMBOSS) set to default parameters, such as matrix (DNAfull), gap opening (10), gap extension (0.5), end gap penalty (false), end gap opening (10), and end gap extension (0.5). In another example, the identity between two nucleic acid sequences can be calculated using the service Matcher (EMBOSS) set to default parameters, such as matrix (DNAfull), gap opening (16), gap extension (4), alternative matching (1).
[0246] All features described herein (including any accompanying claims, abstract, and drawings) and / or all steps of any method or process so disclosed may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.
[0247] In order to better understand the present invention and to show how its embodiments may be carried into effect, reference will now be made to the examples, which are not intended to limit the invention in any way.
[0248] Some embodiments of the invention may be defined by the following terms.
[0249] 1. A cell, which:
[0250] comprising a genome in which at least a first type of sense codon has been recoded such that a first endogenous tRNA is dispensable;
[0251] The first endogenous tRNA is not expressed;
[0252] expressing a first modified tRNA capable of decoding a first type of sense codon, wherein the first modified tRNA carries a first amino acid that is not a naturally occurring cognate amino acid for the first type of sense codon; and
[0253] The cell comprises a gene required for viability, wherein the gene comprises at least one occurrence of a first type of sense codon, and the cell is viable when the first type of sense codon in the gene is decoded as a first amino acid.
[0254] 2. The cell of clause 1, wherein the cell is non-viable if the first type of sense codons in the gene required for viability are decoded according to the canonical genetic code, or wherein the first type of sense codons in the gene required for viability at least partially contributes to the loss of viability if decoded according to the canonical genetic code.
[0255] 3. The cell of Item 1 or Item 2, wherein the gene required for viability is an essential gene or a positive selectable marker.
[0256] 4. The cell of any one of clauses 1-3, wherein the first amino acid is a naturally occurring amino acid.
[0257] 5. The cell of any one of clauses 1 to 4, wherein the second type of sense codon has been re-encoded within the genome; optionally wherein the second endogenous tRNA is dispensable and the cell does not express the second endogenous tRNA; and optionally wherein the cell expresses a second modified tRNA capable of decoding the second type of sense codon, wherein the second modified tRNA carries a second amino acid that is not the natural cognate amino acid for the second type of sense codon.
[0258] 6. The cell of clause 5, wherein the gene required for viability comprises at least one occurrence of the second type of sense codon, and the cell is viable when the second type of sense codon in the gene decodes to the second amino acid.
[0259] 7. The cell of clause 6, wherein the cell is non-viable if the second type of sense codon in a gene required for viability is decoded according to the canonical genetic code, or wherein the second type of sense codon contributes at least in part to the loss of viability if decoded according to the canonical genetic code.
[0260] 8. The cell of any one of clauses 5-7, wherein the second amino acid is a naturally occurring amino acid.
[0261] 9. The cell of any one of clauses 1 to 8, wherein the cell is viable when its genes are decoded by the modified tRNA(s) and is non-viable when its genes are decoded at least in part according to the canonical genetic code.
[0262] 10. A cell having increased resistance to horizontal gene transfer or mobile genetic elements, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid that is not associated with a sense codon in the canonical genetic code, and the cell comprises a gene required for viability that is functional when decoded according to the reassigned genetic code and is not functional when decoded according to the canonical genetic code.
[0263] 11. A cell, which:
[0264] comprising a genome in which a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable;
[0265] The first endogenous tRNA and the second endogenous tRNA are not expressed;
[0266] expressing a first anticodon-exchanged tRNA derived from a naturally occurring first parent tRNA, wherein the first anticodon-exchanged tRNA carries a first amino acid and the first parent tRNA is an isoacceptor for the first amino acid, and wherein the first amino acid is not the naturally cognate amino acid for a first type of sense codon; and
[0267] expressing a second anticodon-exchanged tRNA derived from a naturally occurring second parent tRNA, wherein the second anticodon-exchanged tRNA carries a second amino acid and the second parent tRNA is an isoacceptor for the second amino acid, and wherein the second amino acid is not the naturally cognate amino acid for the second type of sense codon;
[0268] wherein the first and / or second modified tRNA is incapable of decoding any type of codon other than the first type of sense codon and / or the second type of sense codon.
[0269] 12. The cell of clause 11, wherein:
[0270] Due to wobble base pairing, the first and second types of sense codons will be canonically decoded by the same tRNA or overlapping tRNAs, wherein the first anticodon-exchanged tRNA cannot decode any codon type other than the first type of sense codons and / or the second anticodon-exchanged tRNA cannot decode any codon type other than the second type of sense codons; and / or
[0271] The first type of sense codon and the second type of sense codon belong to the formula XXN, and wherein the tRNA exchanged for the first anticodon cannot decode the second type of sense codon, and the tRNA exchanged for the second anticodon cannot decode the first type of sense codon.
[0272] 13. The cell of clause 11 or clause 12, wherein the first amino acid and the second amino acid are different types of amino acids.
[0273] 14. The cell of any one of clauses 11 to 13, wherein the first and second parent tRNAs are derived from the same cell type as in clause 11; optionally wherein the first and / or second anticodon-exchanged tRNA comprises an identity element recognized by an aminoacyl-tRNA synthetase endogenous to the cell.
[0274] 15. The cell of any one of clauses 11-14, wherein the first and second types of sense codons canonically encode serine, the first and second types of sense codons canonically encode alanine, or the first and second types of sense codons canonically encode leucine.
[0275] 16. The cell of any one of clauses 11 to 15, wherein the first and / or second anticodon-exchanged tRNA does not decode a TCC or TCT codon.
[0276] 17. The cell of any one of clauses 11 to 16, wherein the first and / or second amino acid is a naturally occurring amino acid; optionally wherein the first amino acid is any one of alanine, histidine, leucine and proline; and / or the second amino acid is any one of alanine, histidine, leucine and proline.
[0277] 18. A cell, which:
[0278] comprising a genome wherein a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable;
[0279] The first endogenous tRNA and the second endogenous tRNA are not expressed;
[0280] expressing a first modified tRNA capable of decoding a first type of sense codon, wherein the first modified tRNA carries a first amino acid that is not a naturally occurring cognate amino acid for the first type of sense codon; and
[0281] expressing a second modified tRNA capable of decoding a second type of sense codon, wherein the second modified tRNA carries a second amino acid that is not a naturally occurring cognate amino acid for the second type of sense codon;
[0282] in:
[0283] i) the first amino acid is alanine and the second amino acid is alanine;
[0284] ii) the first amino acid is alanine and the second amino acid is histidine;
[0285] iii) the first amino acid is alanine and the second amino acid is leucine;
[0286] iv) the first amino acid is alanine and the second amino acid is proline;
[0287] v) the first amino acid is histidine and the second amino acid is alanine;
[0288] vi) the first amino acid is histidine and the second amino acid is histidine;
[0289] vii) the first amino acid is histidine and the second amino acid is leucine;
[0290] viii) the first amino acid is histidine and the second amino acid is proline;
[0291] ix) the first amino acid is leucine and the second amino acid is alanine;
[0292] x) the first amino acid is leucine and the second amino acid is histidine;
[0293] xi) the first amino acid is leucine and the second amino acid is proline;
[0294] xii) the first amino acid is proline and the second amino acid is alanine;
[0295] xiii) the first amino acid is proline and the second amino acid is histidine;
[0296] xiv) the first amino acid is proline and the second amino acid is leucine; or
[0297] xv) The first amino acid is proline and the second amino acid is proline.
[0298] 19. The cell of clause 18, wherein
[0299] The first modified tRNA is unable to decode a sense codon of the second type, and / or the second modified tRNA is unable to decode a sense codon of the first type; and / or
[0300] The first modified tRNA is incapable of decoding any type of codon except a first type of sense codons, and / or the second modified tRNA is incapable of decoding any type of codon except a second type of sense codons.
[0301] 20. The cell of clause 18 or clause 19, wherein the first modified tRNA is a tRNA exchanged for the anticodon canonically associated with the first amino acid, and / or the second modified tRNA is a tRNA exchanged for the anticodon canonically associated with the second amino acid.
[0302] 21. The cell of any one of clauses 18-20, wherein
[0303] The first modified tRNA is derived from a tRNA that is endogenous to the cell and is an isoacceptor for the first amino acid, or is derived from a tRNA found in a mobile genetic element and is an isoacceptor for the first amino acid; and / or
[0304] The second modified tRNA is derived from a tRNA that is endogenous to the cell and is an isoacceptor for the second amino acid, or is derived from a tRNA found in a mobile genetic element and is an isoacceptor for the second amino acid; and / or
[0305] The first and or second modified tRNA comprises an identity element that is recognized by an aminoacyl-tRNA synthetase endogenous to the cell.
[0306] 22. The cell of any one of clauses 18-21, wherein the first and second types of sense codons canonically encode serine.
[0307] 23. The cell of any one of clauses 18-22, wherein the first type of sense codon is TCA and / or the second type of sense codon is TCG.
[0308] 24. The cell of any preceding clause, wherein the cell is a prokaryotic cell, a bacterial cell, or an Escherichia coli cell.
[0309] 25. A method of increasing resistance of a cell to mobile genetic elements or horizontal gene transfer, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid that is not associated with a sense codon in the canonical genetic code, the method comprising:
[0310] Modifying a gene required for viability to include reassigning at least one occurrence of a sense codon, wherein
[0311] If the reassigned sense codon in the gene decodes to the reassigned amino acid, the cell is viable, and
[0312] The cell is non-viable if the reassigned sense codon in the gene is decoded according to the canonical genetic code, or wherein the reassigned sense codon in the gene, if decoded according to the canonical genetic code, contributes at least in part to the loss of viability.
[0313] Example
[0314] Overview
[0315] The nearly universal genetic code defines the correspondence between codons in genes and amino acids in proteins. It has been widely hypothesized that reconfiguring the genetic code would allow the creation of organisms with novel properties and that genetic firewalls could be created to limit the escape of genetic information from synthetic organisms. However, testing these hypotheses has been elusive. Here, we create reconfigured genetic code / decoder systems that, unlike code-compressed organisms, exhibit semantic and functional orthogonality relative to code / decoder systems of the canonical code. We thereby create orthogonal and mutually orthogonal horizontal gene transfer systems, the latter allowing transfer of genetic information between organisms using the same genetic code but restricting transfer between cells using different genetic codes. Furthermore, we show that locking an orthogonal code into a synthetic organism completely blocks the intrusion of mobile genetic elements that have successfully invaded code-compressed organisms.
[0316] In more detail, we show that codon-compressed genes are read in natural cells, rendering codon compression a poor constraint on gene escape from engineered organisms into the biosphere. Furthermore, we demonstrate that mobile genetic elements using the canonical genetic code and carrying the tRNA decoder necessary to complement tRNAs not present in recipient cells can invade Syn61Δ3 cells. We reassigned sense codons to alternative canonical amino acids in Syn61Δ3, thereby remodeling the structure of the genetic code; we created 16 remodeled codons with features not found in nature. We demonstrated that codon reassignment allows the creation of synthetic genes written in the new codon that are correctly read in synthetic organisms with cognate decoders but incorrectly read in natural cells. We also show that genes written in the canonical codon that are correctly read in natural organisms are incorrectly read in synthetic organisms. The genetic code-decoder system in synthetic organisms exhibits semantic and functional orthogonality relative to the encoder-decoder system of the canonical codon. We exploit this orthogonality to create orthogonal and mutually orthogonal horizontal gene transfer systems, the latter of which allows horizontal transfer of genetic information between cells using the same genetic code, but restricts horizontal transfer between cells using different genetic codes. Furthermore, we show that locking the orthogonal code into the synthetic organism completely blocks the invasion of mobile genetic elements that successfully invade the code-compressed organism.
[0317] Example 1 - Compression Cipher is Non-Orthogonal
[0318] In cells containing a complete complement of tRNAs to read the canonical code, the spectinomycin resistance gene (SpecR WT) written in the canonical genetic code was correctly read and conferred spectinomycin resistance on WT cells. However, consistent with previous observations (18), SpecR WT did not confer spectinomycin resistance on Syn61Δ3 cells ( Figure 1 ), because Syn61Δ3 does not read all codons in the canonical genetic code.
[0319] We created a spectinomycin resistance gene (recSpecR(ΔTCG,TCA)) written using the compressed genetic code that we used to create Syn61 (TCG and TCA codons were replaced by AGC and AGT, respectively, and the TAG stop codon was replaced by TAA). recSpecR(ΔTCG,TCA) confers spectinomycin resistance to Syn61Δ3 cells, which use the same codon compression scheme as recSpecR(ΔTCG,TCA) in their genome ( Figure 1 The recSpecR (ΔTCG, TCA) gene also conferred spectinomycin resistance in cells that read the canonical genetic code; this was expected because the compressed genetic code uses a subset of codons used in the canonical genetic code. We made similar observations for codon-compressed and wt hygromycin-resistance genes (Fig. S1).
[0320] These experiments show that genetic information written in a canonical code can be read in WT cells, but not in cells with full-genome code compression and cognate tRNA deletion. However, code-compressed genes can be read in both cells with genomic code compression and cognate tRNA deletion as well as in WT cells. The codons used in the compressed genetic code are not orthogonal to the tRNA decoders in WT cells. Therefore, there are no obstacles restricting the genetic information from engineered biological cells using a compressed genetic code to be read by natural life forms using a canonical code. Creating an orthogonal genetic code with active barriers that limit the transfer of genetic information from engineered biological systems to natural systems is an important and unresolved challenge.
[0321] Example 2 - tRNA enables accessible codon compaction in organisms
[0322] The WT F plasmid (F(WT)) using the canonical genetic code was efficiently transferred to WT cells. In contrast, as expected, F(WT) was not transferred to Syn61Δ3( Figure 1). However, after selecting to join F(WT) from WT cells to Syn61Δ2 cells (Syn61 cells that lack SerU and serT but contain prfA), we obtained two viable colonies in which the recipient cells had accepted F(WT) (Figure S2). These colonies correspond to rare events, with a frequency 106 times lower than that of colonies caused by F(WT) joining into WT cells. Sequencing of the two clones revealed that they had acquired sequences containing serT from the donor cells. This provides direct experimental evidence that selection for the transfer of mobile genetic elements using the canonical code into cells allows selection of tRNA genes required for reading the canonical genetic code.
[0323] To track the effects of introducing serT into recipient cells in a reproducible system, we created a mobile genetic element, F(WT+serT), which is a variant of F(WT) containing serT. We demonstrated that F(WT+serT) can be transferred to Syn61Δ3 cells and that this transfer is dependent on the presence of serT. Figure 1 We conclude that acquisition of serT is sufficient to circumvent the genetic isolation provided by codon compression and cognate tRNA depletion in Syn61Δ3. These experiments highlight the important challenge of creating systems that actively prevent invasion by mobile genetic elements carrying their own decoders.
[0324] Example 3 - Reconstructing the Password-Structure
[0325] Encoding tRNA CGA Ser serU and encoding tRNA UGA Ser The serT of Syn61Δ3 can decode both TCG and TCA codons and incorporate serine into the protein in Syn61Δ3 (Figure S3). To reassign TCG and TCA codons to different natural amino acids in Syn61Δ3, we created variants of canonical amino acid isoacceptor tRNAs; for each isoacceptor, we changed the anticodon to CGA or UGA. We measured the activity of these chimeric tRNAs in decoding TCG or TCA codons at position 3 of the sfGFP gene (a known permissive site) in Syn61Δ3 (Figure S3). We found that Ala (tRNA CGA Ala , tRNA UGA Ala )、His(tRNA CGA His , tRNA UGA His )、Leu(tRNA CGA Leu , tRNA UGA Leu) and Pro(tRNA CGA Leu , tRNA UGA Leu ) chimeric tRNAs specifically direct the incorporation of the amino acid defined by the parental isoacceptor tRNA in response to their cognate codons (TGC or TCA) at position 3 in sfGFP or position 11 in ubiquitin in Syn61Δ3 ( Figure 2 , Figures 8, 12-21), and produced good protein yields (Figure S3). We noticed that tRNA UGA Leu The fidelity is lower than that of other tRNAs ( Figure 13 Alanyl and leucyl-tRNAs were studied because their anticodons are not identity elements of their cognate aminoacyl-tRNA synthetases and therefore they are expected to be permissive to anticodon mutations; other tRNAs were identified by screening ( Figure 22 We also found that tRNA CGA Ser and tRNA UGA Ser In contrast, our chimeric tRNAs specifically decode the Watson-Crick complement of their anticodon sequence; e.g., tRNA CGA Ala Decodes TCG codons in preference to TCA codons, and tRNA UGA Ala TCA codons were decoded preferentially over TCG codons (Figure S3, Figure 13 These tRNAs are also specific for other TCN codons ( Figure 14 , Figure 15 ), and the growth of most reassigned strains was comparable to that of the parent strain ( Figure 23 ). Therefore, we could independently reassign TCA and TCG codons to Ala, His, Leu, or Pro in Syn61Δ3, thereby creating 16 new genetic codes (Fig. S3, Data File S1, Figure 2 B). In each new genetic code, we varied the identity of the canonical amino acid encoded at a particular sense codon relative to both the canonical codon and the 15 other codons we created ( Figure 2 C).
[0326] In summary, we have restructured the genetic code. Our new genetic code expands the number of codons used to encode Ala and Pro (from 4 to 6), doubles the number of codons used to encode His from 2 to 4, and increases the number of codons used to encode Leu from 6 to 8; this is more codons than are used to encode any amino acid using the canonical code. These experiments also show that the UCN codon box that encodes serine using the canonical code can be split to encode additional canonical amino acids.
[0327] Example 4 - Orthogonal Cipher-Orthogonal Decoder Pair
[0328] Genes written using the canonical genetic code, in which TCG and TCA codons encode serine, will make the correct protein product in native cells that read these codons as serine. However, these genes will produce incorrect (and possibly nonfunctional) protein products in cells that decode these codons to incorporate amino acids other than serine.
[0329] Similarly, synthetic genes in which we use the Syn61 recoding protocol to condense the genetic code and replace the codons for specific natural amino acids with TCG and TCA codons will make the correct protein products in cells that decode the TCG and TCA codons to incorporate the correct amino acids. However, these synthetic genes will produce incorrect (and possibly non-functional) protein products in cells that read the natural genetic code ( Figure 3 ).
[0330] In recSpecR, we converted all 27 GCN codons (which encode alanine as the canonical codon) to TCG codons, and all 6 CAT / C codons (which encode histidine as the canonical codon) to TCA codons (ΔTCG, TCA). This created the orthogonal resistance gene O-SpecR (TCG-Ala, TCA-His). We demonstrated that O-SpecR (TCG-Ala, TCA-His) can be expressed in the presence of Syn61Δ3 (tRNA CGA Ala , tRNA UGA His ) cells, in which TCG is read as Ala and TCA is read as His. We further demonstrated that O-SpecR (TCG-Ala, TCA-His) does not confer spectinomycin resistance in Syn61WT cells, in which TCG and TCA are decoded as Ser, as in the canonical genetic code. Finally, we demonstrated that SpecR WT, in which serine is encoded using TCG and TCA codons, does not confer spectinomycin resistance in Syn61Δ3(tRNA CGA Ala , tRNA UGAHis ) cell spectinomycin resistance ( Figure 3 We extended this approach to five other reassignment schemes, as well as to other genes (Fig. 9). We obtained similar results for the wt hygromycin resistance gene and a hygromycin resistance gene in which all Ala codons were converted to TCG and all histidine codons were converted to TCA (O-HygR (TCG-Ala, TCA-His)) (Fig. S4B).
[0331] These experiments demonstrate that we can create a genetic code-decoder pair for synthetic genes that is functionally orthogonal to the canonical genetic code-decoder pair of natural genes. The orthogonal codons (TCG-Ala, TCA-His) written into the synthetic gene are converted by the orthogonal decoders (tRNA CGA Ala , tRNA UGA His ) is read correctly, but cannot be read by the canonical decoder (tRNA UGA Ser ) read. The canonical code (TCG-Ser, TCA-Ser) written in the natural gene is correctly read by the canonical decoder but cannot be read by the orthogonal decoder.
[0332] The functional orthogonality of genes in cells with altered decoders will depend on the frequency of codon reallocation and the functional consequences of codon reallocation. The results of amino acid substitutions, which are the result of codon reallocation, can be generally and roughly related to differences in amino acid polarity and hydrophobicity (6, 21). Computational methods using evolved sequence and / or structural information can predict the results of amino acid substitutions at specific sites in proteins (22-25). Although the composition of natural genes is fixed, the codon usage in synthetic genes (written in standard codes or any orthogonal codes) can be simply designed to maximize the number of codons that are reallocated, thereby maximizing the functional orthogonality of synthetic genes.
[0333] Example 5 - Orthogonal Horizontal Gene Transfer
[0334] Next, building on the orthogonal genetic code-decoder pair, we created an orthogonal horizontal gene transfer (O-HGT) system consisting of an orthogonal decoder and a mobile genetic element using the orthogonal genetic code. WT cells can transfer the WT mobile genetic element between themselves but cannot transfer the WT mobile genetic element to cells containing the orthogonal decoder. Cells containing the O-HGT system can transfer their mobile genetic element to cells containing a compatible orthogonal decoder but cannot transfer their mobile genetic element to cells containing an incompatible orthogonal decoder or to WT cells.
[0335] As expected, a mobile genetic element (F plasmid, F(WT)) using the canonical genetic code was transferred to WT cells (Syn61 WT). We also showed that F(WT) could not be transferred to Syn61Δ3 (tRNA CGA Ala , tRNA UGA His ) cells, in which TCG codons are read as Ala and TCA codons are read as His ( Figure 4 ).
[0336] Next, we studied the horizontal gene transfer of the mobile genetic element with the genetic code that changes.We synthesized the mobile genetic element O-F1 (TCG-Ala, TCA-His).The genetic code in all the annotated open reading frames of this F plasmid is compressed using the Syn61 scheme, and the GCN codon (it encodes alanine in the canonical code) and the CAT / C codon (it encodes histidine in the canonical code) are converted to TCG and TCA codons respectively in the trfA gene, and the trfA gene is crucial for the replication of the mobile genetic element.
[0337] O-F1 (TCG-Ala, TCA-His) was horizontally transferred to Syn61Δ3 (tRNA CGA Ala , tRNA VGA His ) cells. We further demonstrated that O-F1(TCG-Ala, TCA-His) was not horizontally transferred to cells that read the canonical genetic code ( Figure 4 ). These experiments showed that we can create an O-HGT system.
[0338] Next, we created mutually orthogonal HGT systems that are orthogonal to the natural genetic system and to each other. We created new mobile genetic elements O-F2 (TCG-His, TCA-Ala). The genetic codes in all annotated open reading frames of the F plasmid were compressed using the Syn61 scheme, and GCN codons (which encode alanine in the canonical code) and CAT / C codons (which encode histidine in the canonical code) were converted to TCA and TCG codons in the trfA gene, respectively.
[0339] We demonstrated that O-F2 (TCG-His, TCA-Ala) is able to transfer to Syn61Δ3 (tRNA) in which TCG is decoded as His and TCA is decoded as Ala. CGA His , tRNA UGA Ala) cells. In contrast, O-F2(TCG-His, TCA-Ala) was not transferred into Syn61 cells, which use the canonical genetic code to decode TCG and TCA codons as serine. O-F2(TCG-His, TCA-Ala) was not transferred into Syn61Δ3(tRNA CGA Ala , tRNA UGA His ) cells. In addition, we demonstrated that neither the WT mobile genetic element (F(WT; TCG-Ser, TCA-Ser) nor the O-F1 (TCG-Ala, TCA-His) was transferred to Syn61Δ3(tRNA CGA His , tRNA UGA Ala )Intracellular( Figure 4 ). Further experiments demonstrating the HGT system are shown in Figure 24. In summary, we demonstrated the scalability of our approach by creating five mutually orthogonal horizontal gene transfer systems.
[0340] These experiments demonstrate that we can create orthogonal and mutually orthogonal HGT systems.
[0341] Example 6 - Orthogonal password lock to prevent intrusion password
[0342] We hypothesized that replacing codons for specific natural amino acids in essential genes with TCA and TCG codons and adding tRNAs that reassign these codons to specific natural amino acids would prevent the serT-mediated HGT we observed into Syn61Δ3 ( Figure 1 ).
[0343] In the absence of spectinomycin, F(WT+serT) was transferred to Syn61Δ3(tRNA CGA Ala , tRNA UGA His , O-SpecR (TCG-Ala, TCA-His)) transfer was blocked (10 4 times), because tRNA CGA Ala and tRNA UGA His tRNA in the recipient cell UGA SerCompetition reduces the production of functional proteins from mobile genetic elements. However, this blockade is not sufficient to completely eliminate the transfer of mobile genetic elements. After adding spectinomycin to make O-SpecR (TCG-Ala, TCA-His) an essential gene in the cell, decoding TCG codons into alanine and TCA codons into histidine becomes essential. Under these conditions, the transfer of mobile genetic elements is completely eliminated ( Figure 5 Similar results were obtained for other reconstructed codes and other essential genes ( Figure 25 , Figure 26 ).
[0344] To expand our approach to viral infection, we identified pools of phage from River Cam that could infect Syn61Δ3 (Methods). From these pools, we isolated two individual phages (12 and 06, both T4-like phages) that carry the same tRNA UGA Ser gene and infected with Syn61Δ3 ( Figure 27 ); some viruses are known to carry their own tRNAs and other translation factors to enhance the cellular pool of translation factors and aid in translating codons within their own genes. As expected, expression of this tRNA in Syn61Δ3 was sufficient to confer susceptibility to infection by the (otherwise non-infectious) T4 phage ( Figure 28 We demonstrated that, unlike Syn61Δ3, several reconstructed codon-locked strains were completely resistant to infection with phage 6 and phage 12 ( Figure 5 , Figure 29 ).
[0345] Our results suggest that writing essential genes in an orthogonal code and reading these genes with a cognate orthogonal decoder creates cells that are locked into the orthogonal code. These cells resist invasion by mobile genetic elements using a competing code.
[0346] Example 7 - Genetic code locking to achieve stable phage resistance in synthetic organisms
[0347] The experiments discussed in the previous examples identified a gene encoding a seryl-tRNA (tRNA Ser UGA ) phage and showed that this phage can infect Syn61Δ3. These experiments also showed that we can eliminate infection by these phages by code locking ( Figure 34 However, we also show that, in contrast to conjugative transfer, redistribution is sufficient to abrogate plaque formation. Further experiments were performed to determine why this was the case.
[0348] We speculate that genomic differences may explain the phenotypic differences; one possible explanation is the number of TCG and TCA codons in each genome. The genome of the phage studied here is much larger than that of F'(WT+serT), and the total number of target codons is more than three times ( Figure 30 a) As more positions are affected by amino acid misincorporation, the chance of a deleterious effect is higher.
[0349] It could also be the genomic frequency of the target codon. Compared to F'(WT+serT), the phage genome showed an approximately 25% increase in the frequency of the target codon in its genome. As the frequency of amino acid misincorporation increases, the probability of the open reading frame being non-functionally translated is higher.
[0350] These two effects could be additive. Not only are more genes affected in the phage genome, but these genes are also affected to a greater extent on average. Therefore, we would expect codon reallocation to have a greater impact on phage infection than on conjugative transfer. Although it is possible in principle that the relative usage of TCG and TCA codons in the genomes of phages 12, 6, and F'(WT+serT) could also be responsible, we consider this unlikely, as the reallocation of these two codons to leucine drives the differences between phage infection and conjugative transfer.
[0351] It is possible that mechanistic differences could explain the phenotypic differences. The two modes of horizontal gene transfer, conjugative transfer and phage infection, are fundamentally different processes. For successful phage infection and plaque formation, the entire phage life cycle must be completed. This includes attachment to cells, injection of the viral genome, production of viral proteins, phage genome replication, and maturation of phage particles ( Figure 30 c) Plaque formation occurs only when mature phage particles form and successfully infect neighboring cells. Completion of this entire life cycle is a highly complex process, requiring many components to work together under strict time constraints. For T4-like phages, such as phages 12 and 6, multiple genes are crucial for this process. Defects in any of these genes lead to abolition of plaque formation.
[0352] In contrast, conjugative transfer and subsequent colony formation is a much simpler process. After the donor cell attaches to the recipient, the ssDNA is transferred to the recipient through the mating channel. The DNA is then recycled to form a stable dsDNA plasmid in the recipient cell. Importantly, all proteins involved in this process are either expressed from the recipient cell genome or transferred from the donor along with the DNA. Subsequently, for successful colony formation, it is only necessary to ensure the replication of the plasmid and its proper segregation during cell division ( Figure 30 d) This process requires very few genes from the conjugative elements to be functionally expressed.
[0353] Since more genes need to be functionally expressed from the horizontally transferred DNA for successful plaque formation, ambiguous decoding by codon reallocation is expected to be more detrimental to phage infection. If an essential gene is disrupted enough to render its product nonfunctional, plaque formation can be eliminated.
[0354] A further explanation is the dominant negative effect caused by ambiguous decoding of certain genes. Proteins that form complex interactions are more likely to produce dominant negative effects. Some mutations in viral envelope proteins are known to show dominant negative phenotypes. Mutations in envelope subunits are likely to interrupt oligomerization and correct particle assembly. In T4-like phages, the major capsid protein (gp23) forms a hexamer, which is the basis for particle assembly. In phages 06 and 12, there are three surface-exposed serine residues encoded by TCA ( Figure 32 Misincorporation of an amino acid at one of these positions in a subset of gp23 can disrupt capsid assembly and thereby have a dominant negative effect on plaque formation.
[0355] We recognize that the advantage of codon lock may be to maintain the alternative genetic code over time. The tRNA responsible for codon reconstruction can be inactivated by a variety of mechanisms, such as mutation, deletion, or silencing. This will essentially revert cells with reconstructed genetic codes back to codon-compressed cells with decoder deletions (such as Syn61Δ3) and make them susceptible to infection by phages carrying appropriate tRNA genes. However, if the codon is locked, the tRNA responsible for reconstruction is crucial and cannot be inactivated. This ensures the temporal stability of the reconstructed codon and acute phage resistance.
[0356] We modeled the stability of alternative genetic codes in the presence and absence of codon lock. tRNAs responsible for alternative decoding of TCG and TCA codons were encoded on a low-copy plasmid carrying hygromycin resistance. A second plasmid encoding a variant of the spectinomycin resistance gene (SpecR); for cells without codon lock: recSpec R (ΔTCG, ΔTCA), for cells with code lock, respectively: oSpec R (TCG: Ala, TCA: His) and oSpec R (TCG:Leu, TCA:Leu). Cells were serially passaged in the presence of spectinomycin and the absence of hygromycin (no inherent pressure to maintain the plasmid). At each passage, we measured the fraction of cells that retained the tRNA-encoding plasmid. We found that codon locking stabilizes the alternative codon and acts to maintain the codon ( Figure 31 a).
[0357] Thus, codon-locked cells maintain phage resistance over time, while non-locked cells lose resistance. We exposed cells with and without locked codons from the above time course to phages 12 and 06. We observed that cells with locked genetic codes remained resistant to phage infection, while cells that had lost the tRNA responsible for codon remodeling were susceptible to phage infection ( Figure 31 b).
[0358] These experiments also showed that plasmids can be stably maintained by making them essential to the host based on the genetic code. For example, due to the tRNA on the plasmid and its necessity for decoding essential genes. This can be used to avoid antibiotic remittance in biotechnology applications.
[0359] Example 8 - Phage reproduction assay
[0360] Dilute cells from overnight cultures to an OD 600 The cells were inoculated with phage 12 (MOI = 0.001) and incubated for 24 h in a 3 mL (2xty) volume. Phage titers were assessed by serial dilution (7.5 uL spotting on a layer of top agar) and plaque assay against a permissive strain (Syn61 WT). The control was a blank culture medium (2xty) in which no cells were present. The limit of detection for this assay was 1 plaque per 7.5 uL (133.3 PFU / mL).
[0361] The results are shown in Figure 33 The results showed that phage 12 successfully replicated in Syn61WT and Syn61Δ3 cells, but not in cells with a remodeled and locked genetic code (Syn61Δ3(alaTCGA, hisRtga) and Syn61Δ3(leuQcga, leuQtga)). Phage propagation in code-locked cells resulted in lower phage titers compared to the no-cell control (ctrl.). This is likely because the phage adsorbed to those cells but was unable to replicate. Experiments were performed in biological triplicate.
[0362] discuss
[0363] We created 16 synthetic genetic codes; in each new code, a subset of the sense codons was reassigned to amino acids that differed from the canonical code. Codon reassignment restructured the genetic code and directly altered the number and type of amino acids accessible through point mutations (Figs. S5, S6). Previous experimental work has shown that selecting synonymous codons in individual genes and viruses can alter their robustness and evolvability (26-28), but such approaches have been limited to exploring a subset of the canonical code. Although a large body of theoretical work and a limited number of in vitro experiments have considered the relationship between the structure of the genetic code and its robustness and evolvability (19, 20), the resulting hypotheses cannot be investigated experimentally in living cells. Reconstructing the structure of the genetic code provides new opportunities to experimentally test how the altered code affects the robustness and evolvability of proteins and cellular function. In future work, we aim to use genetic code reconfiguration to accelerate directed evolution.
[0364] We have experimentally exemplified the creation of semantic orthogonality between organisms that use different reassigned codes. We have clearly shown that semantic orthogonality creates functional orthogonality for the genes tested; mismatches between the genetic code used to write a gene and the decoder used to read it lead to the missynthesis of nonfunctional proteins.
[0365] We have created multiple, mutually orthogonal HGT systems in which genes can only be correctly read and transferred by cells with the cognate decoder. Each cell type, with a different code decoder system, implements a different, reconstructed genetic code. These systems allow experimental investigation of the role of HGT in repairing the universal genetic code through competition between pools of genotypes written with different codes (4).
[0366] Protecting synthetic organisms from environmental genetic elements could be valuable for industrial-scale biotechnology applications, where contamination by mobile genetic elements (including viruses) can lead to economic losses and disrupt important supply chains (29). Resistance to natural gene horizontal transfer into organisms with genomic code compression and tRNA deficiency could be circumvented by retrieving missing tRNAs and by using mobile genetic elements that carry these tRNAs. Indeed, mobile genetic elements (including viruses) carry their own tRNAs and other translation factors, which enhance the cellular pool of translation factors and contribute to the translation of codons within their own genes (9, 30, 31).
[0367] Synthetic organisms with essential genes written in an orthogonal genetic code, and a decoder that correctly reads the orthogonal code, confer complete resistance to the transfer of mobile genetic elements written in the canonical code, even when the mobile genetic elements contain tRNAs that would allow the cell to correctly read the canonical code. This defines a paradigm for creating organisms that actively resist invasion by foreign codes.
[0368] New strategies to limit the transfer of genetic information from synthetic organisms to natural organisms could form the basis of genetic firewalls that isolate synthetic genetic systems from the environment. This is an important challenge that complements the challenge of controlling the survival and growth of synthetic organisms for biocontainment, especially when considering the use of engineered organisms outside the laboratory (32). All compressed genetic codes are subsets of the natural code and are correctly read by decoders of the full code; genetic systems written in compressed genetic codes are correctly read by natural organisms. Therefore, compressed genetic codes cannot be used to genetically isolate synthetic organisms from natural organisms. The ability to restructure the genetic code and write genes that are correctly read in synthetic organisms but incorrectly read in natural organisms provides the basis for a powerful strategy to prevent the transfer of genetic information from synthetic organisms to natural organisms. Importantly, this strategy is globally applicable to any gene or genetic system added to a synthetic organism. Because the genetic code is almost universally conserved, we anticipate that the principles we have established will be applicable to a wide range of other organisms.
[0369] References and Notes
[0370] 1.FHCrick, L.Barnett, S.Brenner, RJWatts-Tobin, General nature of the genetic code for proteins. Nature 192, 1227-1232 (1961).
[0371] 2.MWNirenberg, JHMatthaei, The dependence of cell-free proteins synthesis in E.coli upon naturally occurring or synthetic polyribonucleotides. Proc Natl Acad Sci USA 47, 1588-1602 (1961).
[0372] 3.R.J.Hall,F.J.Whelan,J.O.McInerney,Y.Ou,M.R.Domingo-Sananes,Horizontal Gene Transfer as a Source of Conflict and Cooperation inProkaryotes.Front Microbiol 11,1569(2020).
[0373] 4.K.Vetsigian,C.Woese,N.Goldenfeld,Collective evolution and thegenetic code.Proc Natl Acad Sci U S A 103,10696-10701(2006).
[0374] 5.D.de la Torre,J.W.Chin,Reprogramming the genetic code.Nat Rev Genet22,169-184(2021).
[0375] 6.E.V.Koonin,A.S.Novozhilov,Origin and evolution of the genetic code:the universal enigma.IUBMB Life 61,99-111(2009).
[0376] 7.M.Kollmar,S.Muhlhausen,Nuclear codon reassignments in the genomicsera and mechanisms behind their evolution.Bioessays 39,(2017).
[0377] 8.J.Ling et al.,Natural reassignment of CUU and CUA sense codons toalanine in Ashbya mitochondria.Nucleic Acids Res 42,499-508(2014).
[0378] 9.A.L.Borges et al.,Widespread stop-codon recoding in bacteriophagesmay regulate translation of lytic genes.Nat Microbiol 7,918-927(2022).
[0379] 10.M.A.Santos,A.C.Gomes,M.C.Santos,L C.Carreto,G.R.Moura,The geneticcode of the fungal CTG clade.C R Biol 334,607-611(2011).
[0380] 11.D.J.Taylor,M.J.Ballinger,S.M.Bowman,J.A.Bruenn,Virus-host co-evolution under a modified nuclear genetic code.PeerJ 1,e50(2013).
[0381] 12.Y.Shulgina,S.R.Eddy,A computational screen for alternative geneticcodes in over 250,000 genomes.Elife 10,(2021).
[0382] 13.D.G.Gibson et al.,Complete chemical synthesis,assembly,and cloningof a Mycoplasma genitalium genome.Science 319,1215-1220(2008).
[0383] 14.D.G.Gibson et al.,One-step assembly in yeast of 25 overlapping DNAfragments to form a complete synthetic Mycoplasma genitalium genome.Proc NatlAcad Sci U S A 105,20404-20409(2008).
[0384] 15.J.Fredens et al.,Total synthesis of Escherichia coli with arecoded genome.Nature 569,514-+(2019).
[0385] 16.F.J.Isaacs et al.,Precise manipulation of chromosomes in vivoenables genome-wide codon replacement.Science 333,348-353(2011).
[0386] 17.M.J.Lajoie et al.,Genomically recoded organisms expand biologicalfunctions.Science 342,357-360(2013).
[0387] 18.W.E.Robertson et al.,Sense codon reassignment enables viralresistance and encoded polymer synthesis.Science 372,1057-1062(2021).
[0388] 19.G.Pines,J.D.Winkler,A.Pines,R.T.Gill,Refactoring the Genetic Codefor Increased Evolvability.mBio 8,(2017).
[0389] 20.J.Calles,I.Justice,D.Brinkley,A.Garcia,D.Endy,Fail-safe geneticcodes designed to intrinsically contain engineered organisms.Nucleic AcidsRes 47,10439-10451(2019).
[0390] 21.M.Schmidt,V.Kubyshkin,How To Quantify a Genetic Firewall?APolarity-Based Metric for Genetic Code Engineering.Chembiochem 22,1268-1284(2021).
[0391] 22.D.S.Marks,S.W.Michnick,Democratizing the mapping of gene mutationsto protein biophysics.Nature 604,47-48(2022).
[0392] 23.S.Teng,A.K.Srivastava,C.E.Schwartz,E.Alexov,L.Wang,Structuralassessment of the effects of amino acid substitutions on protein stabilityand protein protein interaction.Int J Comput Biol Drug Des 3,334-349(2010).
[0393] 24.V.Parthiban,M.M.Gromiha,D.Schomburg,CUPSAT:prediction of proteinstability upon point mutations.Nucleic Acids Res 34,W239-242(2006).
[0394] 25.P.C.Ng,S.Henikoff,Predicting the efiects of amino acidsubstitutions on protein function.Annu Rev Genomics Hum Genet 7,61-80(2006).
[0395] 26.B.A.Renda,M.J.Hammerling,J.E.Barrick,Engineering reducedevolutionary potential for synthetic biology.Mol Biosyst 10,1668-1678(2014).
[0396] 27.G.Moratorio et al.,Attenuation of RNA viruses by redirecting theirevolution in sequence space.Nat Microbiol 2,17088(2017).
[0397] 28.J.R.Coleman et al.,Virus attenuation by genome-scale changes incodon pair bias.Science 320,1784-1787(2008).
[0398] 29.P.W.Barone et al.,Viral contamination in biologic manufacture andimplications for emerging therapies.Nat Biotechnol 38,563-572(2020).
[0399] 30.P.Alamos et al.,Functionality of tRNAs encoded in a mobile geneticelement from an acidophilic bacterium.RNA Biol 15,518-527(2018).
[0400] 31.T.Tuller et al.,Association between translation efficiency andhorizontal gene transfer within microbial communities.Nucleic Acids Res 39,4743-4755(2011).
[0401] 32. JW Lee, CTY Chan, S. Slomovic, JJ Collins, Next-generation biocontainment systems for engineered organisms. Nat Chem Biol 14, 530-537 (2018).
[0402] Materials and methods
[0403] strains
[0404] Throughout the text, Syn61 WT refers to Syn61(ev2) (18) and Syn61Δ3 refers to Syn61Δ3(ev4) (18).
[0405] Gene recoding
[0406] For all genes and plasmids used in experiments with Syn61A3-derived cells, it was necessary to condense the genetic code according to the recoding rules of Syn61 (TCG and TCA codons were replaced by AGC and AGT, respectively, and the TAG stop codon was replaced by TAA). We recoded the open reading frames as previously described for Syn61 (15). The plasmids used in this study are provided in Zürcher et al., “Refactored genetic code enables bidirectional genetic isolation,” data file S2, Science, and are provided herein as Table 1.
[0407] Construction of tRNA plasmids for decoding TCG and TCA codons in Syn61Δ3
[0408] To incorporate amino acids in response to TCG and TCA codons, we used the pSC101-Kan and pSC101-Hyg plasmids (conferring resistance to kanamycin and hygromycin, respectively) into which we cloned the genes encoding the relevant tRNA or tRNAs. No exogenous aaRS was used, as all tRNAs used in this study were acylated by endogenous E. coli aaRSs. For tRNAs that incorporate amino acids other than serine in response to TCG and TCA codons, we designed genes in which the anticodons of the relevant isoacceptor tRNAs were replaced by CGA and UGA, respectively.
[0409] We constructed pSC101-based tRNA plasmids using HiFi assembly of multiple fragments. Two different tRNA plasmid constructs were used: i) tRNA was expressed under the lpp promoter, and the pSC101 backbone included a pheS-HygR dual selection cassette expressed under the EM7 promoter, and ii) tRNA was expressed using the native expression environment of serT in the E. coli genome, and the pSC101 backbone included a kan expressed under the T3 promoter. R In cases where two tRNAs were expressed from a single plasmid, they were expressed as an operon using the intergenic region between the alaX and alaW tRNA genes in the E. coli genome. Backbone fragments were generated via PCR. tRNAs and tRNA operons were sequenced as oligonucleotides (Merck) or gBlocks (IDT). All cloning was performed in Syn61Δ3.
[0410] Construction of recoded antibiotic plasmids for evaluating genetic code orthogonality
[0411] To evaluate the functionality of antibiotic genes encoded according to different genetic codes, we used a pMB1-based plasmid containing codon-squeezed antibiotic resistance genes into which we cloned genes encoding hygromycin or spectinomycin resistance (data file S2). We constructed a pMB1-based tRNA plasmid using HiFi assembly of multiple fragments. Backbone fragments were generated via PCR. The recoded spectinomycin and hygromycin resistance genes were sequenced as gBlocks (IDT).
[0412] Construction of recoded mobile genetic elements
[0413] Intermediate F(ΔTCG, TCA, TAG) was constructed from synthetic DNA (TWIST Bioscience) via yeast assembly (33). F(ΔTCG, TCA, TAG) was designed by recoding all annotated open reading frames in the RK2 conjugative plasmid as previously described (15). Recoding of trfA was performed by lambda red recombination to produce O-F1(TCG-Ala, TCA-His) and O-F2(TCG-His, TCA-Ala) (34). The recoded versions of trfA were synthesized as gBlocks (IDT). All modifications were performed in E. coli Dh10B. In order to make F plasmids lacking the trfA gene or encoding trfA in a genetic code that cannot be decoded by Dh10B replicable, a pMB1-based helper plasmid was used that expresses WT trfA in its endogenous environment in the RK2 conjugative plasmid. This plasmid contains the amp R amp expressed under the promoter Rgenes and assembled by HiFi assembly from fragments generated by PCR.
[0414] sfGFP expression measurement
[0415] We expressed the sfGFP-His6 gene carrying a single TCG or TCA codon at position 3 in Syn61Δ3 cells containing plasmids encoding tRNA or tRNA operons. We electroporated 50 μL of Syn61Δ3 cells with pBAD_sfGFP reporter plasmid (100 ng) and recovered the cells in 1 mL of SOB for 90 min while shaking at 1050 rpm at 37 ° C. Subsequently, we inoculated the recovered culture (1 mL) into 5 mL of preheated 2xYT medium containing 50 μg / mL apramycin and incubated the cells overnight at 37 ° C while shaking at 220 rpm before preparing electrocompetent cells. We electroporated pSC101-based tRNA plasmids (100 ng) into Syn61Δ3 cells with pBAD_sfGFP and recovered the cells in 500 μL SOB in deep well 96-well plates for 90 min. Subsequently, we inoculated one-tenth of the recovered culture (50 μL) into 450 μL of preheated 2xYT medium supplemented with 200 μg / mL hygromycin and 50 μg / mL apramycin. After 36 h of recovery, 37°C, and 750 rpm, we set up the expression in a 96-well microtiter plate format, inoculating the overnight culture into 500 μL of preheated 2xYT containing hygromycin (200 ng / μL), apramycin (50 ng / μL) and L-arabinose (0.2%) at a ratio of 1:50. The expression was incubated at 37°C for 16 h while shaking at 750 rpm. The plate was centrifuged at 3200 g for 10 min. We resuspended the cell pellet in 150 μL of PBS and transferred 100 μL of it to a Costar transparent 96-well flat-bottom plate. In this plate, we recorded the OD on a PHERAstar FS plate reader (BMG LABTECH). 600 and GFP fluorescence (λ ex :485nm;λ em : 520nm) measurement value (gain set to 0, focus adjusted to 00mm).
[0416] To determine protein yield, sfGFP (WT) was expressed in Syn61Δ3 (16 h at 37°C in 5 mL of 2xTY + 0.2% arabinose) and purified as described below (three elutions, 100 μL each). The protein concentration after purification was determined by nanodrop (elution 1: 0.77 mg / mL; elution 2: 0.09 mg / mL; elution 3: 0.0 mg / mL). The total amount measured was 0.086 mg of protein extracted from 5 mL of culture, equivalent to a protein yield of 17 mg / L of culture.
[0417] Purification of sfGFP-His6x and ubiquitin-His6x proteins
[0418] Syn61Δ3 cells containing pSC101-based tRNA plasmids and pBAD_sfGFP (or ubiquitin) plasmids were cultured in 5 mL (20 mL for ubiquitin) 2xYT medium containing 200 μg / mL hygromycin, 50 μg / mL apramycin, and 0.2% L-arabinose at 37°C for 16 h while shaking at 220 rpm. After expression, the cells were centrifuged, resuspended in 1 mL lysis buffer (1x Bugbuster protein extraction reagent (Novagen), 1x PBS, 50 μg / mL DNase1, 20 mM imidazole, and 100 μg / mL lysozyme), and incubated at 4°C for 1 h. The resulting lysate was centrifuged (16,000 x g) for 30 min at 4°C. The supernatant was then transferred to a 50 μL Ni 2+ -NTA slurry (Qiagen) in 1.5 mL microcentrifuge tubes and incubated at 4 °C for 1 h while mixing by inversion. Ni was collected by gravity filtration on a sintered column. 2+ -NTA beads and washed three times in 500 μL wash buffer (1× PBS, 40 mM imidazole). Finally, the protein was eluted in 100 μL elution buffer (1× PBS, 300 mM imidazole, pH 8) and collected in a fresh microcentrifuge tube via centrifugation (100×g, 4° C., 1 min).
[0419] Complete protein profile
[0420] The protein (ubiquitin) was characterized using a Waters Xevo G2 mass spectrometer coupled to a modified nanoAcquity LC system. Figure 13ESI-MS analysis was performed for the purified samples (as described above) on a BEH C4 UPLC column (1.7 μm; 1.0x100 mm; Waters) over 20 min with a flow rate of 50 μL / min and a water / acetonitrile gradient of 2% vol / vol to 80% vol / vol. Subsequently, the eluted sample was connected to a mass spectrometer (Waters) via a Zspray electrospray ionization source. Data were acquired in positive ion mode with a range of 300-2000 m / z and an applied cone voltage of 30 V. The spectrum was deconvoluted using the MaxEnt1 function in MassLynx software (Waters). To calculate the expected molecular weight, the expected mass of the wild-type protein was determined using GPMAW (LighthouseData) and then manually edited to accommodate the encoded amino acid changes.
[0421] The protein (sfGFP) was characterized using a Waters Vion IMS Qtof mass spectrometer coupled to a modified nanoAcquity LC system. Figure 2 B. Figure 14 、 Figure 15 ) were subjected to ESI-MS analysis. The purified samples (as described above) were separated on an Acuity UPLC protein BEHC4 column (1.7 μm; 2.1 x 50 mm; Waters) over 7 min with a flow rate of 200 μL / min and a water / acetonitrile gradient of 5% vol / vol to 100% vol / vol. The eluted samples were then connected to a mass spectrometer (Waters) via a Zspray electrospray ionization source. Data were acquired in positive ion mode with a range of 100-2000 m / z. The spectra were deconvoluted using the MaxEnt1 function within Unifi software (Waters). To calculate the expected molecular weight, the expected mass of the wild-type protein was determined using GPMAW (LighthouseData) and then manually edited to accommodate the encoded amino acid changes.
[0422] Mass spectra of ubiquitin in the screen of anticodon-modified tRNAs were acquired on an Agilent 1200 LC-MS system equipped with a 6130 quadrupole spectrometer ( Figure 18 The protein was eluted from a Phenomenex Jupiter C4 column (150 x 2 mm, 5 μm). Reverse-phase HPLC was performed using buffer A (0.2% formic acid in water) and buffer B (0.2% formic acid in acetonitrile (MeCN). Mass spectra were acquired in positive mode and analyzed using MS Chemstation software (Agilent Technologies). The entire mass spectrum was acquired using the deconvolution program provided in the software.
[0423] Calculating signal-to-noise ratios and fidelity measurements in ESI-MS spectra
[0424] For the ESI spectrum of sfGFP, we calculated the average signal intensity and standard deviation of the intensity between 20,000 Da and 27,000 Da. We defined noise as the average signal intensity plus twice the standard deviation of this mass window. The limit of the fidelity measurement was calculated as: (1-(N / S)) x 100. (Note: For the spectra in Figure S3, the baseline signal was determined between 20,000 Da and 26,500 Da due to the presence of a peak at approximately 27,000 Da from degradation products). This calculation defines the maximum fidelity that can be obtained from the spectrum, and we note that true biological fidelity can be higher.
[0425] To determine the presence of tRNA CGA XXX and tRNA UGA YYY To decode the specificity of the TCG codon in both cases, where XXX and YYY are different amino acids, we divided the signal intensity at the peak resulting from the incorporation of XXX at TCG by the signal intensity at the expected mass for the incorporation of YYY at TCG. The intensity at the expected mass for the incorporation of YYY at TCG was determined as the maximum signal in a 2 Da window around the theoretically calculated mass.
[0426] To determine the presence of tRNA CGA XXX and tRNA UGA YYY To decode the specificity of the TCA codon in both cases, where XXX and YYY are different amino acids, we divided the signal intensity at the peak resulting from the incorporation of YYY at TCA by the signal intensity at the expected mass for the incorporation of XXX at TCA. The intensity at the expected mass for the incorporation of XXX at TCA was determined as the maximum signal in a 2 Da window around the theoretically calculated mass.
[0427] Western blot of cell lysates from ubiquitin expression experiments
[0428] Syn61Δ3 cells containing a pSC101-based tRNA plasmid and a pBAD_ubiquitin plasmid were grown in 20 mL of 2xTY medium containing 200 μg / mL hygromycin, 50 μg / mL apramycin, and 0.2% L-arabinose at 37°C for 16 h while shaking at 220 rpm. The culture was normalized to an OD600 of 1.0. 500 μL of the normalized culture was lysed with sample buffer (Nupage buffer, 10% β-mercaptoethanol, PMSF) and vortexed vigorously to shear the DNA. The sample was separated by SDS-PAGE (NuPAGE 4-12% in MES buffer) and transferred to a polyvinylidene fluoride (PVDF) membrane using an iBlot 2 dry blotting system (Thermo Fisher Scientific). The membrane was blocked with Odyssey blocking buffer (Cat. No. 927-40000, Li-Cor) in PBS for 30 min at room temperature. The membrane was incubated overnight at 4°C with an anti-His-tag primary antibody (Abcam, catalog number ab18184) in primary antibody solution (diluted 1:1000 in Odyssey T20 (PBS) antibody diluent (927-75001, Li-Cor). All incubations were performed on a platform shaker. The membrane was washed three times with PBST (PBS supplemented with 0.1% Tween-20 (v / v)) and incubated with the secondary antibody goat anti-mouse IRDye680RD 925-68070 (1:15,000 (v / v) in PBS blocking buffer supplemented with 0.2% Tween-20 (v / v) and 0.01% SDS at room temperature for 1 hour. After washing three times with PBST and once with PBS, immunoreactive proteins were visualized on a Typhoon Trio phosphorimager (GE Life Sciences). Samples analyzed by Western blotting were also separated by SDS-PAGE, and the gel was stained with Instant Blue (Expedeon) for 30 min and then washed with water.
[0429] MS / MS of ubiquitin variants
[0430] Solution samples were reduced with dithiothreitol at 37°C and alkylated with chloroacetamide at room temperature in the dark. The samples were digested with LysC (Promega) at 37°C for 4 h, followed by trypsin (Promega) digestion at 37°C overnight. The peptide mixture was acidified and desalted using a homemade C18 (3M Impore) stage tip containing 3 μl of Poros Oligo R3 (Thermo Fisher Scientific) resin. Bound peptides were eluted from the stage tip with 30-80% acetonitrile (MeCN) and partially dried in a Speed Vac (Savant).
[0431] Peptides were separated on an Ultimate 3000 RSLC nano system (Thermo Scientific) equipped with a 75 μm × 25 cm nanoEase C18 T3 column (Waters) using mobile phases of buffer A (2% MeCN, 0.1% formic acid) and buffer B (80% MeCN, 0.1% formic acid). The eluted peptides were directly introduced into a Q Exactive Plus hybrid quadrupole-Orbitrap mass spectrometer (Thermo Fisher Scientific) via a nanospray ion source. The mass spectrometer was operated in data-dependent mode. MS1 spectra were acquired from 380–1600 m / z at a resolution of 70,000, followed by MS2 spectra of the 15 most intense ions at a resolution of 17,500 and an NCE of 27%. An MS target value of 1e6 and an MS2 target value of 1e5 were used. Dynamic exclusion was set to 30 seconds.
[0432] The acquired raw data files were searched against the E. coli UniProt Fasta database (downloaded September 2022) using MaxQuant with an integrated Andromeda search engine (v.1.6.17.0), along with 20 additional ubiquitin sequences (each with a different canonical amino acid at position 11). Carbamidomethylation of cysteine was set as a fixed modification, while oxidation of methionine was set as a variable modification. Enzyme specificity was set to trypsin / p, and a maximum of two missed cleavages were allowed.
[0433] Preparation and electroporation of electrocompetent Syn61Δ3 cells
[0434] Inoculate 250 mL of pre-warmed 2xYT medium with 5 mL of Syn61Δ3 overnight culture and grow at 37°C with shaking (220 rpm) to an OD of 600The cell of 400 μ L cell is eluted on ice and the cell is stirred for 2 hours.Then, 4 hours are stirred for 2 hours.Then, 4 hours are stirred for 3 hours.Then, 4 hours are stirred for 1 hour.Then, 4 hours are stirred for 2 hours.Then, 4 hours are stirred for 3 hours.Then, 4 hours are stirred for 1 hour.Then, 4 hours are stirred for 2 hours.Then, 4 hours are stirred for 3 hours.Then, 4 hours are stirred for 1 hour.Then, 4 hours are stirred for 2 hours.Then, 4 hours are stirred for 3 hours.Then, 4 hours are stirred for 1 hour.Then, 4 hours are stirred for 2 hours.Then, 4 hours are stirred for 3 hours.Then, 4 hours are stirred for 1 hour.Then, 4 hours are stirred for 2 hours. We then inoculated the recovered culture (1 mL) into 5 mL of pre-warmed 2xYT medium containing appropriate antibiotics and incubated the cells overnight at 37°C while shaking at 220 rpm.
[0435] Conjugation assay
[0436] Donor and recipient cells for conjugation assays were grown overnight in 5 mL 2xYT in the presence of appropriate antibiotics (50 μg / mL kanamycin for recipient; 20 μg / mL chloramphenicol for donor). The OD of the culture was determined. 600 The cultures were normalized to OD 600 =2.0. 400 μL of culture was then transferred to a 2 mL microcentrifuge tube and washed twice with 2xYT. After washing, the pellet was resuspended in a final volume of 200 μL. For the conjugate, 100 μL of donor and 100 μL of acceptor were mixed and spotted on a TYE plate in 5 μL droplets. The plate was then incubated at 37°C for 2 h. Subsequently, 2 mL of 2xYT was used to wash the cells from the plate and transferred to a fresh 2 mL microcentrifuge tube. The cells were pelleted by centrifugation (1 min, 3000 x g), resuspended in 1 mL of H2O, and serially diluted (1:10). 0 -10 -7 Dilutions ranging from 1:1 to 2:1 were spotted (3 μL droplets) on 2xYT agar plates containing 50 μg / mL kanamycin and 20 μg / mL chloramphenicol. The plates were incubated at 37°C for 24-36 h, and colonies were manually counted to determine the number of successful transconjugants. For the code lock experiment, the appropriate antibiotic (200 μg / mL hygromycin or 75 μg / mL spectinomycin) was added to the 2xTY agar plates.
[0437] Doubling time measurement
[0438] In a Costar clear 96-well flat-bottom plate, cells were inoculated from dense overnight cultures (1:100 ratio) into 200 μL 2xTY containing 200 μg / mL hygromycin. Cells were grown in a TECAN infinite M200 Pro at 37°C with shaking (880 rpm). OD values were measured every 5 min over a 24-hour period. 600 Measurements were made to determine cell density. A sliding window of 10 time points was used to determine the region of steepest slope of the growth curve. From this region, the doubling time was determined.
[0439] Phage enrichment from environmental samples
[0440] Water samples were collected from different locations along the banks of the River Cam (Cambridge, UK). After filtering through a 0.22 μm filter, 4 mL of the water sample was mixed with 4 mL of 2x LB and 200 μl of an overnight culture of E. coli and then incubated at 37°C in a rotating wheel for 48 h. The culture was then centrifuged at 4500 x g for 15 min, and the filtered supernatant was stored as a phage enrichment.
[0441] Note: Location
[0442] A: Cambridge Water Treatment Plant Outflow (52°13'55.3"N 0°10'15.3"E); B: Grassy Corner (52°13'21.3"N 0°10'00.0"E); C: Coffee Temple (52°13'07.7"N0°09'01.5"E); D: Green Dragon Bridge (52°13'02.9"N0°08'44.8"E); D: Jesus Green's Lock (52°12'45.8"N 0°07'15.4"E); E: Scudamore's at Granta Place (52°12'04.6"N 0°06'56.8"E).
[0443] Plaque purification and phage lysate preparation
[0444] To purify phage plaques, the phage concentrate was serially diluted (10-fold) in LB, and 10 μl of each dilution was added to a vial with 200 μl of an overnight culture of the bacterial host for evaluation. 4 mL of molten top agar (0.35% agarose) was then added, mixed, and poured onto an LB agar plate containing the appropriate antibiotic as an overlay. The resulting plate was incubated overnight at 37°C. Individual phage plaques were picked using a sterile toothpick and resuspended in 100 μl of LB. The mixture was centrifuged and the supernatant was diluted and used for further purification rounds as described above. This process was repeated 3 times to ensure phage purity.
[0445] Phage lysates were collected from bacterial flats that showed nearly confluent lysis after infection with pure phage isolates. The top agar was scraped into a glass universal bottle containing 3 ml LB and homogenized using a sterile pipette. The suspension was then centrifuged (4500 x g, 4 ° C, 20 min). The supernatant obtained was filtered through a 0.22 μm filter and stored in a vial at 4 ° C. As described above, the phage titer was estimated by counting the number of plaques obtained from the plated phage lysate dilutions.
[0446] Phage DNA extraction
[0447] Using the standard phenol / chloroform method as described by Chen et al. (2017) (43), 450 μL of high-titer phage lysate (~10 10 PFU / mL) to obtain phage genomic DNA.
[0448] Efficiency of plaque assay
[0449] Phage lysate was serially diluted (10-fold) in LB. The dilutions were spotted (7.5 μL per spot) on freshly poured and dried top layers (200 μL overnight culture mixed with 4 mL top agar, which was poured as an overlay onto an LB agar plate containing 200 μg / mL hygromycin and 75 μg / mL spectinomycin) and incubated overnight at 37°C. Spot images were taken on an iPhone 8 and converted to grayscale in Adobe Illustrator. For concentrations where a single plaque was expected, the entire top layer was poured out to better assess the plaque forming units at a given concentration (200 μL overnight culture was mixed with 10 μL phage lysate at the target concentration and 4 mL top agar, which was poured as an overlay onto an LB agar plate containing 200 μg / mL hygromycin and 75 μg / mL spectinomycin). All plaque counts shown in the bar graph are from the entire top layer. For titer lysates (>10 6PFU / mL), pouring the entire top layer as described above to avoid lysis from the outside. The maximum titer for infection with phage 06 and phage 12 was ∼7.5 x 10 9 PFU / mL and ~1.1x10 10 PFU / mL. The strains used in this experiment contain different versions of the spectinomycin resistance gene. Syn61 WT contains SpecR WT, Syn61Δ3 contains recSpecR, Syn61Δ3 (tRNA CGA Ala , tRNA UGA His ) contains O-SpecR (TCG-Ala, TCA-His), Syn61Δ3 (tRNA CGA Ala , tRNA UGA Leu ) contains O-SpecR (TCG-Ala, TCA-Leu), Syn61Δ3 (tRNA CGA Leu , tRNA UGA Leu ) contains O-SpecR (TCG-Leu, TCA-Leu), Syn61Δ3 (tRNA CGA Pro , tRNA UGA Leu ) contains O-SpecR (TCG-Pro, TCA-Leu).
[0450] Electron microscopy
[0451] By adding 10 μL of high titer phage lysate (>10 9 Phage samples were prepared by adsorbing 1000 PFU / mL onto charged copper grids and staining with 2% (w / v) uranyl acetate. Transmission electron micrographs (TEM) images were obtained using a FEI Tecnai G2 series transmission electron microscope (accelerating voltage: 200.0 kV; direct magnification: 50,000x) at the Cambridge Advanced Imaging Centre (CAIC) at the University of Cambridge.
[0452] Phage genome sequencing and de novo assembly
[0453] Purified phage DNA for NGS was prepared using the Nextera XT DNA Library Preparation Kit. Libraries were paired-end sequenced on a MiSeq (Illumina, kit v3 (150 cycles)). De novo assembly of phage genomes was performed using the Unicycler in short-read mode with default options. Sequence coverage of the entire phage genome is expressed as the median sequencing coverage in a 250 bp window.
[0454] tRNA screening
[0455] An overview of the sequences used in the tRNA screening is provided as SEQ ID NOs: 7-68. These sequences represent, in order, ArgX (modified to the anticodon for CGA), ArgX (modified to the anticodon for TGA), ArgW (modified to the anticodon for CGA), ArgW (modified to the anticodon for TGA), ileT (modified to the anticodon for CGA), ileT (modified to the anticodon for TGA), PheU (modified to the anticodon for CGA), PheU (modified to the anticodon for TGA), AspT (modified to the anticodon for CGA), AspT (modified to the anticodon for TGA), AsnT (modified to the anticodon for CGA), AsnT (modified to the anticodon for TGA), GltU (modified to the anticodon for CGA), (modified to the anticodon of TGA), GltU (modified to the anticodon of TGA), ValV (modified to the anticodon of CGA), ValV (modified to the anticodon of TGA), ThrT (modified to the anticodon of CGA), ThrT (modified to the anticodon of TGA), ThrU (modified to the anticodon of CGA), ThrU (modified to the anticodon of TGA), GlyU (modified to the anticodon of CGA), GlyU (modified to the anticodon of TGA), GlyT (modified to the anticodon of CGA), GlyT (modified to the anticodon of TGA), GlnU (modified to the anticodon of CGA), GlnU (modified to the anticodon of TGA), GlnV (modified to the anticodon of CGA), GlnV (modified to the anticodon of TGA), MetV (modified to the anticodon of CGA), MetV (modified to the anticodon of TGA), MetY (modified to the anticodon of CGA), MetY (modified to the anticodon of TGA), ThrV (modified to the anticodon of CGA), ThrV (modified to the anticodon of TGA), valW (modified to the anticodon of CGA), valW (modified to the anticodon of TGA), ArgQ (modified to the anticodon of CGA), ArgQ (modified to the anticodon of TGA), ArgV (modified to the anticodon of CGA), ArgV (modified to the anticodon of TGA ), CysT (modified to the anticodon of CGA), CysT (modified to the anticodon of TGA), HisR (modified to the anticodon of CGA), HisR (modified to the anticodon of TGA), ileX (modified to the anticodon of CGA), ileX (modified to the anticodon of TGA), LysQ (modified to the anticodon of CGA), LysQ (modified to the anticodon of TGA), ProK (modified to the anticodon of CGA), ProK (modified to the anticodon of TGA), ProL (modified to the anticodon of CGA), ProL (modified to the anticodon of TGA), ProM (modified to the anticodon of CGA),ProM (modified to the anticodon of TGA), TrpT (modified to the anticodon of CGA), TrpT (modified to the anticodon of TGA), TyrV (modified to the anticodon of CGA), TyrV (modified to the anticodon of TGA), AlaT (modified to the anticodon of CGA), AlaT (modified to the anticodon of TGA), LeuQ (modified to the anticodon of CGA), LeuQ (modified to the anticodon of TGA).
[0456] Table 1
[0457]
[0458]
[0459] References for Materials and Methods
[0460] 1.FHCrick, L.Barnett, S.Brenner, RJWatts-Tobin, General nature of the genetic code for proteins. Nature 192, 1227-1232 (1961).
[0461] 2.MWNirenberg, JHMatthaei, The dependence of cell-free proteins synthesis in E.coli upon naturally occurring or synthetic polyribonucleotides. Proc Natl Acad Sci USA 47, 1588-1602 (1961).
[0462] 3. RJ Hall, FJ Whelan, JOMcInerney, Y. Ou, MR Domingo-Sananes, Horizontal Gene Transfer as a Source of Conflict and Cooperation in Prokaryotes. Front Microbiol 11, 1569 (2020).
[0463] 4.K.Vetsigian,C.Woese,N.Goldenfeld,Collective evolution and thegenetic code.Proc Natl Acad Sci U S A 103,10696-10701(2006).
[0464] 5.D.de la Torre,J.W.Chin,Reprogramming the genetic code.Nat Rev Genet22,169-184(2021).
[0465] 6.E.V.Koonin,A.S.Novozhilov,Origin and evolution of the genetic code:the universal enigma.IUBMB Life 61,99-111(2009).
[0466] 7.M.Kollmar,S.Muhlhausen,Nuclear codon reassignments in the genomicsera and mechanisms behind their evolution.Bioessays 39,(2017).
[0467] 8.J.Ling et al.,Natural reassignment of CUU and CUA sense codons toalanine in Ashbya mitochondria.Nucleic Acids Res 42,499-508(2014).
[0468] 9.A.L.Borges et all.,Widespread stop-codon recoding in bacteriophagesmay regulate translation of lytic genes.Nat Microbiol 7,918-927(2022).
[0469] 10.M.A.Santos,A.C.Gomes,M.C.Santos,L.C.Carreto,G.R.Moura,The geneticcode ofthe fungal CTG clade.C R Biol 334,607-611(2011).
[0470] 11.D.J.Taylor,M.J.Ballinger,S.M.Bowman,J.A.Bruenn,Virus-host co-evolution under a modified nuclear genetic code.PeerJ 1,e50(2013).
[0471] 12.Y.Shulgina,S.R.Eddy,A computational screen for alternative geneticcodes in over 250,000 genomes.Elife 10,(2021).
[0472] 13.D.G.Gibson et al.,Complete chemical synthesis,assembly,and cloningof a Mycoplasma genitalium genome.Science 319,1215-1220(2008).
[0473] 14.D.G.Gibson et al.,One-step assembly in yeast of 25 overlapping DNAfragments to form a complete synthetic Mycoplasma genitalium genome.Proc NatlAcad Sci U S A 105,20404-20409(2008).
[0474] 15.J.Fredens et al.,Total synthesis ofEscherichia coli with a recodedgenome.Nature 569,514-+(2019).
[0475] 16.F.J.Isaacs et al.,Precise manipulation of chromosomes in vivoenables genome-wide codon replacement.Science 333,348-353(2011).
[0476] 17.M.J.Lajoie et al.,Genomically recoded organisms expand biologicalfunctions.Science 342,357-360(2013).
[0477] 18.W.E.Robertson et al.,Sense codon reassignment enables viralresistance and encoded polymer synthesis.Science 372,1057-1062(2021).
[0478] 19.G.Pines,J.D.Winkler,A.Pines,R.T.Gill,Refactoring the Genetic Codefor Increased Evolvability.mBio 8,(2017).
[0479] 20.J.Calles,I.Justice,D.Brinkley,A.Garcia,D.Endy,Fail-safe geneticcodes designed to intrinsically contain engineered organisms.Nucleic AcidsRes 47,10439-10451(2019).
[0480] 21.M.Schmidt,V.Kubyshkin,How To Quantify a Genetic Firewall?APolarity-Based Metric for Genetic Code Engineering.Chembiochem 22,1268-1284(2021).
[0481] 22.D.S.Marks,S.W.Michnick,Democratizing the mapping of gene mutationsto protein biophysics.Nature 604,47-48(2022).
[0482] 23.S.Teng,A.K.Srivastava,C.E.Schwartz,E.Alexov,L.Wang,Structuralassessment of the effects of amino acid substitutions on protein stabilityand protein protein interaction.Int J Comput Biol Drug Des 3,334-349(2010).
[0483] 24.V.Parthiban,M.M.Gromiha,D.Schomburg,CUPSAT:prediction of proteinstability upon point mutations.Nucleic Acids Res 34,W239-242(2006).
[0484] 25.P.C.Ng,S.Henikoff,Predicting the effects of amino acidsubstitutions on protein function.Annu Rev Genomics Hum Genet 7,61-80(2006).
[0485] 26.B.A.Renda,M.J.Hammerling,J.E.Barrick,Engineering reducedevolutionary potential for synthetic biology.Mol Biosyst 10,1668-1678(2014).
[0486] 27.G.Moratorio et al.,Attenuation of RNA viruses by redirecting theirevolution in sequence space.Nat Microbiol 2,17088(2017).
[0487] 28.J.R.Coleman et al.,Virus attenuation by genome-scale changes incodon pair bias.Science 320,1784-1787(2008).
[0488] 29.P.W.Barone et al.,Viral contamination in biologic manufacture andimplications for emerging therapies.Nat Biotechnol 38,563-572(2020).
[0489] 30.P.Alamos et al.,Functionality of tRNAs encoded in a mobile geneticelement from an acidophilic bacterium.RNA Biol 15,518-527(2018).
[0490] 31.T.Tuller et al.,Association between translation efficienoy andhorizontal gene transfer within microbial communities.Nucleic Acids Res 39,4743-4755(2011).
[0491] 32.J.W.Lee,C.T.Y.Chan,S.Slomovic,J.J.Collins,Next-generationbiocontainment systems for engineered organisms.Nat Chem Biol 14,530-537(2018).
[0492] 33.W.E.Robertson et al.,Creating custom synthetic genomes inEscherichia coli with REXER and GENESIS.Nat Protoc 16,2345-2380(2021).
[0493] 34.K.C.Murphy,lambda Recombination and Recombineering.EcoSal Plus 7,(2016).
[0494] 35.K.Wang et al.,Defining synonymous codon compression schemes bygenome recoding.Nature 539,59-64(2016).
Claims
1. A cell, which: comprises a genome in which at least a first type of sense codon has been recoded such that a first endogenous tRNA is dispensable; does not express the first endogenous tRNA; expresses a first modified tRNA capable of decoding the first type of sense codon, wherein the first modified tRNA carries a first amino acid that is not the natural cognate amino acid of the first type of sense codon; and comprises a gene required for viability, wherein the gene comprises at least one occurrence of the first type of sense codon, and the cell is viable when the first type of sense codon in the gene is decoded as the first amino acid.
2. The cell according to claim 1, wherein the first modified tRNA is an anticodon-swapped tRNA canonically associated with the first amino acid.
3. The cell according to claim 1 or 2, wherein the first modified tRNA is derived from a tRNA endogenous to the cell and is an isoacceptor for the first amino acid, or is derived from a tRNA found in a mobile genetic element and is an isoacceptor for the first amino acid.
4. The cell according to any one of claims 1-3, wherein if the first type of sense codon in the gene required for viability is decoded according to the canonical genetic code, the cell is non-viable, or wherein if decoded according to the canonical genetic code, the first type of sense codon in the gene required for viability at least partially contributes to loss of viability.
5. The cell according to any one of claims 1-4, wherein the gene required for viability is an essential gene or a positive selectable marker.
6. The cell according to any one of claims 1-5, wherein the first amino acid is a naturally occurring amino acid.
7. The cell according to any one of claims 1-6, wherein a second type of sense codon has been recoded within the genome.
8. The cell according to claim 7, wherein a second endogenous tRNA is dispensable and the cell does not express the second endogenous tRNA.
9. The cell according to claim 7 or claim 8, wherein the cell expresses a second modified tRNA capable of decoding the second type of sense codon, wherein the second modified tRNA carries a second amino acid that is not the natural cognate amino acid of the second type of sense codon.
10. The cell according to claim 9, wherein the gene required for viability comprises at least one occurrence of the second type of sense codon, and the cell is viable when the second type of sense codon in the gene is decoded as the second amino acid.
11. The cell according to claim 10, wherein if the second type of sense codon in the gene required for viability is decoded according to the canonical genetic code, the cell is non-viable, or wherein if decoded according to the canonical genetic code, the second type of sense codon at least partially contributes to loss of viability.
12. The cell according to any one of claims 9-11, wherein the second modified tRNA is an anticodon-exchanged tRNA that canonically associates with the second amino acid.
13. The cell according to any one of claims 9-12, wherein the second modified tRNA is derived from a tRNA that is endogenous to the cell and is an isoacceptor for the second amino acid, or is derived from a tRNA found in a mobile genetic element and is an isoacceptor for the second amino acid.
14. The cell according to any one of claims 9-13, wherein the second amino acid is a naturally occurring amino acid.
15. The cell according to any one of claims 9-14, wherein the first and second amino acids are of the same type or of different types.
16. The cell according to any one of claims 7-15, wherein the first type of sense codon is TCA and the second type of sense codon is TCG.
17. The cell according to any one of claims 1-16, wherein the first type of sense codon is TCA or TCG.
18. The cell according to any one of claims 1-17, wherein when the genes of the cell are decoded by the modified tRNA, it is viable, and when the genes of the cell are decoded at least partially according to the canonical genetic code, it is non-viable.
19. A cell, which: comprises a genome in which a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable; does not express the first endogenous tRNA and the second endogenous tRNA; expresses a first anticodon-exchanged tRNA derived from a naturally occurring first parental tRNA, wherein the first anticodon-exchanged tRNA carries a first amino acid and the first parental tRNA is an isoacceptor for the first amino acid, and wherein the first amino acid is not the natural cognate amino acid of the first type of sense codon, and expresses a second anticodon-exchanged tRNA derived from a naturally occurring second parental tRNA, wherein the second anticodon-exchanged tRNA carries a second amino acid and the second parental tRNA is an isoacceptor for the second amino acid, and wherein the second amino acid is not the natural cognate amino acid of the second type of sense codon; wherein the first and / or second modified tRNA cannot decode any type of codon other than the first type of sense codon and / or the second type of sense codon.
20. The cell according to claim 19, wherein the first and second types of sense codons canonically encode the same amino acid.
21. The cell according to claim 19 or claim 20, wherein due to wobble base pairing, the first and second types of sense codons are canonically decoded by the same tRNA or overlapping tRNAs, wherein the first anticodon-exchanged tRNA cannot decode any codon type other than the first type of sense codons, and / or the second anticodon-exchanged tRNA cannot decode any codon type other than the second type of sense codons.
22. The cell according to any one of claims 19-21, wherein the first type of sense codons and the second type of sense codons belong to the formula XXN, and wherein the first anticodon-exchanged tRNA cannot decode the second type of sense codons, and the second anticodon-exchanged tRNA cannot decode the first type of sense codons.
23. The cell according to any one of claims 19-22, wherein the first amino acid and the second amino acid are different types of amino acids.
24. The cell according to any one of claims 19-23, wherein the first and second parental tRNAs are derived from the same cell type as the cell according to claim 19.
25. The cell according to any one of claims 19-24, wherein the first and / or second anticodon-exchanged tRNA contains an identity element recognized by an aminoacyl-tRNA synthetase that is endogenous to the cell.
26. The cell according to any one of claims 19-25, wherein the first and second types of sense codons canonically decode serine, the first and second types of sense codons canonically decode alanine, or the first and second types of sense codons canonically decode leucine.
27. The cell according to any one of claims 19-26, wherein the first and / or second anticodon-exchanged tRNA does not decode the TCC or TCT codons.
28. The cell according to any one of claims 19-27, wherein the first type of sense codons is TCA and / or the second type of sense codons is TCG.
29. The cell according to any one of claims 19-28, wherein the first and / or second amino acid is a naturally occurring amino acid.
30. The cell according to any one of claims 19-29, wherein the first amino acid is any one of alanine, histidine, leucine, and proline; and / or the second amino acid is any one of alanine, histidine, leucine, and proline.
31. The cell according to any one of claims 19-30, wherein the first and / or second anticodon-exchanged tRNA is derived from a parental tRNA encoded by ArgQ, ArgU, GltU, HisR, ProK, ProL, ProM, TrpT, ThrU, ThrT, TyrU, TyrV, AlaT, or LeuQ.
32. The cell according to claim 31, wherein the first and / or second anticodon-swapped tRNA is derived from a parental tRNA encoded by HisR, ProM, AlaT or LeuQ.
33. The cell according to any one of claims 19-32, wherein the parental tRNA is an Escherichia coli tRNA.
34. A cell, which: comprises a genome in which a first type of sense codon and a second type of sense codon have been recoded such that a first endogenous tRNA and a second endogenous tRNA are dispensable; does not express the first endogenous tRNA and the second endogenous tRNA; expresses a first modified tRNA capable of decoding the first type of sense codon, wherein the first modified tRNA bears a first amino acid that is not the natural cognate amino acid of the first type of sense codon; and expresses a second modified tRNA capable of decoding the second type of sense codon, wherein the second modified tRNA bears a second amino acid that is not the natural cognate amino acid of the second type of sense codon; wherein: i) the first amino acid is alanine and the second amino acid is alanine; ii) the first amino acid is alanine and the second amino acid is histidine; iii) the first amino acid is alanine and the second amino acid is leucine; iv) the first amino acid is alanine and the second amino acid is proline; v) the first amino acid is histidine and the second amino acid is alanine; vi) the first amino acid is histidine and the second amino acid is histidine; vii) the first amino acid is histidine and the second amino acid is leucine; viii) the first amino acid is histidine and the second amino acid is proline; ix) the first amino acid is leucine and the second amino acid is alanine; x) the first amino acid is leucine and the second amino acid is histidine; xi) the first amino acid is leucine and the second amino acid is proline; xii) the first amino acid is proline and the second amino acid is alanine; xiii) the first amino acid is proline and the second amino acid is histidine; xiv) the first amino acid is proline and the second amino acid is leucine; or xv) the first amino acid is proline and the second amino acid is proline.
35. The cell according to claim 34, wherein the first modified tRNA cannot decode the second type of sense codon, and / or the second modified tRNA cannot decode the first type of sense codon.
36. The cell according to claim 34 or claim 35, wherein the first modified tRNA cannot decode any type of codon other than the first type of sense codon, and / or the second modified tRNA cannot decode any type of codon other than the second type of sense codon.
37. The cell according to any one of claims 34 - 36, wherein the first modified tRNA is a tRNA with an anticodon exchange that is canonically associated with the first amino acid, and / or the second modified tRNA is a tRNA with an anticodon exchange that is canonically associated with the second amino acid.
38. The cell according to any one of claims 34 - 37, wherein the first modified tRNA is derived from a tRNA endogenous to the cell and is an isoacceptor for the first amino acid, or is derived from a tRNA found in a mobile genetic element and is an isoacceptor for the first amino acid; and / or the second modified tRNA is derived from a tRNA endogenous to the cell and is an isoacceptor for the second amino acid, or is derived from a tRNA found in a mobile genetic element and is an isoacceptor for the second amino acid.
39. The cell according to any one of claims 34 - 38, wherein the first and / or second modified tRNA contains an identity element recognized by an aminoacyl - tRNA synthetase endogenous to the cell.
40. The cell according to any one of claims 34 - 39, wherein the first and second types of sense codons canonically encode serine.
41. The cell according to any one of claims 7 - 40, wherein the first type of sense codon is TCA and / or the second type of sense codon is TCG.
42. The cell according to any one of claims 7 - 41, wherein the essential genes of the genome do not contain a naturally occurring instance of the second type of sense codon, and the second endogenous tRNA is a cognate tRNA for the second type of sense codon.
43. The cell according to any one of claims 7 - 42, wherein the genome contains 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 or no naturally occurring instances of the second type of sense codon, and the second endogenous tRNA is a cognate tRNA for the second type of sense codon.
44. The cell according to any one of claims 7-43, wherein the second type of sense codon is TCG and the second endogenous tRNA is tRNA Ser CGA .
45. The cell according to any of the preceding claims, wherein the essential genes of the genome do not contain a naturally occurring instance of the first type of sense codon, and the first endogenous tRNA is a cognate tRNA for the first type of sense codon.
46. The cell according to any of the preceding claims, wherein the genome contains 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 or no naturally occurring instances of the first type of sense codon, and the first endogenous tRNA is a cognate tRNA for the first type of sense codon.
47. The cell according to any of the preceding claims, wherein the sense codon of the first type is TCA and the first endogenous tRNA is tRNA Ser UGA , or the sense codon of the first type is TCG and the first endogenous tRNA is tRNA Ser CGA .
48. The cell according to any of the preceding claims, wherein a plurality of naturally occurring instances of the TCA codon have been replaced by AGT and / or a plurality of naturally occurring instances of the TCG codon have been replaced by AGC.
49. A cell having increased resistance to horizontal gene transfer or mobile genetic elements, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid not associated with the sense codon in the canonical genetic code, and the cell contains a gene required for viability that has a function when decoded according to the reassigned genetic code and has no function when decoded according to the canonical genetic code.
50. The cell according to any of the preceding claims, wherein the genome of the cell has at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9% or 100% identity to any one of SEQ ID NO: 1-6.
51. The cell according to any of the preceding claims, wherein the cell is a prokaryotic cell, a bacterial cell or an Escherichia coli cell.
52. A method of increasing the resistance of a cell to mobile genetic elements or horizontal gene transfer, wherein the cell has been modified to reassign at least one type of sense codon to an amino acid not associated with the sense codon in the canonical genetic code, the method comprising: modifying a gene required for viability to include at least one occurrence of the reassigned sense codon, wherein if the reassigned sense codon in the gene decodes to the reassigned amino acid, the cell is viable, and if the reassigned sense codon in the gene decodes according to the canonical genetic code, the cell is non-viable, or wherein if decoded according to the canonical genetic code, the reassigned sense codon in the gene at least partially contributes to loss of viability.
53. A kit comprising a first cell recoded according to a first orthogonal coding scheme and a second cell recoded according to a second orthogonal coding scheme, wherein the first and second coding schemes are orthogonal to each other.
54. The kit according to claim 53, wherein the first cell is according to any one of claims 1-51 and / or the second cell is according to any one of claims 1-51.
55. The kit according to claim 53 or claim 54, wherein the first cell utilizes the reassignment scheme shown in Figure 2C and the second cell utilizes a different reassignment scheme shown in Figure 2C.
56. A mobile genetic element recoded according to an orthogonal coding scheme.
57. The mobile genetic element according to claim 56, wherein at least one, a plurality or each instance of a particular type of sense codon in at least one gene required for horizontal transfer of genetic information is replaced by a sense codon that canonically encodes a different amino acid.
58. The mobile genetic element according to claim 57, wherein the replacement of the particular type of sense codon is according to the reassignment scheme shown in Figure 2C.
59. A method of preventing horizontal transfer of genetic information between a mobile genetic element and a first cell, the method comprising incubating the mobile genetic element and the first cell wherein the mobile genetic element is according to any one of claims 56 - 58, and the first cell comprises tRNAs that decode codons according to the canonical genetic code or according to a coding scheme orthogonal to the mobile genetic element.
60. A method of altering the susceptibility of a gene to mutations that alter the encoded amino acid sequence, the method comprising: i) identifying a target gene; and ii) incubating a cell comprising the target gene, wherein the cell comprises tRNAs capable of decoding at least one sense codon as a reassigned amino acid.
61. The method according to claim 60, wherein the cell is according to any one of claims 1 - 51.
62. The method according to claim 60 or claim 61, wherein the reassigned amino acid is as shown in Figure 2C.
63. The method according to any one of claims 60 - 62, wherein the genome of the cell has at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, 99.9% or 100% identity to any one of SEQ ID NOs: 1 - 6, and the cell comprises tRNAs with anticodon swapping.
64. Use of a cell according to any one of claims 1 - 51 for the production of a polymer.
65. A method for manufacturing a polymer, the method comprising: culturing a cell according to any one of claims 1 - 51, providing a cell having a nucleic acid sequence encoding the polymer, and obtaining the polymer.
Citation Information
Patent Citations
Synthetic genome
WO2020229592A1