Synthetic genome
By employing REXER and GENESIS with directed conjugation, the method effectively produces viable synthetic prokaryotic genomes with reduced sense codons, addressing inefficiencies in existing genome-wide synonymous codon compression techniques and enabling the biosynthesis of non-canonical biopolymers.
Patent Information
- Application Number
- JP2025125893
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-14
- Filing Date
- 2025-07-28
- Publication Date
- 2025-12-02
AI Technical Summary
Existing methods for genome-wide synonymous codon compression in synthetic genomes are inefficient and unclear regarding their viability, particularly in producing organisms with reduced sense codons for encoding canonical amino acids.
A method involving recombination-mediated genetic modification, such as REXER and GENESIS, combined with directed conjugation, is used to produce viable synthetic prokaryotic genomes with reduced sense codons by replacing up to 99.9% of target codons, utilizing a defined rewriting and refactoring scheme to ensure genome-wide synonymous codon compression.
This approach enables the production of synthetic prokaryotic genomes with significantly reduced sense codons, ensuring viability and allowing for the identification of non-permissible positions with codon-level resolution, facilitating the biosynthesis of genetically encoded non-canonical biopolymers.
Smart Images

Figure 2025175308000017 
Figure 2025175308000018 
Figure 2025175308000019
Abstract
Description
[Technical Field]
[0001] The present invention relates to synthetic genomes and methods for producing same. [Background technology]
[0002] Genome design and synthesis offer a powerful approach to understanding and modifying biology. Genome synthesis has the potential to accelerate metabolic engineering. In particular, genome synthesis may elucidate the function of synonymous codons and facilitate the synthesis of genetically encoded unnatural polymers (Wang, K., et al., 2016. Nature, 539(7627), 59-64).
[0003] The standard genetic code uses 61 sense codons to encode 20 canonical amino acids, with 18 of the 20 amino acids encoded by more than one synonymous codon. Nature selects one sense codon from up to six synonyms to encode each amino acid at each position in a gene. The choice of synonymous codons may affect mRNA folding, transcriptional and translational regulatory sequences, translation rate, cotranslational folding, and protein levels, and may have novel and yet-to-be-understood roles (Wang, K., et al., 2016. Nature, 539(7627), 59-64; and Cambray, G., et al., 2018. Nature biotechnology, 36(10), 1005-1015).
[0004] Genome-wide replacement of target codons with synonymous codons (synonymous codon compression) can provide the basis for reassigning sense codons to non-canonical amino acids (or other monomers) to facilitate the in vivo biosynthesis of genetically encoded non-canonical biopolymers (Chin, JW, 2017. Nature, 550(7674), 53-60).
[0005] Site-directed mutagenesis approaches have been used to replace up to 321 amber stop codons in the E. coli genome (Mukai, T., et al., 2015. Scientific Reports, 5, p. 9699). However, sense codons generally outnumber stop codons by several orders of magnitude, and genome synthesis, rather than mutagenesis, may be the preferred approach to sense codon removal in many cases.
[0006] Genome synthesis allows the creation of mycoplasmas with synthetic genomes (Gibson, DG, et al., 2010. Science, 329(5987), 52-56), which can replicate one or two of the 16 chromosomes. Creation of nine strains of S. cerevisiae in which the DNA has been replaced with synthetic DNA. (Zhang, W., et al., 2017. Science, 355(6329), eaaf3981; and Richardson, SM, et al., 2017. Science, 355(6329), 1040-1044). These experiments were carried out independently. Up to 1 Mb of DNA (0.99 Mb, yeast; 1.08 Mb, mycoplasma) has been replaced in strains. Replicon excision for enhanced genome engineering through programmed recombination (REXER) has been reported to replace over 100 kb of the E. coli genome with synthetic DNA in a single step. Furthermore, it has been shown that REXER can be repeated by genome stepwise interchange synthesis (GENESIS) to replace a 220 kb E. coli genome with 230 kb of synthetic DNA. (Wang, K., et al., 2016. Nature, 539(7627), 59-64; International Publication No. 2018 / 020248 Brochure).
[0007] Genome synthesis involves the identification of synonymous codons in individual genes (Napolitano, MG, et al., 2016. PNAS, 113(38), E5588-E5597), genomic regions, and essential operons (Wang, K., et al., 2016. Nature, 539(7627), 59-64; and Lau, YH, et al. 2017. Nucleic acids research, 4 5(11), 6971-6980). For example, Wang et al. used a defined "rewriting scheme" to replace a 20 kb region of the Escherichia coli genome that is enriched in both essential genes and target codons.
[0008] However, these studies only mutated a small fraction (up to 4.7%) of the targeted sense codons in a single strain's genome. As a result, it is unclear whether applying these methods to genome-wide synonymous codon compression can produce viable genomes. For example, it is unclear whether the defined rewriting scheme tested in Wang et al. can be applied genome-wide to create an organism in which a small number of sense codons are used to encode the 20 canonical amino acids. [Prior art documents] [Patent documents]
[0009] [Patent Document 1] International Publication No. 2018 / 020248 Brochure [Non-patent literature]
[0010] [Non-Patent Document 1] Wang, K., et al., 2016. Nature, 539(7627),59-64 [Non-patent document 2] Cambray, G., et al., 2018. Naturebiotechnology, 36(10), 1005-1015 [Non-patent document 3] Chin, JW, 2017. Nature, 550(7674), 53-60 [Non-patent document 4] Mukai, T., et al., 2015. Scientific reports,5, p.9699 [Non-patent document 5] Gibson, DG, et al., 2010. Science,329(5987), 52-56 [Non-patent document 6] Zhang, W., et al., 2017. Science, 355(6329),eaaf3981 [Non-Patent Document 7] Richardson, SM, et al., 2017. Science,355(6329), 1040-1044 [Non-patent document 8] Napolitano, MG, et al., 2016. PNAS,113(38), E5588-E5597 [Non-Patent Document 9] Lau, YH, et al. 2017. Nucleic acidsresearch, 45(11), 6971-6980 Summary of the Invention [Problem to be solved by the invention]
[0011] Therefore, there is a need for synthetic genomes in which one or more sense codons have been removed, as well as for improved methods for producing synthetic genomes. [Means for solving the problem]
[0012] The present inventors have surprisingly found that viable synthetic prokaryotic genomes can be produced in which one or more sense codons have been removed. In particular, the present inventors have produced viable synthetic genomes in which the number of codons used to encode cellular proteins has been reduced from 64 to 61 by genome-wide rewriting of two sense codons and one stop codon. The present inventors have also produced E. coli host cells containing the synthetic genomes.
[0013] The inventors also surprisingly found that the defined rewriting and refactoring scheme can enable genome-wide synonymous codon compression for over 99.9% of the target codons. The inventors found that alternative rewriting and refactoring at disallowed positions enables genome-wide synonymous codon compression.
[0014] The inventors have also surprisingly found that recombination-mediated genetic modification (e.g., REXER and / or GENESIS) can be combined with directed conjugation to effectively produce synthetic genomes. In particular, the inventors found, for example, that at least about 4 Mb of DNA can be effectively replaced by the method, and that the method allows for the identification of failures (non-permissible positions) in the design of synthetic DNA with codon-level resolution.
[0015] Thus, in one aspect, the invention provides a synthetic prokaryotic genome that comprises five or four or fewer occurrences of one or more sense codons. In some embodiments, the synthetic prokaryotic genome comprises four or three or fewer, three or two or fewer, two or one or fewer, one or zero, or no occurrences of one or more sense codons. In some embodiments, the one or more sense codons consist of one sense codon or two sense codons, preferably two sense codons. In some embodiments, the synthetic prokaryotic genome does not comprise two or more or more, preferably two, occurrences of sense codons, and does not comprise one occurrence of a stop codon, preferably an amber stop codon (TAG).
[0016] The synthetic prokaryotic genome may be a synthetic bacterial genome, preferably a synthetic Escherichia coli genome, a Salmonella enterica genome, or a Shigella dysenteriae genome. In some embodiments, the synthetic prokaryotic genome The synthetic prokaryotic genome may be 100 kb to 10 Mb, or 1 Mb to 10 Mb, or 2 Mb to 6 Mb in size. The synthetic prokaryotic genome may be viable. In some embodiments, the synthetic prokaryotic genome comprises 100 or more, 200 or more, or 1000 or more genes, where the genes may lack one or more occurrences of sense codons, and preferably the genes are essential genes.
[0017] In some embodiments, the one or more sense codons are selected from TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, GCC, CTG, CTA, CTT, CTC, TTG, and TTA; preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA; more preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG, and GCA; and most preferably, the one or more sense codons are TCG and / or TCA.
[0018] In some embodiments, the synthetic prokaryotic genome includes no more than 10 or 9, no more than 5 or 4 occurrences, or no occurrences of the amber stop codon (TAG).
[0019] In a further aspect, the invention provides synthetic prokaryotic genomes comprising 100 or 101 or more, 200 or 201 or more, or 1000 or 1001 or more genes, wherein the genes comprise a total of five or four or fewer occurrences of one or two or more sense codons, preferably the genes are essential genes. In some embodiments, the genes comprise a total of four or three or fewer, three or two or fewer, two or one or fewer, one or zero, or no occurrences of one or two or more sense codons. In some embodiments, the one or more sense codons consist of one sense codon or two sense codons, preferably two sense codons.
[0020] The synthetic prokaryotic genome may be a synthetic bacterial genome, preferably a synthetic Escherichia coli genome, Salmonella enterica genome, or Shigella dysenteriae genome. In some embodiments, the synthetic prokaryotic genome is 100 kb to 10 Mb, or 1 Mb to 10 Mb, or 2 Mb to 6 Mb in size. The synthetic prokaryotic genome may be viable.
[0021] In some embodiments, one or more sense codons are TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, GCC, CTG, CTA, CTT, Preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA, more preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG, and GCA, and most preferably, the one or more sense codons are TCG and / or TCA.
[0022] In some embodiments, the synthetic prokaryotic genome includes no more than 10 or 9, no more than 5 or 4 occurrences, or no occurrences of the amber stop codon (TAG).
[0023] In a further aspect, the invention provides a synthetic prokaryotic genome derived from a parental prokaryotic genome, wherein the synthetic prokaryotic genome comprises less than 10%, 5%, 2%, 1%, 0.5%, 0.1% or no occurrences of one or more sense codons compared to the parental prokaryotic genome. In some embodiments, the one or more sense codons consist of one sense codon or two sense codons, preferably two sense codons.
[0024] The synthetic prokaryotic genome may be a bacterial genome, preferably an Escherichia coli genome, a Salmonella enterica genome, or a Shigella dysenteriae genome. In some embodiments, the synthetic prokaryotic genome is 100 kb to 10 Mb, or 1 Mb to 10 Mb, or 2 Mb to 6 Mb in size. The synthetic prokaryotic genome may be viable.
[0025] In some embodiments, the one or more sense codons are selected from TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, GCC, CTG, CTA, CTT, CTC, TTG, and TTA; preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA; more preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG, and GCA; most preferably, the one or more sense codons are TCG and / or TCA, wherein TCG and / or TCA may be replaced by synonymous sense codons.
[0026] Preferably, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of one or more sense codons in the parental prokaryotic genome are replaced with synonymous sense codons. In some embodiments, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCG and / or TCA in the parental prokaryotic genome are replaced with AGC and / or AGT, and most preferably 90% or more, 95% or more, 98% or more, 99% or more, 99% or more, 90% or more of the occurrences of TCG in the parental prokaryotic genome. 9.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% are replaced with AGC and / or 90%, 95%, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCA in the parental prokaryotic genome are replaced with AGT.
[0027] In some embodiments, the synthetic prokaryotic genome contains 10 amber stop codons (TAG). or 9 or fewer, 5 or 4 or fewer occurrences, or no occurrences, preferably 90% or more, 95% or more, 98% or more, 99% or more, or all occurrences of TAG in the parent prokaryotic genome are replaced with TAA.
[0028] In some embodiments, 99.9% or more, or 100% of the occurrences of two or more sense codons, preferably two sense codons, in the parental prokaryotic genome are replaced with synonymous sense codons, and all occurrences of TAG in the parental prokaryotic genome are replaced with TAA.
[0029] One or more gene pairs that share an overlapping region containing one or more sense codons in the parental prokaryotic genome may be refactored, preferably one or more gene pairs where replacement of one or more of the sense codons with one or more synonymous sense codons alters the encoded protein sequence of both or one of the gene pairs.
[0030] In some embodiments, for a pair of genes in the opposite orientation, a synthetic insert is inserted between the genes, where the synthetic insert comprises an overlapping region, and / or for a pair of genes in the same orientation, a synthetic insert is inserted between the genes, where the synthetic insert comprises (i) a stop codon, (ii) about 20-200 bp upstream of the overlapping region, and (iii) the overlapping region.
[0031] In a further aspect, the invention provides polynucleotides comprising 20 or 21 or more, 30 or 31 or more, 40 or 41 or more, 50 or 51 or more, 100 or 101 or more essential genes that are absent one or more occurrences of a sense codon. In some embodiments, the one or more sense codons consist of one sense codon or two sense codons, preferably two sense codons.
[0032] In some embodiments, the one or more sense codons are selected from TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, GCC, CTG, CTA, CTT, CTC, TTG, and TTA; preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA; more preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG, and GCA; and most preferably, the one or more sense codons are TCG and / or TCA.
[0033] One or more occurrences of a sense codon in the gene may be replaced by a synonymous sense codon, preferably a TCG codon is replaced by AGC and / or a TCA codon is replaced by AGT.
[0034] Essential genes are ribF, lspA, ispH, dapB, folA, imp, yabQ, ftsL, ftsI, murE, murF, mraY, murD, ftsW, murG, murC, ftsQ, ftsA, ftsZ, lpxC, sec M, secA, can, folK, hemL, yadR, dapD, map, rpsB, tsf, pyrH, frr, dxr, ispU, cdsA, yaeL, yaeT, lpxD, fabZ, lpxA, lpxB, dnaE, accA, til S, proS, yafF, hemB, secD, secF, ribD, ribE, thiL, dxs, ispA, dnaX, adk, hemH, lpxH, cysS, folD, entD, mrdB, mrdA, nadD, holA, rlpB, leuS, lnt, glnS, fldA, cydA, infA, cydC, ftsK, lolA, serS, rpsA, msbA, lpxK, kdsB, mukF, mukE, mukB, asnS, fabA, mviN, rne, fabD, fa bG、acpP、tmk、holB、lolC、lolD、lolE、purB、minE、minD、pth、prsA、ispE、l olB、hemA、prfA、prmC、kdsA、topA、ribA、fabI、tyrS、ribC、ydiL、pheT、phe S、rplT、infC、thrS、nadE、gapA、yeaZ、aspS、argS、pgsA、yefM、metG、folE、 yejM, gyrA, nrdA, nrdB, folC, accD, fabB, gltX, ligA, zipA, dapE, dapA r、hisS、ispG、suhB、tadA、acpS、era、rnc、lepB、rpoE、pssA、yfiO、rplS、tr mD、rpsP、ffh、grpE、csrA、ispF、ispD、ftsB、eno、pyrG、chpR、lgt、fbaA、pg k、yqgD、metK、yqgF、plsC、ygiT、parE、ribB、cca、ygjD、tdcF、yraL、yhbV、i nfB、nusA、ftsH、obgE、rpmA、rplU、ispB、murA、yrbB、yrbK、yhbN、rpsI、rpl M、degS、mreD、mreC、mreB、accB、accC、yrdC、def、fmt、rplQ、rpoA、rpsD、rp sK、rpsM、secY、rplO、rpmD、rpsE、rplR、rplF、rpsH、rpsN、rplE、rplX、rplN rpsQ, rpmC, rplP, rpsC, rplV, rpsS, rplB, rplW, rplD, rplC, rpsJ, fusA, r psG、rpsL、trpS、yrfF、asd、rpoH、ftsX、ftsE、ftsY、yhhQ、bcsB、glyQ、gpsA rfaK, kdtA, coaD, rpmB, dfp, dut, gmk, spoT, gyrB, dnaN, dnaA, rpmH, rnpA 、yidC、tnaB、glmS、glmU、wzyE、hemD、hemC、yigP、ubiB、ubiD、hemG、yihA、f tsN、murI、murB、birA、secE、nusG、rplJ、rplL、rpoB、rpoC、ubiA、plsB、lex A、dnaB、ssb、alsK、groS、psd、orn、yjeE、rpsR、chpS、ppa、valS、yjgP、yjgQ、and dnaC.
[0035] In a further aspect, the present invention provides a prokaryotic host cell comprising a synthetic prokaryotic genome according to the invention or a polynucleotide according to the invention.
[0036] The prokaryotic host cell can be viable. The prokaryotic host cell can be a bacterial cell, preferably an Escherichia coli cell, a Salmonella enterica cell, or a Shigella dysenteriae cell. Preferably, the host cell is suitable for use in producing a polypeptide comprising one or more non-proteinogenic amino acids, preferably two or three or more non-proteinogenic amino acids, and most preferably three or four or more non-proteinogenic amino acids.
[0037] In a further aspect, the present invention provides the use of a prokaryotic host cell according to the invention for producing a polypeptide comprising one or more non-proteinogenic amino acids, preferably two or three or more non-proteinogenic amino acids, most preferably three or four or more non-proteinogenic amino acids.
[0038] In a further aspect, the present invention provides a method for producing a synthetic genome, comprising: (a) providing a parent genome; (b) performing one or more rounds of recombination-mediated genetic modification on the parent genome to produce two or more different partial synthetic genomes; (c) performing one or more rounds of directed conjugation with two or more different partial synthetic genomes to produce a synthetic genome; wherein each of the partial synthetic genomes comprises a synthetic region having no more than 50 or 49, no more than 20 or 19, no more than 10 or 9, no more than 5 or 4, or no more than 0 occurrences of each of one or more sense codons, or wherein each of the partial synthetic genomes comprises The method includes a synthetic region having less than 10%, 5%, 2%, 1%, 0.5%, 0.1% occurrence of each of one or more sense codons compared to the corresponding region in the parent genome.
[0039] The synthetic region may collectively represent 90% or more, 95% or more, 99% or more, or 100% of the parental genome, hi some embodiments, the synthetic region is 10-1000 kb, 50-1000 kb, 100-1000 kb, or 100-500 kb in size.
[0040] The method may further comprise testing the viability of the partial synthetic genome after each round of recombination-mediated genetic modification and / or after each round of induced conjugation.
[0041] The two or more different partial synthetic genomes may comprise at least one partial synthetic donor genome and at least one partial synthetic recipient genome. In some embodiments, at least one partial synthetic donor genome comprises a synthetic region and a first selectable marker flanked by two homologous regions immediately downstream of the origin of transfer, and at least one partial synthetic recipient genome comprises a second selectable marker flanked by two corresponding homologous regions, where the first selectable marker may comprise a positive selectable marker and / or the second selectable marker may comprise a negative selectable marker. In some embodiments, the synthetic region present in at least one partial synthetic recipient genome is outside the region flanked by homologous regions. In some embodiments, the method further comprises one or more rounds of selection for the selectable marker.
[0042] The one or more rounds of recombination-mediated genetic modification may include one or more rounds of replicon excision for enhanced genome modification by programmed recombination (REXER).
[0043] The synthetic genome may be a synthetic prokaryotic genome according to the present invention.
[0044] In a further aspect, the present invention provides a synthetic prokaryotic genome produced by the method of the present invention. [Brief explanation of the drawings]
[0045] [Figure 1-1] ~ [Figure 1-3]Figure 1 shows the design of a synthetic genome implementing a defined rewriting scheme for synonymous codon compression. a) Defined rewriting scheme for synonymous codon compression. The synonymous serine codon and three stop codons used in the wild-type E. coli genome are shown. Systematic implementation of the defined rewriting scheme for synonymous codon compression rewrites target codons to their defined synonyms, replacing the amber stop codon TAG with the ochre stop codon TAA. This creates an organism with a rewritten genome that uses a reduced number of serine and stop codons. b) Refactoring of the 3',3' overlaps allows for their independent rewriting. The overlap between the two open reading frames (ORF-1 and ORF-2) is duplicated to create a synthetic insert. This allows for independent rewriting of the ORFs. c) Refactoring of the 5',3' overlap. The overlap plus 20 bp upstream is duplicated to generate a synthetic insert. If the overlap is longer than 1 bp at the end of the upstream ORF, an in-frame TAA is introduced at the beginning of the synthetic insert; this in-frame stop codon ensures termination of translation from the original RBS. Therefore, full-length translation of all downstream ORFs is initiated from the reconstituted RBS in the synthetic insert. d. Map of the synthetic genome design in which all TCG, TCA, and TAG codons have been removed. Outer ring: All 18,218 positions of TCG → AGC, TCA → AGT, and TAG → TAA rewrites. Gray ring: 12 positions of engineered silent mutations in the overlap, 21 refactorings of the 3',3' overlap (b), and 58 refactorings of the 5'5' overlap (c). The two inner rings illustrate genomic segments. Outer ring: Eight genomic segments of the synthetic genome design (A–H). Inner ring: 37 fragments of approximately 100 kb each. Fragment 37 is designated as 37a and 37b to reflect the final assembly. oriC: origin of replication. [Figure 2-1] ~ [Figure 2-2]Figure 1 shows the retrosynthesis of the synthetic genome. a) The genome was cut into eight segments. The synthetic genome was cut into segments A–H, each corresponding to approximately 0.5 Mb (Step 1). The location of the replication origin oriC (orange box) is indicated. The segments were assembled into a completely rewritten genome (in the forward sense, opposite the retrosynthesis arrow) by directed conjugation (Figures 10 and 11). b) The genome segment was cut into 100 kb fragments. The segment was further cut into four to five fragments of approximately 100 kb each. Segment A is indicated; the other segments were treated similarly. Nearly all segments were fully assembled by GENESIS (Figure 4) through successive REXER steps (Figure 3). Each step replaced approximately 100 kb of wild-type genomic sequence with a 100 kb synthetic fragment (Steps 2 and 3). A double selection marker consisting of negative selection marker -1 (rpsL), positive selection marker +1 (KanR), negative selection marker -2 (SacB), and positive selection marker +2 (CmR) was used in alternating rounds of REXER to achieve GENESIS. c, each 100 kb synthetic fragment was cleaved into 10 kb synthetic stretches. Each 100 kb synthetic fragment was further cleaved into 9–14 short synthetic stretches approximately 10 kb long (step 4). BACs carrying the 100 kb synthetic fragments were assembled by homologous recombination in yeast. Each BAC contains a Cas9 cleavage site (black triangle) to allow excision of the synthesized DNA in vivo, homologous regions (HR1 and HR2) for targeting recombination, an appropriate double selection cassette (designated +2, -2) for selecting during REXER and GENESIS, a negative selection marker (designated -1) to allow loss of the backbone after REXER, a BAC YAC origin, and a URA3 marker for maintenance in E. coli and S. cerevisiae. [Figure 3]This figure shows the use of a 100 kb fragment of synthetic DNA to replace the corresponding region in the genome by REXER. REXER (Replicon Excision for Enhanced Genome Modification by Programmed Recombination) utilizes CRISPR / Cas9 and lambda-Red-mediated recombination to replace genomic DNA with synthetic DNA provided from an episome (BAC). This allows for the replacement of large regions of the genome (>100 kb) with synthetic DNA (Wang, K., et al., 2016. Nature, 539(7627), 59-64; WO 2018 / 020248). The black triangle indicates the position of the CRISPR protospacer, which is cleaved by Cas9 to release the synthetic DNA (pink) cassette from the BAC flanked by homology regions (HR). Homology regions 1 and 2 (HR1 and HR2) program the location of recombination into the E. coli genome. The -1 / +1 selection cassettes ensure integration of the synthetic DNA, while the -2 / +2 selection cassettes on the genome ensure excision of the corresponding wild-type DNA. In the example shown, +1 is KanR, -1 is rpsL, +2 is CmR, and -2 is sacB. [Figure 4]Figure 3 shows how GENESIS enables the stepwise replacement of genomic DNA with synthetic DNA to generate rewritten segments. Repeated cycles of REXER (see Figure 3), alternating between positive and negative selection cassettes, enable genome stepwise exchange synthesis (GENESIS) (Wang, K., et al., 2016. Nature, 539(7627), 59-64). This allows for the assembly of large segments of synthetic genomes by iterative addition of fragments that replace the corresponding genomic sequence in a clockwise direction. The first REXER of a 100 kb synthetic fragment of DNA leaves a -1 / +1 selection cassette on the genome that serves as a landing site for downstream integration of a second fragment of synthetic DNA carrying a -2 / +2 selection cassette. In the example shown, +1 is KanR, -1 is rpsL, +2 is CmR, and -2 is sacB; however, the same logic can be used with different permutations of markers on the genome and BAC. [Figure 5]Figure 1. Rewriting of ftsI-murE and map in fragment 1. a, Rewriting landscape of fragment 1. We sequenced six clones after REXER. Each dot represents the frequency of rewriting (y-axis) in the sequenced clone for the target codon at the indicated position in the genome (x-axis). Black dots indicate positions where we observed no rewriting. Four codons in ftsI-murE and the refactoring and one codon in map were rejected. b, Refactoring of a 14-bp ftsI-murE duplication. Codons and duplications are in gray, scaled by their post-REXER replacement frequency in sequenced clones. Using our initial refactoring scheme (1), which duplicated the duplication plus 20 bp of upstream sequence, we did not observe replacement of the duplication with synthetic DNA (in the six clones sequenced after REXER). Refactoring scheme 2, which duplicates the duplication plus 182 bp of upstream sequence, resulted in complete rewriting of this region in 12 of the 16 sequenced post-REXER clones. c, Testing of alternative codons at Ser4 in map. The pheS*-HygR double selectable marker on the constitutive EM7 promoter was introduced upstream of map, followed by an RBS. We replaced the cassette using linear double-stranded DNA introducing alternative codons at position 4 (as indicated) by lambda Red recombination and negative selection for loss of pheS*. DNA with AGC and AGT did not integrate (0 / 16 clones); we recovered one clone for AGC, but sequencing revealed that it contained a mutant AAC (Asn) codon. TCT (6 / 8), TCC (6 / 16), ACA (6 / 8), and TTA (4 / 8) were tolerated. d, Rewriting of the landscape across the genomic region shown in (a) after refactoring scheme 2 for the ftsI-murE duplication and REXER on a BAC containing TCT at position 4 of the map. 2 / 7 post-REXER clones were completely refactored and rewritten, with each target codon replaced in at least 5 / 7 clones. Data from (a) is shown for comparison. [Figure 6-1] ~ [Figure 6-2]Rewriting of rne and yceQ in fragment 9. a) Rewriting landscape of fragment 9. Our designed synthetic sequence for fragment 9 was integrated into the genome using REXER, and 19 clones were fully sequenced by NGS. The rewriting landscape graph shows the rewriting frequency of each target codon across 19 clones. While most codon substitutions were accepted, rewriting of a 26 kb region was consistently rejected; codon positions with zero rewriting frequency in all sequenced clones are indicated by black dots. To pinpoint the problematic sequence, a 10 kb stretch of the genome (G2–7) was deleted in the presence of an episomal copy of synthetic fragment 9. The synthetic sequence was sufficient to support deletion of all stretches except G4 (dark gray box), suggesting the underlying problem resides within this stretch. 0 / 19 clones were fully rewritten. b) Rewriting landscape of stretch G4. After sequencing 10 clones with REXER across the 10 kb stretch "G4," the rewriting landscape shown was generated. This revealed minimal, unambiguous rewriting in yceQ, a predicted protein-coding "gene," for which there is no evidence of transcription, protein synthesis, or homology (Pundir, S., et al., 2017. Methods Mol Biol, 1558, 41-55). All target codons in yceQ were rewritten at least once in individual clones, but never simultaneously; therefore, the minimal rewriting landscape is nonzero, with 0 / 10 clones completely rewritten. This is consistent with epistasis between targeted positions. In the map below the rewriting landscape, sequences annotated as essential and target codons are indicated. Sequence positions (x-axis) are relative to panel a. c, Design modifications of the region surrounding rne in fragment 9. Top, original design of the yceQ rewrite and rne (encoding RNAse E) regulatory sequences. Target codons are indicated. Prne1, 2, and 3 are promoters for the essential gene rne; they are found in and around the hypothetical gene yceQ.The -10 sequence of the major promoter P1rne was mutated according to our initial design. The sequence containing hairpin 1 (hp1) and hairpin 2 (hp2), which bind RNAse E to mediate transcript degradation, is shown; this sequence encompasses the remaining target codons and was also mutated according to our initial design. Bottom: The second codon in yceQ was replaced with a stop codon, and the remaining target codons retained their original sequence. Sequence positions (x-axis) are relative to panel a. d, This modified fragment 9 from c was integrated into the genome, and 4 / 5 sequenced clones were completely rewritten. The graph axes are the same as those in panel a. The rewriting landscape for modified fragment 9 from the five sequenced clones is shown in purple. Data from panel a is replicated for comparison. [Figure 7-1] ~ [Figure 7-3]Figure 1 shows the rewriting of yaaY in fragment 37a. a) Rewriting landscape of fragment 37a. Our designed synthetic sequence for fragment 37a was integrated into the genome by REXER, and six clones were fully sequenced by NGS. While most codon substitutions were accepted, rewriting of the 6.5 kb region was consistently rejected. Target codon positions that were never rewritten in the six sequenced clones are indicated by black dots. b) Identification of problematic target codons. Within the identified 6.5 kb problematic region, we first focused on codons in essential genes (dark gray arrows) rather than non-essential genes (light gray arrows). Sanger sequencing of 24 clones (black bars) showed that two clones were rewritten at all six target codons within the essential gene subsection. Sanger sequencing of the remaining target codons in the essential gene of these two clones revealed that one clone was rewritten at all 17 target codons. This clone was fully sequenced by NGS and used to generate a rewriting landscape in which each target codon was either rewritten or not. This, combined with the rewriting landscape in (a), allowed us to identify a problematic region 1.8 kb upstream of ribF. Here, we focused on four target codons in genes rpsT and yaaY as the closest codons to the essential ribF gene. Sanger sequencing of 33 clones spanning this sequence revealed only one codon that was never rewritten: the codon for Ser70 in the hypothetical gene yaaY (sequencing results are shown as gray, scaled on the genetic maps of rspT and yaaY). Therefore, we investigated alternative codon substitutions in yaaY. c, Alternative codon substitutions in hypothetical gene yaaY. Replacement of TCA with AGT at Ser70 in this gene was unsuccessful.To investigate alternative codon replacement schemes, we introduced a double selectable marker, pheS*-HygR, followed by a RBS on the constitutive EM7 promoter into yaaY, 12 bp upstream of the codon for Ser70. A negative selectable marker was then used to select clones that replaced the cassette by lambda Red recombination using linear double-stranded DNA introducing an alternative codon at position 70. Linear double-stranded DNA with AGT did not integrate (0 / 16 clones), but integration of dsDNA with TCC (2 / 16), TCG (2 / 16), TCT (6 / 16), and AGC (9 / 16) proved viable. d, Rewrite landscape of REXER with a BAC containing the correct fragment 37a carrying AGC at position Ser70 in the hypothetical gene yaaY. When integrated by REXER, we identified 1 / 7 fully rewritten clones. AGC at Ser70 in yaaY was introduced into 4 / 7 clones. [Figure 8] Figure 1 shows the replacement of a duplication of the hypothetical gene yceQ with a regulatory element in rne, which encodes the essential protein RNAse E. (a) In our original design, a programmed substitution of TCA to AGT in the hypothetical gene yceQ results in a mutation in the -10 promoter element of P1rne (boxed). The transcription start site (TSS) of this promoter for rne transcription is indicated by an arrow; this is the major promoter for rne transcription. (b) The replacement of the target codon overlaps with and may disrupt important regulatory hairpins hp2 and hp3 in the long 5' UTR of the rne transcript. hp2 and hp3 mediate a regulatory feedback loop by which RNAse E is recruited to mRNA to promote the degradation of its own transcript. A schematic diagram of the wild-type secondary structure of the rne 5' UTR is shown (Diwa, A., et al., 2000 Genes Dev 14, 1249-1260). The target codon for the synonymous substitution is highlighted. [Figure 9-1] ~ [Figure 9-2]Completion of compartments A-B and H. a) GENESIS started with fragment 4 and proceeded smoothly to fragment 9, where we were unable to rewrite yceQ. Identification and correction of a problem with our initial design of fragment 9 was performed as described in Figure 6 by introducing a stop codon at the start of the predicted yceQ ORF. After exchanging the sacB-CmR(sC) double selection cassette at the end of fragment 9 for the pheS*-HygR(pH) double selection cassette, this strain was primed to act as a recipient for conjugation to assemble a strain in which fragments 4-13 (compartments A+B) were fully rewritten. In parallel, we continued to rewrite the strain containing rewritten fragment 4 to incomplete fragment 9 by GENESIS; this generated a second strain for assembly in which fragments 4-8 and 10-13 were fully rewritten and fragment 9 was partially rewritten. We then integrated oriT 3 kb upstream of the start of fragment 10 in the second strain to generate a donor for conjugation to assemble a strain in which segments 4–13 (segment A+B) were completely rewritten. Conjugation of the donor and recipient strains resulted in a strain in which segments A and B were completely rewritten. b, Individual rewrites of fragments 37a and 1 resulted in incomplete rewrites. We performed both troubleshooting runs independently (Figures 5 and 7). Repair is shown. Each strain then served as the starting point for two independent sets of GENESIS runs: one generating segments 37a–37b (left) and terminating with the rpsL-KanR (rK) cassette, and the other generating segments 1–3 (right) and terminating with the sacB-CmR cassette. We integrated oriT 3 kb upstream of the start of fragment 1, and this strain served as a donor for the directed conjugation of 1 to 3 to 37a to 37b. Correct producers were selected by the acquisition of CmR and loss of rpsL, completing compartment H in a single strain. [Figure 10]Figure 1 shows the assembly of organisms with complete synthetic genomes by conjugation of rewritten genome segments. Synthetic genome segments from multiple individual, partially rewritten genomes were assembled into a single, completely rewritten genome by conjugation (Ma, NJ, et al., 2014. Nat Protoc 9, 2285-2300). The donor (d) and recipient (r) strains possess unique rewritten genome segments; overlapping rewritten homologous regions (3 kb to 400 kb) were utilized to seamlessly recombine the strains. Small homologous regions ranging from 3 to 5 kb are indicated by an asterisk (*). Conjugations in which we used homology (HR) greater than 5 kb are indicated by letters. For assembly, the rewritten genome content from the donor was conjugated clockwise to replace the corresponding wild-type genome segment in the recipient. The origins of strains AB and H are detailed in Figure 9, while all other individual synthetic genomes were generated by GENESIS (Figure 4). Conjugation followed by recombination proceeded until the final fully rewritten strains A through H were assembled, and the sequences were verified by NGS sequencing. [Figure 11]Figure 1 shows the assembly of rewritten genome segments into fully rewritten organisms. a) Schematic of the assembly of partial synthetic donor and recipient genomes into a full synthetic genome by conjugation. In recipient cells, the rewritten genome segment is extended with rewritten DNA, typically 3-4 kb, by lambda Red-mediated recombination and positive and negative selection; this step utilizes genomic markers at the ends of the rewritten sequence introduced by GENESIS, providing regions of homology with the ends of the rewritten fragment in the donor strain. The donor strain is prepared by integrating an origin of transfer (oriT) at the end of the rewritten DNA. Positive and negative selection, as shown, ensures viability of the recipient strain and selects recipients that have successfully integrated the synthetic DNA from the donor. Conjugation of the donor genome to the recipient was facilitated using an F' plasmid containing a mutation in the oriT sequence that precludes transfer. +2, CmR; -2, SacB; +3, HygR; -3, pheS*; +4, gentamicin R; +5, tetracycline R. b, Synthetic genome segments from multiple individual partially rewritten genomes were assembled into a single fully rewritten genome by the indicated sequence of conjugation. Donor (d) and recipient (r) strains carry unique rewritten genome segments. Rewritten genome content from the donor was conjugated clockwise to replace the corresponding WT genome segment in the recipient. Conjugation proceeded until the final fully rewritten strains A through H were assembled. Figure 10 shows the process in more detail, including all homologous regions. [Figure 12]Figure 1 shows the functional results of synonymous codon compression in Syn61. a) Synonymous codon compression and deletion of prfA, serU, and serT. Gray boxes indicate the serine and stop codons along with the tRNAs and their corresponding termination factors in wild-type E. coli (WT genome). The tRNA anticodons and termination factors are attached to the codons they decode with black lines. The tRNA and termination factor genes are shown in black boxes. serT is the only tRNA that decodes the TCA codon in wild-type E. coli and is essential. Synonymous codon compression (Syn.Codon.Comp., Synonymous codon compression) results in a rewritten genome in which i) tRNAs with CGA anticodons should not have cognate codons, and ii) serT should be non-essential. All factors that read the target codon should be non-essential in Syn61. b, Cotranslational incorporation of the noncanonical amino acid (ncAA) Nε-(((2-methylcycloprop-2-en-1-yl)methoxy)carbonyl)-L-lysine (CYPK, Nε-(((2-methylcycloprop-2-en-1-yl)methoxy)carbonyl)-L-lysine) using the orthogonal MmPylRS / tRNAPyl CGA pair was toxic in MDS42 but not in Syn61. When CYPK was provided, this pair incorporated the ncAA in response to the TCG codon in a dose-dependent manner. In MDS42, this incorporation resulted in proteome missynthesis and toxicity. However, in Syn61, which does not contain a TCG codon, it was nontoxic. The lines follow the average of three biological replicates (each shown as a point) at each [CYPK] concentration (0 mM, 0.5 mM, 1 mM, 2.5 mM, and 5 mM). "% Maximum growth" was determined by the final OD600, obtained by dividing the indicated concentration of CYPK by the final OD600 in the absence of CYPK. The final OD600 was determined after 600 min. c, Synonymous codon compression allows deletion of serT in Syn61. PCR flanking the serT locus before (-) and after (clones 1 and 2) replacement with the PheS*-HygR cassette.See also Figure 14. Figure 16. Complete gel. [Figure 13]Figure 1 shows the characterization of organisms with completely synthetic genomes. a) Doubling times for Syn61 and MDS42. Our completely synthetic, engineered E. coli strain, Syn61, has a doubling time 1.6-fold greater than that of the parent strain, MDS42 (Posfai, G. et al., 2006. Science 312, 1044-1046), when grown under standard media conditions (90.1 min vs. 57.6 min in LB + 2% glucose). The ratio of growth rates of Syn61 to MDS42 is 1.7 in LB (reduced carbon catabolite repression) at 37°C, 1.7 in M9 minimal medium, 1.4 in richer medium (2XTY), 2.5 in LB at 25°C, and 1.3 in LB at 42°C. The doubling times for MDS42 and Syn61 in different media conditions are listed: LB at 37°C, 58.3 min and 100.6 min; LB + 2% glucose, 57.6 min and 90.1 min; M9 minimal medium, 130.5 min and 221.1 min; 2XTY, 68.2 min and 92.6 min; LB at 25°C, 86.3 min and 218.4 min; and LB at 42°C, 77.4 min and 99.7 min. Syn61 carrying the serV plasmid without (-) or with (+) exhibited a growth rate ratio of 0.99 (138.3 min vs. 136.2 min). Doubling times represent the mean ± standard deviation from the mean of 10 independently grown biological replicates of each strain (see Methods). b, Representative microscopic images of E. coli strains MDS42 and Syn61. Samples were imaged on an upright Zeiss Axiophot phase-contrast microscope using a 63X 1.25NA Plan Neofluar phase objective (see Methods). c, Histogram of cell length quantified from microscopic images of strains MDS42 and Syn61. The average cell length for MDS42 was 1.97 ± 0.57 μm and for Syn61 was 2.3 ± 0.74 μm. Images of n = 500 cells were taken during the exponential growth phase for both strains. Cell length measurements were performed with Nikon NIS Elements software (see Methods). d, Label-free quantification of the MDS42 and Syn61 proteomes. Each strain was grown in three biological replicates. Each biological replicate was analyzed by tandem mass spectrometry in technical replicates.Technical replicates of biological replicates were merged. A total of 1,084 proteins were quantified across samples. P values for differences in abundance were calculated by two-sample T-test for proteins quantified in at least two biological replicates. Data showed that the abundance of three proteins significantly (P = 0.01) differed between strains: aminopeptidase N (P04825) and peptidase T (P29745) were overrepresented in Syn61, whereas 30S ribosomal protein S20 (P0A7U7) was underrepresented. Protein abundance did not differ by more than 1.14-fold between strains, as judged by LFQ values. [Figure 14-1] ~ [Figure 14-4]Figure 12 shows the results of synonymous codon compression in Syn61. a) Synonymous codon compression and deletion of prfA, serU, and serT in E. coli. Gray boxes show the E. coli serine and stop codons, along with the tRNA and terminator that decode them in wild-type E. coli (WT genome). The tRNA anticodon and terminator are attached to the codons from which they are read with black lines. The tRNA and terminator genes are shown in black boxes. Synonymous codon compression (Syn.Codon.Comp.) results in Syn61 cells with a rewritten genome in which the TCG and TCA codons have been removed. The abundance of each codon is listed within its box. b) Same as Figure 12b, except for the MmPylRS / tRNAPyl anticodon, UGA, shown. Because there are fewer cognate codons for this tRNA in Syn61 than in MDS42, CYPK addition, as observed, might be predicted to be less toxic in Syn61. c, Same as Figure 12b, except for the MmPylRS / tRNAPyl anticodon, GCU, shown. Because there are more cognate codons for this tRNA in Syn61 than in MDS42, CYPK addition, as observed, would be predicted to be more toxic in Syn61. d, serT (dark gray) was deleted by insertion of the PheS*-HygR cassette (black) via lambda-Red-mediated recombination. Recombination generates new junctions 1 and 2, as shown. For each recombination, both junctions were sequence-verified by Sanger sequencing. Arrows above the Sanger chromatogram indicate the exact location of the junction, the sequence corresponding to the selection cassette, and bars correspond to genomic sequence flanking the selection cassette. Primers used to generate selection cassettes with appropriate homology to serU, serT, and prfA for recombination are provided in Figure 23. e, prfA (dark gray) is deleted by insertion of rpsL-KanR (black) via lambda-Red-mediated homologous recombination. The agarose gel is annotated as described in Figure 12c, and the remainder of the data is annotated as described in panel d.The full gel is available in Figure 16. f, serU (dark grey) has been deleted by insertion of a PheS*-HygR cassette (black) by lambda-Red-mediated recombination. The agarose gel is annotated as described in Figure 12c, and the remainder of the data is annotated as described in panel d. The full gel is available in Figure 16. [Figure 15]Figure 1 shows the scale of genome synthesis and the scale and fidelity of rewriting. a, Genome and chromosome synthesis. Sizes (Mb) of synthetic genomes produced for M. genitalium and M. mycoides (Gibson, D.G. et al., 2008. Science 319, 1215-1220; and Gibson, D.G. et al., 2010. Science 329, 52-56) and several S. cerevisiae chromosomes (Shen, Y. et al., 2017. Science 355, aaf4791; Annaluru, N. et al., 2014. Science 344, 55-58; Xie, Z.X. et al., 2017. Science 355, aaf4704; Mitchell, L.A. et al., 2017. Science 355, aaf4831; Dymond, J.S. et al., 2011. Nature 477, 471-476; Wu, Y. et al., 2017. Science 355, aaf4706; Zhang, W. et al., 2017. Science 355, aaf3981; and Richardson, SM et al., 2017. Science 355, 1040-1044) are shown in light gray. The size of the synthetic E. coli genome presented here is shown in dark gray. b, Genome rewriting attempt.Attempts to rewrite the target codons TTA and TTG in S. typhimurium (Lau, Y. et al., 2017. Nucleic Acids Res 45, 6971-6980); AGC, AGT, TTG, TTA, AGA, AGG, and TAG in E. coli (Ostrov, N. et al., 2016. Science 353, 819-822); AGA and AGG in E. coli (Napolitano, M. G. et al., 2016. Proc Natl Acad Sci USA 113, E5588-5597), and the rewriting of all TAG in E. coli (Lajoie, M. J. et al., 2013. Science 342, 357-360) are shown in light grey. A comparison with total removal of TCA, TCG, and TAG in E. coli is presented here (dark gray). The total number of codons rewritten in a single strain is graphed, along with the maximum percentage of target codons rewritten in a single strain for each trial. c, Number of reported unprogrammed mutations and indels as a function of the number of target codons rewritten for the experiments shown in b. [Figure 16] Figure 13 shows the complete gel for Figure 12. The complete gel is shown in the corresponding figure panel. Molecular size standards are annotated and the area shown in the associated figure is outlined in white. [Figure 17-1] ~ [Figure 17-2]This diagram illustrates codon and anticodon interactions in the E. coli genome. The 28 sense codons, along with the amber stop codon, are highlighted in gray. Genome-wide removal of these sense codons, but not other sense codons, allows for the deletion of all their cognate tRNAs without eliminating the ability to decode one or more remaining sense codons in the genome. This is necessary, but not sufficient, for the reassignment of sense codons to unnatural monomers. The codon boxes for serine, leucine, and alanine are highlighted because the endogenous aminoacyl-tRNA synthetases for these amino acids do not recognize the anticodons of their cognate tRNAs. This can facilitate the assignment of codons within these boxes to new amino acids by introducing tRNAs bearing cognate anticodons that do not lead to erroneous aminoacylation by the endogenous synthetases. The total number of all 64 triplet codons in the MDS42 genome (GenBank accession number AP012306), all known codon-anticodon interactions by both Watson-Crick base pairing and wobble, base modifications of tRNA anticodons, tRNA genes, and tRNA relative abundances measured in vivo are reported. This analysis identified 10 codons from the serine, leucine, and alanine groups (serine codons TCG, TCA, AGT, AGC; leucine codons CTG, CTA, TTG, TTA; and alanine codons GCG, GCA) that satisfy both the codon-anticodon interaction and the aminoacyl-tRNA synthetase recognition criteria for codon reassignment. [Figure 18] FIG. 1 shows a designed synthetic E. coli genome (SEQ ID NO: 1). A version of the E. coli MDS42 genome in which the serine codons TCG and TCA and the stop codon TAG within open reading frames (ORFs) have been systematically replaced by their synonyms AGC, AGT, and TAA, respectively. Defined rules for synonymous codon compression and refactoring are used to design a genome in which all 18,218 target codons have been rewritten to their target synonyms. [Figure 19]Figure 2 shows the final synthetic E. coli genome (Syn61) (SEQ ID NO: 2). The sequence of E. coli Syn61, in which all 1.8 x 10 target codons in the genome have been rewritten. Our synthesis of the rewritten genome introduced only eight unprogrammed mutations (Table 6); four of these mutations arose during the preparation of the 100 kb BAC and four arose during the rewriting process. [Figure 20-1] ~ [Figure 20-13] Figure 1 shows the BACs for assembling synthetic genomes. A, BAC-sacB-CmR-rpsL. Nucleotide sequence for the annotated BAC vector carrying a sacB-CmR selection cassette flanked upstream by a 5' homology region (HR) and a CRISPR / Cas9 protospacer sequence (spacer 1). The sacB-CmR cassette is flanked downstream by a 3' homology region, a CRISPR / Cas9 protospacer sequence (spacer 2), and an rpsL selection marker. B, - BAC-rpsL-KanR-sacB. Nucleotide sequence for the annotated BAC vector carrying a rpsL-KanR selection cassette flanked upstream by a 5' homology region (HR) and a CRISPR / Cas9 protospacer sequence (spacer 1). The rpsL-KanR cassette is flanked downstream by a 3' homology region, a CRISPR / Cas9 protospacer sequence (spacer 2), and a sacB selection marker. C, BAC-rpsL-KanR-pheS*-HygR. Nucleotide sequence for an annotated BAC vector carrying the rpsL-KanR selection cassette flanked upstream by a 5' homology region (HR) and a CRISPR / Cas9 protospacer sequence (spacer 1). The rpsL-KanR cassette is flanked downstream by a 3' homology region, a CRISPR / Cas9 protospacer sequence (spacer 2), and a pheS*-HygR selection marker. D, Table of BAC construction. Oligonucleotides used to construct the BAC using synthetic DNA for REXER and homology regions between the synthetic DNA fragments and selection marker. The second tab lists the plasmid backbone and protospacer sequences used for REXER. [Figure 21-1] ~ [Figure 21-2]
[0023] Figures 1A-1C show exemplary spacer plasmid maps. A, Spacer plasmid map. Exemplary map of pKW1_MB1amp_spacer_REXER2 containing a CRISPR insert with a spacer sequence used as a linear or circular spacer for REXER. B, Second generation spacer plasmid map. Exemplary map of pKW3_MB1amp_spacer_REXER2 containing a CRISPR insert with a spacer sequence used as a circular second generation spacer for REXER. [Figure 22-1] ~ [Figure 22-4] Figure 1 shows the constructs for conjugation. A, Gentamicin resistance OriT cassette. B, Primers for the conjugation construct. Oligonucleotide primers used for conjugation. C, pJF146. Non-self-transmissible F' plasmid. [Figure 23]
[0033] Figure 1 shows primers for deletion experiments. Oligonucleotide primers used for deletion of tRNAs serT and serU and release factor prfA in Syn61. DETAILED DESCRIPTION OF THE INVENTION
[0046] Detailed Description As used herein, "comprising," "comprises," and " The term "comprised of" does not necessarily mean "including" or "Includes" is synonymous with "containing" or "contains" and is inclusive or open-ended and does not exclude additional, unrecited components, elements, or steps. and the term "comprised of" should be interpreted as meaning "consisting of " is also included.
[0047] Synthetic Genome genome As used herein, a "genome" is the genetic material of an organism, including both genes and non-coding DNA. As used herein, a "synthetic genome" is a synthetically constructed genome. Typically, a synthetic genome is produced by genetic modification of an existing (i.e., "parent") genome. Thus, a synthetic genome can be derived from a parent genome, i.e., can be identical to the parent genome except for containing one or more genetic modifications. One of skill in the art would readily be able to identify the parent genome on which a synthetic genome is based and the genetic modifications performed. As used herein, a "parent genome" can be any naturally occurring, commercially available, deposited, cataloged, or otherwise known genome, or a derivative thereof.
[0048] The synthetic genomes of the present invention are synthetic prokaryotic genomes. Prokaryotes are unicellular organisms that lack a membrane-bound nucleus, mitochondria, or any other membrane-bound organelles. Prokaryotes are divided into two domains: Archaea and Bacteria. Prokaryotic genomes are generally circular double-stranded pieces of DNA, multiple copies of which may exist at any time.
[0049] Preferably, the synthetic genome of the present invention is a synthetic bacterial genome. Preferably, the synthetic bacterial genome is suitable for heterologous protein production, particularly for the production of polypeptides comprising one or more non-proteinogenic amino acids (e.g., those described in Ferrer-Miralles, N. and Villaverde, A., 2013. Microbial Cell Factories, 12:113). Suitable bacterial genomes include Escherichia (e.g., Escherichia coli), Caulobacteria (e.g., Caulobacter crescentus), photosynthetic bacteria (e.g., Rhodobacter sphaeroides), cold-adapted bacteria (e.g., Bacillus subtilis), and the like. Type bacteria (e.g., Pseudoalteromona shaloplanktis, Shewanella sp. strain Ac10), pseudomonads (e.g., Pseudomonas fluorescens), , Pseudomonas putida, Pseudomonas aeruginosa), halophilic bacteria (e.g., Halomonas elongate, Chromohalobacter salexigens) , streptomycetes (e.g., Streptomyces lividans, Streptomyces griseus) ), Nocardia (e.g., Nocardia lactamjurans lactamdurans), mycobacteria (e.g., Mycobacterium Mycobacterium smegmatis), coryneform bacteria (e.g., Corynebacterium glutamicum, Corynebacterium ammoniagenes, Brevibacterium lactofermentum), bacilli (e.g., Bacillus subtilis, Bacillus brevis, Bacillus megaterium, Bacillus rickettsii), Lactic acid bacteria (e.g., Lactococcus lactis, Lactobacillus plantarum, Lactobacillus niformis, Bacillus licheniformis, Bacillus amyloliquefaciens), and lactic acid bacteria (e.g., Lactococcus lactis, Lactobacillus plantarum, Lactobacillus Examples include the genomes of Lactobacillus casei, Lactobacillus reuteri, and Lactobacillus gasseri. In some embodiments, the synthetic genome is a synthetic Gram-negative bacterial genome.
[0050] Bacterial genomes can range in size from about 130 kb to greater than 14 Mb. Thus, in some embodiments, the synthetic prokaryotic genomes of the invention are 100 kb to 20 Mb, or 130 kb to 15 Mb, or 200 kb to 15 Mb, or 300 kb to 15 Mb, or 500 kb to 15 Mb, or 1 Mb to 15 Mb, or 1 Mb to 10 Mb, or 1 Mb to 8 Mb, or 1 Mb to 6 Mb, or 2 Mb to 6 Mb, or 2 Mb to 5 Mb, or 3 Mb to 5 Mb, or about 4 Mb in size. The synthetic prokaryotic genome may comprise 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, or 2000 or more genes, preferably 1000 or more. The synthetic prokaryotic genome may comprise 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, or 2000 or more genes for which there is evidence of translation and / or predicted protein products, preferably 1000 or more genes. Preferably, the synthetic prokaryotic genome comprises 100 or more, 200 or more, 300 or more, 400 or more, 500 or more essential genes, preferably 300 or more essential genes.
[0051] Preferably, the synthetic genome of the present invention is a synthetic Escherichia coli genome, Salmonella enterica genome, or Shigella dysenteriae genome, as described in Lukjancenko, O., et al., 2010. Microbial ecology, 60(4), pp.708-720; and Karberg, KA, et al., 2011. PNAS, 108(50), It is a phylogenetically related species as disclosed on pp. 20154-20159.
[0052] More preferably, the synthetic genome of the present invention is a synthetic E. coli genome. The parent genome may be any suitable E. coli genome, including MDS42, K-12, MG1655, BL21, BL21(DE3), AD494, Origami, HMS174, BLR(DE3), HMS174(DE3), Tuner(DE3), Origami2(DE3), Rosetta2(DE3), Lemo21(DE3), NiCo21(DE3), T7 Express, SHuffle Express, C41(DE3), C43(DE3), and m15 pREP4 or derivatives thereof (Rosano, GL and Ceccarelli, EA, 2014. Frontiers in microbiology, 5, p. 172). Most preferably, the parent genome is MDS42, MG1655, or BL21 or derivatives thereof. MG1655 is considered a wild-type strain of E. coli. The GenBank ID for the genome sequence of this strain is U00096. BL21 is widely available commercially. For example, it can be purchased from New England BioLabs under catalog number C2530H (https: / / www.neb.com / products / c2530-bl21-competent-e-coli).
[0053] In some embodiments, the synthetic genome is a reduced synthetic genome or a minimal synthetic genome. A "reduced genome" is one in which the size of a parent genome has been reduced by removing non-essential genes and / or non-coding regions. A "minimal genome" is a genome that has been reduced to its minimum size while maintaining viability, for example, by deleting all non-essential regions of the genome.
[0054] The synthetic genomes of the invention can be viable genomes. As used herein, a "viable genome" refers to a genome that contains sufficient nucleic acid sequences to cause and / or maintain cell viability, e.g., a genome that encodes molecules required for replication, transcription, translation, energy production, transport, production of membrane and cytoplasmic components, and cell division.
[0055] Preferably, one or more tRNAs or terminators may be deleted from the synthetic genome, and the synthetic genome may remain viable. For example, tRNAs that decode only the one or more sense codons that have been replaced (or deleted) may be non-essential. Similarly, tRNAs that decode the one or more sense codons that have been replaced (or deleted) may be non-essential if the remaining sense codons that the tRNA decodes can also be decoded by alternative tRNAs. For example, tRNAs Ser UGA serT, which encodes serT, is normally essential because it is the only tRNA that decodes TCA codons in E. coli. However, if the synthetic genome does not contain a TCA codon, serT may be non-essential.
[0056] Sense codon The present invention provides synthetic prokaryotic genomes that comprise five or four or fewer occurrences of one or more sense codons; and / or synthetic prokaryotic genomes derived from a parental genome, wherein the synthetic prokaryotic genome comprises less than 10%, 5%, 2%, 1%, 0.5%, 0.1% occurrences of one or more sense codons compared to the parental genome; and / or synthetic prokaryotic genomes that comprise 100 or 101 or more, 200 or 201 or more, or 1000 or 1001 or more genes that have no occurrences of one or more sense codons.
[0057] The one or more sense codons may consist of 1, 2, 3, 4, 5, 6, 7, or 8 sense codons. Preferably, the one or more sense codons consist of 1 sense codon or 2 sense codons, most preferably 2 sense codons.
[0058] The synthetic prokaryotic genome may include no more than five or four (e.g., five, four, three, two, one) or no occurrences of one or two or more (e.g., one, two, three, four, five, six, seven, or eight) sense codons. In some embodiments, the synthetic prokaryotic genome includes no more than five or four (e.g., five, four, three, two, one, zero) occurrences of each of one or two or more (e.g., one, two, three, four, five, six, seven, or eight) sense codons. In other embodiments, the synthetic prokaryotic genome comprises five or four or fewer (e.g., five, four, three, two, one, zero) of a total (i.e., total) of one or two or more (e.g., one, two, three, four, five, six, seven, or eight) sense codons. In preferred embodiments, the synthetic prokaryotic genome does not comprise one occurrence of a sense codon. In other preferred embodiments, the synthetic prokaryotic genome does not comprise two occurrences of a sense codon.
[0059] A synthetic prokaryotic genome may be derived from a parental genome and include no, or no, five or four (e.g., five, four, three, two, one) occurrences of one or two or more (e.g., one, two, three, four, five, six, seven, or eight) naturally occurring sense codons. In some embodiments, a synthetic prokaryotic genome includes no, or no, five or four (e.g., five, four, three, two, one, zero) occurrences of each of one or two or more (e.g., one, two, three, four, five, six, seven, or eight) naturally occurring sense codons. In other embodiments, the synthetic prokaryotic genome comprises five or four or fewer (e.g., five, four, three, two, one, zero) of the combined (i.e., total) one or more (e.g., one, two, three, four, five, six, seven, or eight) natural sense codons. In one embodiment, the synthetic prokaryotic genome is derived from the parent genome and does not contain one occurrence of a natural sense codon. In another preferred embodiment, the synthetic prokaryotic genome is derived from the parent genome and does not contain two occurrences of a natural sense codon.
[0060] In some embodiments, the synthetic prokaryotic genome comprises 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, or 2000 or more genes, preferably 1000 or more. In some embodiments, the genes are ones for which there is evidence of translation and / or predicted protein products. For example, the synthetic prokaryotic genome may comprise 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 1500 or more, or 2000 or more genes, preferably 1000 or more genes, for which there is evidence of translation and / or predicted protein products. Preferably, the synthetic prokaryotic genome comprises 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, preferably 300 or more. Preferably, the (essential) genes do not have one or more occurrences of a sense codon.
[0061] The synthetic prokaryotic genome may comprise less than 10%, 5%, 2%, 1%, 0.5%, 0.1% occurrences of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, or 8) sense codons compared to the parental genome. In some embodiments, the synthetic prokaryotic genome comprises less than 10%, 5%, 2%, 1%, 0.5%, 0.1% occurrences of each of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, or 8) sense codons compared to the parental genome. In other embodiments, the synthetic prokaryotic genome comprises less than 10%, 5%, 2%, 1%, 0.5%, 0.1% occurrences of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, or 8) sense codons combined compared to the parental genome. In preferred embodiments, the synthetic prokaryotic genome contains 10%, 5%, 2%, 1%, 0.5%, 0.1% less than one sense codon compared to the parental genome, hi other preferred embodiments, the synthetic prokaryotic genome contains 10%, 5%, 2%, 1%, 0.5%, 0.1% less than two sense codons compared to the parental genome.
[0062] A synthetic prokaryotic genome may comprise 100 or 101 or more, 200 or 201 or more, or 1000 or 1001 or more genes that lack one or two or more (e.g., one, two, three, four, five, six, seven, or eight) occurrences of a sense codon. Preferably, all or substantially all genes in the synthetic prokaryotic genome lack one or two or more (e.g., one, two, three, four, five, six, seven, or eight) occurrences of a sense codon. In preferred embodiments, all or substantially all genes in the synthetic prokaryotic genome lack one occurrence of a sense codon. In other preferred embodiments, all or substantially all genes in the synthetic prokaryotic genome lack two occurrences of a sense codon. Substantially all means that all but ten or nine (e.g., ten, nine, eight, seven, six, five, four, three, two, one, or zero) genes contain one or more occurrences of a sense codon.
[0063] The synthetic prokaryotic genome may be a 100 or 101 genome in which one or more (e.g., one, two, three, four, five, six, seven, or eight) occurrences of natural sense codons are absent. The synthetic prokaryotic genome may comprise 1 or more, 200 or 201 or more, or 1000 or 1001 or more genes. Preferably, all or substantially all genes in the synthetic prokaryotic genome do not have one or more (e.g., 1, 2, 3, 4, 5, 6, 7, or 8) occurrences of a natural sense codon. In preferred embodiments, all or substantially all genes in the synthetic prokaryotic genome do not have one occurrence of a natural sense codon. In other preferred embodiments, all or substantially all genes in the synthetic prokaryotic genome do not have two occurrences of a natural sense codon. By substantially all, it is meant that all but 10 or 9 or fewer (e.g., 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0) genes contain one or more occurrences of a natural sense codon.
[0064] Preferably, the genes encode proteins (e.g., the genes are ones for which there is evidence of translation and / or predicted protein products), and / or the genes are essential genes. Thus, in more preferred embodiments, the synthetic prokaryotic genome comprises 100 or more, 200 or more, or 1000 or more protein-encoding genes and / or 100 or more, 200 or more, or 300 or more essential genes that are missing one or two occurrences of sense codons. In other more preferred embodiments, all or substantially all protein-encoding genes and / or essential genes in the synthetic prokaryotic genome do not comprise one or two occurrences of sense codons.
[0065] In preferred embodiments, no protein is translated from any of the one or more remaining occurrences of the sense codon, and / or the gene containing the one or more remaining occurrences of the sense codon is a putative or non-coding gene. In some embodiments, translation of a gene containing one or more remaining occurrences of the sense codon is reduced and / or prevented (e.g., the gene may contain a stop codon in the 5' sequence).
[0066] Any remaining occurrences of sense codons may be necessary to ensure the viability of the synthetic prokaryotic genome. For example, one or more, preferably all, of the remaining occurrences of one or more sense codons in the synthetic prokaryotic genome may be in regulatory elements of essential genes, and / or one or more, preferably all, of the remaining occurrences of one or more sense codons may be in genes for which there is no evidence for translation or predicted protein products (i.e., putative or non-coding genes).
[0067] As used herein, a "sense codon" is a nucleotide triplet that encodes an amino acid. Sense codons can therefore be identified within a genome by gene prediction, i.e., by identifying regions of the genome that encode proteins (i.e., genes) and corresponding open reading frames (ORFs). Typically, the genome naturally contains 61 sense codons: GCT, GCC, GCA, GCG, CGT, CGC, CGA, CGG, AGA, AGG, AAT, AAC, GAT, GAC, TGT, TGC, CAA, CAG, GAA, GAG, GGT, GGC, GGA, GGG, CAT, CAC, ATT, ATC, ATA, TTA, TTG, CTT, CTC, CTA, CTG, AAA, AAG, ATG, TTT, TTC, CCT, CCC, CCA, CCG, TCT, TCC, TCA, TCG, AGT, AGC, ACT, ACC, ACA, ACG, TGG, TAT, TAC, GTT, GTC, GTA, and GTG (reading 5' to 3' on the coding strand of DNA). The standard genetic code uses 61 triplet codons to encode the 20 canonical amino acids. Eighteen of the 20 amino acids are coded for by more than one synonymous codon (see Figure 17). One or more of the sense codons are one or more natural sense codons, i.e., the sense codons present in the parent genome. obtain.
[0068] The 61 sense codons in DNA are transcribed into corresponding mRNA, which is then decoded by one or more tRNAs. The tRNA carries amino acids to the ribosome as directed by the sense codons in the mRNA. The tRNA can recognize one or more sense codons through complementary anticodons. The sequence of sense codons is then translated into a polypeptide (i.e., a sequence of amino acids). The codon and anticodon interactions in the E. coli genome are shown in Figure 17.
[0069] Preferably, genome-wide removal of one or more sense codons, but not other sense codons, can delete all cognate tRNAs corresponding to said one or more sense codons without removing the ability to decode one or more remaining sense codons in the genome. Thus, the one or more sense codons can be selected from TCG, TCA, AGT, AGC, GCG, GCA, GTG, GTA, CTG, CTA, TTG, TTA, ACG, ACA, CCG, CCA, CGG, CGA, CGT, CGC, AGG, AGA, GGG, GGA, GGT, GGC, ATT, and ATC.
[0070] Aminoacyl-tRNA synthetases for serine, leucine, and alanine do not recognize the anticodons of their cognate tRNAs. This can facilitate the assignment of codons within these boxes to new amino acids by introducing tRNAs bearing cognate anticodons that do not lead to incorrect aminoacylation by endogenous synthetases. Therefore, one or more sense codons can be selected from TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, GCC, CTG, CTA, CTT, CTC, TTG, and TTA.
[0071] Preferably, one or more sense codons satisfy both of these criteria, and therefore, one or more sense codons may be selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA. More preferably, one or more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG, and GCA. Most preferably, one or more sense codons are TCG and / or TCA.
[0072] Preferably, one or more sense codons are removed so that the genome is compatible with codon reassignment to non-proteinogenic amino acids.Therefore, one or more sense codons may include one or more of TCA, CTA, or TTA.Alternatively, two or more sense codons are removed, and these two or more sense codons include one or more of the sense codon pairs selected from the group consisting of GCG and GCA; GCT and GCC; TCG and TCA; AGT and AGC; TCT and TCC; CTG and CTA; TTG and TTA; and CTT and CTC.Preferably, two or more sense codons are removed, and these two or more sense codons include one or more of the sense codon pairs selected from the group consisting of GCG and GCA; TCG and TCA; AGT and AGC; CTG and CTA; and TTG and TTA. More preferably, the two or more sense codons include TCG and TCA.
[0073] To achieve the removal of sense codons, they can be replaced with synonymous sense codons. This is preferred to ensure that the encoded protein sequence is unchanged. For example, the present invention provides a method for determining whether 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of one or more sense codons in the parent genome are synonymous. Synthetic prokaryotic genomes are provided in which synonymous codons have been substituted. Those skilled in the art can deduce appropriate synonymous codon substitutions. For example, in E. coli, typically, TCG, TCA, TCT, TCC, AGT, and AGC all encode serine, typically, GCG, GCA, GCT, and GCC all encode alanine, and typically, CTG, CTA, CTT, CTC, TTG, and TTA all encode leucine.
[0074] In some embodiments, the substitution is a defined substitution, i.e., one sense codon is replaced with a single synonymous sense codon. Preferably, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of one or more sense codons in the parental genome are replaced with a defined (i.e., single) synonymous sense codon.
[0075] For example, the defined substitutions may be: GCG replaced with either GCT or GCC; GCA replaced with either GCT or GCC; TCG replaced with any one of TCT, TCC, AGT, or AGC; TCA replaced with any one of TCT, TCC, AGT, or AGC; AGT replaced with any one of TCG, TCA, TCT, or TCC; AGC replaced with any one of TCG, TCA, TCT, or TCC; CTG replaced with any one of CTT, CTC, TTG, or TTA; CTA replaced with any one of CTT, CTC, TTG, or TTA; TTG replaced with any one of CTG, CTA, CTT, or CTC; or TTA replaced with any one of CTG, CTA, CTT, or CTC. Preferably, the one or more defined sense codon substitutions are selected from one or more of: GCG to either GCT or GCC; GCA to either GCT or GCC; TCG to either AGT or AGC; TCA to either AGT or AGC; AGT to either TCA or TCT; AGC to either TCG, TCC, or TCA; TTG to CTT; and TTA to CTC. More preferably, TCG and / or TCA are substituted with AGC and / or AGT. Most preferably, TCG is substituted with AGC, and / or TCA is substituted with AGT.
[0076] Preferably, the defined substitutions are such that the genome is compatible with codon reassignment to non-proteinogenic amino acids, for example, (i) GCG may be substituted with either GCT or GCC, and GCA may be substituted with either GCT or GCC; (ii) TCG may be substituted with either TCT, TCC, AGT, or AGC, and TCA may be substituted with either TCT, TCC, AGT, or AGC; (iii) AGT may be substituted with either TCG, TCA, TCT, or TCC. (iv) CTG may be replaced by either CTT, CTC, TTG or TTA, and CTA may be replaced by either CTT, CTC, TTG or TTA; or (v) TTG may be replaced by either CTG, CTA, CTT or CTC, and TTA may be replaced by either CTG, CTA, CTT or CTC.
[0077] Preferably, the defined substitution scheme is one or more of those listed in the table below:
[0078] [Table 1] JPEG2025175308000002.jpg18965 JPEG2025175308000003.jpg7162
[0079] Preferably, none of these codon substitutions affect the ribosome binding site (AGGAGG), a highly conserved regulatory sequence in E. coli. Selected codon substitutions may be tested in a small test region (e.g., a 20 kb region of the genome that is rich in both essential target genes and target codons) to assess viability. If the codon substitutions are not viable in the small test region, they may be ignored.
[0080] If the replacement of one or more sense codons in parent genome with defined synonymous sense codons does not produce a viable genome, alternative synonymous sense codons can be used.For example, 99.9% of the occurrences of one or more sense codons in parent genome can be replaced with defined (i.e., single) synonymous sense codons, and the remaining 0.1% can be replaced with alternative synonymous sense codons.For example, 99.9% of the occurrences of TCG can be replaced with AGC, and 0.1% can be replaced with TCT, TCC, AGT or AGC; and / or 99.9% of the occurrences of TCA can be replaced with AGT, and 0.1% can be replaced with TCT, TCC, AGT or AGC.
[0081] As used herein, a "stop codon" is a nucleotide triplet that codes for the termination of translation into a protein. Typically, genomes contain three stop codons: TAA ("ochre"), TGA ("opal" or "umber"), and TAG ("annular"). Naturally contains hydroxybenzoates (bar).
[0082] In some embodiments, the synthetic prokaryotic genome further comprises no more than 10 or 9, no more than 5 or 4 occurrences of one or two stop codons, and preferably no more than 10 or 9, no more than 5 or 4 occurrences of the amber stop codon (TAG). Preferably, 90% or more, 95% or more, 98% or more, 99% or more, or all occurrences of TAG in the parental prokaryotic genome are replaced with TAA (ochre stop codon). In preferred embodiments, the synthetic prokaryotic genome comprises no occurrences of the amber stop codon (TAG), and optionally all occurrences of TAG in the parental prokaryotic genome are replaced with TAA (ochre stop codon).
[0083] Thus, in a preferred embodiment, the synthetic prokaryotic genome of the invention does not contain one or more, or two or more, occurrences of sense codons and does not contain one occurrence of a stop codon, preferably an amber stop codon (TAG). In a more preferred embodiment, the synthetic prokaryotic genome of the invention does not contain two occurrences of sense codons, preferably TCG and TCA. The parental prokaryotic genome does not include occurrences of the amber stop codon (TAG), and TCG, TCA, and TAG in the parental prokaryotic genome may be replaced with synonymous codons, for example, 99.9% or more of the occurrences of TCG in the parental prokaryotic genome are replaced with AGC, 99.9% or more of the occurrences of TCA in the parental prokaryotic genome are replaced with AGT, and all occurrences of TAG in the parental prokaryotic genome are replaced with TAA.
[0084] In some embodiments, the synthetic prokaryotic genome comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, or 99.9% identical to SEQ ID NO:1 or SEQ ID NO:2.
[0085] The present invention provides synthetic prokaryotic genomes that are at least 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.95% or 100% identical to SEQ ID NO:1 or SEQ ID NO:2.
[0086] Sequence comparisons may be performed by eye, or more usually, with the aid of readily available sequence comparison programs. These publicly and commercially available computer programs can calculate the sequence identity between two or more sequences.
[0087] Sequence identity can be calculated over a continuous sequence, that is, one sequence is aligned with another sequence, and each amino acid in one sequence is directly compared with the corresponding amino acid in the other sequence, one residue at a time.This is called "ungap" alignment.Typically, such ungap alignment is only carried out over a relatively small number of residues (for example, less than 50 consecutive amino acids).
[0088] While this is a very simple and consistent method, it does not take into account that, for example, in an otherwise identical sequence pair, a single insertion or deletion may result in the subsequent amino acid residue being excluded from the alignment, and therefore may result in a significant drop in percent homology when a global alignment is performed. As a result, most sequence comparison methods are designed to produce optimal alignments that take into account possible insertions and deletions without unduly penalizing the overall homology score. This is achieved by inserting "gaps" in the sequence alignment to attempt to maximize local homology.
[0089] However, these more complex methods assign a "gap penalty" to each gap that occurs in the alignment, so that sequence alignments with as few gaps as possible (reflecting a higher relatedness between the two compared sequences) score higher than those with many gaps, for the same number of identical amino acids. Typically, an "affine gap cost" is used, which imposes a relatively high cost on the existence of a gap and a smaller penalty for each subsequent residue in the gap. This is the most commonly used gap scoring system. High gap penalties naturally produce optimized alignments with fewer gaps. Most alignment programs allow the gap penalty to be modified. However, when using such software for sequence comparison, it is preferable to use the default values. For example, when using the GCG Wisconsin Bestfit package (see below), the default gap penalties for amino acid sequences are set to 0. Nulty is -12 for one gap and -4 for each extension.
[0090] Therefore, calculation of maximum % sequence identity first requires the generation of an optimal alignment, taking into account gap penalties. A suitable computer program for performing such an alignment is the GCG Wisconsin Bestfit package (University of Wisconsin, USA; Devereux et al., 1984, Nucleic Acids Research 12:387). Other examples of software capable of performing sequence comparisons include, but are not limited to, the BLAST package (Ausubel et al., 1999 ibid - see Chapter 18), FASTA (Atschul et al., 1990, J. Mol. Biol., 403-410) and the GENEWORKS comparison tool suite. Both BLAST and FASTA are available for offline and online searching (see Ausubel et al., 1999 ibid, pages 7-58 to 7-60). However, it is preferred to use the GCG Bestfit program.
[0091] Suitably, sequence identity may be determined across the entire sequence. Suitably, sequence identity may be determined across the entire candidate sequence that is compared to the sequences listed herein.
[0092] Although the final sequence identity can be measured in terms of identity, the alignment process itself is typically not based on an all-or-nothing pairwise comparison. Instead, a scaled similarity score matrix is generally used which assigns a score to each pairwise comparison based on chemical similarity or evolutionary distance. An example of such a matrix commonly used is the BLOSUM62 matrix (the default matrix for the BLAST suite of programs). GCG Wisconsin programs generally use either the public default values or a custom symbol comparison table, if supplied (see user manual for further details). Preferably, GCG The public default values for the package are used, or in the case of other software, the default matrix, e.g., BLOSUM62.
[0093] Once the software has produced an optimal alignment, it is possible to calculate percent sequence identity. The software typically does this as part of the sequence comparison and generates a numerical result.
[0094] Refactoring Genomes contain multiple overlapping open reading frames (ORFs), which can be classified as 3',3' (between ORFs in opposite orientation) or 5',3' (between ORFs in the same orientation). One or more sense codons (i.e., those that are substituted) can be found within both classes of overlap in the parent genome.
[0095] If the substitution of one or more sense codons in each ORF within the duplication can be accomplished without changing the encoded protein sequence of either ORF (i.e., by introducing synonymous codons), it may not be necessary to edit (e.g., refactor) the parent genome. However, if the encoded protein sequence is changed by the substitution of one or more sense codons (i.e., one or more synonymous sense codons are not introduced into one or both of the ORFs), it may be necessary to edit (e.g., refactor) the parent genome.
[0096] Therefore, in some embodiments, one or more gene pairs that share an overlapping region containing one or more sense codons in the parent genome are refactored. By "refactored" is meant that the genes are rearranged to prevent changes to the encoded protein sequence. Preferably, the gene pairs are such that a sense codon substitution (e.g., a defined synonymous codon substitution) changes the encoded protein sequence of both or either of the gene pairs. Most preferably, all gene pairs that share an overlapping region containing one or more sense codons in the parent genome are refactored, and the gene pairs are such that a sense codon substitution (e.g., a defined synonymous codon substitution) changes the encoded protein sequence of both or either of the gene pairs.
[0097] For 3',3' duplications (i.e., inverted gene pairs), the synthetic insert may be inserted between the genes. For 3',3' duplications, the synthetic insert may include overlapping regions.
[0098] For 5',3' overlaps (i.e., a pair of genes in the same orientation, including an upstream gene and a downstream gene), a synthetic insert may be inserted between the genes. For 5',3' overlaps, the synthetic insert may include: (i) a stop codon; (ii) about 20 to 200 bp, or 20 to 100 bp, or 20 to 50 bp upstream of the overlapping region; and (iii) the overlapping region. Preferably, the synthetic insert contains (i) a stop codon; (ii) approximately 20 bp upstream of the overlapping region; and (iii) the overlapping region, thereby defining the RBS sequence for the downstream ORF. The sequence and the distance between this RBS and its start codon are conserved.
[0099] In a preferred embodiment, the stop codon is in frame with the original start site for the downstream gene. Preferably, the stop codon is TAA.
[0100] Apart from the specific mutations mentioned above, i.e., mutations aimed at reducing the amount of one or more sense codons (e.g., replacement and / or refactoring of one or more sense codons) and mutations aimed at reducing the amount of amber stop codons, the synthetic prokaryotic genome may comprise up to 1000 or 999, up to 100 or 99, up to 50 or 49, up to 20 or 19, up to 10 or 9 additional (i.e., non-programmed) mutations compared to the parental genome. Preferably, the synthetic prokaryotic genome has 2 x 10 mutations per target codon (i.e., per occurrence of one or more sense codons in the parental genome). -4 It contains one or fewer additional or unprogrammed mutations.
[0101] Polynucleotides The present invention provides polynucleotides comprising one or more genes that lack one or more occurrences of a sense codon. The polynucleotides may comprise 2 or more, 3 or more, 4 or more, 5 or more, 10 or 11, 20 or more, 30 or more, 40 or more, 50 or more, 100 or 101, 200 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or 1001, 1500 or more, or 2000 or 2001 genes that lack one or more occurrences of a sense codon. Preferably, the polynucleotide comprises 100 or 101 or more genes without one or more occurrences of a sense codon, more preferably, the polynucleotide comprises 1000 or 1001 or more genes without one or more occurrences of a sense codon.
[0102] The one or more sense codons may consist of one, two, three, four, five, six, seven, or eight sense codons. Preferably, the one or more sense codons consist of one sense codon or two sense codons, most preferably two sense codons. Thus, in a preferred embodiment, the polynucleotide comprises 100 or 101 or more genes in which one or two sense codons are absent. In another preferred embodiment, the polynucleotide comprises 1000 or 1001 or more genes in which one or two sense codons are absent.
[0103] The one or more sense codons may be selected from TCG, TCA, AGT, AGC, GCG, GCA, GTG, GTA, CTG, CTA, TTG, TTA, ACG, ACA, CCG, CCA, CGG, CGA, CGT, CGC, AGG, AGA, GGG, GGA, GGT, GGC, ATT, and ATC. Alternatively, the one or more sense codons may be selected from TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, The one or more sense codons may be selected from GCC, CTG, CTA, CTT, CTC, TTG, and TTA. Preferably, the one or more sense codons are selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA. More preferably, the one or more sense codons are selected from TCG, TCA, TTG, TTA, GCG, and GCA. Most preferably, the one or more sense codons are TCG and / or TCA.
[0104] One or more sense codons of a gene may be replaced with synonymous sense codons. Preferably, the replacement is a defined replacement, i.e., one sense codon is replaced with a single synonymous sense codon.
[0105] For example, GCG may be replaced with GCT or GCC; GCA may be replaced with GCT or GCC; TCG may be replaced with TCT, TCC, AGT, or AGC; TCA may be replaced with TCT, TCC, AGT, or AGC; AGT may be replaced with TCG, TCA, TCT, or TCC; AGC may be replaced with TCG, TCA, TCT, or TCC; CTG may be replaced with CTT, CTC, TTG, or TTA; CTA may be replaced with CTT, CTC, TTG, or TTA; TTG may be replaced with CTG, CTA, CTT, or CTC; or TTA may be replaced with CTG, CTA, CTT, or CTC. Preferably, the one or more defined sense codon substitutions are selected from GCG to GCT or GCC; GCA to GCT or GCC; TCG to AGT or AGC; TCA to AGT or AGC; AGT to TCA or TCT; AGC to TCG or TCC or TCA; TTG to CTT; and TTA to CTC. More preferably, TCG and / or TCA are substituted with AGC and / or AGT. Most preferably, TCG is substituted with AGC and / or TCA is substituted with AGT.
[0106] In some embodiments, the gene is one for which there is evidence of translation and / or a predicted protein product.
[0107] In a preferred embodiment, the gene is an essential gene. The essential genes are ribF, lspA, ispH, dapB, folA, imp, yabQ, ftsL, ftsI, murE, murF, mraY, murD, ftsW, murG, murC, ftsQ, ftsA, ftsZ, lpxC, secM, secA, can, folK, hemL, yadR, dapD, map, rpsB, tsf, pyrH, frr, dxr, ispU, cdsA, yaeL, yaeT, lpxD, fabZ, lpxA, lpxB, dnaE, accA, tilS, proS, yafF, hemB, secD, secF, ribD, ribE, thiL, dxs, ispA, dnaX, adk, hemH, lpxH, cysS, folD, entD, mrdB, mrdA, nadD, holA, rlpB, leuS, lnt, glnS, fldA, cydA, infA, cydC, ftsK, lolA, serS, rpsA, msbA, lpxK, kdsB, mukF, mukE, mukB, asnS, fabA, mviN, rne, fabD, fabG, acpP, tmk, holB, lolC, lolD, lolE, purB, minE, minD, pth, prsA, ispE, lolB, hemA, prfA, prmC, kdsA, topA, ribA, fabI, tyrS, ribC, ydiL, pheT, pheS, rplT, infC, thrS, nadE, gapA, yeaZ, aspS, argS, pgsA, yefM, metG, folE, yejM, gyrA, nrdA, nrdB, folC, accD, fabB, gltX, ligA, zipA, dapE, dapA, der, hisS, ispG, suhB, tadA, acpS, era, rnc, lepB, rpoE, pssA, yfiO, rplS, trmD, rpsP, ffh, grpE, csrA, ispF One or more selected from the list consisting of ispD, ftsB, eno, pyrG, chpR, lgt, fbaA, pgk, yqgD, metK, yqgF, plsC, ygiT, parE, ribB, cca, ygjD, tdcF, yraL, yhbV, infB, nusA, ftsH, obgE, rpmA, rplU, ispB, murA, yrbB, yrbK, yhbN, rpsI, rplM, degS, mreD, mreC, mreB, accB, accC, yrdC, def, fmt, rplQ, rpoA, rpsD, rpsK, rpsM, secY, rplO, rpmD, rpsE, rplR, rplF, rpsH, rpsN, rplE, rplX, rplN, rpsQ, rpmC, rplP, rpsC, rplV, rpsS, rplB, rplW, rplD, rplC, rpsJ, fusA, rpsG, rpsL, trpS, yrfF, asd, rpoH, ftsX, ftsE, ftsY, yhhQ, bcsB, glyQ, gpsA, rfaK, kdtA, coaD, rpmB, dfp, dut, gmk, spoT, gyrB, dnaN, dnaA, rpmH, rnpA, yidC, tnaB, glmS, glmU, wzyE, hemD, hemC, yigP, ubiB, ubiD, hemG, yihA, ftsN, murI, murB, birA, secE, nusG, rplJ, rplL, rpoB, rpoC, ubiA, plsB, lexA, dnaB, ssb, alsK, groS, psd, orn, yjeE, rpsR, chpS, ppa, valS, yjgP, yjgQ, and dnaC may be selected.
[0108] RibF, lspA, ispH, dapB, folA, imp, yabQ, lpxC, secM, secA can、folK、hemL、yadR、dapD、map、rpsB、tsf、pyrH、frr、dxr、ispU、cdsA、yae L、yaeT、lpxD、fabZ、lpxA、lpxB、dnaE、accA、tilS、proS、yafF、hemB、secD、 secF, ribD, ribE, thiL, dxs, ispA, dnaX, adk, hemH, lpxH, cysS, folD, entD mrdB、mrdA、nadD、holA、rlpB、leuS、lnt、glnS、fldA、cydA、infA、cydC、fts K、lolA、serS、rpsA、msbA、lpxK、kdsB、mukF、mukE、mukB、asnS、fabA、mviN、r ne、fabD、fabG、acpP、tmk、holB、lolC、lolD、lolE、purB、minE、minD、pth、p rsA、ispE、lolB、hemA、prfA、prmC、kdsA、topA、ribA、fabI、tyrS、ribC、ydiL 、pheT、pheS、rplT、infC、thrS、nadE、gapA、yeaZ、aspS、argS、pgsA、yefM、m etG、folE、yejM、gyrA、nrdA、nrdB、folC、accD、fabB、gltX、ligA、zipA、dapE 、dapA、der、hisS、ispG、suhB、tadA、acpS、era、rnc、lepB、rpoE、pssA、yfiO 、rplS、trmD、rpsP、ffh、grpE、csrA、ispF、ispD、ftsB、eno、pyrG、chpR、lgt、 fbaA、pgk、yqgD、metK、yqgF、plsC、ygiT、parE、ribB、cca、ygjD、tdcF、yraL 、yhbV、infB、nusA、ftsH、obgE、rpmA、rplU、ispB、murA、yrbB、yrbK、yhbN、rp sI、rplM、degS、mreD、mreC、mreB、accB、accC、yrdC、def、fmt、rplQ、rpoA、rp sD、rpsK、rpsM、secY、rplO、rpmD、rpsE、rplR、rplF、rpsH、rpsN、rplE、rplX、rplN, rpsQ, rpmC, rplP, rpsC, rplV, rpsS, rplB, rplW, rplD, rplC, rpsJ, fusA, rpsG, rpsL, trpS, yrfF, asd, rpoH, ftsX, ftsE, ftsY, yhhQ, bcsB, glyQ, gpsA, rfaK, kdtA, coaD, rpmB, dfp, dut, gmk, spoT, gyrB, dnaN, dnaA, rpmH, rnpA, yidC, tnaB, glmS, glmU, wzyE, hemD, hemC, yigP, ubiB, ub, iD, hemG, yihA, ftsN, murI, murB, birA, secE, nusG, rplJ, rplL, rpoB, rpoC, ubiA, plsB, lexA, dnaB, ssb, alsK, groS, psd, orn, yjeE, rpsR, chpS, ppa, valS, yjgP, yjgQ, and dnaC.
[0109] It is also a TCG certificate / TCA certificate 1. 1. 2. 1. 1. 1. 1. 1. 1. 1. 2. 1. 1. 1. One of the 1 and 2 of these are ribF, lspA, ispH, dapB, folA, imp, yabQ, lpxC, and secM. secA, can, folK, hemL, yadR, dapD, map, rpsB, tsf, pyrH, frr, dxr, ispU, cd sA, yaeL, yaeT, lpxD, fabZ, lpxA, lpxB, dnaE, accA, tilS, proS, yafF, hemB. secD, secF, ribD, ribE, thiL, dxs, ispA, dnaX, adk, hemH, lpxH, cysS, folD entD, mrdB, mrdA, nadD, holA, rlpB, leuS, lnt, glnS, fldA, cydA, infA, cyd C, ftsK, lolA, serS, rpsA, msbA, lpxK, kdsB, mukF, mukE, mukB, asnS, fabA. mviN, rne, fabD, fabG, acpP, tmk, holB, lolC, lolD, lolE, purB, minE, minD. pth、prsA、ispE、lolB、hemA、prfA、prmC、kdsA、topA、ribA、fabI、tyrS、rib C, ydiL, pheT, pheS, rplT, infC, thrS, nadE, gapA, yeaZ, aspS, argS, pgsA. yefM、metG、folE、yejM、gyrA、nrdA、nrdB、folC、accD、fabB、gltX、ligA、zi pA、dapE、dapA、der、hisS、ispG、suhB、tadA、acpS、era、rnc、lepB、rpoE、pss A, yfiO, rplS, trmD, rpsP, ffh, grpE, csrA, ispF, ispD, ftsB, eno, pyrG, ch pR, lgt, fbaA, pgk, yqgD, metK, yqgF, plsC, ygiT, parE, ribB, cca, ygjD, tdc F, yraL, yhbV, infB, nusA, ftsH, obgE, rpmA, rplU, ispB, murA, yrbB, yrbK. yhbN, rpsI, rplM, degS, mreD, mreC, mreB, accB, accC, yrdC, def, fmt, rplQ.rpoA, rpsD, rpsK, rpsM, secY, rplO, rpmD, rpsE, rplR, rplF, rpsH, rpsN, rplE, rplX, rplN, rpsQ, rpmC, rplP, rpsC, rplV, rpsS, rplB, rplW, rp lD, rplC, rpsJ, fusA, rpsG, rpsL, trpS, yrfF, asd, rpoH, ftsX, ftsE, ftsY, yhhQ, bcsB, glyQ, gpsA, rfaK, kdtA, coaD, rpmB, dfp, dut, gmk, spo Selected from the list consisting of T, gyrB, dnaN, dnaA, rpmH, rnpA, yidC, tnaB, glmS, glmU, wzyE, hemD, hemC, yigP, ubiB, ubiD, hemG, yihA, ftsN, murI, murB, birA, secE, nusG, rplJ, rplL, rpoB, rpoC, ubiA, plsB, lexA, dnaB, ssb, alsK, groS, psd, orn, yjeE, rpsR, chpS, ppa, valS, yjgP, yjgQ, and dnaC. Preferably, the polynucleotide comprises two or more, three or more, four or more, five or more, five or more, ten or eleven, twenty or more, thirty or more, forty or more, fifty or more, one hundred or more, or two hundred or more essential genes that lack the TCG and / or TCA codon.
[0110] In some embodiments, the polynucleotide is a sequence similar to or similar to SEQ ID NO: 1 or SEQ ID NO: 2. comprises a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, 99.5%, 99.8%, or 99.9%, or 100% identical to a fragment of either SEQ ID NO:1 or SEQ ID NO:2, preferably the fragment is at least 10 kb, 20 kb, 50 kb, 100 kb, or 500 kb in length.
[0111] Preferably, the polynucleotide is viable. That is, the polynucleotide can be incorporated into a genome such that the genome is a viable genome. Preferably, the polynucleotide can replace the corresponding region of the parent genome and retain the viability of the genome. As used herein, "viable genome" refers to a genome that contains sufficient nucleic acid sequences to cause and / or maintain cellular viability, e.g., a genome that encodes molecules required for replication, transcription, translation, energy production, transport, production of membrane and cytoplasmic components, and cell division. Thus, the present invention also provides a viable synthetic prokaryotic genome (e.g., a viable synthetic E. coli genome) comprising a polynucleotide of the invention.
[0112] The present invention provides polynucleotides which are at least 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.95% or 100% identical to SEQ ID NO:1 or SEQ ID NO:2, or to a fragment of any of SEQ ID NO:1 or SEQ ID NO:2, preferably the fragment is at least 10kb, 20kb, 50kb, 100kb or 500kb in length.
[0113] Host cells and uses thereof host cell The present invention also provides a host cell comprising the synthetic prokaryotic genome or polynucleotide of the present invention. The host cell may be an isolated host cell.
[0114] The host cell of the present invention is a prokaryotic cell. More preferably, the host cell is a bacterial cell. Preferably, the bacterial host cell is suitable for heterologous protein production, particularly for the production of polypeptides containing one or more non-proteinogenic amino acids (e.g., those described in Ferrer-Miralles, N. and Villaverde, A., 2013. Microbial Cell Factories, 12:113). Suitable bacterial host cells include Escherichia (e.g., Escherichia coli), Caulobacteria (e.g., Caulobacter crescentus), photosynthetic bacteria (e.g., Rhodobacter sphaeroides), cold-adapted bacteria (e.g., Pseudoalteromonas haloplanktis, Shewanella sp. strain Ac10), Pseudomonas (e.g., Pseudomonas fluorescens, Pseudomonas putida, Pseudomonas aeruginosa), halophilic bacteria (e.g., Halomonas elongata, Chromohalobacter salexigens), Streptomycetes (e.g., Streptomyces lividans, Streptomyces griseus), Nocardia, and the like. Examples of suitable bacterial host cells include: Aerobic bacteria (e.g., Nocardia lactamdurans), Mycobacteria (e.g., Mycobacterium smegmatis), Coryneform bacteria (e.g., Corynebacterium glutamicum, Corynebacterium ammoniagenes, Brevibacterium lactofermentum), Bacillus (e.g., Bacillus subtilis, Bacillus brevis, Bacillus megaterium, Bacillus licheniformis, Bacillus amyloliquefaciens), and Lactic acid bacteria (e.g., Lactococcus lactis, Lactobacillus plantarum, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri). In some embodiments, the bacterial host cell is a Gram-negative bacterium.
[0115] Preferably, the host cell is Escherichia coli, Salmonella enterica, or Shigella dysenteriae. More preferably, the host cell is Escherichia coli. Suitable E. coli host cells include MDS42, K-12, MG1655, BL21, BL21(DE3), AD494, Origami, HMS174, BLR(DE3), HMS174(DE3), Tuner(DE3), Origami2(DE3), Rosetta2(DE3), and Lemo21(DE3). , NiCo21(DE3), T7 Express, SHuffle Express, C41(DE3), C43(DE3), and m15 pREP4 or their derivatives (Rosano, GL and Ceccarelli, EA, 2014. Frontiers in microbiology, 5, (p. 172). Most preferably, the host cell is MDS42, MG1655, or BL21 or a derivative thereof. MG1655 is considered the wild-type strain of E. coli. The GenBank ID for the genome sequence of this strain is U00096. BL21 is widely available commercially. For example, it can be purchased from New England BioLabs under catalog number C2530H. This can be done.
[0116] The host cell may preferably be the same as the one from which (or derived from) the synthetic prokaryotic genome or polynucleotide was present. For example, if the synthetic prokaryotic genome is a synthetic E. coli genome, the host cell is preferably E. coli. If the parent genome of a cell has been modified to produce a synthetic prokaryotic genome of the invention, the host cell is preferably the same cell, i.e., the host cell comprising the synthetic prokaryotic genome is preferably the same as the host cell of the parent genome (parent host cell).
[0117] The host cell may be viable, ie, capable of growing and replicating.
[0118] When a cell's genome is modified to produce a synthetic prokaryotic genome of the present invention, the synthetic prokaryotic genome preferably does not substantially reduce the growth rate when present in a parent host cell. Thus, preferably, a host cell comprising the synthetic prokaryotic genome does not substantially reduce the growth rate compared to a host cell comprising the parent genome. In some embodiments, a host cell comprising the synthetic prokaryotic genome has a doubling time that is 4-fold, 3-fold, 2-fold, or less than about 1.6-fold slower than a host cell comprising the parent genome. Doubling time can be determined by any method known to those of skill in the art. In some embodiments, doubling time is determined in LB medium at 37°C, 25°C, or 42°C.
[0119] When a cell's genome is modified to produce a synthetic prokaryotic genome of the present invention, the synthetic prokaryotic genome preferably does not cause any substantial phenotypic change when present in a parent host cell. Thus, preferably, host cells comprising the synthetic prokaryotic genome do not have any substantial phenotypic change compared to host cells comprising the parent genome. In some embodiments, host cells comprising the synthetic prokaryotic genome have an average cell length that is 100%, 50%, or less than about 20% longer than host cells comprising the parent genome. For example, the cell length may be about 1.5 to 3 microns. Cell length can be determined by any method known to those of skill in the art. In some embodiments, host cells comprising the synthetic prokaryotic genome have a proteome that is not substantially different from the proteome of a host cell comprising the parent genome. The proteome can be determined by any method known to those of skill in the art.
[0120] Reassignment to alternative canonical amino acids In some embodiments, one or more sense codons (i.e., those removed from the parent genome) are reassigned to encode alternative canonical amino acids. For example, if TCG and TCA are removed, one or both can be reassigned to encode a canonical amino acid other than serine (e.g., alanine).
[0121] For example, the synthetic prokaryotic genome of the present invention substantially or completely lacks one or more sense codons. Therefore, one or more tRNAs or terminators may be deleted from the synthetic genome. For example, tRNAs that decode the one or more sense codons that have been replaced (or deleted) may be deleted from the synthetic prokaryotic genome. tRNAs that decode the one or more sense codons that have been replaced (or deleted) may be deleted, such that the tRNAs decode only the one or more sense codons that have been replaced (or deleted). Alternatively, if the tRNA decodes one or more sense codons that have been replaced (or deleted) and one or more sense codons that have not been replaced (or deleted), and if the tRNA is non-essential for one or more sense codons that have not been replaced (or deleted) (i.e., one or more of the sense codons decoded by the tRNA are decoded by one or more alternative tRNAs), the synthetic prokaryotic genome will remain viable. For example, if the synthetic prokaryotic genome lacks a TCA sense codon, the tRNA Ser UGA serT encoding tRNA may be deleted, and / or if the synthetic prokaryotic genome lacks a TCG sense codon, Ser CGA The serU encoding serU may be deleted. Deletion of one or more tRNAs can be used, for example, in combination with reassigned endogenous tRNAs or orthogonal aminoacyl-tRNA synthetase / tRNA pairs to reassign one or more sense codons to alternative amino acids.
[0122] For example, if TCG and TCA have been removed from a synthetic prokaryotic genome, tRNA Ser UGA serT, and tRNA encoding Ser CGA serU, which encodes either tRNA , may be deleted from the synthetic prokaryotic genome. CGA (e.g., tRNA Ala CGAcan be reassigned to an orthogonal aminoacyl-tRNA synthetase / tRNA CGA The pairs can be introduced into a host cell (e.g., by heterologous nucleic acid or by incorporation into a synthetic prokaryotic genome) to reassign TCG to an alternative canonical amino acid. Thus, in some embodiments, a host cell of the invention contains one or more heterologous nucleotides encoding one or more reassigned tRNAs and / or an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. In some embodiments, the host cell of the invention further comprises a plasmid encoding an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. Alternatively, the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair can be introduced into the host cell by incorporation into a synthetic prokaryotic genome. Thus, in some embodiments, the synthetic prokaryotic genome encodes the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair, preferably with the gene encoding the native tRNA deleted from the parent prokaryotic genome. In preferred embodiments, the host cell of the invention further comprises one or more reallocated tRNAs. Methods for reallocating tRNAs are well known to those skilled in the art.
[0123] Reassignment to encode alternative canonical amino acids can increase biosafety. Thus, in some embodiments, the host cells of the invention have increased biosafety. Accordingly, the invention provides host cells with improved biosafety.
[0124] For example, reassignment to encode alternative canonical amino acids can render host cells containing the synthetic prokaryotic genome resistant to bacteriophage infection. Because one or more bacteriophage genes typically contain one or more sense codons, when the one or more bacteriophage genes are translated, the alternative canonical amino acids can be incorporated into the corresponding bacteriophage proteins. Incorporation of the alternative canonical amino acids can destabilize, destroy, or reduce the activity of the proteins, thereby reducing bacteriophage infectivity and rendering the host cell resistant to bacteriophage infection.
[0125] Thus, in some embodiments, the host cells of the invention are resistant to phage infection. For example, when a cell's genome is modified to produce a synthetic prokaryotic genome of the invention, the synthetic prokaryotic genome may increase resistance to phage infection when present in a parent host cell. Thus, in some embodiments, a host cell comprising a synthetic prokaryotic genome The cells have increased resistance to the phage compared to host cells containing the parental genome.
[0126] Thus, the present invention provides phage-resistant host cells and host cells with increased phage resistance.
[0127] Reassignment to encode alternative canonical amino acids can also allow genetic material, e.g., antibiotic resistance genes, to be designed so that they are functional in the engineered strain but not in the wild-type strain. For example, genetic material can be incorporated into host cells of the invention (e.g., by heterologous nucleic acid or by incorporation into a synthetic prokaryotic genome) such that the host cells grow under certain conditions (e.g., in the presence of an antibiotic), but other host cells (e.g., parent host cells) do not. Thus, in some embodiments, the host cells of the invention can render compositions comprising the host cells more resistant to contamination by other host cells (e.g., other prokaryotes).
[0128] Reassignment to non-proteinogenic amino acids In some embodiments, one or more sense codons (ie, those removed from the parent genome) are reassigned to encode a non-canonical amino acid (a non-proteinogenic amino acid).
[0129] Therefore, the present invention provides the use of a host cell according to the present invention for producing a polypeptide comprising one or more non-proteinogenic amino acids, preferably two or three or more non-proteinogenic amino acids, most preferably three or four or more non-proteinogenic amino acids.
[0130] The present invention also provides polypeptides obtained or obtainable by using the host cells according to the present invention. In some embodiments, the polypeptides comprise one or more non-proteinogenic amino acids, preferably two or more non-proteinogenic amino acids, and most preferably three or more non-proteinogenic amino acids. Thus, the present invention also provides polypeptides comprising two or more non-proteinogenic amino acids and polypeptides comprising three or more non-proteinogenic amino acids.
[0131] As used herein, a "non-proteinogenic amino acid" (also known as a "non-encoded amino acid" or "non-canonical amino acid") is an amino acid that is not naturally encoded or found in the genetic code. Despite the use of only 22 amino acids (the proteinogenic amino acids, i.e., the 20 in the standard genetic code and two additional ones that can be incorporated by specialized translation machinery) by the translation machinery to assemble proteins, over 140 amino acids are known to occur naturally in proteins, and thousands more may occur naturally or can be synthesized in the laboratory. Thus, non-proteinogenic amino acids may include any amino acid except L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan and L-tyrosine, and optionally L-pyrrolysine and L-selenocysteine.
[0132] In some embodiments, the non-proteinogenic amino acid is an unnatural amino acid (UAA).
[0133] The non-proteinogenic amino acid or UAA is not particularly limited. Suitable non-proteinogenic amino acids and UAAs are well known to those skilled in the art, for example, those disclosed in Neumann, H., 2012. FEBS letters, 586(15), pp. 2057-2064; and Liu, CC and Schultz, PG, 2010. Annual review of biochemistry, 79, pp. 413-444. In some embodiments, the non-proteinogenic amino acid or UAA is The metabolic amino acids and / or UAAs include p-acetylphenylalanine, m-acetylphenylalanine, O-allyltyrosine, phenylselenocysteine, p-propargyloxyphenylalanine, p-azidophenylalanine, p-boronophenylalanine, O-methyltyrosine, p-aminophenylalanine, p-cyanophenylalanine, m-cyanophenylalanine, p-fluorophenylalanine, p-iodophenylalanine, p-bromophenylalanine, p-nitrophenylalanine, L-DOPA, 3-aminotyrosine, 3-iodotyrosine, p-isopropylphenylalanine, 3-(2-naphthyl)alanine, biphenylalanine, homoglutamic acid, and the like. The amino acid sequence may be selected from one or more of 2-nitrobenzyl lysine, D-tyrosine, p-hydroxyphenyllactic acid, 2-aminocaprylic acid, bipyridylalanine, HQ-alanine, p-benzoylphenylalanine, o-nitrobenzylcysteine, o-nitrobenzylserine, 4,5-dimethoxy-2-nitrobenzylserine, o-nitrobenzyllysine, o-nitrobenzyltyrosine, 2-nitrophenylalanine, dansylalanine, p-carboxymethylphenylalanine, 3-nitrotyrosine, sulfotyrosine, acetyllysine, methylhistidine, 2-aminononanoic acid, 2-aminodecanoic acid, pyrrolidine, Cbz-lysine, Boc-lysine, and allyloxycarbonyllysine.
[0134] Prokaryotes, such as E. coli, typically cannot incorporate most eukaryotic post-translational modifications, such as ubiquitination, glycosylation, and phosphorylation, and they typically cannot perform other eukaryotic maturation processes, as well as proteolytic protein maturation. Furthermore, correct disulfide bond formation and lipopolysaccharide contamination can be challenging (see Ovaa, H., 2014. Frontiers in chemistry, 2, p.15). However, the use of anti- Therapeutic proteins, such as enzymes, cytokines, and the like, typically retain post-translational modifications and disulfide bonds and often require proteolytic maturation to achieve their correctly folded state. Therefore, the majority of therapeutic proteins are produced in eukaryotic and mammalian cell systems. However, expression in prokaryotic host cells, such as E. coli, is generally inexpensive, amenable to genetic modification, versatile for developing mutation libraries, and suitable for industrial-scale fermentation (Ovaa, H., 2014. Frontiers in chemistry, 2, p. 15 ).
[0135] Thus, in some embodiments, the polypeptide is a therapeutic polypeptide, preferably having a mammalian protein modification introduced by one or more non-proteinogenic amino acids. For example, amber codon suppression has previously been used to incorporate one or more non-proteinogenic amino acids (i.e., mammalian protein modifications) into therapeutic polypeptides. The present invention allows for the incorporation of two or more non-proteinogenic amino acids. Thus, the present invention provides therapeutic polypeptides comprising two or more non-proteinogenic amino acids.
[0136] Because the synthetic prokaryotic genomes of the present invention substantially or completely lack one or more sense codons, one or more tRNAs or terminators may be deleted from the synthetic genome. For example, tRNAs that decode only the one or more sense codons that have been replaced (or deleted) may be deleted from the synthetic prokaryotic genome. For example, if the synthetic prokaryotic genome lacks a TCA sense codon, then tRNAs Ser UGA setT encoding tRNA may be deleted, and / or if the synthetic prokaryotic genome lacks a TCG sense codon, Ser CGA The serU gene encoding the nucleotide sequence may then be deleted. The synthetic prokaryotic genome may then be used (in conjunction with an orthogonal aminoacyl-tRNA synthetase-tRNA pair) to direct the incorporation of non-proteinogenic amino acids into proteins.
[0137] Genetic code expansion generates non-proteinaceous amino acids in response to an unassigned codon (e.g., an amber stop codon, UAG) introduced at a desired site in a desired gene. Orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pairs are used to direct the incorporation of amino acids into proteins. The orthogonal synthetase does not recognize endogenous tRNAs and specifically aminoacylates the orthogonal cognate tRNA (which is not an effective substrate for endogenous synthetases) with non-proteinogenic amino acids provided to (or synthesized by) the cell (Chin, JW, 2017. Nature, 550(7674), 53-60). Those skilled in the art can identify and / or generate suitable orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pairs (e.g., Elliott, TS et al., 2014. Nat Biotechnol 32, 465-472; Elliott, TS, et al., 2016. Cell Chem Biol 23, 805-815; and Krogager, TP et al., 2018. Nat Biotechnol 36, 156-159). Thus, in some embodiments, the host cells of the invention further comprise one or more heterologous nucleotides (e.g., a plasmid) encoding an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. In preferred embodiments, the host cells of the invention further comprise a plasmid encoding an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. Alternatively, the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair can be introduced into the host cell by incorporation into a synthetic prokaryotic genome. Thus, in some embodiments, the synthetic prokaryotic genome encodes the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair, preferably in which the gene encoding the native tRNA has been deleted from the parent prokaryotic genome.
[0138] Thus, in some embodiments, the host cell of the invention further comprises one or more heterologous nucleotides (e.g., a plasmid) comprising one or more genes comprising said sense codons. In preferred embodiments, the host cell further comprises a plasmid comprising a gene comprising said sense codons. The one or more sense codons may be present at desired sites in the gene, preferably, the desired sites allowing for the incorporation of one or more non-proteinogenic amino acids (i.e., mammalian protein modifications) into a polypeptide, preferably a therapeutic polypeptide.
[0139] In other embodiments, the sense codons can be present in one or more genes in the synthetic prokaryotic genome (e.g., heterologous nucleotides can be incorporated into the synthetic prokaryotic genome). The one or more sense codons can be present at desired sites in the gene, preferably sites that allow for the incorporation of one or more non-proteinogenic amino acids (i.e., mammalian protein modifications) into a polypeptide, preferably a therapeutic polypeptide.
[0140] For example, if TCG and TCA have been removed from a synthetic prokaryotic genome, tRNA Ser UGA serT, and tRNA encoding Ser CGA serU, which encodes an orthogonal aminoacyl-tRNA synthetase / tRNA, may be deleted from the synthetic prokaryotic genome. CGA The pair may be used in combination with a (heterologous) gene containing a TCG codon, such that the pair encodes a polypeptide containing one or more non-proteinogenic amino acids. Thus, the host cell of the present invention may, for example, be a host cell that contains (i) an orthogonal aminoacyl-tRNA synthetase / tRNA CGA and (ii) a plasmid containing a gene containing one or more TCG codons. Similarly, if AGT and AGC are removed, tRNA Ser GCUserV, which encodes an orthogonal aminoacyl-tRNA synthetase / tRNA, may be deleted from the synthetic prokaryotic genome. ACU Paired and / or orthogonal aminoacyl-tRNA synthetases / tRNAs GCU Similarly, if CTG and CTA are removed, tRNA Leu CAG leuP, Q, T, V, and tRNA encoding Leu UAG The leuW encoding the orthogonal aminoacyl-tRNA synthetase / tRNA may be deleted from the synthetic prokaryotic genome. CAG Similarly, if TTG and TTA are removed, tRNA Leu CAA leuX, which encodes Leu UAA The leuZ encoding the orthogonal aminoacyl-tRNA synthetase / tRNA may be deleted from the synthetic prokaryotic genome. CAA Paired and / or orthogonal axes Aminoacyl-tRNA synthetase / tRNA UAA Similarly, if GCG and GCA are removed, tRNA Ala UGC The alaT, U, and V encoding the orthogonal aminoacyl-tRNA synthetase / tRNA may be deleted from the synthetic prokaryotic genome. CGC Pairs may be used.
[0141] In some embodiments, the synthetic prokaryotic genome lacks a gene encoding a release factor (e.g., RF1), and / or the host cell lacks a release factor (e.g., RF1) to increase the efficiency of incorporation of non-proteinogenic amino acids.
[0142] Methods for producing synthetic genomes In one aspect, the present invention provides a method for producing a synthetic genome, comprising: (a) providing a parent genome; (b) performing one or more rounds of recombination-mediated genetic modification on the parent genome to produce two or more different partial synthetic genomes; (c) performing one or more rounds of directed conjugation with two or more different partial synthetic genomes to produce a synthetic genome; The present invention provides a method comprising:
[0143] Genetic modification via recombination Preferably, one or more rounds of recombination-mediated genetic modification are used to edit 10-1000 kb, 50-1000 kb, 100-1000 kb, or 100-500 kb of the parent genome to provide two or more different partial synthetic genomes. Thus, in preferred embodiments, each round of recombination-mediated genetic modification inserts or replaces 10 kb or more, 50 kb or more, 100 kb or more, or about 100 kb of DNA from the parent genome.
[0144] As used herein, the term "recombination-mediated genetic modification" (also known as "recombineering") refers to a method for genetic modification (i.e., genome editing) based on the homologous recombination system. Typically, recombineering is based on homologous recombination in E. coli mediated by bacteriophage proteins RecE / RecT from Rac prophage or Redαβδ from bacteriophage lambda. Any suitable method of recombination-mediated genetic modification may be used. Methods for recombination-mediated genetic modification are well known to those skilled in the art.
[0145] In "classical recombination" (exemplified by lambda Red-mediated recombination in E. coli), short regions of synthetic DNA can be inserted into a genome or used to replace genomic DNA in a two-step process: (i) transformation of cells with linear double-stranded DNA (dsDNA) containing a stretch of synthetic DNA, linked to a positive selectable marker, and flanked by homology regions (HR) at each end of the target region in the genome, and (ii) recombination mediated by the homology regions, followed by selection for genomic integration by the positive selectable marker. This approach can be used to insert or replace 2-3 kb of genomic DNA. Therefore, when classical recombination is used, many rounds of recombination-mediated genetic modification are required to edit 100-500 kb of the parent genome.
[0146] Therefore, in a preferred embodiment, the one or more rounds of recombination-mediated genetic modification include one or more rounds of replicon excision for enhanced genome modification by programmed recombination (REXER).
[0147] REXER is described in WO 2018 / 020248, which is incorporated herein by reference. Each round of REXER involves the recombination of parental genomic DNA. It can be used to insert or replace about 50 kb to 250 kb, or about 100 kb.
[0148] Therefore, one or more rounds of genetic modification via recombination may i) providing a host cell (e.g., E. coli), the host cell comprising an episomal replicon (e.g., a plasmid or a bacterial artificial chromosome) and a target nucleic acid (e.g., a genome), the episomal replicon comprising a donor nucleic acid sequence (i.e., a synthetic region), the donor nucleic acid sequence comprising, in order, 5'-homologous recombination sequence 1-desired sequence-homologous recombination sequences 2-3', the desired sequence comprising a positive selectable marker, and the target nucleic acid comprising, in order, 5'-homologous recombination sequence 1-negative selectable marker-homologous recombination sequences 2-3'; ii) providing a helper protein (e.g., lambda red protein) capable of supporting nucleic acid recombination in said host cell; iii) helper proteins capable of supporting nucleic acid excision in said host cell; and and / or providing RNA (e.g., CRISPR / Cas9 protein / RNA); iv) inducing excision of the donor nucleic acid sequence; v) incubating to allow recombination between the excised donor nucleic acid and the target nucleic acid; and vi) selecting recombinants that have integrated the donor nucleic acid into the target nucleic acid may include:
[0149] Suitably, the step of selecting recombinants that have integrated the donor nucleic acid into the target nucleic acid comprises selecting for the gain of a positive selectable marker on the donor nucleic acid and the loss of a negative selectable marker on the target nucleic acid. Suitably, the gain of the positive selectable marker on the donor nucleic acid and the loss of the negative selectable marker on the target nucleic acid are carried out simultaneously. Suitably, the desired sequence comprises both a positive selectable marker and a negative selectable marker. Suitably, the negative selectable marker is sacB (sucrose sensitivity), rpsL (S12 ribosomal protein - streptomycin sensitivity), or phe ST251A_A294G (4-chlorophenylalanine sensitive). Suitably the positive selectable marker is selected from the group consisting of Cm R (chloramphenicol resistant), Kan R (kanamycin resistance), Hyg R (hygromycin resistant), gentamicin R (gentamicin resistant), or tetracycline R (tetracycline resistance). Suitably, the step of selecting for recombinants comprises sequential selection for said positive and negative markers, or sequential selection for said negative and positive markers. Suitably, the step of selecting for recombinants comprises simultaneous selection for said positive and negative markers.
[0150] Suitably, the method above further comprises inducing at least one double-stranded break in a target nucleic acid sequence, wherein the double-stranded break is between the homologous recombination sequence 1 and the homologous recombination sequence 2. Suitably, at least two double-stranded breaks are induced in the target nucleic acid sequence, each double-stranded break being between the homologous recombination sequence 1 and the homologous recombination sequence 2.
[0151] Suitably, the excised donor nucleic acid starts at the homologous recombination sequence 1 and ends at the homologous recombination sequence 2.
[0152] Suitably, the episomal replicon comprises a negative selectable marker independent of the donor nucleic acid sequence. Suitably, the method comprises the further step of selecting for loss of the episomal replicon by selecting for loss of the negative selectable marker independent of the donor nucleic acid sequence. Suitably, the episomal replicon comprises, in order: excision cleavage site 1 - donor nucleic acid sequence - excision cleavage site 2. Suitably, the target nucleic acid is capable of functioning in the host cell. Suitably, said episomal replicon is a plasmid nucleic acid. Suitably, said episomal replicon is a bacterial artificial chromosome (BAC). Suitably, said target nucleic acid is the host cell genome. do.
[0153] Episomal replicons (e.g., BACs) can be used to express homologous sequences in S. cerevisiae, as described, for example, in Kouprina, N., et al., 2004. Methods Mol Biol 255, 69-89. Recombination sequences can be assembled by recombinant DNA technology. Assembly can combine 7-14 stretches of synthetic DNA, each 6-13 kb long; a selection construct (containing a negative and / or positive selection marker); and a BAC shuttle vector backbone. The stretches of synthetic DNA can correspond entirely to the donor nucleic acid sequences (i.e., synthetic regions) in the episomal replicon, each containing 80-200 bp of overlapping DNA sequence, with the overlapping regions not including any of the targets being rewritten. The stretches can be provided in pSC101 or pST vectors flanked by appropriate restriction sites (e.g., BsaI, AvrII, SpeI, or XbaI). Therefore, during assembly, the synthetic DNA stretches can be excised by digestion with the corresponding restriction enzymes. Assembly of the episomal replicon can be verified by sequencing.
[0154] Suitably, the two homologous regions may be 30 to 100 bp, or 40 to 50 bp, or about 50 bp in length.
[0155] CRISPR / Cas9 machinery may be used for excision. In some embodiments, the CRISPR / Cas9 machinery comprises Cas9, a tracrRNA, and two spacer RNAs, where the spacer RNAs target two homologous regions for excision. In preferred embodiments, the spacer RNAs are linear double-stranded spacers. In other embodiments, the CRISPR / Cas9 machinery comprises Cas9 and two sgRNAs, where the sgRNAs target two homologous regions for excision.
[0156] Lambda Red recombination machinery may be used for recombination. The lambda Red recombination machinery may include lambda alpha / beta / gamma.
[0157] The method may include performing one or more rounds of REXER, i.e., the above steps with a first donor nucleic acid sequence, selecting additional donor sequences contiguous with the first donor nucleic acid sequence, and repeating the steps with the additional donor nucleic acid sequences until a partial synthetic genome is assembled. This is known as genome-level exchange synthesis (GENESIS), as described in Wang, K. et al., 2016. Nature 539, 59-64, and is shown schematically in Figure 4.
[0158] In a preferred embodiment, the donor sequence corresponds to a region of a synthetic genome according to the invention and / or a polynucleotide according to the invention.
[0159] Thus, the donor sequence (i.e., the synthetic region) may contain 20 or 19 or fewer occurrences of one or more sense codons, and / or the donor sequence may contain 10 or 11 or more, 20 or 21 or more, or 100 or 101 or more genes without one or more occurrences of a sense codon.
[0160] Donor sequences (i.e., synthetic regions) may be selected such that they have no more than 50 or 49, no more than 20 or 19, no more than 10 or 9, no more than 5 or 4, or no more than 0 occurrences of each of one or more sense codons, and / or have less than 10%, 5%, 2%, 1%, 0.5%, 0.1% or less occurrences of one or more sense codons compared to the corresponding region in the parent genome. or two or more occurrences of each of the sense codons, and / or one or more missing occurrences of the sense codons, or 10 or more, 20 or more, or 100 or more genes (i.e., non-synthetic regions).
[0161] The donor sequence (i.e., synthetic region) may also be refactored relative to the sequence of the parent genome (i.e., non-synthetic region). For 3',3' overlaps (i.e., reverse-oriented gene pairs), the synthetic insert may be inserted between genes. For 3',3' overlaps, the synthetic insert may include the overlapping region. For 5',3' overlaps (i.e., same-oriented gene pairs), the synthetic insert may be inserted between genes. For 5',3' overlaps, the synthetic insert may include (i) a stop codon; (ii) approximately 20-200 bp, or 20-100 bp, or 20-50 bp upstream of the overlapping region; and (iii) the overlapping region. Preferably, the synthetic insert comprises (i) a stop codon; (ii) approximately 20 bp upstream of the overlapping region; and (iii) the overlapping region. In a preferred embodiment, the stop codon is located downstream of the gene Preferably, the stop codon is TAA.
[0162] Preferably, the donor sequences (i.e., synthesis regions) are 50 to 10,000 kb, 100 to 5,000 kb, 100 to 2,000 kb, 100 to 1,000 kb, or 100 to 500 kb in total size. Preferably, each donor sequence is 50 to 300 kb, 100 to 200 kb, or about 100 kb in size.
[0163] Thus, the donor sequences may each be approximately 100 kb in size and may be identical to the corresponding sequences in the parent genome, except that they do not contain one or more occurrences of sense codons, and all gene pairs that share overlapping regions containing one or more sense codons in the parent genome are refactored, where the sense codon substitution alters the encoded protein sequence of both or one of the gene pairs.
[0164] In preferred embodiments, the viability of the genome is tested after each round of recombination-mediated genetic modification. In some embodiments, the sequence of the genome is verified after each round of recombination-mediated genetic modification.
[0165] Partially Synthetic Genomes The present invention provides two or more different partial synthetic genomes.
[0166] As used herein, a "partially synthetic genome" is a genome in which one or more contiguous regions of a parent genome have been edited (i.e., the partial synthetic genome comprises one or more synthetic regions), where the one or more contiguous (synthetic) regions do not occupy the entirety of the parent genome. Preferably, a partially synthetic genome of the invention has one contiguous (synthetic) region. In contrast, a "synthetic genome" may include genome edits that occupy substantially all of the parent genome.
[0167] The partial synthetic genome of the present invention may be a prokaryotic genome. Preferably, the partial synthetic genome of the present invention is a bacterial genome. More preferably, the partial synthetic genome of the present invention is an Escherichia coli, Salmonella enterica, or Shigella dysenteriae genome. Most preferably, the partial synthetic genome of the present invention is an E. coli genome. In some embodiments, the partial synthetic genome is a small or minimal partial synthetic genome. In a preferred embodiment, the partial synthetic genome is a viable genome.
[0168] In some embodiments, the partial synthetic genomes of the invention are between 100 kb and 20 Mb, or between 130 kb and 15 Mb, or between 200 kb and 15 Mb, or between 300 kb and 15 Mb, or between 500 kb and 15 Mb, or between 1 Mb and 15 Mb, or between 1 Mb and 10 Mb, or between 1 Mb and 8 Mb, or between 1 Mb and 6 Mb, or between 2 Mb and 6 Mb, or between 2 Mb and 5 Mb, or between 3 Mb and 5 Mb, or about 4 Mb in size.
[0169] A partial synthetic genome may comprise a synthetic region having no more than 50 or 49, no more than 20 or 19, no more than 10 or 9, no more than 5 or 4, or no more than 0 occurrences of each of one or more sense codons, or a partial synthetic genome may comprise a synthetic region having less than 10%, 5%, 2%, 1%, 0.5%, 0.1% occurrences of each of one or more sense codons compared to the corresponding region in the parent genome.
[0170] Preferably, the synthetic region is 50 to 10,000 kb, 100 to 5,000 kb, or 100 to 500 kb in size.
[0171] Thus, a partial synthetic genome may comprise one or more contiguous regions of 100-5000 kb having 10 or 9 or less, 5 or 4 or less, or zero occurrences of each of one or more sense codons, and / or a partial synthetic genome may comprise one or more contiguous regions of 100-5000 kb having less than 10%, 5%, 2%, 1%, 0.5%, 0.1% or less occurrences of each of one or more sense codons compared to the corresponding region in the parent genome, and / or a partial synthetic genome may comprise one or more contiguous regions of 100-5000 kb having 10 or 11 or more, 20 or 21 or more, or 100 or 101 or more genes with no occurrences of one or more sense codons.
[0172] The remainder of the partial synthetic genome (i.e., non-synthetic regions) may have unaltered sense codons. Thus, a partial synthetic genome may include one or more non-synthetic regions with 100% or 99% occurrence of each sense codon compared to the corresponding region in the parent genome, and / or a partial synthetic genome may include one or more non-synthetic regions with 100 or 101 or more genes with occurrence of each sense codon. The non-synthetic regions may be 500 kb to 20 Mb, or 500 kb to 10 Mb, or 500 kb to 5 Mb, or about 3.5 Mb in size.
[0173] For example, a partially synthetic genome may include one contiguous region of 100 to 5,000 kb having 10 or more, 20 or more, or 100 or more genes that lack one or more occurrences of a sense codon (i.e., the synthetic region), and one contiguous region of 500 kb to 10,000 kb having 100 or more genes that have occurrences of each sense codon (i.e., the non-synthetic region).
[0174] The two or more different partial synthetic genomes may be derived from the same parent genome, i.e., may contain substantially the same sequences, for example, the two or more different partial synthetic genomes may share 90%, 95%, 99%, or 99.5% sequence identity.
[0175] The two or more different partial synthetic genomes may contain one or more synthetic regions such that the synthetic regions collectively account for 90% or more, 95% or more, 99% or more, or 100% of the parent genome. Preferably, each of the two or more different partial synthetic genomes contains one or more synthetic regions, and the synthetic regions do not substantially overlap (e.g., the overlap between the synthetic regions is 10 kb or less, preferably about 3-4 kb). Thus, each of the two or more different partial synthetic genomes contains one or more synthetic regions, such that the synthetic regions collectively account for 90% or more, 95% or more, 99% or more, or 100% of the parent genome. Preferably, each of the two or more different partial synthetic genomes contains one or more synthetic regions, and the synthetic regions do not substantially overlap (e.g., the overlap between the synthetic regions is 10 kb or less, preferably about 3-4 kb). Thus, each of the two or more different partial synthetic genomes contains one or more synthetic regions. The invention may also include unique or substantially unique synthetic regions of:
[0176] Thus, in a preferred embodiment, each of the two or more different partial synthetic genomes comprises one contiguous synthetic region of 100 to 5000 kb having 10 or 11 or more, 20 or 21 or more, or 100 or 101 or more genes that are absent from one or more occurrences of a sense codon, and one non-synthetic contiguous region of 500 kb to 10000 kb having 100 or 101 or more genes that have occurrences of each sense codon, wherein the synthetic regions collectively occupy substantially all of the parent genome, and the synthetic regions are substantially non-overlapping.
[0177] Two or more different partial synthetic genomes may be suitable for induced conjugation. Thus, in a preferred embodiment, the two or more different partial synthetic genomes comprise at least one partial synthetic donor genome and at least one partial synthetic recipient genome. The method of the present invention may further comprise one or more rounds of recombination-mediated genetic modification, preferably lambda red-mediated genetic modification (before induced conjugation), to provide at least one partial synthetic donor genome and at least one partial synthetic recipient genome. The method may further comprise one or more rounds of selection for at least one partial synthetic donor genome and at least one partial synthetic recipient genome.
[0178] At least one partial synthetic donor genome may comprise a first selectable marker flanked by two regions of homology immediately downstream of the synthetic region and the origin of transfer, and at least one partial synthetic recipient genome may comprise a second selectable marker flanked by two corresponding regions of homology, wherein the first selectable marker may comprise a positive selectable marker and / or the second selectable marker may comprise a negative selectable marker.
[0179] Suitably, the negative selectable marker is sacB (sucrose sensitivity), rpsL (S12 ribosomal protein - streptomycin sensitivity), or phe ST251A_A294G (4-chlorophenylalanine sensitive). Suitably the positive selectable marker is selected from the group consisting of Cm R (chloramphenicol resistant), Kan R (kanamycin resistance), Hyg R (hygromycin resistant), gentamicin R (gentamicin resistant), or tetracycline R (tetracycline resistance). The selectable marker may be different in one or more steps of the recombination-mediated genetic modification.
[0180] Preferably, the synthetic regions present in at least one partial synthetic recipient genome are outside of the regions flanked by homologous regions, i.e., the synthetic regions do not substantially overlap. Preferably, the homologous regions are between 3 kb and 500 kb in length, most preferably about 3 to 5 kb.
[0181] Induced conjugation One or more rounds of directed conjugation may be performed on two or more different partial synthetic genomes of the invention to produce a synthetic genome.
[0182] Each round of induced conjugation can be used to provide a partial synthetic genome with a larger contiguous synthetic region. For example, after one or more rounds of recombination-mediated genetic modification, there can be eight partial synthetic genomes, each with a contiguous synthetic region of about 500 kb. After the first round of induced conjugation, two of the partial synthetic genomes can be combined to provide six partial synthetic genomes, each with a contiguous synthetic region of about 500 kb, and one partial synthetic genome with a contiguous synthetic region of about 1 Mb. A second round can be used to provide six partial synthetic genomes, each with a contiguous synthetic region of about 500 kb. Five partial synthetic genomes each having a continuous synthetic region of about 1.5 Mb can be provided, and one partial synthetic genome having a continuous synthetic region of about 1.5 Mb; or four partial synthetic genomes each having a continuous synthetic region of about 500 kb, and two partial synthetic genomes each having a continuous synthetic region of about 1 Mb. After several rounds of guided conjugation, a complete synthetic genome (i.e., one having a continuous synthetic region of about 4 Mb) can be provided. Examples are shown schematically in Figures 10 and 11b.
[0183] Any suitable method of induced conjugation may be used. Methods of induced conjugation are well known to those skilled in the art and are described, for example, in Ma, NJ, Moonan, DW and Isaacs, FJ, 2014. Nature Protocols, 9(10), p.2285. The pathway to synthetic genomes is not limited.
[0184] Therefore, one or more rounds of induced conjugation may be i) providing a first host cell comprising a partial synthetic recipient genome and a second host cell comprising a partial synthetic donor genome and a conjugated plasmid; ii) conjugation of the partial synthetic recipient genome and the partial synthetic donor genome; and iii) Combinations in which synthetic regions of the donor genome are incorporated into a partially synthetic recipient genome. Recombinant selection step may include:
[0185] The partially synthetic donor genome may comprise a first selectable marker flanked by two regions of homology immediately downstream of the synthetic region and the origin of transfer, and the partially synthetic recipient genome may comprise a second selectable marker flanked by two corresponding regions of homology, the first selectable marker may comprise a positive selectable marker and / or the second selectable marker may comprise a negative selectable marker. Thus, step (iii) may comprise: The method may involve selection for a functional marker, i.e., selection for the gain of a first selectable marker and the loss of a second selectable marker.
[0186] Suitably, the negative selectable marker is sacB (sucrose sensitivity), rpsL (S12 ribosomal protein - streptomycin sensitivity), or phe ST251A_A294G (4-chlorophenylalanine sensitive). Suitably the positive selectable marker is selected from the group consisting of Cm R (chloramphenicol resistant), Kan R (kanamycin resistance), Hyg R (hygromycin resistant), gentamicin R (gentamicin resistant), or tetracycline R(tetracycline resistance). The selectable marker may be different in one or more steps of the recombination-mediated genetic modification.
[0187] Preferably, the homology region is 3 kb to 500 kb in length, most preferably about 3 to 5 kb. Preferably, if the induced conjugation step is the final step of the induced conjugation, the homology region is 50 kb to 500 kb.
[0188] Step (ii) may include incubating the first and second host cells. For example, the first and second host cells may be mixed, transferred to an appropriate medium (e.g., an agar plate), and incubated at about 37°C for about 1 to 3 hours.
[0189] The conjugated plasmid may be an F plasmid, and preferably the conjugated plasmid does not contain an origin of transfer (e.g., Figure 22c).
[0190] In a preferred embodiment, the viability of genome is tested after each round of induction conjugation.Advantageously, this allows verifying that genome editing (for example, sense codon substitution) results in viable genome, and correcting unauthorized editing.In some embodiments, the sequence of genome is verified after each round of induction conjugation.
[0191] Those skilled in the art will understand that they can combine all features of the invention disclosed herein without departing from the scope of the invention as disclosed.
[0192] Preferred features and embodiments of the present invention will now be described by way of non-limiting example.
[0193] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of chemistry, biochemistry, molecular biology, microbiology, and immunology, which are within the capabilities of those skilled in the art. Such techniques are explained in the literature, e.g., Sambrook, J., Fritsch, EF, and Maniatis, T. (1989) Molecular Cloning: A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory. Press; Ausubel, FM et al. (1995 and periodic supplements) Current Protocolsin Molecular Biology, Ch. 9, 13 and 16, John Wiley & Sons; Roe, B.,Crabtree, J. and Kahn, A. (1996) DNA Isolation and Sequencing: EssentialTechniques, John Wiley & Sons; Polak, JM and McGee, J.O'D. (1990) In Situ Hybridization: Principles and Practice, Oxford University Press; Gait, MJ (1984) Oligonucleotide Synthesis: A Practical Approach, IRL Press; and Lilley, DM and Dahlberg, JE (1992) Methods in Enzymology: DNA Structures Part A: Synthesis and PhysicalAnalysis of DNA, Academic Press. [Example] [Example]
[0194] Designing genomes using synonymous codon compression We first designed a version of the E. coli MDS42 genome (Uniprot accession number AP012306.1) in which the serine codons TCG and TCA and the stop codon TAG in the open reading frame (ORF) were systematically replaced with their synonyms AGC, AGT, and TAA, respectively (Fig. 1a, Fig. 18, SEQ ID NO: 1). We previously demonstrated that this defined rewriting scheme for synonymous codon compression is feasible in a 20 kb region of the E. coli genome rich in essential genes (Wang, K. et al., 2016. Nature 539, 59-64). However, this region accounts for only 0.46% of the target codons in the genome.
[0195] E. coli contains many overlapping open reading frames (ORFs), and we classify overlaps as 3',3' (between ORFs in opposite orientations) or 5',3' (between ORFs in the same orientation). The targeted codons are found within both classes of overlaps. If rewriting each ORF within a 3',3' overlap could be achieved without changing the encoded protein sequence of either ORF, i.e., by introducing synonymous codons, the overlap structure was maintained and the sequence was directly rewritten. However, if this was not possible, we duplicated the overlapping region and rewritten each ORF individually (Figure 1b, Table 1).
[0196] For the 5', 3' overlaps, we separated the ORFs by duplicating both the region of overlap between the ORFs and a 20-bp sequence upstream of the overlap. This refactoring allowed us to rewrite each ORF independently (Fig. 1c, Table 1). Our strategy preserved the sequence of the RBS for the downstream ORF and the distance between this RBS and its start codon.
[0197] Using defined rules and refactoring for synonymous codon compression, we designed a genome in which all 18,218 target codons were rewritten to their target synonyms (FIG. 1d).
[0198] [Table 2] JPEG2025175308000005.jpg252170 JPEG2025175308000006.jpg250143 [Example]
[0199] Composition of rewritten compartments We performed a retrosynthesis on the designed genome, similar to that commonly used to design synthetic routes to small molecules (Figure 2). We cleaved the genome into eight segments, A through H, of approximately 0.5 Mb each (Figure 1d, Figure 2a, Figure 18, SEQ ID NO: 1), and then cleaved each segment into four to five fragments (Figure 2b). This yielded 37 fragments ranging from 91 kb to 136 kb (Figure 1d, Table 2). We placed boundaries between fragments and segments in the intergenic regions between non-essential genes. The fragments were further cleaved into nine to fourteen stretches of approximately 10 kb each (Figure 2c, Table 2).
[0200] We constructed BACs for REXER (Fig. 2c, Fig. 20) containing each fragment by homologous recombination in S. cerevisiae (Wang, K. et al., 2016. Nature 539, 59-64; and Kouprina, N., et al., 2004. Methods Mol Biol 255, 69-89). For fragment 36, BAC assembly proceeded smoothly (Table 3). Fragment 37 was difficult to assemble, so we split it into two 50 kb fragments (37a and 37b) and straightened them out for assembly (Table 3).
[0201] We initiated genome replacement in seven different strains using REXER. The starting point for REXER in each strain corresponded to the beginning of compartment A, C, D, E, F, G, or H (Fig. 1d, 2b, Fig. 3), with compartment B later being established in compartment A, as described below. We marked the starting point of genome replacement in each strain by introducing cassettes carrying positive and negative selection markers. We used Cas9 (Jiang, W., et al., 2013. Nat Biotechnol 31, 233-239), lambda Red recombination machinery (Datsenko, KA & Wanner, BL, 2000. Proc Natl Acad Sci USA 97, 6640-6645), and for each compartment, a BAC containing the first transcribed fragment was introduced into the relevant strain, encoding the relevant Cas9 spacer (Jiang, W., et al., 2013. Nat Biotechnol 31, 233-239). The addition of NA to cells initiated replacement of genomic DNA. Cas9-mediated excision of the rewritten DNA from the BAC and lambda Red-mediated recombination of this DNA into the genome resulted in replacement of a section of genomic DNA with the rewritten DNA, removal of the positive and negative selectable markers from the genome, and introduction of new orthogonal positive and negative selectable markers. Clones that had recombined across the target region were selected based on the loss of the negative selectable marker from the genome and the acquisition of the positive selectable marker from the BAC.
[0202] In each strain, the positive and negative selectable markers introduced in the first REXER provided templates for the next round of REXER, enabling genome-level exchange synthesis (GENESIS) (Figures 2b and 4). We used spacer-encoding plasmids for the initial rounds of REXER (Table 4, Figures 20d and 21). However, we subsequently found that REXER could be initiated by electroporation of linear double-stranded spacers generated by PCR (Table 4, Figure 21a). Because these spacers do not propagate through cell division, this allowed cells from one REXER step to be used more quickly for the next REXER step. This progress accelerated GENESIS. For sections A, C, D, E, F, and G, we proceeded with GENESIS in a clockwise direction for four to five REXER steps until we replaced approximately 0.5 Mb of genomic DNA with synthetic DNA. Since section A was started first and completed before the other sections, we progressed GENESIS through section B once we reached the end of section A.
[0203] After each REXER, we sequenced the resulting genome to identify cells that were completely rewritten across the targeted region of the genome (Table 4). In parallel, we performed multiple single-step REXERs (Table 4) to rapidly identify 100 kb regions of the genome that may be difficult to rewrite, and then we sequenced them via GENESIS. We reached them by using synthetic DNA. For 35 of the 38 steps, including all of compartments A, C, D, E, F, and G, we were able to completely rewrite the targeted genomic sequence with GENESIS. We observed only incomplete replacement of the corresponding genomic region with synthetic DNA for fragment 9 in compartment B and for fragments 37a and 1 in compartment H (Table 4).
[0204] [Table 3] JPEG2025175308000008.jpg255165 JPEG2025175308000009.jpg25593
[0205] [Table 4]
[0206] [Table 5]
[0207] [Table 6] [Example]
[0208] Identifying and repairing design flaws By sequencing several clones after REXER, we were able to score the frequency with which each target codon was rewritten, thereby aggregating the rewriting landscape for the genomic region. From the rewriting landscape in fragment 1, we directly identified the fourth codon (Ser4, TCA) in map, an essential gene encoding methionine aminopeptidase, as being difficult to rewrite using our defined rewriting scheme (Fig. 5a). We also identified a second region encompassing a 14-bp overlap of the essential genes ftsI and murE, as well as several serine codons in ftsI and murE, which were not replaced by our rewritten and refactored sequence. Because we had previously rewritten this region with the same rewriting scheme, we added 182 bp to the overlap instead of the 20 bp used here. When the addition was replicated (Wang, K. et al., 2016. Nature 539, 59-64) (Fig. 1c), we conclude that the defect in the synthetic DNA for this region is in its refactoring, not its rewriting. REXER using the new fragment 1 BAC, which contained both the extended refactoring (Fig. 5b) and the TCA to TCT mutation at Ser4 in map (Fig. 5c, Table 5), enabled the complete rewriting of the targeted 100 kb region of the genome (Fig. 5d).
[0209] From the post-REXER rewriting landscape for fragment 9, we identified a 26-kb genomic region that had not been rewritten (Figure 6). Attempts to delete a 10-kb region of the genome within and around this region in the presence of a BAC containing rewritten fragment 9 narrowed the region that was difficult to rewrite to 10 kb of the genome. REXER across the 10-kb genomic region revealed a minimum within the resulting rewriting landscape at yceQ. This identified five target codons within yceQ as problematic for rewriting. Similarly, post-REXER rewriting landscapes for fragment 37a, followed by further sequencing, allowed us to identify a single codon at the 3' end of yaaY that had not been rewritten (Figure 7).
[0210] Both yceQ and yaaY encode "predicted proteins," multiple insertions in yceQ are viable, and there is no evidence of mRNA production and / or protein synthesis from these predicted genes (Pundir, S., et al., 2017. Methods Mol Biol 1558, 41-55). Notably, all of the hard-to-rewrite codons in yceQ and yaaY are located within the 5' untranslated region (UTR) of the essential gene. We have demonstrated that yceQ and yaaY are highly reproducible. These findings suggest that the sequence changes introduced by recoding aY negatively affect the regulation of neighboring essential genes. Indeed, we mapped the target codon in yceQ to RNA secondary structures and promoter elements within the 5'UTR of rne (encoding the essential ribonuclease RNase E) (Figure 8), and these sequences are essential for controlling RNAse E homeostasis (Schuck, A., et al. 2009. Mol Microbiol 72, 470-478).
[0211] We modified fragment 9 by introducing a stop codon into the 5' sequence of yceQ, thereby minimizing any potential translation but retaining the native sequence for regulating rne transcription (Figure 6, Table 5). REXER on this new BAC completely rewrote the corresponding 100 kb genomic region (Figure 6, Table 5). REXER on a new BAC containing fragment 37a, which replaced the problematic codon in yaaY with TCA to AGC, completely rewrote the corresponding region of the genome (Figure 7, Table 5).
[0212] By identifying and correcting all initially problematic sequences, we completed the assembly of a strain in which compartments A and B were completely rewritten (Figure 9), and a strain in which compartment H was completely rewritten (Table 5, Figure 9), completing the assembly of all compartments in seven different strains.
[0213] [Table 7] JPEG2025175308000014.jpg25589 [Example]
[0214] Assembly of the rewritten genome We developed a conjugation-based strategy to assemble the rewritten segments into a single genome (Isaacs, FJ et al., 2011. Science 333, 348-353; Ma, NJ, et al., 2014. Nat Protoc 9, 2285-2300; and Lederberg, J. & Tatum, EL, 1946. Nature 158, 558). Our strategy assembles the rewritten genome clockwise by conjugating rewritten "donor" segments containing the origin of transfer (oriT) to their adjacent rewritten "recipient" segments that have been extended to provide homology with the donor (Figure 10, Figure 11a, Figure 22a, b). This generates a new genome containing both the donor and recipient rewritten segments. The cells containing this new genome can then be used as a recipient for the next rewritten donor, and the process can be repeated to gradually add rewritten segments to the rewritten recipient, allowing the rewritten genome to be assembled (Figure 10, Figure 11a, b). The donor cells contained a form of F' plasmid that facilitates the transfer of the donor genome to recipient cells, but unlike standard F' plasmids, it does not have the ability to transfer itself to recipient cells (Figure 22c). As a result, this F' plasmid does not need to be lost from recipient cells after every conjugation. This accelerated our workflow.
[0215] We initiated conjugation by mixing donor and recipient cells and varied the conjugation time and conditions to control the degree of genome transfer from the donor to the recipient. After conjugation between donor and recipient cells, we selected recipient cells and then selected those recipients that acquired positive markers at the end of the rewritten sequence from the donor and lost negative markers at the end of the recipient extension (Figure 11a).
[0216] We performed convergent synthesis of the rewritten genome through sections A through E (Figure 10, Figure 11b). We then used strains A through E as recipients for F to generate rewritten strains A through F. A through F were then used as recipients for F through G to generate A through G; this conjugation used a fairly long shared rewritten sequence (0.4 Mb) between the donor and recipient strains to increase conjugation efficiency.
[0217] To create a fully rewritten genome, we first created a recipient strain by introducing 37a and 37b into A through G to create A through G-37ab (providing a 115 kb homologous region in the final donor). We created the final donor strain by conjugation between the H and AB strains, resulting in the HA-09 strain, in which H, A, and fragment 9 from section B are rewritten (Figures 10, 11b). By adding additional sequences from A and B to H, we ensured that we did not erase the rewriting of A in the final conjugation. The final conjugation between the HA-09 donor strain and the A through G-37ab recipient strain resulted in the synthesis of an E. coli strain, which we named E. coli Syn61, in which 1.8 x 10 4 All target codons are rewritten (Figure 19, SEQ ID NO: 2). Our synthesis of the rewritten genome introduced only eight unprogrammed mutations (Table 6), four of which occurred during the preparation of the 100 kb BAC and four during the rewriting process.
[0218] [Table 8] JPEG2025175308000016.jpg25550 [Example]
[0219] Consequences of synonymous codon compression in Syn61 Syn61 doubled only 1.6-fold slower than MDS42 in LB plus glucose at 37°C, and this rate increased at 25°C and decreased at 42°C (Fig. 13a). Syn61 contains 65% more AGT and AGC codons than MDS42, yet providing additional copies of serV, the tRNA that decodes these codons (Fig. 12a), did not increase growth (Fig. 13a), suggesting that serV is not limiting. Imaging of Syn61 cells suggests that they are slightly longer than MDS42 (Fig. 13b, c). The Syn61 proteome was comparable to that of MDS42 (Fig. 13d). Orthogonal aminoacyl-tRNA synthetase / tRNA targeting the TCG codon. CGA Co-translational incorporation of non-canonical amino acids using the pair was highly toxic in MDS42 but completely non-toxic in Syn61, providing phenotypic validation for the removal of TCG codons in Syn61 (Fig. 12b). This approach also provided further insights (Fig. 14a, b, c). Ser UGA serT, encoding serU, is essential because it is the only tRNA that decodes TCA codons in E. coli. Because Syn61 does not contain a TCA codon, serT should be non-essential in our strain. Indeed, we demonstrated that serT (Figures 12c, 14d, and 23) as well as serU and prfA (Figures 14e, f, and 23) can be easily removed in Syn61. These data provide functional confirmation that we removed the target codon from the genome, show that the tRNA and the release factor that decodes the target codon can be removed in Syn61, and demonstrate the unique properties of Syn61 resulting from rewriting. [Example]
[0220] Consideration We generated E. coli with an entire 4 Mb genome replaced with synthetic DNA, and the scale of genome replacement in our experiments is approximately four times larger than previously reported for genome replacement in mycoplasma or chromosomal replacement in a single strain of S. cerevisiae (Figure 15a).
[0221] We have identified all known 1.8 × 10 4 We demonstrated genome-wide removal of 20 targeted codons (two sense codons, TCG and TCA, and an amber codon, TAG). Our study removes 60-fold more codons than experiments removing amber stop codons by site-directed mutagenesis (Figure 15b). Furthermore, this demonstrates complete, genome-wide rewriting of all targeted sense codons (Figure 15b). Therefore, we created a synthetic organism that uses 61 codons instead of the usual 64. The new organism uses a reduced number of sense codons to encode the 20 canonical amino acids.
[0222] Our synthetic genome contains 2 × 10 genes per target codon. -4 This contains only 1.05 unprogrammed mutations per target codon (Figure 15c), which compares favorably with the 1.05 unprogrammed mutations reported for replacing amber codons by site-directed mutagenesis (Lajoie, MJ et al., 2013. Science 342, 357-360) (Figure 15c).
[0223] Our final synthetic genome was rewritten using a refactoring and rewriting scheme we defined, using rewriting rules we previously determined for only 83 (0.43%) of the genome's target codons (Wang, K. et al. 2016. Nature 539, 59-64 ) The rewriting rules are 1.8 × 10 of the genome. 4 99.9% of the target codons were And the refactoring rule worked on 99% of the duplications.
[0224] Our initial rewriting scheme corrections resulted in 1.8 × 10 4 Only seven of the target codons were required. One of these codons was in an essential gene, while the other six were within the 5'UTR of an essential gene. Therefore, all but one of the changes in our defined rewriting scheme correct unintended modifications to the 5'UTR of an essential gene, rather than the direct effect of the altered synonyms on translation.
[0225] The strategy we developed to cleave engineered genomes into sections, fragments, and stretches and achieve design through the convergent, seamless, and robust integration of REXER, GENESIS, and guided conjugation provides a blueprint for future genome synthesis. In future studies, we will further characterize the consequences of synonymous codon compression in E. coli Syn61 and test additional rewriting schemes in E. coli and other organisms. Furthermore, we will test sense codon reassignment for noncanonical biopolymer synthesis. [Example]
[0226] method Rewritten genome design We based our synthetic genome design on the sequence of the E. coli MDS42 genome (accession number AP012306.1, published October 7, 2016), which has 3547 annotated CDSs. We manually curated the annotation of the starting genome to remove three CDSs and add another 12. The three predicted CDSs we removed were htgA, ybbV, and yzfA. There is no evidence that these sequences encode proteins (Pundir, S., et al., 2017. Methods Mol Biol 1558, 41-55). These sequences completely or largely overlap with well-characterized genes, making them difficult to rewrite without disrupting the overlapping genes or creating large repetitive regions. Conversely, the pseudogenes ydeU, ygaY, pbl, yghX, yghY, agaW, yhiK, yhjQ, rph, ysdC, glvG, and cybC were recommended for CDS. To enable negative selection with rpsL, we inserted a genomic copy of rpsL into the rpsL K43R Finally, our in-house deep sequencing of MDS42 revealed a 51-bp insertion between mrcB and heML that had not been reported in AP012306.1. We manually introduced and annotated this insertion in our starting genome sequence.
[0227] We created a custom Python script to i) identify and rewrite all target codons and ii) identify and resolve overlapping gene sequences containing the target codon. From our curated MDS42 starting sequences, we used the script to generate a new synthetic genome in which all TCG, TCA, and TAG codons were replaced with AGC, AGT, and TAA, respectively. The script reported 91 CDSs with overlaps containing the target codon. In 33 cases, genes were duplicated tail-to-tail (3', 3') (Table 1). Twelve of these could be rewritten by introducing silent mutations into the overlapping genes, while the remaining 21 were duplicated to separate the genes (Figure 1b). Fifty-eight examples of head-to-tail (5', 3') overlapping genes were resolved by duplicating the overlap plus 20 bp of upstream sequence to allow endogenous expression of the downstream gene (Figure 1c). For overlaps longer than 1 bp, an in-frame TAA was introduced to terminate expression from the original RBS for downstream genes. Because prfB (release factor RF-2) was not annotated as a CDS in our starting MDS42 genome due to its regulatory internal stop codon, we introduced an in-frame TAA to terminate expression from the original RBS for downstream genes. All target codons were manually rewritten, thereby maintaining internal stop codons. The resulting genome design contained 3556 CDSs with 1,156,625 codons, of which 18,218 were rewritten (Figure 18, SEQ ID NO: 1).
[0228] Retrosynthesis of rewritten stretches We divided the designed genome into 37 fragments ranging from 91 to 136 kb. We determined that i) the border sequences consisted of 5'-NGG-3' PAM so that REXER4 could be used for integration if necessary, and ii) the PAM was located within 50 bp of the target codon. iii) the PAM is located between non-essential genes; iv) the PAM is located in any of the promoters The boundary sequences separating these fragments were chosen so as not to interfere with any of the annotated features. We called the regions approximately 50-100 bp upstream and downstream of these boundaries "landing sites" and annotated them as Lxx, where xx is the number of the upstream fragment; e.g., L01 is the landing site between fragments 1 and 2. In our design, landing site sequences are contained at the 3' end of the fragment and the next 5' end; as a result, all 37 fragments contain 54-155 bp of overlapping homology with their adjacent fragments.
[0229] Each fragment was further decomposed into 7-14 stretches of 4-15 kb. We designed the stretches to contain 80-200 bp of overlap with each other, and the overlapping region was defined as an intergenic region that did not contain any rewriting targets. A total of 409 stretches were synthesized (GENEWIZ, USA) and delivered to pSC101 or pST vectors flanked by BsaI, AvrII, SpeI, or XbaI restriction sites. The synthetic stretches naturally contained at least one of these restriction sites. Neither contained one.
[0230] Construction of selection cassettes and plasmids for REXER / GENESIS The cloning procedures described in this section were carried out in E. coli DH10b, which is resistant to streptomycin due to the rpsLK43R mutation. The plasmid pKW20_CDFtet_pAraRedCas9_tracrRNA used throughout this study was prepared as previously described. It encodes Cas9 and lambda-Red recombination components alpha / beta / gamma under the control of an arabinose-inducible promoter, and tracrRNA under its native promoter (Wang, K. et al., 2016. Nature 539, 59-64).
[0231] The protospacer for REXER is located in the plasmid pKW1_MB1 AmpThe plasmid pKW3_MB1 is encoded in the pKW3_MB1 spacer (Fig. 21a), which contains the pMB1 origin of replication, an ampicillin resistance marker, and a protospacer array under the control of its endogenous promoter, as previously described (Wang, K. et al., 2016. Nature 539, 59-64). From this plasmid, we created the derivative pKW3_MB1 Amp _Tracr K A tracrRNA-spacer was constructed (Table 5), which further contains tracrRNA upstream of the protospacer array. To this end, we transformed a PCR product containing tracrRNA with its modified endogenous promoter into pKW1_MB1 by Gibson assembly using NEBuilder HiFi Master Mix. Amp _Space From this plasmid, the fragment was also synthesized by Gibson assembly. A derivative encoding additional Cas9 was constructed, pKW5_MB1 Amp _Tracr K This was named _Cas9_spacer.
[0232] For each REXER step, a derivative of one of these three plasmids was constructed to carry a protospacer / direct repeat array containing two (REXER2) or four (REXER4) protospacers corresponding to the target sequences for cleaving the BAC and genome. Different protospacer arrays were constructed from overlapping oligos by multiple rounds of PCR, and the products were then transformed into pKW1_MB1. Amp _Spacer, pKW3_MB1 Amp _Tracr K _spacer or pKW5_MB1 Amp _Tracr K Restriction site AccI in the backbone of the Cas9 spacer The resulting promoters were inserted between the EcoRI and EcoRI sites by Gibson assembly. The spacer arrays were verified to be mutation-free by Sanger sequencing.
[0233] The positive-negative selection cassette used in REXER and GENESIS was -1 / +1 (rpsL-Kan R ), -2 / +2(sacB-Cm R ) and -3 / +3(pheS T251A_A294G -Hyg R ) -1 / +1 and -2 / +2 are as previously described (Wang, K. et al., 2016. Nature 539, 59-64). -3 / +3 is pheS T251A_A294G is dominant lethal in the presence of 4-chlorophenylalanine, and Hyg R confers resistance to hygromycin. Both proteins are expressed polycistronically under the control of the EM7 promoter. The -3 / +3 cassette was synthesized de novo. The -3 / +3 cassette is encoded by the pheS promoter. * / Hyg R It is also called.
[0234] Construction of E. coli strains containing dual selection cassettes in the genomic landing site. According to our design, each region of the genome targeted for replacement with a synthetic fragment is flanked by an upstream landing site and a downstream landing site, and these genomic landing site sequences are the same as those described above. Initiation of REXER / GENESIS requires the insertion of a double selection cassette into the upstream genomic landing site. We inserted the double selection cassette into the landing site by lambda-Red-mediated recombination. Briefly, the sacB-Cm R or rpsL-Kan ROne of the cassettes was PCR amplified using primers containing homologous regions to the desired genomic landing site. For recombination experiments, we prepared electrocompetent cells as previously described (Wang, K. et al., 2016. Nature 539, 59-64) and transfected 3 μg of purified PCR product into 10 cells carrying the pKW20_CDFtet_pAraRedCas9_tracrRNA plasmid expressing the Lambda Red alpha / beta / gamma genes. 0 μL MDS42 rpsLK43R Cells were electroporated with OD 6000 under the control of the arabinose promoter (pAra). 600 The recombination machinery was induced by adding 0.5% L-arabinose for 1 hour starting at pH 0.2. Pre-induced cells were electroporated and then allowed to recover in 4 mL of super optimal broth (SOB) medium at 37°C for 1 hour. Cells were then resuspended in 10 μg / mL tetrahydrofuran. The cells were diluted in 100 mL of LB medium containing tetracycline and grown at 200 rpm at 37°C for 4 hours. Afterwards, the cells were spun down, resuspended in 4 mL of HO, serially diluted, and plated in a 10 μg / mL tetracycline, 18 μg / mL chloramphenicol (sacB-Cm R for 50 μg / mL of kanamycin (rpsL-Kan R The mixture was incubated overnight at 37°C on an LB agar plate containing 100% ethanol (for use).
[0235] BAC assembly and delivery We constructed bacterial artificial chromosome (BAC) shuttle vectors containing 97–136 kb of synthetic DNA. At the 5′ end, the synthetic DNA was flanked by a region of homology to the genome (HR1) and a Cas9 cleavage site. At the 3′ end, the synthetic DNA was flanked by a double selection cassette, a region of homology to the genome (HR2), and a second Cas9 cleavage site. The BAC also contained a negative selection marker, a BAC origin, a URA marker, and a YAC origin (CEN6 centromere fused to an autonomously replicating sequence (CEN / ARS)) (Figure 2c, Figures 20a–c).
[0236] BACs were assembled by homologous recombination in S. cerevisiae. Each assembly combines i) 7-14 stretches of synthetic DNA, each 6-13 kb in length, with ii) a selection construct (see below) and iii) a BAC shuttle vector backbone. (Figure 20a-c, Wang, K. et al., 2016. Nature 539, 59-64).
[0237] Synthetic DNA stretches were prepared using their source vectors provided by GENEWIZ. The fragments were excised from the 5'-terminal end of the 5'-terminal fragment by digestion with BsaI, AvrII, SpeI, or XbaI restriction sites. In the case of AvrII, SpeI, and XbaI, restriction digestion was followed by mung bean nuclease treatment to remove sticky ends. did.
[0238] The selection construct contains a homology region to the 3'-most stretch of the fragment, a double selection cassette (sacB-Cm R or rpsL-Kan R ), a homology region (HR2) to the targeted genomic locus, a negative selection marker (rpsL, sacB or pheS * -Hyg R) and YAC. See Figure 20d for the specific double selection cassette, negative selection marker, and homology region sequences. We assembled an episomal version of the selection construct in a pSC101 backbone from the three PCR fragments using NEBuilder HiFi DNA Assembly Master Mix. This episomal version was designed so that restriction digestion with BsaI would generate DNA fragments for BAC assembly.
[0239] The BAC backbone containing the BAC origin and URA3 marker was amplified by PCR using a previously described BAC (Wang, K. et al., 2016. Nature 539, 59-64) as a template, and the PCR product was used for BAC assembly. The primers used for these PCR assemblies are listed in Figure 20d.
[0240] To assemble the stretches, selection constructs, and BAC backbones, 30–50 fmol of each piece of DNA was transformed into S. cerevisiae spheroplasts, which were prepared as previously described (Kouprina, N., et al., 2004. Methods Mol Biol 255, 69–89). After assembly, we identified yeast clones potentially harboring the correctly assembled BAC by colony PCR at the junctions of the overlapping fragments and the vector insertion junction. Clones that appeared correct by colony PCR were sequence verified by next-generation sequencing after transformation into E. coli, as described below.
[0241] The assembled BAC was extracted from yeast using the Gentra Puregene Yeast / Bact.Kit (Qiagen) according to the manufacturer's instructions. 42 rpsLK43RThe assembled BACs were transformed into cells. Due to the large size of the BACs, we sometimes observed inefficient electroporation into target cells. As a result, we introduced the oriT-apramycin cassette, provided as a PCR product with a 50-bp homology region by lambda-Red-mediated recombination (as described above), into some BACs after assembly (Figures 20a-c). This facilitated the transfer of BACs from successfully transformed E. coli to other strains by conjugation.
[0242] Synthesis of rewritten compartments with REXER and GENESIS We used various genomic and plasmid selection markers for the sequential REXER experiments (GENESIS) (Table 4). We used rpsL-Kan at the genomic landing site for selection. R (-1 / +1) or sacB-Cm R We used the (-2 / +2) cassette. We used the rpsL-Kan as an episomal selection marker. R -sacB(-1 / +1, -2), rpsL-Kan R -pheS * -Hyg R (-1 / +1, -3 / +3) or sacB-Cm R A -rpsL(-2 / +2, -1) cassette was used.
[0243] For each REXER, a pKW20_CDFtet_pAraRedCas9_tracrRNA and MDS42 containing a dual selection cassette at the relevant upstream genomic landing site were inserted. rpsLK43R The cells were transformed with the relevant BAC. We cultured the cells in 2% glucose, 5 μg / ml tetracycline and the antibiotic of choice for the BAC (i.e., 18 μg / ml tetracycline). Cells were plated onto LB agar supplemented with 50 μg / ml chloramphenicol or 50 μg / ml kanamycin. We inoculated individual colonies into LB medium containing 5 μg / ml tetracycline and the BAC-specific antibiotic and grew the cells overnight at 37°C and 200 rpm. The overnight culture was diluted to an OD of 0.05 in LB medium containing 5 μg / ml tetracycline and the BAC-specific antibiotic and grown at 37°C with shaking for approximately 2 hours until the OD was approximately 0.2. To induce Lambda Red expression, we added arabinose powder to the culture to a final concentration of 0.5% and incubated the culture for an additional hour at 37°C with shaking. We harvested the cells at OD600≈0.6 and made them electrocompetent as previously described (Wang, K. et al., 2016. Nature 539, 59-64).
[0244] For each REXER experiment, a linear dsDNA protospacer array was PCR amplified from pKW1_MB1Amp_spacer using universal primers (Figure 21a). Approximately 5-10 μg of the resulting DpnI-digested and purified PCR product was transformed into 100 μL of electrocompetent and induced cells. Cells were allowed to recover in 4 ml of SOB medium at 37°C for 1 h, then diluted with 100 mL of LB supplemented with 5 μg / mL tetracycline and the antibiotic selected for the BAC, and incubated with shaking at 37°C for an additional 4 h. Alternatively, electrocompetent and induced cells were transformed with 5 μg of the circular protospacer array (pKW1_MB1Amp_spacer or pKW3_MB1Amp_spacer plasmid) and allowed to recover in SOB medium at 37°C for 1 hour, then transferred to 100 mL of LB supplemented with 100 μg / mL ampicillin for an additional 4 hours at 37°C with shaking (Figure 21a, b). When REXER2 was not sufficient, we performed REXER4 using the pKW5_MB1Amp_spacer plasmid as previously described (Wang, K. et al., 2016. Nature 539, 59-64).
[0245] We spun down the culture, resuspended it in 4 ml of Milli-Q filtered water, and Serial dilutions were plated onto selective LB agar plates containing 50 μg / ml tetracycline, a drug that selected for negative selection markers and an antibiotic that selected for the positive marker derived from the BAC. The plates were incubated overnight at 37°C. Multiple colonies were picked, resuspended in Milli-Q filtered water, and incubated with 50 μg / ml kanamycin, 18 μg / ml chloramphenicol, and 10 μg / ml tetracycline. The colonies were plated onto several LB agar plates supplemented with phenicol, 200 μg / ml streptomycin, 7.5% sucrose, or 2.5 mM 4-chlorophenylalanine. Colony PCR was also performed from the resuspended colonies using primer pairs flanking both the genomic locus of the landing site and the location of the newly integrated selection cassette from the BAC. REXER-mediated recombination generated an approximately 500 bp band at the upstream genomic locus, similar to the control MDS42 rk / MDS42 sC A 2.5 kb (rk-landing site) or 3.5 kb (sC-landing site) band for the strain indicates successful removal of the landing site from the genome. Primer pairs flanking the 3' end of the replaced DNA yielded bands of approximately 2.5 kb (rk selection cassette on pBAC) or 3.5 kb (sC selection cassette on pBAC) and the control MDS42, indicating successful integration of the selection marker. rk / MDS42 sC This produces a 500 bp band for the strain.
[0246] If a plasmid-based circular protospacer array was used in a previous REXER experiment, the plasmid had to be lost before the next experiment. Therefore, successful clones from the first REXER experiment were grown to high-density cultures at 37°C with shaking in LB supplemented with 2% glucose, 5 μg / mL tetracycline, and an antibiotic that selected for positive markers in the genome. Two μL of the culture was then streaked onto an LB agar plate containing the same supplements and incubated overnight at 37°C. Several colonies were replica plated onto LB agar plates and LB agar plates supplemented with 100 μg / mL ampicillin to screen for loss of the plasmid.
[0247] BACEdit When a loss-of-function mutation was encountered in the selection cassette on the BAC in E. coli, the defective cassette was replaced with the appropriate double selection cassette provided as a PCR product flanked by 50 bp homologous regions and integrated by lambda-Red-mediated recombination (Figure 20d).
[0248] Changes in the synthetic rewritten sequence of the BAC were introduced by a two-step replacement approach, either to correct natural mutations or to change the rewritten codon. For BACs containing selection cassettes -2 / +2 and -1 at the ends of the rewritten sequence, the -3 / +3 cassette was provided as a PCR product flanked by 50-bp homologous regions targeting the desired locus and integrated by lambda Red-mediated recombination, followed by selection for +3. Due to homology between the rewritten DNA and the genome, some of the resulting clones contained -3 / +3 on the BAC and some on the genome. To identify clones with the cassette on the BAC, clones were replica plated on agar plates, selecting for (1) +3, (2) against -3, and (3) +2 and against -3. Only clones that survived on plates (1) and (2) but not (3) had the -3 / +3 cassette integrated into the BAC. The location of the cassette was determined using a QIAprep Spin BAC purification using a Miniprep Kit followed by validation by genotyping In the second step, the -3 / +3 cassette was replaced by providing a PCR product of the desired sequence flanked by 50 bp homology regions and integrated by lambda-Red-mediated recombination, followed by selection for +2 and against -3. The BAC was genotyped as above, and the sequence was verified by NGS.
[0249] Preparation of nontransferable F' plasmids and episomal conjugate transfer We engineered a version of the F' plasmid used for conjugation of genomic DNA and transfer of BACs between strains, enabling the transfer of sequences carrying oriT without transferring the F' plasmid itself (Figure 22c). We achieved this by deleting the nick site of the origin of transfer (oriT) within the F' plasmid itself, a related approach that has been previously reported (Strand, TA, et al., 2014. PLoS One 9, e90372). The F' plasmid derivative, pRK24 (addgene #51950), was cloned with a 50 bp homologous The desired marker is incorporated as a PCR product flanked by regions modified by Tet R Instead of Kan R A variant of pKW20 carrying lambda- First, the ampicillin resistance was conferred in pRK24 by PCR. The β-lactamase gene in T5 was replaced with an artificial T5-luxABCDE operon (Bryksin, AV & Matsumura, I., 2010. PLoS One 5, e13244) that produces bioluminescence, allowing visual identification of infected bacterial cells. R The cells were then cultured for selection with 50 μg / mL apramycin. To achieve this, the oriT gene was replaced with T3-aac3, which produces aminoglycoside 3-N-acetyltransferase IV. Finally, a 24-bp deletion of the nick site in oriT was performed by incorporating EM7-bsd, which expresses blasticidin-S deaminase, allowing selection with 50 μg / mL blasticidin in low-salt TYE / LB. The resulting F' plasmid, designated pJF146 (Figure 22c), was extracted using a QIAprep SpinMiniprep Kit (QIAgen) and transformed by electroporation into the donor strain for subsequent conjugation.
[0250] The transfer of episomal DNA containing oriT was carried out by conjugation (Isaacs, FJ et al., 2011. Science 333, 348-353; and Ma, NJ, et al. 2014. Nat Protoc 9, 2285-2300). The donor strain was double transformed with pJF146 and the assembled BAC containing oriT (see above). The recipient strain was transformed with pKW20. 1 ml of donor and recipient cultures were grown to saturation overnight in selective LB medium and then The cells were then washed three times with antibiotic-free LB medium. The resuspended donor and recipient strains were combined at a 4:1 ratio and spotted onto TYE agar plates and incubated at 37°C for 1 hour. The cells were washed off the plates and serially diluted onto LB agar plates containing 2% glucose, 5 μg / ml tetracycline (selective for the recipient strain), and the antibiotic (selective for the BAC). Successful BAC transfer was confirmed by colony PCR of the BAC-vector insertion junction.
[0251] Assembling synthetic genomes from rewritten compartments The transfer of genomic DNA, combined with subsequent recBCD-mediated recombination, resulted in the assembly of a partial synthetic E. coli genome into a synthetic genome. In preparing the donor and recipient strains, rpsL-HygR-oriT or Gm RThe pheS-oriT cassette was supplied as a PCR product and integrated into the donor strain genome by lambda-Red-mediated recombination (Fig. 22a, b). * -Hyg R The cassette was integrated approximately 3 kb downstream of the synthetic DNA of the donor strain. * -Hyg R Genomic DNA served as a template for PCR amplification of a 3 kb synthetic DNA segment carrying a selection cassette. This PCR product was introduced into a recipient strain to replace the WT DNA by lambda-Red-mediated recombination, thereby replacing the selection marker at the 3' end of the synthetic segment and generating a 3 kb region of homology to the donor synthetic DNA. This strategy resulted in the synthesis of 3 kb of synthetic segments with pheS-Hyg at the 3' end. R Furthermore, pJF146 was transformed into the donor strain and its sensitivity to tetracycline was confirmed. In contrast, pKW20 was maintained in the donor strain to induce tetracycline sensitivity. It confers resistance to lacycline.
[0252] For conjugation, donor and recipient strains were grown to saturation overnight in LB medium containing 2% glucose, 5 μg / ml tetracycline and 50 μg / ml kanamycin or 20 μg / ml chloramphenicol (donor) and 50 μg / ml apramycin and 200 μg / ml hygromycin B (recipient). The overnight cultures were diluted 1:10 in the same selective LB medium and analyzed at OD 600The donor and recipient cultures were grown to a pH of 0.5. 50 ml of both the donor and recipient cultures were washed three times with LB medium containing 2% glucose, then each was resuspended in 400 μl of LB medium containing 2% glucose. 320 μl of the donor was mixed with 80 μl of the recipient, spotted onto a TYE agar plate, and incubated at 37°C. The incubation time varied from 1 to 3 hours, depending on the length of the transferred synthetic DNA and the doubling time of the recipient strain. The cells were washed from the plate, transferred to 100 ml of LB medium containing 2% glucose and 5 μg / ml tetracycline, and incubated with shaking at 37°C for 2 hours. Subsequently, 50 μg / ml kanamycin or 20 μg / ml chloramphenicol (selection for the donor-transferred positive selectable marker) was added, and then incubated at 37°C for an additional 2 hours. The cultures were spun down, resuspended in 4 ml of Milli-Q filtered water, and incubated for 2 hours. Serial dilutions were plated onto selective plates of LB agar containing 5% glucose, 5 μg / ml tetracycline, 2.5 mM 4-chlorophenylalanine, and 50 μg / ml kanamycin or 20 μg / ml chloramphenicol. Successful DNA transfer and recombination were confirmed by the pheS * -Hyg R Loss of the cassette, integration of the donor selection cassette and absence of the Gm-oriT cassette were determined by colony PCR.
[0253] Preparation of whole genome and BAC libraries for next-generation sequencing Using the DNEasy Blood and Tissue Kit (QIAgen) according to the manufacturer's instructions. E. coli genomic DNA was purified. BAC was extracted from cells using the QIAprep Spin Miniprep Kit (QIAgen) according to the manufacturer's instructions. We have confirmed that this kit is We found that this method was suitable for purifying BACs over 10 kb in size. We avoided vigorous shaking of the sample throughout the purification to reduce DNA shearing.
[0254] Paired-end Illumina sequencing libraries were prepared using the Illumina Nextera XT Kit according to the manufacturer's instructions. Sequencing data were obtained on an Illumina MiSeq using the MiSeq Reagent kit v3 for 2 × 300 or 2 × 75 cycles.
[0255] Sequencing data analysis The standard workflow for sequence analysis in this study is summarized in the iSeq package. Briefly, sequencing reads were aligned to the reference rewritten or wild-type genome using bowtie2 with soft clipping activated. (Langmead, B. & Salzberg, SL, 2012. NatMethods 9, 357-359). Aligned reads were classified and indexed using samtools (Li, H. et al., 2009. Bioinformatics 25, 2078-2079). A customized Python script was used with samtools. This script was used to generate variant calling summaries in conjunction with functions from the igvtools and igvtools tools. Protein and structural variations were assessed (Thorvaldsdottir, H., et al., 2013. BriefBioinform 14, 178-192).
[0256] We created a custom Python script to generate rewrite landscapes across target genomic regions. Briefly, the script takes a BAM alignment file, a fasta reference, and a GenBank annotation file as input. It identifies target codons for rewriting and aggregates reads that align to these target codons in an alignment file. It then outputs the rewriting frequency at each target codon and plots these frequencies over the length of the desired genomic region.
[0257] Measurement and analysis of proliferation rate Bacterial colonies were grown overnight at 37°C in LB containing 2% glucose and 100 μg / mL streptomycin. The overnight cultures were diluted 1:50 and monitored for growth under varying temperatures (25°C, 37°C, or 42°C) and media conditions (LB, LB with 2% glucose, M9 minimal media, 2XTY). OD 600 Measurements were taken every 5 minutes for 18 hours on a Biomek automated workstation platform with high speed linear shaking.
[0258] To determine doubling time, growth curves were log2 transformed. During the linear phase of the curve during exponential growth, the first derivative was determined (d(log2(x)) / dt), and the 10 consecutive time points with the largest log2 derivative were used to calculate the doubling time for each replicate. A total of 10 independently grown biological replicates were used for the engineered Syn61 strain and the wt MDS42 strain. rpsLK43R The mean doubling time and standard deviation from the mean were calculated for all replicates, n=10.
[0259] Microscopy and cell size measurement Cells were grown with shaking in LB supplemented with 100 μg / mL streptomycin to approximately OD 600 The bacteria were grown to a pH of 0.2. A thin layer of bacteria was sandwiched between an agarose pad and a coverslip. Standard microscope slides were prepared using 1% agarose pads (Sigma-Aldrich A4018-5G). A 2-4 μl sample of the bacterial culture was placed on top of the pad. This was covered with a #1 coverslip supported on both sides by glass spacers adapted to the height of the pad approximately 1 mm. The sample was imaged on an upright Zeiss Axiophot phase-contrast microscope using a 63X 1.25NA Plan Neofluar phase objective (Zeiss UK, Cambridge, UK). Images were taken with an IDS ueye monochrome camera under the control of ueye cockpit software (IDS Imaging Development Systems GmbH, Obersulm, Germany). Ten fields of view were photographed for each sample. Images were loaded into Nikon NIS Elements software (Nikon Instruments, Surrey, UK) for further quantification. General analysis An intensity threshold was applied to segment the bacteria using the tool. A size limit of 1 micron was imposed to remove background particulates and dust. Length measurements were then performed on the segmented bacteria using the general analysis and quantification tool.
[0260] mass spectrometry Three biological replicates were performed for each strain. Proteins from each E. coli lysate were solubilized in a buffer containing 6 M urea in 50 mM ammonium bicarbonate, reduced with 10 mM DTT, and alkylated with 55 mM iodoacetamide. After alkylation, the proteins were diluted to 1 M urea in 50 mM ammonium bicarbonate and digested with Lys-C (Promega, UK) at a protein-to-enzyme ratio of 1:50 for 2 hours at 37°C, followed by trypsin (Promega, UK) at a protein-to-enzyme ratio of 1:100 for 12 hours at 37°C. The collected peptide mixture was acidified by adding formic acid to a final concentration of 2% v / v. Nanoscale capillary LC-MS / MS was performed using an Ultimate U3000 HPLC (ThermoScientific Dionex, San Jose, USA) to deliver a flow rate of approximately 300 nL / min. Digests were analyzed in duplicate (1 μg starting protein / injection) by S. C18 Acclaim PepMap100 3 μm, 75 μm × 250 mm nanoViper (ThermoScientific Dionex, San Jose, CA). Peptides were captured by a C18 AcclaimPepMap100 5 μm, 100 μm x 20 mm nanoViper (ThermoScientific Dionex, San Jose, USA) prior to separation on a nanoViper column (ThermoScientific Dionex, San Jose, USA). Peptides were eluted with a 100-minute gradient of acetonitrile (2% to 60%). The analytical column outlet was directly coupled to a hybrid dual-pressure linear ion trap mass spectrometer (Orbitrap Velos, ThermoScientific, San Jose, USA) via a nanoflow electrospray ionization source. Data-dependent analysis was performed using a resolution of 30,000 for the complete MS spectrum, followed by 10 MS / MS spectra in the linear ion trap. MS spectra were collected over an m / z range of 300-2000. MS / MS scans were collected using a threshold energy of 35 for collision-induced dissociation. All raw files were processed in MaxQuant 1.5.5.1 using standard settings and searched against E. coli strain K-12 using the Andromeda search engine built into the MaxQuant software suite. Enzyme searches were performed. Specificity was trypsin / P for both endoproteinases. A maximum of two false cleavages were allowed for each peptide. Carbamidomethylation of cysteine was set as a fixed modification with oxidized methionine, and protein N-acetylation was considered a variable modification. The search was performed with an initial mass tolerance of 6 ppm for precursor ions and 0.5 Da for CID MS / MS spectra. The false discovery rate was fixed at 1% at the peptide and protein level. Statistical analysis was performed using the Perseus (1.5.5.3) module in MaxQuant. Prior to statistical analysis, peptides mapping to known contaminants, reverse hits, and protein groups identified only by site were removed. Only protein groups identified with at least two peptides, one of which was unique, were considered for data analysis. For proteins quantified at least once in each strain, the mean abundance of each protein across Syn61 replicates was divided by its abundance across MDS42 replicates and then log2-transformed. P values for differences in abundance between strains were calculated by a two-sample t-test (Perseus). .
[0261] Orthogonal aminoacyl-tRNA synthetase tRNA xxx CYPK uptake toxicity using s (Elliott, TS et al., 2014. Nat Biotechnol 32,465-472; Elliott, TS, et al., 2016. Cell Chem Biol 23, 805-815; and Krogager, TP et al., 2018. Nat Biotechnol 36, 156-159) Electrocompetent MDS42 and Syn61 cells were transfected with PylRS and tRNA. Pyl xxx Plasmid pKW1_MmPylS_PylT for expression XXX Transform, where XXX is the anticodon shown. tRNA Pyl The anticodon of pKW1_MmPylS_Py lT CGA ), UGA(pKW1_MmPylS_PylT UGA ) or GCU(pKW1_MmPylS_PylT GCU ) Three variants of this plasmid were used. Cells were grown overnight in LB medium containing 75 μg / ml spectinomycin. The overnight culture was diluted 1:100 in LB supplemented with Nε-(((2-methylcycloprop-2-en-1-yl)methoxy)carbonyl)-L-lysine (CYPK) at 0 mM, 0.5 mM, 1 mM, 2.5 mM, and 5 mM, and growth was measured as described above. "% Maximum Growth" was calculated as the final OD in the absence of CYPK. 600 Final OD in the presence of the indicated concentrations of CYPK divided by 600 The final OD was determined as 600 was decided after 600 minutes.
[0262] Deletion of prfA, serU, and serT by homologous recombination To ensure that expression of the selected protein is independent of decoding by serU or serT, we reprogrammed pheS according to the reprogramming scheme described in Figure 1a. * -Hyg R and rpsL-Kan R A rewritten version of the cassette was synthesized de novo. To delete prfA, the rewritten rpsL-Kan R was amplified with an oligo containing approximately 50 bp of homology to the prfA flanking genomic sequence. * -Hyg R The oligonucleotide sequences are provided in Figure 23. Syn61 cells harboring the plasmid pKW20_CDFtet_pAraRedCas9_tracrRNA were cultured in LB Cells were made competent as above, using 2xTY instead of 1xTY. Cells were electroporated with approximately 8 μg of PCR product, allowed to recover in 4 mL of SOB for 1 hour, and then transferred to 100 mL of 2xTY supplemented with 5 μg / mL tetracycline. After 4 hours, cells were spun down, resuspended in 500 μL of HO, and resuspended in 5 μg / mL tetracycline and 200 μg / mL hygromycin B (pheS). * -Hyg R for 50 μg / ml of kanamycin (rpsL-Kan RSerial dilutions were plated onto 2xTY agar plates supplemented with 100% ethanol (for 10 min). In each case, deletions were verified by colony PCR using primers flanking the desired locus.
[0263] All publications mentioned in the above specification are incorporated herein by reference. Various modifications and variations of the disclosed methods, cells, compositions, and uses of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been disclosed in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the disclosed methods for carrying out the invention that are obvious to those skilled in the art are intended to be within the scope of the appended claims.
Claims
1. A synthetic prokaryotic genome comprising at least one of the following (a) to (e): (a) five or less than four occurrences of one or more sense codons; (b) containing four or three or fewer, three or two or fewer, two or one or fewer, one or zero occurrences, or no occurrences of one or more sense codons; (c) 100 or more, 200 or more, or 300 or more genes, which contain a total of no more than five or four occurrences of one or more sense codons; (d) a gene of (c) that contains a total of no more than four or three, no more than three or two, no more than two or one, no more than one or zero occurrences, or no occurrences, of one or more sense codons; (e) occurrence of one or more sense codons at less than 10%, 5%, 2%, 1%, 0.5%, 0.1% compared to the parental prokaryotic genome from which the synthetic prokaryotic genome is derived;
2. 2. The synthetic prokaryotic genome of claim 1, wherein (c) 100 or more, 200 or more, or 300 or more genes that contain a total of five or four or fewer occurrences of one or more sense codons are essential genes.
3. 3. The synthetic prokaryotic genome of claim 1 or 2, which is a synthetic bacterial genome.
4. 4. The synthetic prokaryotic genome of claim 3, wherein the synthetic bacterial genome is a synthetic Escherichia coli genome, a synthetic Salmonella enterica genome, or a synthetic Shigella dysenteriae genome.
5. 5. The synthetic prokaryotic genome of claim 1, wherein the one or more sense codons consist of one sense codon or two sense codons.
6. 5. The synthetic prokaryotic genome of claim 1, wherein the one or more sense codons consist of two sense codons.
7. 7. A synthetic prokaryotic genome according to any one of claims 1 to 6, which does not contain two or more occurrences of sense codons and does not contain one occurrence of a stop codon.
8. 7. The synthetic prokaryotic genome of any one of claims 1 to 6, which does not contain any occurrences of two sense codons and does not contain any occurrences of an amber stop codon (TAG).
9. 8. The synthetic prokaryotic genome of any of claims 1 to 7, wherein one or more sense codons are selected from TCG, TCA, TCT, TCC, AGT, AGC, GCG, GCA, GCT, GCC, CTG, CTA, CTT, CTC, TTG, and TTA.
10. 8. The synthetic prokaryotic genome of any of claims 1 to 7, wherein one or more sense codons are selected from TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA.
11. 8. The synthetic prokaryotic genome of any preceding claim, wherein one or more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG and GCA.
12. 8. The synthetic prokaryotic genome of any one of claims 1 to 7, wherein one or more sense codons are TCG and / or TCA.
13. 13. The synthetic prokaryotic genome of any of claims 1 to 12, comprising (e) and wherein 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of one or more sense codons in the parental prokaryotic genome are replaced with synonymous sense codons, and / or 90%, 95%, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCA in said parental prokaryotic genome are replaced with AGT.
14. 13. The synthetic prokaryotic genome of any of claims 1 to 12, comprising (e) and wherein 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCG and / or TCA in the parental prokaryotic genome are replaced with AGC and / or AGT, and / or 90%, 95%, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCA in said parental prokaryotic genome are replaced with AGT.
15. (e) and 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more of the occurrence of TCGs in the parent prokaryotic genome; 13. The synthetic prokaryotic genome of any of claims 1 to 12, wherein 99.8% or more, 99.9% or more, or 100% of occurrences of TCA in the parental prokaryotic genome are replaced with AGC and / or 90%, 95%, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of occurrences of TCA in the parental prokaryotic genome are replaced with AGT.
16. 16. The synthetic prokaryotic genome of any of claims 1-15, comprising (e) and wherein 99.9% or more, or 100% of occurrences of two or more sense codons in the parental prokaryotic genome are replaced with synonymous sense codons, and all occurrences of TAG in said parental prokaryotic genome are replaced with TAA.
17. 16. The synthetic prokaryotic genome of any of claims 1-15, comprising (e) and wherein 99.9% or more, or 100% of the occurrences of two sense codons in a parental prokaryotic genome are replaced with synonymous sense codons, and all occurrences of TAG in said parental prokaryotic genome are replaced with TAA.
18. 18. The synthetic prokaryotic genome of any one of claims 1 to 17, comprising (e) and wherein one or more pairs of genes sharing an overlapping region containing one or more sense codons in the parental prokaryotic genome are refactored.
19. 18. The synthetic prokaryotic genome of any of claims 1-17, comprising (e) and wherein one or more gene pairs that share an overlapping region containing one or more sense codons in a parental prokaryotic genome, wherein substitution of said sense codons with one or more synonymous sense codons alters the encoded protein sequence of both or one of the gene pairs.
20. 20. The synthetic insert of claim 18 or 19, wherein for a pair of genes in an inverted orientation, a synthetic insert is inserted between the genes, said synthetic insert comprising an overlapping region, and / or for a pair of genes in the same orientation, a synthetic insert is inserted between the genes, said synthetic insert comprising (i) a stop codon, (ii) about 20-200 bp upstream of the overlapping region, and (iii) the overlapping region. Prokaryotic genomes.
21. A synthetic prokaryotic genome according to any one of claims 1 to 20, which is viable.
22. A polynucleotide comprising 20 or 21 or more, 30 or 31 or more, 40 or 41 or more, 50 or 51 or more, 100 or 101 or more essential genes, wherein one or more sense codons are absent.
23. The essential genes are ribF, lspA, ispH, dapB, folA, imp, yabQ, ftsL, ftsI, murE, murF, mraY, murD, ftsW, murG, murC, ftsQ, ftsA, fts Z, lpxC, secM, secA, can, folK, heml, yadR, dapD, map, rpsB, tsf, pyrH, frr, dxr, ispU, cdsA, yaeL, yaeT, lpxD, fabZ, lpx A, lpxB, dnaE, accA, tilS, proS, yafF, hemB, secD, secF, ribD, ribE, thiL, dxs, ispA, dnaX, adk, hemH, lpxH, cysS, folD, entD, mrdB, mrdA, nadD, holA, rlpB, leuS, lnt, glnS, fldA, cydA, infA, cydC, ftsK, lolA, serS, rpsA, msbA, lpxK, kdsB, m ukF、mukE、muk3、asn3、fabb。、mvii、 e、fabD、fabbァ、acp / 、tmk、hol3、 lolD、lolE、pur3、minn、min、、pth、pr s。、ispE、lolB、hem。、prf。、prm3、kdd A、top。、ribb。、fabゥ、tyrr″ 、、、、、、、、、、、、、、、、、、 、、、、、、、、 gap。、yeaaコaasppウ、aarggウ、pgss| et1、foll・、yejjュ、gyrr。、nrdd。、nrrd3、ffol C、accD、faab3、glt8、liggA、ziip。、dappE 、dap。、der、his3、isp1、suh3、tadd。、a cpa、era、rnc、lep3、rpoE、psssss。y 、rplm、trm、、rps.、ffh、grrpE、cssrr。、i pヲ、isppD、ftss、eno、pyr1、chhp2、lggt、 fbaa。、pgkkayqgg、、mmttォ、yqggヲ、pls3、yg iエ、parE、ribb3、cca、ygjD、tdcP、yrraャ、yhb6、inf3、nuss|、ftssィ、obgg・、rpmm| plオ、issp、murr。、yrbb「、yrbbォ、yhbb、 sゥ、rplュ、deggウ、mreD、mreC、mree3、accc 、accC、yrdC、def、fmt、rplアアrpo。、r psD、rpsss、rpsss、seccケ、rplッ、rpmm、、rrpss Earplイ、rpllヲ、rpss 2、rpsョ、rppl・、rrpl️ 、Sh1201、Sh1213,Sh1213、Sh1213,Sh166、 rps「、rpl、、rplW、rplD、rplC、rpss4、ff uss。、rps1、rpsss、ttrp3、yrfPaassdrrpoィ、five. gps。。rfaaa、kdtt。、coaa、、rpm3、dfpadu t、gmk、spoo4、gyyr3、dnaョ、dna!| rnppA、yidC、tna3、glmm3、glmm5、wzyyE、he mD、hemC、yigP、ubiB、ubii、、hem1、yiih A、ftsNmmurゥ、mur™、birr| 、Sh144、Shmfl,、Shmfl3、Shmfl3、Shiff!! lex。、dna、ssb、alss,、groo、psd、orrn、23. The polynucleotide of claim 22, comprising an essential gene selected from one or more of the list consisting of yjeE, rpsR, chpS, ppa, valS, yjgP, yjgQ, and dnaC.
24. A prokaryotic host cell comprising the synthetic prokaryotic genome of any one of claims 1 to 21 or the polynucleotide of claim 22 or 23.
25. 25. Use of the prokaryotic host cell of claim 24 for producing a polypeptide comprising one or more non-proteinogenic amino acids.
26. 25. Use of the prokaryotic host cell of claim 24 for producing a polypeptide comprising two or more non-proteinogenic amino acids.
27. 25. Use of the prokaryotic host cell of claim 24 for producing a polypeptide comprising three or more non-proteinogenic amino acids.
28. 1. A method for producing a synthetic prokaryotic genome, comprising: (a) providing a parent prokaryotic genome; (b) performing one or more rounds of recombination-mediated genetic modification on the parental prokaryotic genome to produce two or more different partially synthetic prokaryotic genomes; (c) performing one or more rounds of directed conjugation with said two or more different partial synthetic prokaryotic genomes to produce a synthetic prokaryotic genome; wherein each of said partially synthetic prokaryotic genomes comprises one or more sense codons or wherein each of said partially synthetic prokaryotic genomes comprises a synthetic region having less than 50 or 49, 20 or 19, 10 or 9, 5 or 4, or 0 occurrences of each of one or more sense codons compared to the corresponding region in said parent prokaryotic genome.
29. 1. A method for producing a synthetic prokaryotic genome, comprising: (a) providing a parent prokaryotic genome; (b) performing one or more rounds of recombination-mediated genetic modification on the parental prokaryotic genome to produce two or more different partially synthetic prokaryotic genomes; (c) performing one or more rounds of directed conjugation with said two or more different partial synthetic prokaryotic genomes to produce a synthetic prokaryotic genome; wherein each of said partial synthetic prokaryotic genomes comprises a synthetic region having 50 or 49 or less, 20 or 19 or less, 10 or 9 or less, 5 or 4 or less, or 0 occurrences of each of one or more sense codons, or each of said partial synthetic prokaryotic genomes comprises a synthetic region having less than 10%, 5%, 2%, 1%, 0.5%, 0.1% of said occurrences of each of one or more sense codons compared to the corresponding region in said parental prokaryotic genome; the composite regions collectively cover 90% or more, 95% or more, 99% or more, or 100% of the parent prokaryotic genome; the synthetic region is 10-1000 kb, 50-1000 kb, 100-1000 kb, or 100-500 kb in size; the viability of the partially synthetic prokaryotic genome is tested after each round of recombination-mediated genetic modification and / or after each round of induced conjugation; the two or more different partially synthetic prokaryotic genomes comprise at least one partially synthetic donor prokaryotic genome and at least one partially synthetic recipient prokaryotic genome; the at least one partially synthetic donor prokaryotic genome comprises a first selectable marker flanked by two regions of homology immediately downstream of the synthetic region and the origin of transfer, and the at least one partially synthetic recipient prokaryotic genome comprises a second selectable marker flanked by two corresponding regions of homology; or 29. The method of claim 28, wherein the synthetic prokaryotic genome is a synthetic prokaryotic genome according to any one of claims 1 to 21.
30. 1. A method for producing a synthetic prokaryotic genome, comprising: (a) providing a parent prokaryotic genome; (b) performing one or more rounds of recombination-mediated genetic modification on the parental prokaryotic genome to produce two or more different partially synthetic prokaryotic genomes; (c) performing one or more rounds of directed conjugation with the two or more different partial synthetic prokaryotic genomes to produce synthetic prokaryotic genomes, wherein each of the partial synthetic prokaryotic genomes comprises a synthetic region having no more than 50 or 49, no more than 20 or 19, no more than 10 or 9, no more than 5 or 4, or no more than 0 occurrences of each of one or more sense codons, or each of the partial synthetic prokaryotic genomes has, compared to a corresponding region in the parental prokaryotic genome: a synthetic region having less than 10%, 5%, 2%, 1%, 0.5%, 0.1% of said occurrences of each of one or more sense codons; the composite regions collectively cover 90% or more, 95% or more, 99% or more, or 100% of the parent prokaryotic genome; the synthetic region is 10-1000 kb, 50-1000 kb, 100-1000 kb, or 100-500 kb in size; the viability of the partially synthetic prokaryotic genome is tested after each round of recombination-mediated genetic modification and / or after each round of induced conjugation; the two or more different partially synthetic prokaryotic genomes comprise at least one partially synthetic donor prokaryotic genome and at least one partially synthetic recipient prokaryotic genome; the at least one partially synthetic donor prokaryotic genome comprises a first selectable marker flanked by two regions of homology immediately downstream of the synthetic region and the origin of transfer, the at least one partially synthetic recipient prokaryotic genome comprises a second selectable marker flanked by two corresponding regions of homology, and the first selectable marker comprises a positive selectable marker; and / or the second selectable marker comprises a negative selectable marker; and / or the method further comprises one or more rounds of selection for the selectable marker; Alternatively, the synthetic prokaryotic genome is a synthetic prokaryotic genome according to any one of claims 1 to 21; the method according to claim 29.
31. 1. A method for producing a synthetic prokaryotic genome, comprising: (a) providing a parent prokaryotic genome; (b) performing one or more rounds of recombination-mediated genetic modification on the parental prokaryotic genome to produce two or more different partially synthetic prokaryotic genomes; (c) performing one or more rounds of directed conjugation with said two or more different partial synthetic prokaryotic genomes to produce a synthetic prokaryotic genome; wherein each of said partial synthetic prokaryotic genomes comprises a synthetic region having 50 or 49 or less, 20 or 19 or less, 10 or 9 or less, 5 or 4 or less, or 0 occurrences of each of one or more sense codons, or each of said partial synthetic prokaryotic genomes comprises a synthetic region having less than 10%, 5%, 2%, 1%, 0.5%, 0.1% of said occurrences of each of one or more sense codons compared to the corresponding region in said parental prokaryotic genome; the composite regions collectively cover 90% or more, 95% or more, 99% or more, or 100% of the parent prokaryotic genome; the synthetic region is 10-1000 kb, 50-1000 kb, 100-1000 kb, or 100-500 kb in size; the viability of the partially synthetic prokaryotic genome is tested after each round of recombination-mediated genetic modification and / or after each round of induced conjugation; the two or more different partially synthetic prokaryotic genomes comprise at least one partially synthetic donor prokaryotic genome and at least one partially synthetic recipient prokaryotic genome; The at least one partially synthetic donor prokaryotic genome comprises a first selectable marker flanked by two regions of homology immediately downstream of the synthetic region and the origin of transfer, the at least one partially synthetic recipient prokaryotic genome comprises a second selectable marker flanked by two corresponding regions of homology, and the first selectable marker comprises a positive selectable marker. and / or said second selectable marker comprises a negative selectable marker; said synthetic region present in said at least one partially synthetic recipient prokaryotic genome is outside of the region flanked by said homologous regions; and / or said method further comprises one or more rounds of selection for said selectable marker; or The synthetic prokaryotic genome is a synthetic prokaryotic genome according to any one of claims 1 to 21; 31. The method of claim 30.
Citation Information
Patent Citations
Genome editing
WO2018020248A1