Methods of editing nucleic acid sequences

WO2026202389A1PCT designated stage Publication Date: 2026-10-01CONSTRUCTIVE BIOLOGY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/059039
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure EP2026059039_01102026_PF_FP_ABST
    Figure EP2026059039_01102026_PF_FP_ABST
Patent Text Reader

Abstract

In an aspect, the present invention relates to methods of introducing a sequence of interest into a target nucleic acid. The invention also relates to methods of assembling nucleic acid sequences comprising iterating the methods of introducing a sequence of interest into a target nucleic acid, as well as assembly of replicons encoding larger amounts heterologous nucleic acid, using systems free of CRISPR / Cas9.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS OF EDITING NUCLEIC ACID SEQUENCES

[0002] FIELD OF THE INVENTION

[0003] In an aspect, the present invention relates to methods of introducing a sequence of interest into a target nucleic acid. The invention also relates to methods of assembling nucleic acid sequences comprising iterating the methods of introducing a sequence of interest into a target nucleic acid.

[0004] BACKGROUND OF THE INVENTION

[0005] Strategies for replacing genomic DNA with synthetic DNA1-12enable genome engineering and provide a basis for powerful technologies to create entirely synthetic genomes that cannot be created by editing methodologies alone. Genome synthesis has been used to create synthetic genomes including in E. coli (4Mb), where it has been used to create a recoded organism13. The work in E. coli removed over 18,000 synonymous codons, to create an organism with a compressed genetic code which uses just 61 codons to encode the canonical amino acids. The recoded E. coli with a compressed genetic code provides a foundation for the creation of virus resistant cells, and for sense codon reassignment for non-canonical amino acid incorporation and encoded non-canonical polymer synthesis14.

[0006] The E. coli genome synthesis was based on REXER (Replicon Excision Enhanced Recombination)5. In REXER (Fig. la) a bacterial artificial chromosome (BAC) - containing an insert composed of the synthetic DNA of interest flanked by a double selection cassette (composed of a positive and negative selection marker) and 50-152 bp regions of homology to the genome - is transformed into cells containing a helper plasmid that encodes CRISPR / Cas9 and the lambda red recombination machinery. A single clone containing the correct BAC is isolated and made competent and then a spacer array is introduced into the competent cells to activate Cas9 mediated excision of the insert containing the synthetic DNA from the BAC. The spacer sequences are designed such that the spacer RNAs base pair with the unique sequences in the homology region of the insert and precisely cleave at the junctions between the insert and the BAC backbone (Fig. lb). The precisely excised insert is then used, by the lambda red recombination machinery, to insert the synthetic DNA into the genome in place of the corresponding genomic DNA. Distinct double selection cassettes in the genome and synthetic DNA are used to select for the desired integration (Fig. la). It takes four days to go from cells with an appropriately marked genome to having clonal colonies on a post-REXER agar plate.

[0007] A single step of REXER has been used to replace up to 136kb of the genome with synthetic, recoded DNA. The genome produced from one step of REXER provides a template for the next round of REXER, and iteration of REXER (Genome interchange stepwise synthesis, GENESIS) enables largersections of the E. col genome to be replaced with synthetic DNA. 38 REXER steps - each requiring the design, synthesis, cloning and validation of bespoke spacer pairs - were used to replace the entire E. coli genome (across seven strains) with synthetic recoded DNA. The recoded DNA was then compiled, by conjugation, into a single strain to create a recoded organism.

[0008] REXER uses an RNA-guided DNA endonuclease to introduce a double strand break and excise the target DNA. The excision site is located in the homologous recombination region, as noted above. Since the homologous recombination sequence is defined, the scope for introducing recognition sequences for other nucleases is very limited, and although various nucleases are mentioned in WO2018020248, none are employed.

[0009] A development of REXER, CONEXER, used conjugative transfer to iterate the steps of REXER and facilitate creation of larger episomes comprising recoded DNA (Zurcher et al., Nature. 2023 Jun 28;619(7970): 555- 562). CONEXER also placed the excision sites of the insert in the backbone of the vector, such that the excised DNA always includes one or more nucleotides derived from backbone sequence. This permits excision of the insert from the episomal replicon using the same pair of CRISPR spacers - ‘universal spacers’ - which can be used to perform REXER at any target locus with a given episomal replicon backbone. This massively simplifies the introduction of synthetic DNA into a target nucleic acid, such as an E. coli genome, and enables accelerated methods for large-scale genome engineering and whole genome synthesis. The arrangement of spacers that bind within the backbone and minimize the distance of the cleavage site from the junction between the backbone and the insert leads to the excision of an insert flanked by 6 base pairs of the backbone on each end (Fig. 1c). These 6 bp sequences are not commonly homologous to the region of the target nucleic acid flanking the regions complementary to the insert. The presence of short, 6 bp non-homologous blunt ends on the excised synthetic DNA does not impede recombination and integration efficiency, and scarless integration is still achieved.

[0010] However, CONEXER is dependent on CRISPR / Cas9 and requires the provision of spacer RNAs and Cas9 protein to enable the process. Spacer arrays are typically separated by repetitive regions, which increases the likelihood of construct collapse by homologous recombination, adding technical uncertainty to the process. A system not dependent on spacer RNAs would be simpler to implement and more robust in operation.

[0011] SUMMARY OF THE INVENTION

[0012] The invention provides a method which introduces a sequence of interest into a target nucleic acid using a host cell system. The host cell includes an episomal replicon comprising a backbone sequence and a donor nucleic acid sequence. The donor nucleic acid sequence contains, in order from 5' to 3': a first homologous recombination sequence, the sequence of interest, and a second homologous recombinationsequence. The backbone sequence includes first and second excision sites positioned adjacent to the first and second homologous recombination sequences respectively.

[0013] The method involves providing helper proteins capable of supporting nucleic acid recombination in the host cell, along with an endonuclease having a recognition sequence which is not present in the host cell system apart from the first and second excision sites. The endonuclease recognizes the first and second excision sites. The donor nucleic acid sequence can be excised by the endonuclease, and the excised donor nucleic acid may then recombine with the target nucleic acid during an incubation period. In an aspect of the invention, there is provided a method of introducing a sequence of interest into a target nucleic acid, the method comprising

[0014] a) providing a host cell

[0015] said host cell comprising an episomal replicon,

[0016] said episomal replicon comprising a backbone sequence and a donor nucleic acid sequence, wherein said donor nucleic acid sequence comprises in order: 5 ’ - homologous recombination sequence 1 - sequence of interest - homologous recombination sequence 2 - 3’,

[0017] wherein the backbone sequence comprises a first excision site positioned adjacent to homologous recombination sequence 1 and a second excision site positioned adjacent to homologous recombination sequence 2,

[0018] said host cell further comprising a target nucleic acid;

[0019] b) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;

[0020] c) providing an endonuclease having a recognition sequence which is not present in the host cell system apart from the first and second excision sites and which recognises the first and second excision sites;

[0021] e) inducing excision of said donor nucleic acid sequence by the endonuclease; and f) incubating to allow recombination between the excised donor nucleic acid and said target nucleic acid.

[0022] Typically, such an endonuclease will have a recognition sequence at least 12 to 18 nucleotides in length, and up to 20 to 40 nucleotides in length.

[0023] The endonuclease may be selected from various types including meganucleases or homing endonucleases, zinc finger nucleases (ZFNs), and transcription activator-like effector nucleases (TALENs). In some implementations, the endonuclease may be a LAGLID ADG homing endonuclease (LHE), such as I-Scel, I-Crel, I-Dmol or I-Anil. In certain cases, I-Scel may be selected as theendonuclease. The excised nucleic acid may include nucleic acid sequence derived from the backbone sequence at each terminus.

[0024] In an embodiment, each terminus of the excised nucleic acid comprises nucleic acid sequence derived from the backbone sequence. The excised donor nucleic acid may preferably include 18, 17, 16, 15, 14, 13 or fewer base pairs of nucleic acid sequence derived from the backbone sequence at each terminus. The ends may be blunt ends, or comprise staggered ends with overhangs of 6, 5, 4, 3, 2 or 1 nucleotides. The excision sites can be selected from endonuclease cut sites, which may be either the same or different from each other. These excision sites may be arranged in the same orientation or in opposite orientations relative to each other. In one embodiment, the backbone sequence may comprise 9 and 13 bases, with a 4bp overhang.

[0025] In an embodiment, the episomal replicon is a bacterial artificial chromosome. The episomal replicon may be delivered to the host cell by conjugative transfer.

[0026] The target nucleic acid may be the genome of the host cell. The host cell may be a prokaryotic cell, such as Escherichia coli.

[0027] In an aspect of the invention, there is provided a method of assembling a nucleic acid sequence, the method comprising:

[0028] (i) performing the steps of a method of introducing a sequence of interest into a target nucleic acid of the invention, to introduce a first donor nucleic acid sequence into a first target nucleic acid in order to create a second target nucleic acid; and

[0029] (ii) performing the steps of a method of introducing a sequence of interest into a target nucleic acid of the invention, to introduce a second donor nucleic acid sequence into the second target nucleic acid in order to create a third target nucleic acid. Part (i) and part (ii) may be iterated.

[0030] Excision of the sequence of interest from the episomal replicon to provide the donor nucleic acid is preferably performed by endonuclease digestion using an endonuclease with a recognition sequence of at least 16 nucleotides, as set forth in the first aspect of the invention.

[0031] Integration into the target nucleic acid is performed by homologous recombination.

[0032] In an embodiment, the method further comprises:

[0033] (iii) performing the steps of a method of introducing a sequence of interest into a target nucleic acid of the invention, to introduce a third donor nucleic acid sequence into the third target nucleic acid in order to create a fourth target nucleic acid;

[0034] iterating parts (i), (ii), and (iii), and wherein

[0035] the sequence of the first recognition site for part (iii) is the same for each iteration and / or the sequence of the second recognition site for part (iii) is the same for each iteration.In a particular embodiment, part (i) comprises the use of a donor-nucleic-acid-sequence-encoding episomal replicon comprising a first backbone sequence, and part (ii) comprises the use of a donor-nucleic-acid-sequence-encoding episomal replicon comprising a second backbone sequence, wherein the first backbone sequence comprises a first marker or set of markers, the first excision site within said first backbone sequence, and encodes the second excision site within said first backbone sequence; and

[0036] the second backbone sequence comprises a second marker or set of markers, the first excision site within said second backbone sequence, and the second excision site within said second backbone sequence; wherein

[0037] the first marker or set of markers is different from the second marker or set of markers.

[0038] In a further aspect of the invention, there is provided a method for constructing an episomal replicon, comprising the steps of:

[0039] a) providing a donor episomal replicon, said replicon comprising:

[0040] a first homology region HRn which is specific for an integration step n, and a second, universal, homology region uHR,

[0041] a first excision site positioned adjacent to HRn and a second excision site positioned adjacent to uHR;

[0042] a donor nucleic acid DNAn positioned between HRn and uHR; and

[0043] a double selection cassette, comprising positive and negative selection markers;

[0044] b) providing a host cell comprising an assembly episomal replicon comprising a double selection cassette comprising positive and negative selection markers, flanked by HRn and uHR, the double selection cassette in the assembly replicon comprising different markers to the selection cassette in the donor replicon;

[0045] c) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;

[0046] d) providing an endonuclease specific for the first excision site and the second excision site, having a recognition sequence which is not present in the host cell system apart from the first and second excision sites;

[0047] e) inducing excision of said donor nucleic acid sequence DNAn by the endonuclease in the host cell; and

[0048] f) incubating to allow recombination between the excised donor nucleic acid and said assembly replicon to form a second assembly replicon, which comprises the nucleic acid DNAn.The endonuclease may be selected from various types including meganucleases or homing endonucleases, zinc finger nucleases (ZFNs), and transcription activator-like effector nucleases (TALENs). In some implementations, the endonuclease may be a LAGLID ADG homing endonuclease (LHE), such as I-Scel, I-Crel, I-Dmol or I-Anil. In certain cases, I-Scel may be selected as the endonuclease. The excised nucleic acid may include nucleic acid sequence derived from the backbone sequence at each terminus.

[0049] In an embodiment, each terminus of the excised nucleic acid comprises nucleic acid sequence derived from the backbone sequence. The excised donor nucleic acid may comprise 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 or fewer base pairs of nucleic acid sequence derived from the backbone sequence at each terminus. Preferably, the excised donor nucleic acid comprises 13 or fewer base pairs of nucleic acid sequence derived from the backbone sequence at each terminus.

[0050] In an embodiment, the episomal replicon is a bacterial artificial chromosome. The episomal replicon may be delivered to the host cell by conjugative transfer.

[0051] In a preferred aspect of the invention, the donor episomal replicon is comprised in a donor host cell and the cell assembly replicon is comprised in a recipient host cell. Conjugation between the donor and recipient host cells can advantageously transfer the donor episomal replicon to the recipient host cell. The donor host cell preferably comprises a non-transferrable F’ plasmid, such that the F’- plasmid is not transferred to the recipient host cell. Preferably, the F’ plasmid is non-transferable through oriT deletion. The donor episome is transferrable, and may contain an oriT.

[0052] Selection for the host cell can be accomplished as described, but employing positive and negative selection markers present in recombined donor and assembly replicon DNA.

[0053] The target nucleic acid may be the genome of the host cell. The host cell may be a prokaryotic cell, such as Escherichia coli.

[0054] In one embodiment, the donor nucleic acid comprises a homology region HRn+1, and the method further comprises a further step (g) of introducing into the host cell a further donor episomal replicon comprising a donor nucleic acid DNAn+1, and inducing excision of said donor nucleic acid sequence DNAn+1 by the endonuclease in the host cell; and incubating to allow recombination between the excised donor nucleic acid DNAn+1 and said second assembly replicon to form a third assembly replicon, which comprises the nucleic acid DNAn and nucleic acid DNAn+1.

[0055] Said method steps can be iteratively repeated.

[0056] The replicon provided in this aspect of the invention may be used in the foregoing aspects which describe methods of assembling a nucleic acid sequence.In an advantageous embodiment of all aspects of the present invention, the host cell is lacking competent recA and / or recO. Preferably, the host cell lacks recA (ArecA).

[0057] BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Fig. 1 | The steps in REXER-mediated integration of ~100 kb of synthetic DNA into the E. coli genome using homology region (HR)-specific spacers, and the relationship between the cuts directed by HR-specific spacers and universal spacers.

[0059] a, REXER allows integration of more than 100 kb of synthetic DNA (pink) into the genome, either through replacement of the genomic DNA (as shown here) or by insertion into the genome. A bacterial artificial chromosome (BAC) containing the synthetic DNA of interest is electroporated into competent cells with a suitably marked genome, the cells also contain a helper plasmid encoding the Cas9 protein and the lambda red recombination components. The cell is then induced with arabinose to express the helper plasmid genes and made electrocompetent again. HR-specific spacer arrays (either plasmid based as shown, or as linear DNA) are then electroporated into the cell, leading to CRISPR / Cas9 mediated in vivo excision of the synthetic DNA flanked by a double selection cassette and HRs to the genome, from the BAC. The lambda red recombination machinery then uses the HRs to direct the integration of the excised DNA into the genome. Triangles denote the Cas9 cleavage sites at the HRs (grey boxes) flanking the synthetic DNA. +1, blue is karf1,' -1, yellow is rpsL,' +2, green is cat, -2, pink is sacB,' +3, dark blue is tefi,' -3, purple is pheS* +4, orange is ampP,' b, Previously, we designed spacer RNAs specific for each HR flanking the genomic locus for recombination. The BAC sequence flanking the synthetic DNA insert contains a constant PAM sequence (black box). Directing Cas9 with HR-specific spacer RNA allows precise excision of the synthetic DNA at the ends of HR1 and HR2. c, Universal spacer RNAs direct Cas9 to the constant sequence of the BAC backbone. To create a universal cut site for any BAC independent of the REXER locus, Cas9 is directed to the PAM sequences directly flanking the HRs. This adds an additional 6 bp non- homologous sequences to both ends of the excised DNA fragment.

[0060] FIG. 2 | Methods for introducing DNA sequences into a target nucleic acid using bacterial artificial chromosomes, according to an embodiment.

[0061] BAC stepwise insertion synthesis (BASIS) for iterative assembly of large DNA in BACs.

[0062] a, The donor BAC contains HRn and uHR, homologous to the recipient BAC. The BAC backbone contains oriT and universal spacers. uHR is a universal homology region for all steps of insertion. HRn is specific for the nth step of insertion. The BAC insert contains HRn+1, which serves as HRn for the (n+l)th step. The BAC contains a double selection cassette, -3, +3 shown. The assemblyBAC contains a distinct double selection cassette, -1,+1 shown, flanked by HRn and uHR. This DNA insert is excised from the donor BAC and inserted into the assembly BAC in the recipient cell. Green triangles indicate cut sites for Cas9 excision. Note in the main text HRn is described as HR1. In the example shown, the selectable markers are +1, blue, kanR’. -1, yellow, rpsl . +3, purple, hy roR’. -2, orange, PheS +4, petrol, GentamycinR.

[0063] b, BASIS workflow. The donor BAC is delivered by conjugation to the recipient cell

[0064] containing the assembly BAC and expressing Cas9 and lambda red components. The insert is excised from the donor BAC and inserted into the recipient BAC, as shown in (a). Iteration of this process, using alternating sets of markers, allows for the insertion of n DNA fragments into the assembly BAC.

[0065] FIG. 3 | Schematic of CONEXER-mediated replacement of genomic fragment 100k24 with a synthetic recoded counterpart, and different molecular arrangements for the donor BAC and helper plasmids.

[0066] a, Schematic depiction of a CONEXER experiment for the replacement of -100 kbp of wild-type DNA in the E. coli genome with a synthetic counterpart, where protein coding genes have been recoded by systematically replacing TCG, TCA and TAG codons by synonymous AGC, AGT and TAA codons, respectively. The donor cell harbours a BAC encoding -100 kbp of synthetic recoded DNA spanning LS23 to LS24, flanked by homology regions and a dual positive / negative selection cassette. The BAC contains an additional, distinct negative selection marker in the BAC backbone, and an oriT sequence to enable its conjugative transfer. The recipient cell contains a genomically integrated dual positive / negative selection marker at position LS23, and a helper plasmid encoding an endonuclease for cleaving the BAC and lambda red recombination components. Conjugative transfer of the donor BAC into the recipient cell is followed by induction of expression of the endonuclease and recombination factors encoded in the helper plasmid, to promote recombination. Cells that undergo successful recombination are selected for using various selective pressures as indicated in the figure, b, Schematic depiction of the different donor BAC arrangements tested in this study. BAC_100k24_vl is as described in Zurcher 2023. BAC_100k24_v2 and BAC_100k24_v3 do not contain spacer sequences in the backbone, and the protospacer sites for Cas9 cleavage are replaced with I-Scel recognition sites, either in the same or in the opposite orientation, c, Schematic depiction of the different helper plasmids used in this study. pHelper_Cas9 is as described in Zurcher 2023. pHelper l-Scel contains lambda red recombination components and an I-Scel gene, all under control of an arabinose-inducible promoter, d, Graphical depiction of the various selectable markers used in the CONEXER process.FIG. 4 | Recoding landscape plots showing the frequency of recoding at genomic loci, according to aspects of the present disclosure

[0067] Clonal recoding landscapes for post-CONEXER clones in the different conditions described in this study, across the genomic region spanning from LS23 to LS24. For each clone, the sequencing depth across the region of interest is shown (Read depth). The Recoding status indicates whether, for all targeted codons for synonymous replacement in the region, the alleles detected in sequencing correspond to the recoded or the wild-type (WT) codon. Fully recoded clones show recoded alleles across the entire region of interest.

[0068] DETAILED DESCRIPTION

[0069] The methods disclosed in the prior art, such as REXER, require a new set of homology region (HR)-specific spacers to be cloned for each locus that is targeted; these spacers can be challenging to clone, and spacer cloning can be expensive and time consuming. For instance, the recent E. coli genome synthesis required the cloning of 78 unique spacers. Each new set of spacers must be designed to avoid undesired cutting of the target nucleic acid. Additionally, varying the spacer sequence can affect the excision efficiency and this may contribute to variation in the efficiency of REXER at distinct genomic loci. The requirement for HR-specific spacer RNA complicates the workflow and may limit the scalability of REXER.

[0070] CONEXER avoids this drawback by placing the excision sites in the episome backbone. The use of universal spacers specific for the backbone sequence demonstrated that these spacers can be used for scarless integration of synthetic DNA into a target nucleic acid. The presence of short non-homologous ends on the excised synthetic DNA did not impede recombination and integration efficiency, and scarless integration was achieved. It was suggested that the non-homologous ends of the DNA in the BAC may be removed by exonucleases prior to recombination, or by flap endonucleases such as EcoIX during recombination, similar to the mechanism described for FEN1 in eukaryotes. However, the CONEXER method is limited to the use of RNA-guided DNA exonucleases (CRISPR), and thus still requires the provision of specific RNA spacer sequences in the host cell. Such sequences comprise repeats and can introduce undesired recombination events.

[0071] The term "meganuclease" as used herein refers to an endonuclease that recognizes and cleaves DNA at specific recognition sequences typically ranging from 14 to 40 base pairs in length. Meganucleases include, but are not limited to, LAGLID ADG homing endonucleases (LHEs) such as I-Scel, I-Crel, I-Dmol, and I-Anil. These enzymes are characterized by their ability to recognize and cleave long DNA target sites with high specificity, making them useful tools for precise genetic modifications in various organisms.The present disclosure relates to methods for introducing a sequence of interest into a target nucleic acid and methods for assembling nucleic acid sequences. These methods provide efficient and flexible approaches for genetic engineering and genome synthesis.

[0072] A method for introducing a sequence of interest into a target nucleic acid comprises providing a host cell containing an episomal replicon. The episomal replicon includes a backbone sequence and a donor nucleic acid sequence. The donor nucleic acid sequence comprises, in order from 5' to 3': a first homologous recombination sequence, the sequence of interest, and a second homologous recombination sequence. The backbone sequence includes first and second excision sites positioned adjacent to the first and second homologous recombination sequences, respectively.

[0073] The method further involves providing helper proteins capable of supporting nucleic acid recombination in the host cell. An endonuclease having a recognition sequence not present in the host cell system apart from the first and second excision sites is also provided. The endonuclease recognizes and cleaves at the first and second excision sites.

[0074] Excision of the donor nucleic acid sequence is induced by the endonuclease. The excised donor nucleic acid may then recombine with a target nucleic acid during an incubation period. This allows for precise integration of the sequence of interest into the target nucleic acid.

[0075] The method may be iterated to assemble larger nucleic acid sequences. Multiple rounds of introducing donor nucleic acid sequences into progressively modified target nucleic acids allow for the stepwise assembly of complex synthetic DNA constructs.

[0076] In some examples, the host cell may be a reduced-genome Escherichia coli strain. The use of a reduced-genome strain can provide advantages such as increased genomic stability and simplified genetic manipulation.

[0077] The endonuclease employed in the method may be selected from various types, including meganucleases, zinc finger nucleases, or transcription activator-like effector nucleases. The choice of endonuclease can be tailored to the specific requirements of the genetic engineering application. The episomal replicon used to deliver the donor nucleic acid sequence may be a bacterial artificial chromosome or other suitable vector. In some examples, the episomal replicon may be delivered to the host cell by conjugative transfer.

[0078] These methods provide powerful tools for genome engineering, enabling the precise introduction of synthetic DNA sequences and the assembly of large DNA constructs. The approaches described herein may be applied to create custom synthetic genomes, engineer biosynthetic pathways, or study gene function through targeted modifications.

[0079] The presence of short non-homologous ends on the excised synthetic DNA did not impede recombination and integration efficiency, and scarless integration was achieved. It was suggested thatthe non-homologous ends of the DNA in the BAC may be removed by exonucleases prior to recombination, or by flap endonucleases such as EcoIX during recombination, similar to the mechanism described for FEN1 in eukaryotes. However, the CONEXER method is limited to the use of RNA-guided DNA exonucleases (CRISPR), and thus still requires the provision of specific RNA spacer sequences in the host cell.

[0080] The term "meganuclease" as used herein refers to an endonuclease that recognizes and cleaves DNA at specific recognition sequences typically ranging from 14 to 40 base pairs in length. Meganucleases include, but are not limited to, LAGLID ADG homing endonucleases (LHEs) such as I-Scel, I-Crel, I-Dmol, and I-Anil. These enzymes are characterized by their ability to recognize and cleave long DNA target sites with high specificity, making them useful tools for precise genetic modifications in various organisms.

[0081] The present disclosure relates to methods for introducing a sequence of interest into a target nucleic acid and methods for assembling nucleic acid sequences. These methods provide efficient and flexible approaches for genetic engineering and genome synthesis.

[0082] A method for introducing a sequence of interest into a target nucleic acid comprises providing a host cell containing an episomal replicon. The episomal replicon includes a backbone sequence and a donor nucleic acid sequence. The donor nucleic acid sequence comprises, in order from 5' to 3': a first homologous recombination sequence, the sequence of interest, and a second homologous recombination sequence. The backbone sequence includes first and second excision sites positioned adjacent to the first and second homologous recombination sequences, respectively.

[0083] The method further involves providing helper proteins capable of supporting nucleic acid recombination in the host cell. An endonuclease having a recognition sequence not present in the host cell system apart from the first and second excision sites is also provided. The endonuclease recognizes and cleaves at the first and second excision sites.

[0084] Excision of the donor nucleic acid sequence is induced by the endonuclease. The excised donor nucleic acid may then recombine with a target nucleic acid during an incubation period. This allows for precise integration of the sequence of interest into the target nucleic acid.

[0085] The method may be iterated to assemble larger nucleic acid sequences. Multiple rounds of introducing donor nucleic acid sequences into progressively modified target nucleic acids allow for the stepwise assembly of complex synthetic DNA constructs.

[0086] In some examples, the host cell may be a reduced-genome Escherichia coli strain. The use of a reduced-genome strain can provide advantages such as increased genomic stability and simplified genetic manipulation.The endonuclease employed in the method may be selected from various types, including meganucleases, zinc finger nucleases, or transcription activator-like effector nucleases. The choice of endonuclease can be tailored to the specific requirements of the genetic engineering application. The episomal replicon used to deliver the donor nucleic acid sequence may be a bacterial artificial chromosome or other suitable vector. In some examples, the episomal replicon may be delivered to the host cell by conjugative transfer.

[0087] These methods provide powerful tools for genome engineering, enabling the precise introduction of synthetic DNA sequences and the assembly of large DNA constructs. The approaches described herein may be applied to create custom synthetic genomes, engineer biosynthetic pathways, or study gene function through targeted modifications.

[0088] The host cell system for introducing a sequence of interest into a target nucleic acid may comprise various types of cells. Prokaryotic cells such as Escherichia coli are commonly used as host cells due to their rapid growth and ease of genetic manipulation. Other bacterial species may also serve as suitable host cells. In some examples, eukaryotic cells such as yeast or mammalian cells may be employed as host cells.

[0089] The episomal replicon used in the method comprises a backbone sequence and a donor nucleic acid sequence. The backbone sequence provides the structural framework and replication elements for the episomal replicon. The donor nucleic acid sequence contains the genetic elements to be introduced into the target nucleic acid.

[0090] Various types of episomal replicons may be utilized, including bacterial artificial chromosomes (BACs) and plasmids. BACs can accommodate large DNA inserts, including inserts of over 1 Mbp (Zurcher et al., 2023) making them suitable for introducing long sequences. Plasmids are smaller circular DNA molecules that can replicate independently of the host chromosome.

[0091] The donor BAC may be derived from natural F plasmids or Pl phage backbones, and may contain molecular components that facilitate their maintenance and stable propagation across cell divisions. The BAC may contain partitioning systems that ensure each daughter cell receives at least one copy. BACs may contain the parABS system, which enables active segregation. ParB proteins bind to parS centromere-like sequences on the plasmid, forming nucleoprotein complexes. ParA is a Walker-type ATPase that forms dynamic gradients on the nucleoid, pulling ParB-bound plasmids to opposite cell poles before division. BACs derived from F plasmids may contain the sopABC system (an equivalent mechanism), where SopB binds sopC sites and SopA provides the motor function.

[0092] The BAC may contain toxin-antitoxin systems that provide post-segregational killing to enforce plasmid maintenance. Examples of toxin-antitoxin systems that may be present in BACs and related F plasmids include the ccdAB system and the parED system. In some cases, ccdA and ccdB genes maybe mutated. Mutation of these genes may alter their function to enhance stability or reduce potential toxicity to the host cell.

[0093] An origin of transfer (oriT) sequence may be included in the donor BAC to enable transfer by conjugation. The oriT sequence serves as the initiation site for DNA transfer during bacterial conjugation, allowing the BAC to be transferred between cells.

[0094] The BAC may contain an alternative origin (oriS) for rolling-circle replication during conjugative transfer.

[0095] The BAC may contain copy-control mechanisms involving antisense RNAs, such as the copA / copB system, ensuring the plasmid doesn't over-replicate and destabilize.

[0096] The BAC may contain multimer resolution systems to facilitate the distribution of plasmid copies across daughter cells, or encode sequences that facilitate multimer resolution by host-encoded systems. For example, the BAC backbone may encode a psi site, which may be recognized by host XerCD dimer resolution systems. A dimer resolution system comprising a resD gene with an rfsF cut site may be included in the donor BAC. The resD gene encodes a site-specific recombinase that can resolve plasmid dimers, while the rfsF site serves as the recognition sequence for this recombinase. This system can help maintain plasmid stability by preventing the accumulation of multimeric forms.

[0097] The backbone sequence of the episomal replicon includes first and second excision sites positioned adjacent to the homologous recombination sequences of the donor nucleic acid. These excision sites serve as recognition sequences for the endonuclease used to excise the donor nucleic acid sequence. The donor nucleic acid sequence comprises, in order from 5' to 3': a first homologous recombination sequence, the sequence of interest, and a second homologous recombination sequence. The homologous recombination sequences facilitate integration of the sequence of interest into the target nucleic acid through homologous recombination.

[0098] In some examples, the episomal replicon may contain selectable markers to allow for identification and selection of cells containing the replicon. Antibiotic resistance genes or auxotrophic markers may be used for this purpose.

[0099] The structure and components of the episomal replicon can be tailored to the specific requirements of the genetic engineering application. Factors such as insert size, host cell compatibility, and desired genetic manipulations may influence the choice and design of the episomal replicon.

[0100] The donor nucleic acid sequence comprises three main components arranged in a specific order: a first homologous recombination sequence, the sequence of interest, and a second homologous recombination sequence. This arrangement facilitates the precise integration of the sequence of interest into the target nucleic acid through homologous recombination.The homologous recombination sequences are regions of DNA that share sequence similarity with specific locations in the target nucleic acid. These sequences serve as recognition sites for the cellular recombination machinery, enabling the alignment and exchange of genetic material between the donor and target nucleic acids. The length and composition of the homologous recombination sequences can affect the efficiency of the recombination process.

[0101] In designing effective homologous recombination sequences, several factors may be considered: 1. Sequence length: Homologous recombination sequences typically range from 30 to 100 base pairs in length. Longer sequences may increase recombination efficiency but can be more challenging to synthesize and manipulate.

[0102] 2. Sequence uniqueness: The homologous recombination sequences may be selected to be unique within the target nucleic acid to promote specific integration at the desired location.

[0103] 3. GC content: A balanced GC content in the homologous recombination sequences may help maintain DNA stability and facilitate efficient recombination.

[0104] 4. Absence of repetitive elements: Avoiding repetitive DNA sequences in the homologous recombination regions may reduce the likelihood of non-specific recombination events.

[0105] 5. Proximity to target site: Positioning the homologous recombination sequences close to the intended integration site in the target nucleic acid may enhance recombination efficiency.

[0106] The sequence of interest, located between the two homologous recombination sequences, contains the genetic information to be introduced into the target nucleic acid. This sequence may encode proteins, regulatory elements, or other functional genetic components. The size of the sequence of interest can vary depending on the specific application and the capacity of the chosen episomal replicon.

[0107] In some examples, the sequence of interest may include selectable markers or reporter genes to facilitate the identification of successful recombination events. These markers may be flanked by site-specific recombination sequences to allow for their subsequent removal if desired.

[0108] The design of the donor nucleic acid sequence may be optimized to enhance the efficiency of the recombination process and the stability of the integrated sequence. Considerations may include codon optimization for the host organism, removal of unwanted restriction sites, and inclusion of transcriptional terminators or insulator sequences to prevent interference with neighboring genes in the target nucleic acid.

[0109] The excision sites in the episomal replicon backbone are positioned adjacent to the homologous recombination sequences of the donor nucleic acid. This arrangement allows for precise excision of the donor nucleic acid sequence, including the homologous recombination sequences, by the chosenendonuclease. The excised fragment can then participate in homologous recombination with the target nucleic acid.

[0110] In some examples, the donor nucleic acid sequence may be designed to introduce multiple genetic elements simultaneously. This can be achieved by including additional homologous recombination sequences within the donor sequence, separating distinct genetic elements to be integrated at different locations in the target nucleic acid.

[0111] The structure and composition of the donor nucleic acid sequence can be tailored to suit specific genetic engineering applications, such as gene insertion, deletion, or replacement. By carefully designing the homologous recombination sequences and the sequence of interest, researchers can achieve precise and efficient genetic modifications in a wide range of host organisms.

[0112] The host cell system for introducing a sequence of interest into a target nucleic acid may comprise various types of cells. Prokaryotic cells such as Escherichia coli are commonly used as host cells due to their rapid growth and ease of genetic manipulation. Other bacterial species may also serve as suitable host cells. In some examples, eukaryotic cells such as yeast or mammalian cells may be employed as host cells.

[0113] The episomal replicon used in the method comprises a backbone sequence and a donor nucleic acid sequence. The backbone sequence provides the structural framework and replication elements for the episomal replicon. The donor nucleic acid sequence contains the genetic elements to be introduced into the target nucleic acid.

[0114] The design of the donor nucleic acid sequence may be optimized to enhance the efficiency of the recombination process and the stability of the integrated sequence. Considerations may include codon optimization for the host organism, removal of unwanted restriction sites, and inclusion of transcriptional terminators or insulator sequences to prevent interference with neighboring genes in the target nucleic acid.

[0115] The excision sites in the episomal replicon backbone are positioned adjacent to the homologous recombination sequences of the donor nucleic acid. This arrangement allows for precise excision of the donor nucleic acid sequence, including the homologous recombination sequences, by the chosen endonuclease. The excised fragment can then participate in homologous recombination with the target nucleic acid.

[0116] In some examples, the donor nucleic acid sequence may be designed to introduce multiple genetic elements simultaneously. This can be achieved by including additional homologous recombination sequences within the donor sequence, separating distinct genetic elements to be integrated at different locations in the target nucleic acid.The structure and composition of the donor nucleic acid sequence can be tailored to suit specific genetic engineering applications, such as gene insertion, deletion, or replacement. By carefully designing the homologous recombination sequences and the sequence of interest, researchers can achieve precise and efficient genetic modifications in a wide range of host organisms.

[0117] The method employs various types of endonucleases to excise the donor nucleic acid sequence from the episomal replicon. Meganucleases, also known as homing endonucleases, recognize and cleave long DNA target sites, typically 14-40 base pairs in length. Examples of meganucleases include LAGLIDADG homing endonucleases (LHEs) such as I-Scel, I-Crel, I-Dmol, and I-Anil. These enzymes exhibit high specificity for their recognition sequences, making them useful for precise genetic modifications.

[0118] Zinc finger nucleases (ZFNs) represent another class of endonucleases that can be used in the method. ZFNs consist of a DNA-binding domain derived from zinc finger proteins fused to a DNA cleavage domain, typically from the FokI restriction enzyme. The DNA-binding domain can be engineered to recognize specific DNA sequences, allowing for targeted cleavage at desired locations.

[0119] Transcription activator-like effector nucleases (TALENs) function similarly to ZFNs but utilize a different DNA-binding domain. TALENs incorporate a DNA-binding domain derived from transcription activator-like effectors (TALEs) fused to a nuclease domain. The TALE DNA-binding domain can be customized to recognize specific DNA sequences, enabling precise targeting of genomic loci.

[0120] Excision sites in the backbone sequence of the episomal replicon serve as recognition sequences for the chosen endonuclease. These sites are positioned adjacent to the homologous recombination sequences of the donor nucleic acid. The strategic placement of excision sites allows for precise excision of the donor nucleic acid sequence, including the flanking homologous recombination sequences, upon cleavage by the endonuclease.

[0121] A recipient cell may contain a helper plasmid encoding an endonuclease, such as I-Scel, under the control of an inducible promoter. This arrangement allows for controlled expression of the endonuclease, enabling temporal regulation of the excision process. The inducible promoter may respond to specific stimuli, such as chemical inducers or environmental cues, to initiate endonuclease expression at the desired time point.

[0122] To modulate the expression levels of the endonuclease, different ribosome binding site (RBS) sequences may be used upstream of the endonuclease gene, such as I-Scel. The RBS plays a role in translation initiation, and variations in RBS sequence can affect the efficiency of protein synthesis. By employing different RBS sequences, the expression level of the endonuclease can be fine-tuned to optimize the excision process and minimize potential toxicity to the host cell.The choice of endonuclease and the design of excision sites may be tailored to the specific requirements of the genetic engineering application. Factors such as recognition sequence length, cleavage efficiency, and potential off-target effects may influence the selection of an appropriate endonuclease system for a given experimental setup.

[0123] The method involves providing helper proteins capable of supporting nucleic acid recombination in the host cell. These helper proteins play a crucial role in facilitating the integration of the excised donor nucleic acid sequence into the target nucleic acid.

[0124] One commonly used system for supporting recombination is the lambda Red recombination system. This system comprises three main components: Exo, Beta, and Gam. Exo is a 5' to 3' exonuclease that generates single-stranded DNA overhangs. Beta is a single-stranded DNA binding protein that promotes annealing of complementary DNA strands. Gam inhibits the host RecBCD exonuclease, protecting linear DNA fragments from degradation.

[0125] The lambda Red recombination system may be expressed from a helper plasmid under the control of an inducible promoter, such as the arabinose-inducible pBAD promoter. This allows for temporal control of recombination activity, minimizing potential negative effects on host cell growth and viability.

[0126] Alternative recombination systems may also be employed, such as the RecET system from the Rac prophage. The RecET system functions similarly to the lambda Red system, with RecE acting as an exonuclease and RecT as a single-stranded DNA annealing protein.

[0127] The recombination process begins with the excision of the donor nucleic acid sequence from the episomal replicon by the chosen endonuclease. The endonuclease recognizes and cleaves at the excision sites flanking the donor sequence, releasing a linear DNA fragment containing the sequence of interest and the homologous recombination sequences.

[0128] Once excised, the linear donor DNA fragment is protected from degradation by the host cell's nucleases, either through the action of Gam in the lambda Red system or by other mechanisms in alternative recombination systems.

[0129] The helper proteins then facilitate the integration of the excised donor sequence into the target nucleic acid through homologous recombination. This process involves several steps:

[0130] 1. The exonuclease (Exo in lambda Red or RecE in RecET) processes the linear donor DNA, generating 3' single-stranded overhangs at the homologous recombination sequences.

[0131] 2. The single-stranded DNA binding protein (Beta in lambda Red or RecT in RecET) binds to the exposed single-stranded regions, protecting them from degradation and promoting their interaction with complementary sequences in the target nucleic acid.3. The single-stranded overhangs of the donor DNA align with their complementary sequences in the target nucleic acid, forming heteroduplex structures.

[0132] 4. DNA replication and repair mechanisms of the host cell complete the integration process, resolving the heteroduplex structures and incorporating the sequence of interest into the target nucleic acid. The efficiency of the recombination process may be influenced by various factors, including the length and sequence composition of the homologous recombination regions, the size of the sequence of interest, and the expression levels of the helper proteins.

[0133] In some examples, additional proteins may be provided to enhance recombination efficiency or specificity. For instance, single-stranded DNA binding proteins from other organisms or engineered variants of recombination proteins may be employed to optimize the process for specific applications or host organisms.

[0134] The recombination process may be further modulated by controlling the expression of host cell recombination and repair proteins. For example, in some host strains, the recA gene, which encodes a key protein in homologous recombination, may be deleted or inactivated to reduce unwanted recombination events and improve the specificity of the desired integration.

[0135] Following the recombination process, cells that have successfully integrated the sequence of interest may be selected using appropriate markers or screening methods. This may involve the use of antibiotic resistance genes, auxotrophic markers, or fluorescent reporters incorporated into the donor sequence. The described recombination process allows for precise and efficient integration of the donor nucleic acid sequence into the target nucleic acid, enabling a wide range of genetic engineering applications, from single gene modifications to large-scale genome engineering projects.

[0136] The method for introducing a sequence of interest into a target nucleic acid involves several key steps that can be optimized and varied for different applications.

[0137] First, a host cell is provided containing an episomal replicon, such as a bacterial artificial chromosome (BAC), which comprises a backbone sequence and a donor nucleic acid sequence. The donor nucleic acid sequence contains the sequence of interest flanked by homologous recombination sequences. In some examples, the donor BAC is delivered to the recipient cell by conjugative transfer. This process involves direct cell-to-cell contact and transfer of the BAC from a donor cell to the recipient cell. Conjugative transfer can be an efficient method for introducing large DNA constructs into bacterial cells.

[0138] After conjugation, the method may include an additional outgrowth phase before inducing excision of the donor sequence. This outgrowth period allows the recipient cells to recover and begin expressingany necessary genes from the transferred BAC. The duration of the outgrowth phase can be optimized based on the specific host strain and BAC construct being used.

[0139] Next, helper proteins capable of supporting nucleic acid recombination are provided in the host cell. These may include components of recombination systems such as the lambda Red system or the RecET system. The expression of these helper proteins can be controlled using inducible promoters to optimize the timing of recombination events.

[0140] An endonuclease is then provided to recognize and cleave at specific excision sites in the BAC backbone. The choice of endonuclease can be tailored to the specific application. For example, meganucleases like I-Scel may be used due to their long recognition sequences and high specificity. Alternatively, engineered nucleases such as zinc finger nucleases or TALENs could be employed for increased flexibility in target site selection.

[0141] Excision of the donor nucleic acid sequence is induced by activating the endonuclease. This may involve inducing expression of the endonuclease gene or providing the protein directly. The timing and level of endonuclease activity can be optimized to maximize excision efficiency while minimizing potential toxicity to the host cell.

[0142] Following excision, the host cells are incubated under conditions that allow for recombination between the excised donor nucleic acid and the target nucleic acid. The duration and conditions of this incubation period can be adjusted based on the specific recombination system being used and the characteristics of the host cell.

[0143] Throughout the process, selection markers or screening methods may be employed to identify cells that have successfully integrated the sequence of interest. This can involve the use of antibiotic resistance genes, auxotrophic markers, or fluorescent reporters.

[0144] The efficiency and specificity of this method can be further optimized by adjusting factors such as the length and composition of the homologous recombination sequences, the size of the sequence of interest, and the expression levels of helper proteins and endonucleases. Additionally, the use of host strains with modifications to endogenous recombination pathways may enhance the desired integration events while reducing unwanted recombination.

[0145] Figure 1 illustrates some key steps in a related process for integrating synthetic DNA into a bacterial genome. While the specific details differ, this figure demonstrates the general concept of using a bacterial artificial chromosome to deliver synthetic DNA, excising the DNA of interest, and integrating it into a target genome through recombination.

[0146] The method of introducing a sequence of interest into a target nucleic acid can be iterated to assemble larger nucleic acid sequences. This iterative process allows for the stepwise construction of complex synthetic DNA constructs.Figure 2 illustrates key components and steps involved in the assembly of larger nucleic acid sequences. The figure shows a donor BAC 1, an assembly BAC 2, and a recipient cell 3.

[0147] In a first iteration, a donor BAC 1 containing a first donor nucleic acid sequence is introduced into a recipient cell 3 containing an assembly BAC 2. The donor nucleic acid sequence is excised from the donor BAC 1 and recombines with the assembly BAC 2 through homologous recombination sequences. This results in the integration of the first donor sequence into the assembly BAC 2.

[0148] For subsequent iterations, a new donor BAC 1 containing a second donor nucleic acid sequence may be introduced into the recipient cell 3. The second donor sequence is then excised and recombines with the assembly BAC 2 that now contains the first donor sequence. This process can be repeated multiple times to progressively assemble larger nucleic acid constructs within the assembly BAC 2.

[0149] The efficiency of the assembly process may be optimized through several strategies:

[0150] 1. Design of homologous recombination sequences: The length and sequence composition of the homologous regions can be optimized to enhance recombination efficiency. Longer homology arms may increase integration efficiency but may also increase the complexity of BAC construction.

[0151] 2. Selection of endonuclease: The choice of endonuclease for excising donor sequences can impact the efficiency and specificity of the process. Meganucleases with long recognition sequences may provide high specificity.

[0152] 3. Expression control of recombination proteins: Fine-tuning the expression levels of helper proteins involved in recombination may optimize the assembly process. Inducible promoters can be used to control the timing and level of expression.

[0153] 4. Host strain engineering: Modifications to the recipient cell 3, such as deletion of certain recombination genes, may enhance the desired integration events while reducing unwanted recombination.

[0154] 5. Outgrowth conditions: Optimizing the duration and conditions of outgrowth periods between iterations can improve cell recovery and BAC stability.

[0155] 6. Selection strategies: Employing alternating selection markers between iterations can facilitate identification of successful integration events.

[0156] 7. Size of donor sequences: The size of individual donor sequences may be adjusted to balance assembly efficiency with the desired final construct size.

[0157] 8. BAC backbone design: Optimizing the backbone sequences of both donor BAC 1 and assembly BAC 2 can enhance stability and reduce unwanted recombination events.By iteratively applying the method and incorporating these optimization strategies, large synthetic DNA constructs can be assembled in a controlled and efficient manner. This approach enables the creation of complex genetic systems, synthetic genomes, or large-scale genome modifications.

[0158] The methods described herein have numerous applications in the fields of genome engineering and synthetic biology. These techniques enable precise and efficient introduction of synthetic DNA sequences into target genomes, facilitating a wide range of genetic modifications and engineering approaches.

[0159] In genome engineering applications, the methods may be used to introduce large synthetic DNA fragments into bacterial genomes. This allows for extensive genome modifications, including the replacement of multiple genes or entire genomic regions with synthetic sequences. Such capabilities are valuable for studying gene function, optimizing metabolic pathways, or creating organisms with novel properties.

[0160] The techniques described may be applied to create custom synthetic genomes. By iteratively introducing multiple synthetic DNA fragments, researchers can progressively replace large portions of a genome with designed sequences. This approach enables the construction of minimal genomes, recoded genomes with altered genetic codes, or genomes optimized for specific functions.

[0161] In synthetic biology applications, the methods provide tools for assembling and integrating large synthetic DNA constructs. This facilitates the design and implementation of complex genetic circuits, biosynthetic pathways, or artificial chromosomes. The ability to introduce and assemble large DNA sequences in a controlled manner supports the development of engineered organisms for biotechnology applications.

[0162] The techniques may be used for metabolic engineering of microorganisms. By introducing synthetic pathways or modifying existing metabolic networks, researchers can optimize the production of valuable compounds such as biofuels, pharmaceuticals, or industrial chemicals.

[0163] These methods also have potential applications in creating cell factories for the production of proteins or other biomolecules. The ability to precisely integrate large DNA sequences encoding biosynthetic pathways or protein expression systems can enhance the efficiency and yield of recombinant protein production.

[0164] In the field of functional genomics, the techniques allow for systematic genome-wide modifications. This can support large-scale studies of gene function, genetic interactions, or evolutionary processes. An advantage of the described methods is their independence from CRISPR / Cas9 systems. This characteristic provides flexibility in experimental design and may be beneficial in situations where CRISPR / Cas9 components are undesirable or incompatible with the host organism.The use of meganucleases or other site-specific endonucleases for donor DNA excision offers high specificity in targeting predetermined sequences. This can reduce the likelihood of off-target effects compared to some other genome editing approaches.

[0165] The methods described allow for the introduction of large DNA fragments, potentially exceeding 100 kilobases in size. This capability enables the engineering of complex genetic systems or the simultaneous modification of multiple genomic loci in a single step.

[0166] The iterative nature of the assembly process provides a modular approach to constructing large synthetic DNA sequences. This modularity allows for flexibility in design and facilitates the creation of diverse genetic constructs.

[0167] By utilizing bacterial artificial chromosomes (BACs) as vectors for donor DNA, the methods can accommodate large insert sizes while maintaining stability in bacterial hosts. This characteristic supports the manipulation and transfer of extensive genetic elements.

[0168] The techniques described may be adapted for use in various prokaryotic and potentially eukaryotic host organisms. This versatility expands the range of potential applications across different biological systems.

[0169] In summary, the methods presented offer powerful tools for genome engineering and synthetic biology applications. These techniques enable precise and large-scale genetic modifications, supporting diverse research and biotechnology endeavors in fields such as metabolic engineering, functional genomics, and synthetic genome construction.

[0170] Thus, in an embodiment of the invention, there is provided a method of introducing a sequence of interest into a target nucleic acid, the method comprising

[0171] a) providing a host cell

[0172] said host cell comprising an episomal replicon,

[0173] said episomal replicon comprising a backbone sequence and a donor nucleic acid sequence, wherein said donor nucleic acid sequence comprises in order: 5 ’ - homologous recombination sequence 1 - sequence of interest - homologous recombination sequence 2 - 3’,

[0174] wherein the backbone sequence comprises a first excision site positioned adjacent to homologous recombination sequence 1 and a second excision site positioned adjacent to homologous recombination sequence 2,

[0175] said host cell further comprising a target nucleic acid;

[0176] b) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;c) providing an endonuclease having a recognition sequence which is not present in the host cell system apart from the first and second excision sites, and which is specific for the excision sites;

[0177] d) inducing excision of said donor nucleic acid sequence by the endonuclease; and e) incubating to allow recombination between the excised donor nucleic acid and said target nucleic acid.

[0178] The steps a), b), and c) need not be performed in order and need not be performed as distinct steps. For instance, the endonuclease may be provided before the provision of the helper protein(s) capable of supporting nucleic acid recombination in said host cell. However endonuclease must be provided before the induction of excision. In addition, all helper protein(s) capable of supporting nucleic acid recombination are provided before the incubation to allow recombination. Steps a), b) and c) may be performed simultaneously.

[0179] The episomal replicon comprises a backbone sequence and a donor nucleic acid sequence, and the backbone sequence does not overlap with the donor nucleic acid sequence. As such, a constant backbone sequence may be used for the introduction of multiple different donor nucleic acid sequences. In embodiments where multiple donor nucleic acids are introduced into a target by iterating the methods of the invention, more than one type of backbone sequence may be used. For instance, backbones comprising different markers, such as selection markers, may be used such that the successful introduction of the sequence of interest can be identified at each step.

[0180] The first and second excision sites allow for cleavage of the episomal replicon to excise the donor nucleic acid sequence. The excision sites provide all nucleic acid sequence required for recognition and cleavage by the endonuclease, and so no part of homologous recombination sequence 1 or homologous recombination sequence 2 is recognised in order for cleavage to take place.

[0181] The excision sites are adjacent to the homologous recombination sequences of the donor nucleic acid. Preferably each excision site is contiguous with the homologous recombination sequence. In such embodiments, there are no base pairs in between the sequences required for excision and the homologous recombination sequence. In other embodiments 1, 2, 3, or 4 base pairs of intervening sequence may be tolerated.

[0182] Any backbone sequence between the site at which the episomal replicon is cleaved and the homologous recombination sequence will be present as a part of the excised nucleic acid. As such, the excised nucleic acid may comprise: a first portion of the backbone sequence, the donor nucleic acid, and a second portion of the backbone sequence. Thus, in an embodiment, each terminus of the excised nucleic acid comprises nucleic acid sequence derived from the backbone sequence.The first and second portion of backbone sequence may each be 13, 12, 11, 10, 9, 8, 7, 6 or fewer base pairs in length, in addition to any intervening sequence between the excision site and the homologous recombination sequence. In embodiments, the portions of backbone sequence may comprise singlestranded overhangs, or may be blunt ended. I-Scel leaves a 4 base overhang. It is believed that backbone sequences are removed during the homologous integration procedure, and multiple configurations are thus tolerated.

[0183] The methods of the invention employ an endonuclease which does not cleave the genome of the intended host cell. For example, such an endonuclease is an endonuclease having a 16 nucleotide recognition sequence, which statistically will not cut in the human genome (the recognition sequence should occur less than once). A nuclease having a recognition sequence of 12 nucleotides is statistically non-cutting in E. coli. A recognition sequence of 12 or more nucleotides, such as 13, 14, 15, 16, 17 or 18 or more nucleotides, is preferred. However, any length of recognition sequence can be employed, as long as it does not cleave the host cell genome or at unintended locations in the episomal replicon or any helper plasmids.

[0184] Homing endonucleases (also called meganucleases) recognize long DNA target sites (14-40 bp). Beyond I-Scel, examples include I-Crel, I-Dmol, I-Anil, I-Ceul, and others; see Belfort et al., Nucleic Acids Res. 1997;25(17):3379-3388 and Stoddard, Structure. 2011;19(1):7-15. Some homing endonucleases, such as I-Ceul, cleave wild-type E. coli, and thus cannot be used without prior removal of the cleavage sites from the E. coli genome. Others may be less than 100% sequence specific, and may tolerate one or more nucleotide changes in the recognition sequence. Preferably, a homing endonuclease is selected which is highly specific for a sequence which does not appear in the host cell. Homing endonucleases generally fall within one of four families, characterized by the sequence motifs LAGLID ADG, GIY-YIG, H-N-H and His-Cys box. LAGLID ADG homing endonucleases (LHE) are preferred.

[0185] Exemplary LHE nucleases are set forth in the table below:

[0186] Table 3

[0187] Enzyme Recognition Sequence (Top strand 5'— >3' Approx. Cut Site Length Source / Notes unless noted) of Site

[0188] l-Scel 5'-TAGGGATAACAGGGTAAT-3' (SEQ ID Cleaves near 18 bp From

[0189] NO: 11) 3'-ATCCCTATTGTCCCATTAA-5' center Saccharomyces (SEQ ID NO: 18) cerevisiae

[0190] (yeast).Enzyme Recognition Sequence (Top strand 5'— >3' Approx. Cut Site Length Source / Notes unless noted) of Site

[0191] l-Crel Recognizes a 22 bp sequence (consensus Center 22 bp From often cited as Chlamydomona 5'-CAAAACGTCCTGAAGACAGTTT-3') s reinhardtii. (SEQ ID NO: 12) Binds as a homodimer to a symmetric target.

[0192] l-Ppol 5'-CTCTTAAAGGTAGCCTGCC-3' (top Cleaves within site -15-16 From Physarum strand) (Often shown with a break after bp polycephalum. "AAGG") (SEQ ID NO: 13)

[0193] l-Dmol Recognizes a long 31 bp sequence(e.g., Internal cleavage 31 bp From

[0194] 5'-TCTCATGTTTAACTATGCTTTGCTGCC Desulfurococcu GAAT-3') (SEQ ID NO: 14) s mobilis.

[0195] l-Ceul Binds a 19 bp sequence in the 23S rRNA Internal cleavage -19 bp From gene (e.g., Chlamydomona 5'-TAACGGTTTACCTTTGAGT-3') (SEQ ID s eugametos. NO: 15)

[0196] l-Anil Often cited as recognizing -20 bp: Internal cleavage -20 bp From

[0197] 5'-GGGGTCTCTCTTCAAAACGT-3' (SEQ Aspergillus ID NO: 16) nidulans.

[0198] I-Tevl Recognizes -37 bp within a T4 phage Cleaves -22 bp -37 bp Phage T4 intron(e.g., downstream of homing 5'-GTACTTCCTCTGGGTTAAACGGCAAG binding site endonuclease AGATTTTTTTT-3') (SEQ ID NO: 17) (Gl intron).

[0199] l-Tevll Another T4 phage meganuclease Internal cleavage >30 bp Phage T4.

[0200] site(internal 38-40 bp region) Related to I- Tevl

[0201] l-Tevlll Recognizes a segment in T4 phage Internal cleavage -37-38 Third T4

[0202] genome(~37-38 bp) bp meganucleaseEnzyme Recognition Sequence (Top strand 5'— >3' Approx. Cut Site Length Source / Notes unless noted) of Site

[0203] l-Panl Recognizes ~24 bp region(Pyrococcus Internal cleavage 24 bp From species) PyrococcusAmino Acid MHMKNIKKNQVMN MNTKYNKEFLLYL MALTNAQILAVIDS MHNNENVSGISAYL Sequence LGPNSKLLKEYKSQL AGFVDGDGSIIAQIK WEETVGQFPVITHH LGLIIGDGGLYKLKY IELNIEQFEAGIGLIL PNQSYKFKHQLSLT VPLGGGLQGTLHC KGNRSEYRVVITQKS GDAYIRSRDEGKTY FQVTQKTQRRWFL YEIPLAAPYGVGFA ENLIKQHIAPLMQFLI CMQFEWKNKAYMD DKLVDEIGVGYVR KNGPTRWQYKRTI DELNVKSKIQIVKGD HVCLLYDQWVLSPP DRGSVSDYILSEIKP TRYELRVSSKKLYY NQVVHRWGSHTVP HKKERVNHLGNLVI YFANMLERIRLFNM LHNFLTQLQPFLKL FLLEPDNINGKTCT TWGAQTFKHQAFNK REQIAFIKGLYVAEG KQKQANLVLKIIEQ ASHLCHNTRCHNPL LANLFIVNNKKTIPN DKTLKRLRIWNKNK LPSAKESPDKFLEV HLCWESLDDNKGR NLVENYLTPMSLAY ALLEIVSRWLNNLG CTWVDQIAALNDS NWCPGPNGGCVHA WFMDDGGKWDYN VRNTIHLDDHRHGV KTRKTTSETVRAVL VVCLRQGPLYGPG KNSTNKSIVLNTQSF YVLNISLRDRIKFVH DSLSEKKKSSP ATVAGPQQRGSHF TFEEVEYLVKGLRN TILSSHLNPLPPERAG

[0204] (SEQ ID NO: 22) VV (SEQ ID NO: 23) KFQLNCYVKINKNK GYT (SEQ ID NO: 24) PIIYIDSMSYLIFYNL IKPYLIPQMMYKL PNTISSETFLK

[0205] (SEQ ID NO: 21)

[0206] Organism Saccharomyces Chlamydomonas Physarum Desulfurococcus cerevisiae (mt) reinhardtii polycephalum mobilis (Archaea) Name I-Scel I-Crel I-Ppol l-Dinol

[0207]

[0208]

[0209] An alternative to homing endonucleases is Zinc-Finger Nucleases (ZFNs). ZFNs are composed of A DNA-binding domain engineered from zine-finger motifs and a cleavage domain from the FokI restriction enzyme.

[0210] When two ZFNs bind opposite DNA strands, the FokI domains dimerize to induce a double-strand break. ZFNs can be customized to target a specific locus by engineering the zine-finger array. See Kim et al., Proc Natl Acad Sci U S A. 1996;93(3): 1156-1160; Umov et al., Nat Rev Genet.

[0211] 2010;ll(9):636-646.

[0212] Similar to ZFNs, Transcription Activator-Like Effector Nucleases (TALENs) fuse the catalytic domain of FokI to a TALE (transcription activator-like effector) DNA-binding domain. Each TALE repeat recognizes a single base, making the system more modular and often easier to engineer than zine-finger arrays. See Christian et al. Genetics. 2010; 186(2):757-761 ; Miller et al. Nat Biotechnol.

[0213] 2011;29(2): 143-148.

[0214] In some aspects, when employing ZFNs or TALENs in the methods of the invention, the backbone sequence of the episomal replicon may be designed to accommodate the dimerization requirement of the FokI cleavage domain. Because FokI functions as a dimer, two ZFN or TALEN monomers may bind to opposite DNA strands in a tail-to-tail orientation to enable cleavage. Accordingly, each excision site in the backbone sequence may comprise a pair of binding sites positioned on opposite strands, with appropriate spacing between them to permit FokI dimerization and cleavage.

[0215] The spacing between ZFN or TALEN binding sites may be selected to optimize cleavage efficiency. In some cases, a spacer region of approximately 5 to 7 base pairs between the two binding sites may be suitable for ZFNs, while TALENs may function with spacer regions of approximately 12 to 21 base pairs. The optimal spacing may vary depending on the specific ZFN or TALEN architecture and linker length employed.

[0216] In designing excision sites for ZFNs or TALENs, the binding sites may be positioned such that cleavage occurs adjacent to or within a short distance of the junction between the backbone sequence and the homologous recombination sequence. This arrangement may minimize the amount of backbone sequence retained on the excised donor nucleic acid. In some implementations, the excised donor nucleic acid may comprise backbone-derived sequences at each terminus, similar to embodiments employing meganucleases.

[0217] For ZFN-based implementations, each zinc finger module typically recognizes a 3 base pair DNA sequence, and arrays of 3 to 6 zinc fingers may be assembled to recognize target sequences of 9 to 18 base pairs per monomer. When two ZFN monomers are employed, the combined recognition sequence may span 18 to 36 base pairs or more, providing high specificity for the excision sites in the backbone sequence.For TALEN-based implementations, each TALE repeat recognizes a single nucleotide, allowing for modular assembly of DNA-binding domains that recognize target sequences typically ranging from 14 to 20 base pairs per monomer. The combined recognition sequence for a TALEN pair may thus span 28 to 40 base pairs or more, which may provide sufficient specificity to avoid cleavage at unintended sites in the host cell system.

[0218] In some embodiments employing ZFNs or TALENs, the same pair of nucleases may be used to cleave at both the first and second excision sites, provided that identical or substantially similar recognition sequences are incorporated at both sites in the backbone sequence. Alternatively, different ZFN or TALEN pairs may be employed for the first and second excision sites if distinct recognition sequences are desired.

[0219] The expression of ZFNs or TALENs in the host cell may be controlled using inducible promoters, similar to the control of meganuclease expression described herein. This temporal control may help minimize potential toxicity associated with nuclease expression and may allow for coordination of excision with other steps in the method.

[0220] Restriction enzymes may also be engineered to have specific recognition sequences. See Takeuchi et la., Proc Natl Aca d Sci USA. 2014;l 11(11):4061-4066.

[0221] The backbone of the episomal replicon may be engineered such that the recognition site for the endonuclease employed is positioned adjacent to the junction between the backbone and the homologous region of the donor nucleic acid. Generally, the length of overhang or the amount of backbone nucleic acid which is present on the excised DNA fragment is minimised. However, we shown that the presence of backbone nucleic acid does not prevent scarless integration of the donor nucleic acid into the target nucleic acid.

[0222] In some embodiments, the helper protein(s) capable of supporting nucleic acid recombination and / or the at least one endonuclease that is capable of cleaving the first and / or second excision site are encoded on the episomal replicon. In some embodiments, the helper protein(s) capable of supporting nucleic acid recombination and / or the at least one endonuclease that is capable of cleaving the first and / or second excision site are encoded on a separate episomal replicon, such as a plasmid. This separate episomal replicon may be known as a helper episomal replicon or a helper plasmid.

[0223] In some embodiments, one or more integrated stability systems are introduced into the donor BAC to enhance plasmid maintenance and faithful segregation. These systems may include multimer resolution systems, toxin-antitoxin modules, partitioning loci, or combinations thereof.

[0224] The BAC may contain a plasmid-encoded site-specific recombinase system to resolve plasmid dimers or multimers. For example, a resD gene with its cognate rfsF recognition sites may be introduced under a native or constitutive promoter. The ResD protein promotes resolution of plasmid cointegrates ordimers through site-specific, recA-independent recombination at the rfsF sequences. Two rfsF sites are typically required for efficient resolution.

[0225] Alternatively, or additionally, the BAC may contain recognition sites for host-encoded resolution systems, such as psi sites recognized by the chromosomal XerCD recombinase system, or cer-like sequences that facilitate dimer resolution through host machinery.

[0226] When a resD-based system is employed, the resD gene may be cloned together with toxin-antitoxin genes such as ccdA and ccdB, which may be part of the same operon or adjacent loci. The ccdAB genes encode a post-segregational killing system: CcdB is a stable toxin that targets DNA gyrase, while CcdA is a labile antitoxin. The ratio between CcdA and CcdB regulates the repression state of the ccd operon — if the level of CcdA is greater than or equal to CcdB, repression occurs, preventing the harmful effect of CcdB on the cell. Otherwise, derepression occurs, leading to toxicity in plasmid-free segregants.

[0227] In a preferred embodiment, mutated ccdA and / or ccdB sequences may be used. For example, a mutated ccdB sequence may be employed such that the CcdB protein is not expressed at all, or is expressed at reduced levels, ensuring safe propagation of the system while retaining the regulatory architecture of the operon. Alternatively, ccdA may be mutated to alter the CcdA: CcdB ratio or the binding affinity between antitoxin and toxin.

[0228] Other toxin-antitoxin systems may be substituted or used in combination, including parED, mazEF, or relBE systems, each providing post-segregational killing through distinct mechanisms.

[0229] The BAC may additionally or alternatively contain active partitioning systems to ensure proper segregation. For F-plasmid-derived BACs, this may include the sopABC locus, where SopB binds sopC centromere-like sites and SopA provides the motor function for active segregation. For other BAC backbones, the parABS system may be present, wherein ParB binds parS sites and ParA (a Walker-type ATPase) establishes dynamic gradients that pull plasmid copies to opposite cell poles prior to division.

[0230] Copy number regulation may be achieved through antisense RNA systems such as copA / copB, which control replication initiation frequency, or through iteron-based handcuffing mechanisms involving the replication initiator protein (e.g., RepE binding to iterons at oriV).

[0231] The nucleic acid sequences for multimer resolution systems, toxin-antitoxin modules, and partitioning loci may be obtained from publicly available plasmid sequences. For example, the resD gene and rfsF cut sites may be obtained from the complete sequence of the mini-F plasmid (GenBank accession number AP001918), which provides the complete annotated sequence including the dimer resolution locus. Similarly, sequences for parABS, sopABC, ccdAB, parED, and other stability determinants maybe obtained from annotated F plasmid sequences (e.g., GenBank accession AP001918 for mini-F, or other deposited F-factor sequences), Pl phage sequences, or from characterized low-copy-number plasmid backbones. The skilled person can readily identify and extract the relevant coding sequences and recognition sites from these references.

[0232] In a particular embodiment, there is provided a method comprising:

[0233] 1) providing a host cell

[0234] said host cell comprising an episomal replicon, such as a BAC,

[0235] said episomal replicon comprising a backbone sequence and a donor nucleic acid sequence, wherein said donor nucleic acid sequence comprises in order: 5 ’ - homologous recombination sequence 1 - sequence of interest - homologous recombination sequence 2 - 3’,

[0236] wherein the backbone sequence comprises a first excision site positioned adjacent to the homologous recombination sequence 1 and a second excision site positioned adjacent to the homologous recombination sequence 2, wherein the first excision site comprises and the second excision site comprise endonuclease recognition sequences,

[0237] said host cell further comprising a target nucleic acid, for instance the genome of the host cell;

[0238] 2) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;

[0239] 3) providing an endonuclease having a recognition sequence which is not present in the host cell system apart from the first and second excision sites, such as a I-Scel endonuclease;

[0240] 4) inducing excision of said donor nucleic acid sequence by the endonuclease; and 5) incubating to allow recombination between the excised donor nucleic acid and said target nucleic acid.

[0241] The steps 1), 2) and 3) need not be performed in order and need not be performed as distinct steps. For instance, the endonuclease may be provided before the provision of the helper protein(s) capable of supporting nucleic acid recombination in said host cell. However, all components required for excision are provided before the induction of excision. In addition, all helper protein(s) capable of supporting nucleic acid recombination are provided before the incubation to allow recombination. Steps 1), 2) and 3) may be performed simultaneously.

[0242] The episomal replicon may be provided to the host cell by conjugation. Thus, the methods may further comprise the step of delivering the episomal replicon to the host cell by conjugative transfer. The episomal replicon, such as a BAC, may be included in a donor cell for transfer. The episomal replicon may comprise an origin of tranfer (priT). The donor cell may comprise a non-transferable F’ plasmid.The methods of the invention are particularly suitable for iteration in order to assemble large synthetic nucleic acid sequences. For instance, for the construction of artificial genomes. Thus, in an aspect of the invention, there is provided a method of assembling a nucleic acid sequence, the method comprising:

[0243] (i) performing the steps of any of the methods of the invention to introduce a first donor nucleic acid sequence into a first target nucleic acid in order to create a second target nucleic acid; and (ii) performing the steps of any of the methods of the invention to introduce a second donor nucleic acid sequence into the second target nucleic acid in order to create a third target nucleic acid.

[0244] Parts (i) and (ii) may be iterated multiple times. This allows the introduction of a first donor nucleic acid, a second donor nucleic acid, a third donor nucleic acid, and potentially further donor nucleic acids. When the technique is iterated the product of one round of the method of the invention may act as a target nucleic acid sequence for the next round of nucleic acid introduction.

[0245] The nucleases used in the iterations may be the same or different. Since the recognition sequences are in the backbone DNA, the same endonuclease can be used and the recognition site can be constant throughout the procedure.

[0246] The backbone sequence of the episomal replicon may be different during part (i) and during part (ii). For instance, the episomal replicon may comprise a marker or markers to allow identification or selection of the successful introduction of the sequence of interest. In order to allow rounds of nucleic acid introduction that include identification or selection, the episomal replicon of part (i) may comprise a first marker or set of markers and the episomal replicon of part (ii) may comprise a second marker or set of markers. In embodiments where parts (i) and (ii) are iterated, this may mean that a first marker or set of markers is used for every odd numbered selection and a second marker or set of markers is used for every even numbered selection. In other embodiments, further markers, such as a third marker or set of markers, may be used as part of a pattern of iterations. To allow for selection, the marker or markers for each round of nucleic acid introduction should be different from the marker or markers used in the previous round.

[0247] The backbone sequence of the episomal replicon during part (i) may be a first backbone sequence comprising a first marker or set of markers and have a first pair of recognition sequences as described herein. The backbone sequence of the episomal replicon during part (ii) may be a second backbone sequence comprising a second marker or set of markers and a second pair of recognition sequences as described herein. This pattern may be maintained during iterations such that the first backbone sequence is present during every odd-numbered iteration and the second backbone sequence is present during every even-numbered iteration. The pattern of iterations may also include further backbone sequences, such as a third backbone sequence, comprising a further marker or set of markers and furtherpairs of recognition sequences. The first marker or set of markers and the second or further marker or set of markers are different from each other, to allow selection of each successful nucleic acid introduction during the rounds of recombination. The recognition sequences in each backbone sequence allow cleavage of the encoding backbone.

[0248] In a further aspect, the principles established herein can be extended to realize the scarless assembly and cloning, through iterative insertion, of megabases of DNA in episomes in E. coli. The invention thus concerns an assembly episomal replicon in which to iteratively insert and assemble DNA (Fig. 2a).

[0249] In embodiments, this replicon comprises approximately 50 bp of sequence homologous to one end of the next sequence to be inserted (HR1); this is immediately followed by apositive and negative selection cassette and a universal homology region (uHR), which is complementary to the other end of the sequence to be inserted. Longer homologies, as long as multiple 10s of kb of sequence, may also be employed.

[0250] The invention also provides donor episomal replicons with the CONEXER backbone, containing universal spacers and oriT (Fig. 2a). In the donor replicons, HR1 is within one end of the next DNA sequence to be inserted into the recipient BAC; this DNA sequence is followed by a distinct positive and negative selection cassette and a universal homology region. Each step of assembly (Fig. 2a, b) proceeds by conjugation of the donor replicon to recipient cells containing the assembly replicon, nuclease-mediated excision of the sequence from HR1 to uHR from the donor replicon, and recombination mediated insertion of this sequence into the assembly replicon. Selection for loss of the negative selection markers on the assembly replicon, and gain of the positive marker from the sequence excised from the donor replicon selects for cells containing the assembly replicon with the correct insertion. Cells containing the new assembly replicon provide the input for the next step of insertion. This approach was named BAC stepwise insertion synthesis (BASIS).

[0251] Therefore, in a further aspect of the invention, there is provided a method for constructing an episomal replicon comprising a plurality of assembly steps, wherein step n comprises the steps of:

[0252] a) providing a donor episomal replicon, said replicon comprising:

[0253] a backbone, said backbone comprising a first homology region HRn which is specific for an integration step n, and a second, universal, homology region uHR, a first excision site positioned adjacent to HRn and a second excision site positioned adjacent to uHR;

[0254] a donor nucleic acid DNAn, said donor nucleic acid comprising a homology region HRn+1, specific for an assembly step n+1;

[0255] a double selection cassette, comprising positive and negative selection markers;

[0256] b) providing a host cell comprising an assembly episomal replicon comprising a double selection cassette comprising positive and negative selection markers, flanked by HRn and uHR, thedouble selection cassette in the assembly replicon comprising different markers to the selection cassette in the donor replicon;

[0257] c) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;

[0258] c) providing an endonuclease without a recognition sequence in the host cell apart from the first and second excision sites;

[0259] d) inducing excision of said donor nucleic acid sequence DNAn by the endonuclease; and e) incubating to allow recombination between the excised donor nucleic acid and said assembly replicon to form a second assembly replicon, which comprises the nucleic acid DNAn. The assembly replicon carrying the donor nucleic acid can in turn be used as an assembly replicon in a second step, in which a second donor replicon comprising homology regions HRn+1 and uHR and a second donor nucleic acid DNAn+1 is employed to introduce a second donor nucleic acid into the assembly replicon generated in the first step.

[0260] Alternating positive and negative selection marker sets allow an infinite number of steps to be performed iteratively, assembling an episomal replicon of any desired size. Preferably, the number of steps performed may be at least 2, and 100 or less; 50 or less; 25 or less; and most preferably about 10, 9, 8, 7, 6 or 5.

[0261] The size of the replicon which is assembled is preferably between 1 and 100 Mb. Preferably, it is 2 to 50Mb, 3 to 25 Mb, 4 to 15Mb, or 5 to 10Mb.

[0262] Thus, in an embodiment, the invention comprises performing the method of the above aspect of the invention, further comprising the steps of introducing into the host cell a further donor episomal replicon comprising a donor nucleic acid DNAn+1, and inducing excision of said donor nucleic acid sequence DNAn by the RNA-guided DNA endonuclease in the host cell; and incubating to allow recombination between the excised donor nucleic acid DNAn+1 and said assembly replicon to form a second assembly replicon, which comprises the nucleic acid DNAn and nucleic acid DNAn+1.

[0263] The steps of this embodiment of the invention may be performed iteratively, inserting donor nucleic acids DNAn+2, n+3, n+4 etc into the assembly episomal replicon.

[0264] BASIS can be used to generate episomes or other DNA vectors or segments which are useful for continuous genome synthesis (CGS). BASIS can also be itself applied continuously, without sequencing steps, for the continuous production of artificial DNA, whether episomes, genomes, bacterial or other genes and DNA to processes for continuous genome synthesis (CGS).

[0265] BASIS can be used to assemble large sections of human genomic DNA, which includes exonic, intronic and intergenic regions, into a single episome.In order to allow rapid iteration of CONEXER by directly using an un-sequenced pool of clones from one CONEXER as the input for the next CONEXER, the removal of sequencing steps to validate each clone is desirable.

[0266] Continuous genome synthesis can be performed by directly using the output from one round of CONEXER - without identifying an individual, fully recoded clone by sequencing - as the input for the next round of CONEXER.

[0267] The host cell used in aspects of the present invention is advantageously lacking competent recA and / or recO. Preferably, the host cell lacks recA (ArecA).

[0268] EPISOMAL REPLICONS

[0269] Embodiments of the invention comprise an episomal replicon comprising a donor nucleic acid sequence. The donor nucleic acid may be DNA.

[0270] The term “episome” has its ordinary meaning in the art, for example any accessory extrachromosomal replicating genetic element that can exist either autonomously or can become integrated with the chromosome.

[0271] An episomal replicon is an episomal nucleic acid which possesses its own origin of replication capable of functioning within said host cell.

[0272] The episomal replicon may be a plasmid. A plasmid means a small circular nucleic acid (usually DNA, most usually double stranded DNA) molecule. A plasmid within a cell is physically separated from any chromosomal nucleic acid such as DNA and can replicate independently. Considering plasmids, “small” means they are typically no bigger than 10 kb. Suitably a plasmid useful in the invention has the following genetic elements: an origin of replication cognate for the host cell; and at least one selection marker.

[0273] The episomal replicon may be a BAC. The BAC may comprise the following genetic elements: an origin of replication cognate for the host cell; and at least one selection marker.

[0274] BACs and plasmids differ from each other by their replication origin. A BAC has a special replication origin which typically makes the BAC a single copy in each cell and helps the BAC to maintain a bigger size (up to several hundred kb). Plasmids have a plasmid replication origin which typically makes the plasmid multiple copies (ranging from a few copies to a few hundred copies per cell) in each cell and typically of a size up to around 10 kb.

[0275] The episomal replicon may be a yeast artificial chromosome (YAC). The YAC may comprise the following genetic elements: an origin of replication cognate for the host cell; and at least one selection marker.Multiple origins of replication active in the same cell on the same single nucleic acid are not usually desirable. This is especially true for example when a multicopy episomal nucleic acid such as a plasmid is carrying the donor nucleic acid - in this scenario it is clearly not desirable to incorporate the plasmid origin of replication into (for example) a BAC or into the host genome. Thus, suitably said excised linear donor nucleic acid does not comprise an origin of replication. Suitably the target nucleic acid sequence comprises an origin of replication.

[0276] Suitably the origin of replication on the episomal replicon comprising the donor sequence must match with host, e.g. all prokaryotic. Suitably the origins of replication on the episomal replicon comprising the target and on the episomal replicon comprising donor sequence must match with host, e.g. all prokaryotic.

[0277] Suitably the episomal replicon comprising the donor sequence comprises a prokaryotic origin of replication. Suitably the replicon comprising the target sequence comprises a prokaryotic origin of replication. Suitably the replicon comprising the target sequence is an episomal replicon and comprises a prokaryotic origin of replication. Suitably the host cell is prokaryotic. Suitably the synthetic genome is a synthetic prokaryotic genome.

[0278] TARGET NUCELIC ACID

[0279] The target nucleic acid may be any suitable for the introduction of the donor nucleic acid sequence. In particular, the target nucleic acid may be a DNA molecule suitable for the introduction of a donor DNA molecule.

[0280] The target nucleic acid may comprise homologous recombination sequence 1 and homologous recombination sequence 2. The target nucleic acid may comprise a selection marker or selection markers, which may be flanked by the homologous recombination sequences. For instance, the target nucleic acid may comprise a negative selection marker.

[0281] The target nucleic acid may be a region of a nucleic acid that also possess its own origin of replication capable of functioning within the host cell. The target nucleic acid may be a plasmid. The target nucleic acid may be a BAC. The target nucleic acid may be a YAC. In a particular embodiment, the target nucleic acid is the genome of the host cell.

[0282] When the invention is applied to a genome, suitably the genome is a non-human genome, suitably a non-mammalian genome. Suitably the genome is a prokaryotic genome, suitably a bacterial genome. In particular, the genome may be an E. coli genome.

[0283] In a particular embodiment, the episomal replicon comprising a donor DNA molecule is a BAC and the target nucleic acid is the genome of an E. coli cell.HOMOLOGOUS RECOMBINATION

[0284] In theory any nucleotide sequence can be chosen as the site for homologous recombination sequences. The nucleotide sequence for homologous recombination may be unique. For instance, the nucleotide sequence for homologous recombination may be unique within the target sequence into which the donor sequence is being recombined. In other examples, homologous recombination sequence 1 and / or homologous recombination sequence 2 may be unique within the target sequence into which the donor sequence is being recombined.

[0285] Alternatively, homologous recombination sequence 1 and / or homologous recombination sequence 2 may be not unique within the target sequence into which the donor sequence is being recombined. In such examples, selection may be used to identify the successful introduction of the sequence of interest into the desired site. For instance, off-target integration may not result in the removal or disruption of a negative selection marker, or off-target integration may not repair double-strand breaks induced at the site of introduction.

[0286] Suitably the sequence for homologous recombination is non-repetitive.

[0287] Suitably the sequence for homologous recombination is at least 30 nucleotides long. Homologous recombination sequences as short as 30 nucleotides may lead to a low efficiency; thus for high efficiency suitably the homologous recombination sequence is at least 40 nucleotides in length, suitably at least 50 nucleotides, suitably 50 to 100 nucleotides, most suitably 50 to 65 nucleotides.

[0288] The sequence for homologous recombination is selected on the target sequence and introduced into the donor sequence. Therefore, the homologous recombination sequence 1 (HR1) and homologous recombination sequence 2 (HR2) on the donor sequence show 100% sequence identity to the HR1 and HR2 on the target sequence.

[0289] Use of lambda Red recombination permits short nucleotide sequences to be used for homologous recombination, as outlined above. Other recombination support systems may be used. For example, the RecBCD system might be used. When the RecBCD system is used, suitably the step of “providing helper protein(s) capable of supporting nucleic acid recombination in said host cell” consists of inducing or permitting expression of the RecBCD system within the host cell.

[0290] When using the RecBCD system or other recombination support systems, the skilled operator will pay attention to the requirements of those systems on the sequences selected for homologous recombination. For example, the RecBCD system may require longer homologous recombination sequences such as 3 to lOkb in length.In more detail, RecBCD is a natural E. col recombination system consisting of three components RecB, RecC, and RecD. The three subunits make up an ATP -dependent helicase / nuclease complex that is essential for both homologous recombination during the course of transduction and conjugation as well as in repair of double-strand breaks in E. coli. Studies in which double strand breaks are induced in vivo in E. coli DNA show that double-strand break repair (DSBR) can proceed via one of two recombination pathways. Both pathways require RecBCD and RecA, but one depends on the resolvase enzyme, RuvABC, while the other does not and instead relies on RecG. The recB and recD genes form an operon while recC is situated nearby but has its own promoter. The three gene products form a heterotrimer which is also known as Exonuclease V. In case any further guidance is needed, details can be found in the publicly available EcoCyc database under ‘RecBCD’, for example for the K12-MG-1655 strain of E. coli (Keseler et al. (2013), “EcoCyc: fusing model organism databases with systems biology”, Nucleic Acids Research 41: D605-12).

[0291] Thus, to support recombination according to this embodiment, at least RecBCD should be expressed in the host cell.

[0292] It may be that RecA is also required; thus, more suitably to support recombination according to this embodiment, at least RecBCD and RecA should be expressed in the host cell. Most suitably to support recombination according to this embodiment, RecBCD and RecA should be expressed in the host cell. Another alternative to the lambda red system is the RecET system. RecE and RecT are E. coli genes of phage origin. RecE mimics lambda red alpha, and RecT mimics lambda red beta (Muyrers, J.P., Zhang, Y ., Buchholz, F. & Stewart, A. F. RecE / RecT and Redalpha / Redbeta initiate double-stranded break repair by specifically interacting with their respective partners. Genes Dev. 14, 1971-1982 (2000)). The RecET combination performs comparatively to lambda red alpha / beta combination. Lambda red alpha and beta are the actual components that carry out recombination, while lambda red gamma is an inhibitor of the RecBCD system.

[0293] Suitably recombination support is provided via the lambda red system, for example from the commercially available pRed / ET plasmid from Gene Bridges ("Quick & Easy E. coli Gene Deletion Kit" from Gene Bridges GmbH, Im Neuenheimer Feld 584, 69120 Heidelberg, Germany.).

[0294] This system in this setup is first described in Datsenko et al 2000 (Datsenko K. A. & Wanner, B. L. One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc. Natl. Acad. Sci. U.S.A. 97, 6640-6645 (2000)), which is hereby incorporated herein by reference specifically for details of the Lambda Red system.

[0295] The inventors teach that the pRed / ET plasmid is based on the pKD46 plasmid in Datsenko et al 2000 (as judged by sequence identity), and therefore the pKD46 plasmid may be used as a template to perform PCR for the construction of lambda red system.When said helper protein(s) capable of supporting nucleic acid recombination comprise lambda Red proteins, suitably the following proteins are expressed in said host cell:Table 1

[0296] Lambda Essential? Function / Exemplary sequence / accession number

[0297] notes

[0298] Red

[0299] protein

[0300] alpha essential SEQ ID NO: ATGACACCGGACATTATCCTGCAGCGTACCGGGATCGATG 1 TGAGAGCTGTCGAACAGGGGGATGATGCGTGGCACAAATT ACGGCTCGGCGTCATCACCGCTTCAGAAGTTCACAACGTG ATAGCAAAACCCCGCTCCGGAAAGAAGTGGCCTGACATGA AAATGTCCTACTTTCACACCCTGCTTGCTGAGGTTTGCAC CGGTGTGGCTCCGGAAGTTAACGCTAAAGCACTGGCCTGG

[0301] Lambda red GG AAAAC AGT ACGAGAAC GAG GCC AGAACCC T GT T T GAAT TCACTTCCGGCGTGAATGTTACTGAATCCCCGATCATCTA

[0302] recombinatio TCGCGACGAAAGTATGCGTACCGCCTGCTCTCCCGATGGT

[0303] n

[0304] TTATGCAGTGACGGCAACGGCCTTGAACTGAAATGCCCGT TTACCTcccgggATTTCATGAAGTTCCGGCTCGGTGGTTT CGAGGCCATAAAGTCAGCTTACATGGCCCAGGTGCAGTAC AGCATGTGGGTGACGCGAAAAAATGCCTGGTACTTTGCCA ACTATGACCCGCGTATGAAGCGTGAAGGCCTGCATTATGT C GT GAT T GAGC GGGAT GAAAAGT AC AT GGC GAGT T T T GAG GAGATCGTGCCGGAGTTCATCGAAAAAATGGACGAGGCAC TGGCTGAAATTGGTTTTGTATTTGGGGAGCAATGGCGATA

[0305] G

[0306] beta essential SEQ ID NO: ATGAGTACTGCACTCGCAACGCTGGCTGGGAAGCTGGCTG 2 AACGTGTCGGCATGGATTCTGTCGACCCACAGGAACTGAT CACCACTCTTCGCCAGACGGCATTTAAAGGTGATGCCAGC GATGCGCAGTTCATCGCATTACTGATCGTTGCCAACCAGT ACGGCCTTAATCCGTGGACGAAAGAAATTTACGCCTTTCC TGATAAGCAGAATGGCATCGTTCCGGTGGTGGGCGTTGAT

[0307] Lambda red GGCTGGTCCCGCATCATCAATGAAAACCAGCAGTTTGATG

[0308] recombinatio GCATGGACTTTGAGCAGGACAATGAATCCtgtacaTGCCG n GATTTACCGCAAGGACCGTAATCATCCGATCTGCGTTACC GAATGGATGGATGAATGCCGCCGCGAACCATTCAAAACTC GCGAAGGCAGAGAAATCACGGGGCCGTGGCAGTCGCATCC

[0309]

[0310] C AAAC GG AT GT T ACGT C AT AAAGC CAT GAT TCAGTGTGCC CGTCTGGCCTTCGGATTTGCTGGTATCTATGACAAGGATG AAGCCGAGCGCATTGTCGAAAATACTGCATACACTGCAGA ACGTCAGCCGGAACGCGACATCACTCCGGTTAACGATGAA ACCATGCAGGAGATTAACACTCTGCTGATCGCCCTGGATA AAACATGGGAT GACGACT T ATT GCCGCT CT GT T CCCAGAT ATTTCGCCGCGACATTCGTGCATCGTCAGAACTGACACAG GCCGAAGCAGTAAAAGCTCTTGGATTCCTGAAACAGAAAG CCGCAGAGCAGAAGGTGGCAGCATGA

[0311] gamma SEQ ID NO: ATGGATATTAATACTGAAACTGAGATCAAGCAAAAGCATT 3 CACTAACCCCCTTTCCTGTTTTCCTAATCAGCCCGGCATT TCGCGGGCGATATTTTCACAGCTATTTCAGGAGTTCAGCC ATGAACGCTTATTACATTCAGGATCGTCTTGAGGCTCAGA

[0312] inhibiting GCTGGGCGCGTCACTACCAGCAGCTCGCCCGTGAAGAGAA RecBCD AGAGGCAGAACTGGCAGACGACATGGAAAAAGGCCTGCCC C AGC ACC T GT T T G AAT C GC T AT GC AT C GAT C AT T T GC AAC GCCACGGGGCCAGCAAAAAATCCATTACCCGTGCGTTTGA TGACGATGTTGAGTTTCAGGAGCGCATGGCAGAACACATC CGGTACATGGTTGAAACCATTGCTCACCACCAGGTTGATA T T GAT T C AG AGGT AT AA

[0313]

[0314] When said helper protein(s) capable of supporting nucleic acid recombination comprise RecET proteins, suitably the following proteins are expressed in said host cell:

[0315] Table 1A

[0316] RecET

[0317] Essential? Function / notes Exemplary sequence / accession number protein

[0318] SEQ ID NO: 4

[0319] GenBank Accession: AAA97507.1 (RecE 5’ to 3’ exonuclease;

[0320] RecE essential exonuclease VIII from E. coll Rac analogous to lambda red prophage )

[0321] alpha (Exo)

[0322] SEQ ID NO: 5

[0323] GenBank Accession: AAA97506.1 (RecT Single-stranded DNA

[0324] RecT essential recombinase from E. coll Rac annealing protein; analogous prophage )

[0325] to lambda red beta (Beta)

[0326]

[0327] RecET

[0328] Essential? Function / notes Exemplary sequence / accession number protein

[0329] SEQ ID NO: 9

[0330] GenBank Accession: AAA97507.1 (RecE 5' to 3' exonuclease;

[0331] RecE essential exonuclease VIII from E. coll Rac analogous to lambda red prophage )

[0332] alpha (Exo)

[0333] SEQ ID NO: 10

[0334] GenBank Accession: AAA97506.1 (RecT Single-stranded DNA

[0335] RecT essential recombinase from E. coll Rac annealing protein; analogous prophage )

[0336] to lambda red beta (Beta)

[0337]

[0338] Unlike the lambda red system, the RecET system does not require a Gam equivalent protein for inhibiting RecBCD, as RecE and RecT may function in certain host cell contexts without RecBCDinhibition. In some embodiments, the RecET system may be used in host cells where RecBCD activity is reduced or absent, or where the recombination substrate is protected from RecBCD degradation by other means.

[0339] HOMOLOGOUS RECOMBINATION SEQUENCES

[0340] In order to choose a homologous recombination sequence, the following steps may be used:

[0341] Choose 50 to 100 nucleotides in the desired position in the sequence of the nucleic acid being altered (target nucleic acid) such as a bacterial genome or plasmid backbone.

[0342] Perform a BLAST search of the chosen sequence against the target nucleic acid.

[0343] Consider the E- value for the chosen sequence compared to the closest match in the BLAST search - typically an E- value compared to an undesired target site elsewhere in the target nucleic acid of greater than IO-20would be too high; if this is discovered then suitably an alternative homologous recombination sequence is selected.

[0344] Suitably standard BLAST tool is used to calculate the E- value for homologous recombination (HR) sequences. One such online tool is at http: / / biocyc.org / ECOLI / blast.html. Suitably the focus is on how unique a given HR sequence is as judged by / / -value. Suitably it is not necessary to consider / calculate affinity. In principle any sequence that can work with classical recombination, is going to work better with the invention.

[0345] In more detail, if HR sequences can work with classical recombination, they are going to work better in the invention. Suitably the HR sequences for the invention are selected following the exact principle and requirement as for classical recombination using lambda red system. Lor example, the inventors typically design HRs 50-70 bp in length and blast against the E. coli genome for an expected value lower than IO-20, ( / '.'-value, a measurement of how unique a given sequence is; the lower the / / -value is, the more unique the sequence is. Any suitable tool for calculation may be used, for example standard BLAST tool to calculate the / / -value for HR sequences. One such online tool is at http: / / biocyc.org / ECOLI / blast.html. j.Values lower than IO-20 / / -value are not expected to be necessary, although of course sequences with lower values are still useful in the invention.

[0346] The / / -value is a measurement of how unique a given sequence is. Because classical recombination solely relies on the specificity of the homology regions, it requires a relatively stringent / / -value cut off such as IO-20. Because the methods of the invention may boost locus specificity not only by the specificity of the homology regions but also by the simultaneous loss of the negative selection marker and gain of positive selection marker, the methods can in principle tolerate less stringent A'-valuc(s)(e.g. less stringent homology regions). However, it is practically very straightforward to generate homology regions with stringent E-value, so suitably the IO-20E- value cut off is used.

[0347] SELECTABLE MARKERS

[0348] Any of the methods of the invention may comprise the further step of selecting for recombinants having incorporated the donor nucleic acid into the target nucleic acid. This step would be performed after the induction step to allow recombination.

[0349] The sequence of interest may comprise a positive selectable marker. Such markers include any that would allow the identification or selection of a cell comprising the marker.

[0350] The target nucleic acid may comprise in order: 5' - homologous recombination sequence 1 - negative selectable marker - homologous recombination sequence 2 - 3'.

[0351] Selecting for recombinants having incorporated said donor nucleic acid into said target nucleic acid may comprise selection for gain of the positive selectable marker of the donor nucleic acid and loss of the negative selectable marker of the target nucleic acid. Suitably selection for gain of the positive selectable marker of the donor nucleic acid and loss of the negative selectable marker of the target nucleic acid is carried out simultaneously. In other embodiments, the step of selecting for recombinants comprises sequential selection for the positive and negative markers, or sequential selection for the negative and positive markers.

[0352] The sequence of interest may comprise both a positive selectable marker and a negative selectable marker.

[0353] The methods of the invention may further comprise the step of:

[0354] inducing at least one double stranded break in the target nucleic acid sequence, wherein said double stranded break is between said homologous recombination sequence 1 and said homologous recombination sequence 2.

[0355] Suitably at least two double stranded breaks are induced in the target nucleic acid sequence, wherein each said double stranded break is between said homologous recombination sequence 1 and said homologous recombination sequence 2.

[0356] The episomal replicon may comprise a negative selectable marker independent of the donor nucleic acid sequence. Suitably said method comprises the further step of selecting for loss of the episomal replicon by selecting for loss of said negative selectable marker independent of the donor nucleic acid sequence.Some methods of the invention comprise a combinatorial selection approach involving a positive marker and loss of a negative marker. Use of this “double selection” scheme actually also helps with site specificity. For example, if a recombination event takes place at an inappropriate site, it could result in acquisition of the positive selectable marker. However, by using simultaneous selection for the positive marker and loss of the negative marker, even if the nucleic acid has been incorporated into the target nucleic acid at an inappropriate site (thereby conferring the positive marker), such molecules still would not be selected because if they have recombined into an inappropriate site they will not have simultaneously resulted in the loss of the negative marker. Therefore, as well as being a useful selection in its own right, this actually adds to the technical benefit of assisting in the site specificity by selecting not only for acquisition of the donor sequence but also simultaneous deletion of the sequence being removed / replaced.

[0357] Examples of suitable selectable markers are shown in the following table (Table 2). Furthermore, any suitable antibiotic marker may be used. Examples of such antibiotic markers include TetR, AmpR, HyR, and ErmR, which allow for a selection scheme including tetracycline, ampicillin, hygromycin, or erythromycin resistance, respectively.

[0358] Table 2Marker Selection Accession number / nucleic acid

[0359] sacB Sucrose sacB

[0360] sensitivity ATGAACATCAAAAAGTTTGCAAAACAAGCAACAGTATTAACCTTTACTACCGCACTGCTGGCAGGAGGCGCAACTCAAGCGTTTGCGAAAGA AACGAACCAAAAGCCATATAAGGAAACATACGGCATTTCCCATATTACACGCCATGATATGCTGCAAATCCCTGAACAGCAAAAAAATGAAA SEQ ID NO: 4

[0361] AATATAAAGTTCCTGAGTTCGATTCGTCCACAATTAAAAATATCTCTTCTGCAAAAGGCCTGGACGTTTGGGACAGCTGGCCATTACAAAACA CTGACGGCACTGTCGCAAACTATCACGGCTACCACATCGTCTTTGCATTAGCCGGAGATCCTAAAAATGCGGATGACACATCGATTTACATGT TCTATCAAAAAGTCGGCGAAACTTCTATTGACAGCTGGAAAAACGCTGGCCGCGTCTTTAAAGACAGCGACAAATTCGATGCAAATGATTCTA TCCTAAAAGACCAAACACAAGAATGGTCAGGTTCAGCCACATTTACATCTGACGGAAAAATCCGTTTATTCTACACTGATTTCTCCGGTAAAC ATTACGGCAAACAAACACTGACAACTGCACAAGTTAACGTATCAGCATCAGACAGCTCTTTGAACATCAACGGTGTAGAGGATTATAAATCAA TCTTTGACGGTGACGGAAAAACGTATCAAAATGTACAGCAGTTCATCGATGAAGGCAACTACAGCTCAGGCGACAACCATACGCTGAGAGATC CTCACTACGTAGAAGATAAAGGCCACAAATACTTAGTATTTGAAGCAAACACTGGAACTGAAGATGGCTACCAAGGCGAAGAATCTTTATTT AACAAAGCATACTATGGCAAAAGCACATCATTCTTCCGTCAAGAAAGTCAAAAACTTCTGCAAAGCGATAAAAAACGCACGGCTGAGTTAGC AAACGGCGCTCTCGGTATGATTGAGCTAAACGATGATTACACACTGAAAAAAGTGATGAAACCGCTGATTGCATCTAACACAGTAACAGATG

[0362] rpsL (S12 Streptomycin rpsL

[0363] ribosomal sensitivity ATGGCAACAGTTAACCAGCTGGTACGCAAACCACGTGCTCGCAAAGTTGCGAAAAGCAACGTGCCTGCGCTGGAAGCATGCCCGCAAAAACG protein)

[0364] SEQ ID NO: 5 TGGCGTATGTACTCGTGTATATACTACCACTCCTAAAAAACCGAACTCCGCGCTGCGTAAAGTATGCCGTGTTCGTCTGACTAACGGTTTCG AAGTGACTTCCTACATCGGTGGTGAAGGTCACAACCTGCAGGAGCACTCCGTGATCCTGATCCGTGGCGGTCGTGTTAAAGACCTCCCGGGT

[0365] <- T T r <- T T 7\ r r 7\ r 7\ r r <- T 7\ r <- T <- <- T <- r <- r T T <- T\ r T <- r T r r r <- r- <- T T 7\ 7\ 7\ <- T\ r- r- <- T 7\ 7\ <- r- 7\ <- <- r- T r <- T T r r- 7\ 7\ <- T 7\ T <- <- r- <- T <- T\ T\ <- r- <- T r r- T 7\ 7\ PheS with two P-chloro- PheS T251A A294G

[0366] point phenylalanine ATGTCACATCTCGCAGAACTGGTTGCCAGTGCGAAGGCGGCCATTAGCCAGGCGTCAGATGTTGCCGCGTTAGATAATGTGCGCGTCGAATA mutations

[0367] SEQ ID NO: 6 TTTGGGTAAAAAAGGGCACTTAACCCTTCAGATGACGACCCTGCGTGAGCTGCCGCCAGAAGAGCGTCCGGCAGCTGGTGCGGTTATCAACG AAGCGAAAGAGCAGGTTCAGCAGGCGCTGAATGCGCGTAAAGCGGAACTGGAAAGCGCTGCACTGAATGCGCGTCTGGCGGCGGAAACGATT GATGTCTCTCTGCCAGGTCGTCGCATTGAAAACGGCGGTCTGCATCCGGTTACCCGTACCATCGACCGTATCGAAAGTTTCTTCGGTGAGCT TGGCTTTACCGTGGCAACCGGGCCGGAAATCGAAGACGATTATCATAACTTCGATGCTCTGAACATTCCTGGTCACCACCCGGCGCGCGCTG ACCACGACACTTTCTGGTTTGACACTACCCGCCTGCTGCGTACCCAGACCTCTGGCGTACAGATCCGCACCATGAAAGCCCAGCAGCCACCG ATTCGTATCATCGCGCCTGGCCGTGTTTATCGTAACGACTACGACCAGACTCACACGCCGATGTTCCATCAGATGGAAGGTCTGATTGTTGA TACCAACATCAGCTTTACCAACCTGAAAGGCACGCTGCACGACTTCCTGCGTAACTTCTTTGAGGAAGATTTGCAGATTCGCTTCCGTCCTT

[0368]

[0369] CnF Chloramphenicol CmR

[0370] resistance ATGGAGAAAAAAATCACTGGATATACCACCGTTGATATATCCCAATGGCATCGTAAAGAACATTTTGAGGCATTTCAGTCAGTTGCT CAATGTACCTATAACCAGACCGTTCAGCTGGATATTACGGCCTTTTTAAAGACCGTAAAGAAAAATAAGCACAAGTTTTATCCGGCC SEQ ID NO: 7

[0371] TTTATTCACATTCTTGCCCGCCTGATGAATGCTCATCCGGAATTCCGTATGGCAATGAAAGACGGTGAGCTGGTGATATGGGATAGT GTTCACCCTTGTTACACCGTTTTCCATGAGCAAACTGAAACGTTTTCATCGCTCTGGAGTGAATACCACGACGATTTCCGGCAGTTT CTACACATATATTCGCAAGATGTGGCGTGTTACGGTGAAAACCTGGCCTATTTCCCTAAAGGGTTTATTGAGAATATGTTTTTCGTC

[0372] KanR Kanamycin KanR

[0373] resistance ATGATTGAACAAGATGGATTGCACGCAGGTTCTCCGGCCGCTTGGGTGGAGAGGCTATTCGGCTATGACTGGGCACAACAGACAATC GGCTGCTCTGATGCCGCCGTGTTCCGGCTGTCAGCGCAGGGGCGCCCGGTTCTTTTTGTCAAGACCGACCTGTCCGGTGCCCTGAAT SEQ ID NO: 8

[0374] GAACTGCAGGACGAGGCAGCGCGGCTATCGTGGCTGGCCACGACGGGCGTTCCTTGCGCAGCTGTGCTCGACGTTGTCACTGAAGCG GGAAGGGACTGGCTGCTATTGGGCGAAGTGCCGGGGCAGGATCTCCTGTCATCTCACCTTGCTCCTGCCGAGAAAGTATCCATCATG GCTGATGCAATGCGGCGGCTGCATACGCTTGATCCGGCTACCTGCCCATTCGACCACCAAGCGAAACATCGCATCGAGCGAGCACGT ACTCGGATGGAAGCCGGTCTTGTCGATCAGGATGATCTGGACGAAGAGCATCAGGGGCTCGCGCCAGCCGAACTGTTCGCCAGGCTC

[0375]

[0376] AAGGCGCGCATGCCCGACGGCGAGGATCTCGTCGTGACCCATGGCGATGCCTGCTTGCCGAATATCATGGTGGAAAATGGCCGCTTTThus, in an embodiment, the negative selectable marker is selected from the group consisting of sacB (sucrose sensitivity) or rpsL (S12 ribosomal protein - streptomycin sensitivity). In some embodiments, the positive selectable marker is selected from the group consisting of CmR (chloramphenicol resistance) or KanR (kanamycin resistance).

[0377] EXCISION / INTRODUCTION OF DOUBLE STRANDED BREAKS

[0378] The methods of the invention include a mechanism for excision / introduction of double stranded breaks. Suitably excision is performed to generate a linear donor nucleic acid. Suitable endonucleases are described above.

[0379] In embodiments, the expression of the endonuclease is constitutive, and induced under the control of an inducible promoter such as the pBAD promoter from the E. coli arabinose operon.

[0380] Induction of expression is well within the abilities of a skilled worker in the art. For example, the sequence of interest (such as the I-Scel sequence) is placed under the control of an inducible promoter. That promoter activity is induced when desired. For example, the well-known arabinose (pAra) promoter may be used, which is induced in the presence of arabinose. Similarly, the skilled worker may choose inducible or constitutive promoters from a vast array of well-known promoters suitable for inducible or constitutive expression as desired.

[0381] Alternatives to pBAD include the Lad repressor, which is induced by IPTG (Isopropyl P-D-l-thiogalactopyranoside). The Lad repressor binds to the lac operator, preventing transcription in the absence of IPTG. Addition of IPTG relieves repression. See Muller-Hill B. The lac Operon: A Short History of a Genetic Paradigm. Walter de Gruyter & Co 1996;

[0382] the T7-Lac (pET System), in which a T7 promoter drives expression, but A. coli must produce T7 RNA polymerase (often from the DE3 lysogen under lacUV5 control). IPTG induction triggers T7 RNA polymerase production, which then transcribes the target gene from the T7 promoter. See Studier FW, Moffatt BA. JMolBiol. 1986; 189(1): 113-130; Studier FW. Protein Expr Purif 2005;41(l):207-234; Tet-Based Systems, in which the Tet repressor (TetR) binds the operator, blocking transcription in the absence of inducer. Addition of tetracycline (or aTc) releases TetR, inducing expression. “Tet-On” uses a modified TetR that activates in the presence of tetracycline. See Gossen M, Bujard H. Proc Natl Acad Sci USA. 1992;89(12):5547-555; Lutz R, Bujard H. Nucleic Acids Res. 1997;25(6): 1203-1210; Rhamnose-Inducible Promoters (prhaBAD)in which the rha operon is controlled by the RhaS / RhaR activators. Expression is off in the absence of L-rhamnose and on upon inducer addition. See Giacalone et al., Biotechniques. 2006;40(3):355-364; Wegerer et al., Microb Cell Fact. 2008;7:8; Xylose-Inducible Promoters (pxyl) wherein Xylose activates a specific regulator (e.g., XylR) that drives expression of xylose-metabolism genes. See Khlebnikov et al., J Bacteriol. 2000;182(24):7029-7034;the Cumate Switch (Pcmt) wherein the cumate repressor (CymR) binds to the operator in the absence of cumate, blocking transcription. Addition of cumate relieves repression. See Xiao et al. Gene.

[0383] 2008;426(l-2): 16-22; as well as hybrid systems combining more than one system.

[0384] The skilled operator will realise that if alternative endonucleases are employed in the invention then the corresponding alternate recognition sequence should be used in the backbone DNA of the episomal vector. This is well within the ambit of the skilled worker, and exemplary recognition sequences are set forth above and in the literature.

[0385] In one embodiment it may be desirable to induce a cut on the target nucleic acid in order to assist in selection for recombinants. In this embodiment suitably there are 3 cuts - two on the episomal replicon to excise the donor nucleic acid and one on the target nucleic acid to assist in selection. Thus suitably said target nucleic acid comprises in order: 5' - homologous recombination sequence 1 - cut site -homologous recombination sequence 2 - 3'.

[0386] Suitably said target nucleic acid comprises in order:

[0387] a) 5' - homologous recombination sequence 1 - cut site - homologous recombination sequence 2 -3'

[0388] b) 5' - homologous recombination sequence 1 - positive selectable marker - homologous recombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0389] c) 5' - homologous recombination sequence 1 - negative selectable marker - homologous recombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0390] d) 5' - homologous recombination sequence 1 - positive selectable marker - negative selectable marker - homologous recombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0391] e) 5' - homologous recombination sequence 1 - negative selectable marker - positive selectable marker - homologous recombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0392] When applying the invention in multiple rounds, the donor nucleic acid of a first round may contribute / become part of the target nucleic acid in next round. Thus, suitably the sequence of interest may comprise in order:

[0393] a) 5' - homologous recombination sequence 1 - cut site - homologous recombination sequence 2 -3'

[0394] b) 5' - homologous recombination sequence 1 - positive selectable marker - homologousrecombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0395] c) 5' - homologous recombination sequence 1 - negative selectable marker - homologous recombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0396] d) 5' - homologous recombination sequence 1 - positive selectable marker -negative selectable marker - homologous recombination sequence 2 - 3', further comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0397] e) 5' - homologous recombination sequence 1 - negative selectable marker - positive selectable marker - homologous recombination sequence 2 - 3', further

[0398] comprising a cut site between said homologous recombination sequence 1 and homologous recombination sequence 2

[0399] Suitably the cut site on the target nucleic acid or sequence of interest is different from the excision site on the episomal replicon / donor nucleic acid. A different endonuclease may be provided.

[0400] Said cut site may be between said positive / negative selectable markers, or may be within said positive / negative selectable markers. Suitably said target nucleic acid comprises two such cut sites. Suitably said cut site is adjacent to one of said homologous recombination sequences. Suitably said two cut sites comprise a first cut site adjacent to said homologous recombination sequence 1, and a second cut site adjacent to said homologous recombination sequence 2.

[0401] APPLICATIONS

[0402] The invention involves the introduction of a sequence of interest into a target nucleic acid. “Introducing a sequence of interest”, as used herein, means that the sequence of interest is integrated into the target nucleic acid such that the resultant nucleic acid sequence comprises the sequence of interest. This may be referred to as incorporation of the sequence of interest into a target nucleic acid.

[0403] Introducing a sequence of interest into a target nucleic acid may comprise replacing a part of the target nucleic acid sequence with the sequence of interest. Thus, after a replacement, the resultant nucleic acid sequence comprises the sequence of interest and only part of the original sequence of the target nucleic acid. For example, in embodiments wherein a genome is the target nucleic acid sequence, a part of the genome sequence may be replaced by the introduced sequence of interest during an iteration of the method of the invention. The sequence of interest may replace a region within the target nucleic acid that is smaller, the same size, or larger than the sequence of interest.Introducing a sequence of interest may comprise inserting the sequence of interest into the target nucleic acid. As used herein, “inserting” means that that the sequence of interest is introduced into a site within the target nucleic acid such that the resultant nucleic acid sequence comprises all of the original sequence of the target nucleic acid and also includes the inserted sequence of interest.

[0404] Methods of the invention may include multiple steps whereby a donor nucleic acid, for instance encoding selection markers, is introduced into the target nucleic acid and then replaced by a sequence of interest.

[0405] As an example, the methods of the invention may be used to insert a sequence of interest into a genome, or other target nucleic acid, such that all original sequences of the genome remain in addition to the newly inserted sequence of interest. In another example, the methods of the invention may be used to replace part of a genome, or other target nucleic acid, with a sequence of interest such that the overall size of the genome is unaltered. In yet other examples, the methods of the invention may be used to replace part of a genome, or other target nucleic acid, with a sequence of interest that is longer than the replaced region, such that the resultant nucleic acid comprises only part of the original genome sequence and the overall size of the genome is increased.

[0406] The invention is useful in the construction of plasmids. The invention is useful in manipulation of host genomes. The invention is useful in the construction of artificial chromosomes such as BACs.

[0407] The invention finds particular application in the making of large sized nucleic acid constructs. The invention finds particular application in the creation of high diversity libraries. In this regard, a transformation efficiency of approximately 108is achievable using current transformation techniques. However, a transformation efficiency of IO10or beyond is extremely challenging and / or problematic. According to the present invention, a first half-library may be created and transformed into a first host cell (population of host cells). This first half-library is then transformed with nucleic acid encoding the second half-library. By using recombination according to the present invention, those two half-libraries are then combined in vivo resulting in a library having diversity of 1010, which has advantageously been obtained having only ever needing to use a transformation efficiency of 105.

[0408] HOST CELL

[0409] Suitably the host cell is a prokaryotic cell. Suitably the host cell is a bacterial cell.

[0410] In one embodiment, the host cell is in vitro i.e. in the laboratory. In one embodiment, the methods of the invention are in vitro methods. In some embodiments, the methods are not practiced in vivo. Suitably the host cell is not part of a live human or animal body. Suitably the host cell is selected from one of the host cells used in the examples below.The host cell may be any gram-negative bacterium. The host cell may be E. Col . The host cell may be any / i. coli strain (such as MG 1655 orBL21), or cells derived therefrom.

[0411] MG 1655 is considered as the wild type strain of E.coli. The GenBank ID of genomic sequence of this strain is U00096 (U00096.3 as of the date of filing). BL21 is widely available commercially.

[0412] The host organism such as E.coli may be chosen, or may be manipulated, in order to inhibit naturally occurring repair mechanisms to ensure the absence of, or extremely low likelihood of, double stranded repair. For example, the RecBCD system may be mutated or inhibited provided that suitable helper protein(s) capable of supporting nucleic acid recombination in the host cell are present in place of RecBCD, e.g. the lambda Red proteins described herein or other suitable recombination support proteins. For example, in one embodiment RecBCD may be inhibited because it can interfere with lambda red components and reduce the efficiency of recombination using double strand DNA with short homology regions (e.g. around 50 bp)(degraded by RecBCD system) carried out by lambda red components. However, if long homology regions (e.g. around 3-5 kb) are used, RecBCD can be an alternative as recombination support protein(s) to lambda red components as recombination support protein(s).

[0413] OPTIONAL ADDITIONAL STEPS OR FEATURES

[0414] In one embodiment, the invention may involve a first recombination step carried out by conventional techniques. This has the advantage of allowing introduction into the target site of contra-selectable markers.

[0415] Optionally the invention comprises a final step of a final recombination which may be accomplished either by the methods disclosed herein or by conventional recombination. For example, this may be advantageous in removing selectable markers which have served their purpose and are no longer required for further iterations of the methods disclosed herein.

[0416] In some embodiments, the iterative methods disclosed herein begin and continue without a first conventional homologous recombination event.

[0417] In an embodiment, the endonuclease is employed to cut at a site intended to be replaced by recombination event, thereby creating selective pressure against the cut (and not recombined) target nucleic acid i.e. negative selection by double stranded break. In another embodiment, this negative selection by double stranded break in the target sequence is used to improve selection with a 3 -double strand break embodiment (2 double strand breaks for excision of the donor nucleic acid and one double strand break between the HR1 and HR2 sequences on the target nucleic acid making 3 DS breaks / cuts in total).In a particular embodiment, there is provided a method of introducing a sequence of interest into a target nucleic acid comprising

[0418] 1) providing an E. col host cell comprising a genome, the genome comprising a target nucleic acid;

[0419] 2) delivering a BAC to the host cell by conjugative transfer,

[0420] said BAC comprising a backbone sequence and a donor nucleic acid sequence, wherein said donor nucleic acid sequence comprises in order: 5 ’ - homologous recombination sequence 1 - sequence of interest - homologous recombination sequence 2 - 3’,

[0421] wherein the backbone sequence comprises a first excision site positioned adjacent to the homologous recombination sequence 1 and a second excision site positioned adjacent to the homologous recombination sequence 2;

[0422] 3) providing lambda red proteins capable of supporting nucleic acid recombination in said host cell;

[0423] 4) providing an endonuclease which does not recognise a host cell sequence apart from the first and second excision sites;

[0424] 5) inducing excision of said donor nucleic acid sequence by the endonuclease;

[0425] 6) incubating to allow recombination between the excised donor nucleic acid and said target nucleic acid; and

[0426] 7) selecting for recombinants having incorporated said donor nucleic acid into said target nucleic acid. Any of the alternatives for the features of this method that are discussed herein may be substituted for the corresponding features of this method. For instance, the BAC may be any episomal replicon, the target nucleic acid may be any target nucleic acid, the lambda red proteins may be substituted for any suitable for the purpose, etc.

[0427] As discussed herein, steps 1) to 4) may be performed in any order or simultaneously.

[0428] In another embodiment, the method is for the assembly of a nucleic acid sequence, and comprises performing steps 1) to 7) to introduce a first donor nucleic acid sequence into a first target nucleic acid in order to create a second target nucleic acid; and then performing steps 1) to 7) to introduce a second donor nucleic acid sequence into the second target nucleic acid in order to create a third target nucleic acid. This process may be iterated, wherein the product of each iteration is the target of the next iteration. A BAC comprising a first backbone sequence may be used for each odd-numbered iteration and a BAC comprising a second backbone sequence may be used for each even-numbered iteration. The first backbone sequence and the second backbone sequence may encode different spacers and may comprise different selection markers.All of the features described herein (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.

[0429] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made to the Examples, which are not intended to limit the invention in any way.

[0430] EXAMPLES

[0431] Methodology

[0432] Strains and plasmids used in this study

[0433] As a recipient strain, we used a reduced-genome, streptomycin resistant E. coli (MDS42; Scarab Genomics) for the experiments described in this application. A subset of experiments uses a ArecA version of MDS42. The ArecA mutant was generated through methods described in WO2023078997. All recipient strains contained a positive-negative selection marker immediately upstream of the region targeted for replacement, specifically landing site 23 (LS23, Fig 3) as described in WO2023078997. As a donor strain, we used E. coli DHIOb for the conjugative transfer of BACs containing donor DNA sequences.

[0434] All cloning procedures were performed in E. coli DHIOb. We performed BAC assemblies using yeast homologous recombination in .S', cerevisiae strain BY4741.

[0435] We used BAC 100k24 in this study (Ziircher et al., Nature. 2023 Jul;619(7970):555-562). This BAC carries -100 kb of synthetic DNA with a defined synonymous codon compression scheme in which two serine codons (TCG and TCA) and a stop codon (TAG) are replaced through defined recoding rules (TCG to AGC, TCA to AGT, and TAG to TAA). The 100k24 BAC as described in Ziircher et al., (Nature, 2023) is referred to as BAC_100k24_vl in this application. BAC_100k24_vl contains universal protospacer sequences flanking the 100k24 synthetic DNA insert, which are amenable by cleavage by Cas9, as well as universal spacer sequences encoded in the BAC backbone (Fig 3). Using homologous recombination in yeast and methods analogous to those described in WO2023078997, we generated derivatives of this BAC, where the 100k24 synthetic DNA insert is flanked by I-Scel recognition sites, either in the same orientation (BAC_100k24_v2) or in opposite orientations (BAC_100k24_v3). Both BACs containing I-Scel recognition sites do not contain universal spacer sequences in the BAC backbone (Fig 3). All BAC backbones contain an oriT for conjugative transfer.We used the following positive and negative selection markers: sacB (conferring sucrose sensitivity, - 2), cat (chloramphenicol resistance, +2), rpsL (streptomycin sensitivity, -1), kan" (kanamycin resistance, +1), / ?eST251A A294G(pheS*) (4-chloro-phenylalanine (4-CP) sensitivity, -3), and tef (tetracycline resistance, +5)13.

[0436] For Cas9-based experiments, we used the helper plasmid pKW20 previousy described (Wang, K. et al. Nature 539, 59-64, doi: 10.1038 / nature20124 (2016)), referred to in this application as pHelper_Cas9 (Fig 3). We created an alternative version of the helper plasmid, modified to delete the CRISPR-Cas components (Cas9) and introduce the I-Scel gene under control of an arabinose inducible promoter. This modified plasmid is labelled pHelper l-Scel (Fig 3). The I-Scel gene was synthesized de novo, and inserted in replacement of the Cas9 in pHelper_Cas9 by Gibson assembly using primers detailed in Table 6. The I-Scel gene sequence was sourced from a public database, and then further revised by altering a single TCG codon to AGC, ensuring future compatibility with Syn61 systems if required. The complete list of plasmids is detailed in Table 7.

[0437] CONEXER CONEXER requires preparation of a conjugation competent donor cell and a recipient cell. The donor cells, in this case E. coll DHIOb, carry the non-transferable conjugative plasmid pJF146 (accession number MK809154.1, Fredens et al., 2019) and the BAC with the synthetic DNA for integration, an oriT sequence and a universal spacer array. In the Cas9-based experiment, the orientation of the oriT ensures that the spacer array enters the recipient cell last to minimise the risk of partial excision by premature initiation of Cas9 cleavage in the recipient cell. The recipient cell carries a genomic double selection cassette at LS23, marking the upstream end of the integration site, and the helper plasmid pKW20 for inducible expression of lambda-red components and either Cas9 or I-Scel. In BACs used with I-Scel, the spacer array is removed.

[0438] Here, we describe CONEXER with a donor strain carrying a BAC with a 100 kb synthetic DNA insert with rpsL-karf1followed by pheS* on the BAC backbone; and a recipient strain carrying a genomic sacB-cat selection cassette at LS23 (Fig 3A). We grew the donor strain to saturation overnight in 25 ml LB medium with selection for pJF146 (50 pg / mL apramycin) and selection for the BAC (50 pg / mL kanamycin (even numbered 100k24 BACs) or 20 pg / mL chloramphenicol (odd-numbered 100k25 BACs)). We grew the recipient strain to saturation overnight in 25 ml LB medium with selection for the helper plasmid (5 pg / mL tetracycline), the genomic double selection cassette (20 pg / mL chloramphenicol or 50 pg / mL kanamycin) and suppression of endonuclease and lambda-red expression (2% glucose). We harvested the cells from each culture by centrifugation and washed the pellet two times in 1 mb LB medium. After the final wash we resuspended the pellets in 800 pl LB. We mixed

[0439]

[0440] dried, incubated the plates at 37 °C for 1 hour. Following conjugation, we washed cells off the plate and transferred all into 250 mL pre-warmed LB medium with selection for recipient cells carrying the helper plasmid (5 pg / mL tetracycline), and the BAC (50 pg / mL kanamycin or 20 pg / mL chloramphenicol) for 2 hours, and subsequently induced expression of endonuclease and lambda-red (0.5% L-arabinose). After 1.5 hours of incubation at 37 °C with shaking we harvested cells by centrifugation and immediately transferred all into 250 ml pre-warmed LB with 50 pg / mL kanamycin (or 20 pg / mL chloramphenicol), 5 pg / mL tetracycline, and 2% glucose to terminate recombination by supressing expression of endonuclease and lambda-red. After another 2.5 h incubation with shaking at 37°C we spun the culture by centrifugation and resuspended the pellet in 2 mL Milli-Q filtered water. The cell suspension was spread in serial dilutions on LB agar plates with selection for the helper plasmid (5 pg / mL tetracycline), selection for the integration of the double selection cassette at LS24 (50 pg / ml kanamycin), selection for the loss of the double selection cassette at LS23 (7.5% sucrose), and selection for the loss of the BAC backbone (2.5 mM 4CP). In selection plates without sucrose, we added 2 % glucose to supress endonuclease and lambda-red expression.

[0441] For experiments in ArecA hosts, we grew cells for 2-8 h in 250 ml pre-warmed LB with 50 pg / mL kanamycin (or 20 pg / mL chloramphenicol), 5 pg / mL tetracycline, and 2% glucose to allow for cells who received the BAC to expand prior to the 1.5 h induction of endonuclease and lambda-red expression. This increased the number of successful recombinants from CONEXER experiments. After overnight incubation at 37 °C, we picked colonies and resuspended them in 30 pL water. We assessed each clone by colony PCR for the loss of the sacB-cat cassette at LS23 and integration of the rps -katB cassette at LS24. Additionally, we characterized the phenotypes of the post-CONEXER clones by stamping them in selective media, as described in Robertson, W. E. et al. Nat Protoc 16, 2345-2380, doi:10.1038 / s41596-020-00464-3 (2021). Specifically, we considered a clone to have passed phenotypic validation if it displayed resistance to chloramphenicol, 4-chloro-phenylalanine, streptomycin and tetracycline, and sensitivity to kanamycin and sucrose. We selected colonies with verified genotypes and phenotypes for whole-genome sequencing by NGS. Oligonucleotide sequences used in genotyping are provided in Table 5.

[0442] Next-generation sequencing (NGS) and sequencing data analysis

[0443] BACS and genomic DNA (gDNA) were extracted from overnight cultures of E. coll using a QuickExtract Extraction Kit from Bioserarch Technologies. Preparation for NGS has been previously described13;5; Robertson, (2021). Briefly, paired-end sequencing libraries were prepared with the Nextera XT DNA Library Preparation Kit (Illumina) following the manufacturer’s protocol.Libraries were paired-end sequenced on a NextSeq 2000 (Illumina). The downstream sequencing analysis was achieved with a custom Python script that outputs, for each clone, the frequency of recoding at each target codon plotted across the genomic region being analysed.

[0444] Table 5: Oligos used to genotype genomic loci after CONEXER

[0445] Oligo F (5' -> 3') Oligo R (5' -> 3')

[0446] LS23 TGGCATCACGAGACACAT TACAACATGACCAGCGGATTCAA CACAGAG (SEQ ID NO: 25) G (SEQ ID NO: 26)

[0447] LS24 TTACCGCCTGGTTCTCTG TTCTTGTTCCAGTTGACCTTTCT ACTTAACGC (SEQ ID NO: CCCAG (SEQ ID NO: 28)

[0448] 27)

[0449]

[0450] Table 6: Oligos used to clone an l-Scel version of the helper plasmid

[0451] Oligo F (5' -> 3') Oligo R (5' -> 3')

[0452] l-Scel insert caatggcgatagatcctagcaaggagga cag cag cctag gg atcctag catcttcatctaaaat aatcatcATGCAT AT GAAAAACA atactTTATTTCAGGAAAGTTTCGGA TCAAAAAAAACCAGGTAATG GGAGATAGTGTTCGGCAG (SEQ ID AACCTGGGT (SEQ ID NO: 29) NO: 30)

[0453] pHelper CTCCTCCGAAACTTTCCTGA CTGG I I I I I I I I GA I G I I I I I CATAT vector AATAAagtatattttagatgaagatgcta G C ATg atg atttcctccttg ctaggatctatcgcca ggatccctaggctgct (SEQ ID NO: ttgctcccc (SEQ ID NO: 32)

[0454] 31)

[0455]

[0456] Table 7: Plasmids and BACs used

[0457] Plasmid Description Genbank# Reference

[0458] pHelper_Cas9 Contains lambda-red MN927219 Wang et al.

[0459] recombination components 2016 and Cas9 under arabinose

[0460] inducible promoter as well as

[0461] tracrRNA

[0462] pHelper_l-Scel Contains lambda-red N / A This study recombination components

[0463] and l-Scel under arabinose

[0464] inducible promoter as well as

[0465] tracrRNA

[0466]

[0467] Example 1 - A rapid and general method to create custom synthetic genomes in Escherichia coli Summary

[0468] Whole genome synthesis and large-scale genome engineering promise to provide powerful approaches for understanding organism function, wholesale engineering of biosynthetic pathways, and creating organisms with functions beyond those found in nature. Simple, robust, accelerated, and scalable methods for replacing genomic DNA with synthetic DNA will make genome synthesis, and large-scale genome engineering more accessible.

[0469] Here we report an approach that simplifies and accelerates the introduction of more than 100 kb of synthetic DNA into the Escherichia coli genome. Our method accomplishes this using a rapid (one day) protocol, which may be iterated to introduce even larger synthetic DNA sequences. Crucially, the method standardizes and unifies all the necessary components such that the user only needs to clone the synthetic DNA of interest into a bacterial artificial chromosome and then implement a standard protocol.

[0470] The method described herein is free of CRISPR and does not require spacer arrays or Cas9. It thus simplifies CONEXER and REXER, whilst avoiding the practical disadvantages of incorporating long endonuclease cleavage sites into insert nucleic acid. While it was previously shown that the presence of a 6 bp blunt overhang, generated by Cas9 cleavage, still enables scarless integration of synthetic DNA, it was unknown whether the larger, staggered overhangs generated by I-Scel would still enablescarless synthetic DNA integration. The results described here show that I-Scel-promoted recombination is compatible with scarless integration of synthetic DNA.

[0471] Results

[0472] As discussed herein, it was unclear whether the use of homing endonuclease would prove as efficient as CRISPR / Cas9 in cleaving donor episomes, whether nucleases such as I-Scel would be toxic in recoded E. coll such as Syn61, nor whether the cleavage patterns and resulting backbone DNA fragments retained on insert DNA would compromise recombination efficiency.

[0473] To investigate the efficiency and fidelity of CONEXER with I-Scel endonuclease, we used PCR using Cas-9 CONEX BAC as template to replace CRISPR system and Cas9 cut sites with the I-Scel gene and associated cleavage sites. We assembled the following versions of the BAC in yeast:

[0474] • CONEXER BAC with I-Scel cut sites in the same orientation - BAC_100k24_v2

[0475] • CONEXER BAC with I-Scel cut sites in the opposite orientation - BAC_100k24_v3

[0476] In both versions, CRISPR components are removed. Additionally, an integrated dimerresolution system (resD) with rfsF cut site introduced under the native promotor and upstream ccdA and ccdB genes. Both genes are essential for transcription repression of the operon. Ratio between CcdA and CcdB regulates the repression state of the ccd operon - if level of ccdA is greater than or equal to ccdB, repression occurs preventing harmful effect of ccdB on the bacteria. Otherwise, de-repression occurs. In the current assembly, a mutated ccdB sequence is used such that ccdB sequence is not expressed at all.

[0477] Syn61wt genomic DNA (Fredens 2019) was used as template for amplifying the synthetic DNA fragments (100k24). The amplified and purified fragments were assembled into BACs by homologous in yeast, as described in (Robertson 2021). The clones were verified by next generation sequencing (NGS) for correct assembly of all fragments with intact BAC-backbone. In the helper plasmid, we replaced CRISPR-associated sequences with the I-Scel gene. A modified pHelper plasmid was created by replacing the cas9 gene with the I-Scel gene, placing this gene under the control of the pBAD promoter for induction by arabinose, creating pHelper l-Scel.

[0478] I-Scel DNA sequence:

[0479] ATGCATATGAAAAACATCAAAAAAAACCAGGTAATGAACCTGGGTCCGAACTCTAAACT GCTGAAAGAATACAAATCCCAGCTGATCGAACTGAACATCGAACAGTTCGAAGCAGGTA TCGGTCTGATCCTGGGTGATGCTTACATCCGTTCTCGTGATGAAGGTAAAACCTACTGTA TGCAGTTCGAGTGGAAAAACAAAGCATACATGGACCACGTATGTCTGCTGTACGATCAG TGGGTACTGTCCCCGCCGCACAAAAAAGAACGTGTTAACCACCTGGGTAACCTGGTAAT CACCTGGGGCGCCCAGACTTTCAAACACCAAGCTTTCAACAAACTGGCTAACCTGTTCAT CGTTAACAACAAAAAAACCATCCCGAACAACCTGGTTGAAAACTACCTGACCCCGATGTCTCTGGCATACTGGTTCATGGATGATGGTGGTAAATGGGATTACAACAAAAACTCTACCA ACAAATCGATCGTACTGAACACCCAGTCTTTCACTTTCGAAGAAGTAGAATACCTGGTTA AGGGTCTGCGTAACAAATTCCAACTGAACTGTTACGTAAAAATCAACAAAAACAAACCG ATCATCTACATCGATTCTATGTCTTACCTGATCTTCTACAACCTGATCAAACCGTACCTGA TCCCGCAGATGATGTACAAACTGCCGAACACTATCTCCTCCGAAACTTTCCTGAAATAA

[0480] (SEQ ID NO: 33)

[0481] We set out to perform CONEXER experiments as illustrated in Figure 1A, aiming to evaluate the effectiveness of I-Scel based CONEXER in comparison to Cas9 CONEXER, in different arrangements and different genetic backgrounds. In all cases, the donor DHIOb strain carries an orzT-deficient F plasmid, and a -100 kbp synthetic DNA insert corresponding to fragment 100k24, coupled to an rpsL-kanR double selection marker. Donor strains vary in the version of the BAC they contain, which can be in a format suitable for excision of the synthetic DNA insert with CRISPR / Cas9 (BAC_100k24_vl), or with I-Scel (BAC_100k24_v2 and BAC_100k24_v3).

[0482] The recipient strains, derived from MDS42, contain a genomically integrated sacB-cat double selection marker at LS23, immediately upstream of the corresponding 100k24 fragment, as well as either a pHelper_Cas9 or a pHelper l-Scel plasmid.

[0483] The results of performing CONEXER experiments with different combinations of donors and recipients are detailed below. After each experiment, the indicated number of colonies was picked and characterized by genotyping and phenotyping as described in the CONEXER methodology section.

[0484] Table 9

[0485] CONEXER CONEXER CONEXER recipient Genotype / phenotype experiment # donor BAC verification

[0486] Cl BAC_100k24_vl MDS42(ArecA)-sC23 + 5 out of 5 correct pHelper_Cas9

[0487] C2 BAC_100k24_v2 MDS42-sC23 + pHelper_I-SceI 5 out of 5 correct

[0488] C3 BAC_100k24_v2 MDS42(ArecA)-sC23 + pHelper l- 3 out of 5 correct Scel

[0489] C4 BAC_100k24_v3 MDS42-sC23 + pHelper_I-SceI 3 out of 3 correct

[0490]

[0491] C5 BAC_100k24_v3 MDS42(ArecA)-sC23 + pHelper l- 5 out of 5 correct Scel

[0492] C6 BAC_100k24_v2 MDS42(ArecA)-sC23 + No colonies pHelper_Cas9

[0493] C7 BAC_100k24_v3 MDS42(ArecA)-sC23 + No colonies

[0494] pHelper_Cas9

[0495]

[0496] Clones that pass genotypic and phenotyping verification have undergone the expected recombination event as outlined in Figure 3A. The results suggest that, despite the replacement of the CRISPR / Cas9 system with an endonuclease-based excision of the synthetic DNA insert prior to recombination, I-Scel-based CONEXER enables the replacement of the 100k24 fragment of the E. coli MDS42 genome with a synthetic recoded counterpart. This highlights the suitability of the method disclosed here for replacing large sections of the E. coli genome and creating custom synthetic genomes. We observed I-Scel-based CONEXER-mediated replacement of the -100 kbp section in genetic backgrounds with and without a functional recA gene. No colonies were obtained in experiments where the donor BAC contained I-Scel excision sites, but the helper plasmid harboured Cas9, indicating that the recombination event is dependent on I-Scel being expressed and cleaving the donor BAC prior to recombination.

[0497] In the course of these experiments, E. coli cells behaved similarly when harbouring pHelper_Cas9 and pHelper l-Scel, suggetins that there is no significant additional toxicity attributable to I-Scel compared to Cas9.

[0498] While phenotyping and genotyping indicates that the recombination event has occurred, it does not provide direct information about whether the resulting colonies have integrated the entirety of the synthetic DNA insert at the designated locus (in this case, whether the method provides cells that are fully recoded across the entire region of interest), or whether integration of the synthetic DNA into the recipient genome occurs scarlessly.

[0499] We set out to further confirm the extent of recoding in post-CONEXER clones and the fidelity of recombination by next-generation sequencing (see Methodology section). The results showed that the synthetic DNA is integrated into the genome scarlessly despite the non-homologous sequences at both ends resulting from retention of backbone DNA upon excision by I-Scel, and that the rates of complete recoding achieved I-Scel method are comparable to those of CRISPR / Cas9.Table 10

[0500] CONEXER CONEXER donor CONEXER recipient Fully recoded experiment # BAC

[0501] Cl BAC_100k24_vl MDS42(ArecA)-sC23 + pHelper_Cas9 2 out of 5

[0502] C2 BAC_100k24_v2 MDS42-sC23 + pHelper_I-SceI 4 out of 5

[0503] C3 BAC_100k24_v2 MDS42(ArecA)-sC23 + pHelper_I-SceI 3 out of 3

[0504] C4 BAC_100k24_v3 MDS42-sC23 + pHelper_I-SceI 1 out of 3

[0505] C5 BAC_100k24_v3 MDS42(ArecA)-sC23 + pHelper_I-SceI 3 out of 5

[0506] C6 BAC_100k24_v2 MDS42(ArecA)-sC23 + pHelper_Cas9 N / A

[0507] C7 BAC_100k24_v3 MDS42(ArecA)-sC23 + pHelper_Cas9 N / A

[0508]

[0509] Figure 4 shows recoding landscapes for each of the NGS-verified clones in the CONEXER experiments described here. The plots indicate, across the genomic region of the post-CONEXER corresponding to fragment 100k24 which is being targeted for replacement, whether the detected alelle is recoded or nonrecoded. The read coverage across the region of interest is also shown.

[0510] Together, the results show that I-Scel-based CONEXER is a suitable method for large-scale genome replacement workflows and the creation of custom synthetic genomes, including recoded genomes, and that the rates of fully recoded clones obtained by the I-Scel based method are comparable to those of the CONEXER-based method.

[0511] Table 8

[0512] Condition Selection Glucose Induced at Colony Phenotypically Fully OD count correct recoded

[0513] 1A R then D - 0.2 Approx 48 out of 48 9 out of 300 16

[0514]

[0515] 2A R then D - 0.50 Approx 13 out of 48 6 out of 300 13

[0516] 3A R then D - 0.61 Approx 7 out of 48 6 out of 7

[0517] 300

[0518] 4A Dual - 0.09 45 33 out of 45 11 out of (control) selection 24

[0519] 5A Dual - 0.6 >10,000 48 out of 48 0 out of selection 16

[0520] 6A Dual - 0.6 >10,000 48 out of 48 0 out of selection 16

[0521] IB R then D + 0.2 35 35 out of 35 -

[0522] 2B R then D + 0.39 10 No data -

[0523] 3B R then D + 0.56 0 0 -

[0524] 4B Dual + 0.13 1 1 out of 1 - selection

[0525] 5B Dual + 0.44 >1000 No data - selection

[0526] 6B Dual + 0.57 1 No data - selection

[0527]

[0528] Example 2 - Optimization of CONEXER Protocol ParametersSummary

[0529] To improve colony counts and recombination efficiency in CONEXER experiments, we systematically tested different protocol parameters.

[0530] Experimental Design

[0531] We tested the following optimization conditions:

[0532] • Selection strategy: Selection for recipient (R) first then donor (D) versus dual selection (selecting for both simultaneously)

[0533] • Different ODeoo values for arabinose induction (—0.1, ~0.2, -0.4, -0.6). For experiments with induction at ODeoo 0.4 and 0.6, cells were first incubated overnight, and the next day cultures were diluted to ODeoo 0.2 before growing to desired ODeoo.

[0534] • With or without glucose supplementation in the initial recovery phase.

[0535] For all experiments, 50 mb recovery volumes were used and ODeoo of all donor / recipient cultures was normalized to 0.5 before conjugation. We used the 100k24 CONEXER as a model system. The results of the optimization experiments are summarized in the above table 8. Adding glucose in the initial recovery did not improve culture growth, and may interfere with arabinose induction (conditions IB to 6B showed generally poorer results). Selecting for recipient (R) then donor (D) helped the cultures grow faster; however, overall efficiency of marker swap after CONEXER was lower for higher ODs (conditions 2A, 3A).Dual selection and induction at higher OD produced high colony counts (>10,000) and 100% efficiency of marker swap after CONEXER (conditions 5 A, 6A); however, this approach required overnight recovery to reach higher OD and resulted in 0% fully recoded clones.

[0536] For quicker turnover and improved efficiency, condition 1A (selection for R then D, no glucose, induction at OD 0.2) provided an optimal balance of colony count (approximately 300), phenotypic correctness (100%), and proportion of fully recoded clones (9 out of 16, approximately 56%). These results indicate that the selection strategy and timing of arabinose induction may be optimized to balance colony counts with the frequency of obtaining fully recoded clones. In some cases, selecting for recipient first then donor, without glucose supplementation, and inducing at lower ODeoo values may provide improved outcomes for CONEXER-based genome engineering workflows.References

[0537] 1. Santos, C. N., Regitsky, D. D. & Yoshikuni, Y. Implementation of stable and complex biological systems through recombinase-assisted genome engineering. Nat Commun 4, 2503, doi:10.1038 / ncomms3503 (2013).

[0538] 2. Santos, C. N. & Yoshikuni, Y. Engineering complex biological systems in bacteria through recombinase-assisted genome engineering. Nat Protoc 9, 1320-1336, doi: 10.1038 / nprot.2014.084 (2014).

[0539] 3. Krishnakumar, R. et al. Simultaneous non-contiguous deletions using large synthetic DNA and site-specific recombinases. Nucleic Acids Res 42, ell 1, doi: 10.1093 / nar / gku509 (2014).

[0540] 4. Wang, G. et al. CRAGE enables rapid activation of biosynthetic gene clusters in undomesticated bacteria. NatMicrobiol 4, 2498-2510, doi: 10.1038 / s41564-019-0573-8 (2019).

[0541] 5. Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59-64, doi: 10.1038 / nature20124 (2016).

[0542] 6. Itaya, M., Tsuge, K., Koizumi, M. & Fujita, K. Combining two genomes in one cell: stable cloning of the Synechocystis PCC6803 genome in the Bacillus subtilis 168 genome. Proc Natl Acad Sci USA 102, 15971-15976, doi: 10.1073 / pnas.0503868102 (2005).

[0543] 7. Lau, Y. H. et al. Large-scale recoding of a bacterial genome by iterative recombineering of synthetic DNA. Nucleic Acids Res 45, 6971-6980, doi: 10.1093 / nar / gkx415 (2017).

[0544] 8. Lartigue, C. et al. Creating bacterial strains from genomes that have been cloned and engineered in yeast. Science 325, 1693-1696, doi: 10.1126 / science.1173759 (2009).

[0545] 9. Ostrov, N. et al. Design, synthesis, and testing toward a 57-codon genome. Science 353, 819-822, doi: 10.1126 / science.aaf3639 (2016).

[0546] 10. Dymond, J. S. et al. Synthetic chromosome arms function in yeast and generate phenotypic diversity by design. Nature 477, 471-476, doi: 10.1038 / nature 10403 (2011).

[0547] 11. Gibson, D. G. et al. Creation of a bacterial cell controlled by a chemically synthesized genome. Science 329, 52-56, doi: 10.1126 / science.1190719 (2010).

[0548] 12. Wang, K., de la Torre, D., Robertson, W. E. & Chin, J. W. Programmed chromosome fission and fusion enable precise large-scale genome rearrangement and assembly. Science 365, 922-926, doi: 10.1126 / science. aay0737 (2019).

[0549] 13. Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514-518, doi: 10.1038 / s41586-019-l 192-5 (2019).14. de la Torre, D. & Chin, J. W. Reprogramming the genetic code. Nat Rev Genet 22, 169-184, doi:10.1038 / s41576-020-00307-7 (2021).

Claims

-67-CLAIMS1. A method of introducing a sequence of interest into a target nucleic acid, the method comprising a) providing a host cellsaid host cell comprising an episomal replicon,said episomal replicon comprising a backbone sequence and a donor nucleic acid sequence, wherein said donor nucleic acid sequence comprises in order: 5 ’ - homologous recombination sequence 1 - sequence of interest - homologous recombination sequence 2 - 3’,wherein the backbone sequence comprises a first excision site positioned adjacent to homologous recombination sequence 1 and a second excision site positioned adjacent to homologous recombination sequence 2,said host cell further comprising a target nucleic acid;b) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;c) providing an endonuclease having a recognition sequence which is not present in the host cell system apart from the first and second excision sites and which recognises the first and second excision sites;e) inducing excision of said donor nucleic acid sequence by the endonuclease; and f) incubating to allow recombination between the excised donor nucleic acid and said target nucleic acid.

2. The method according to claim 1, wherein the endonuclease is a meganuclease, a zinc finger nuclease (ZFN) or a transcription activator-like effector nuclease (TALEN).

3. The method according to claim 2, wherein the endonuclease is a meganuclease, optionally a LAGLID ADG homing endocuclease (LHE) which can optionally be selected from the group consisting of I-Scel, I-Crel, I-Dmol and I-Anil.

4. The method according to claim 3, wherein the meganuclease is I-Scel.-68-5. The method according to any one of claims 1 to 4, wherein each terminus of the excised nucleic acid comprises nucleic acid sequence derived from the backbone sequence.

6. The method according to claim 5, wherein the excised donor nucleic acid comprises 9 or fewer base pairs of nucleic acid sequence derived from the backbone sequence at each terminus.

7. The method according to any preceding claim, wherein the excision sites are selected from endonuclease cut sites which are the same or different.

8. The method according to claim 7, wherein the excision sites are arranged in the same orientation or in opposite orientations.

9. The method according to any preceding claim, wherein the episomal replicon is a bacterial artificial chromosome.

10. The method according to any preceding claim, wherein the episomal replicon is delivered to the host cell by conjugative transfer.

11. The method according to any preceding claim, wherein the target nucleic acid is the genome of the host cell.

12. The method according to any preceding claim, wherein the host cell is a prokaryotic cell.

13. The method according to any preceding claim, wherein the prokaryotic cell is Escherichia coli.

14. A method of assembling a nucleic acid sequence, the method comprising:(i) performing the steps of any one of claims 1 to 13 to introduce a first donor nucleic acid sequence into a first target nucleic acid in order to create a second target nucleic acid; and(ii) performing the steps of any one of claims 1 to 13 to introduce a second donor nucleic acid sequence into the second target nucleic acid in order to create a third target nucleic acid.-69-15. The method of claim 14, wherein part (i) and part (ii) are iterated.

16. The method of any one of claims 14 or 15, further comprising:(iii) performing the steps of any one of claims 1 to 13 to introduce a third donor nucleic acid sequence into the third target nucleic acid in order to create a fourth target nucleic acid;iterating parts (i), (ii), and (iii).

17. The method of any one of claims 14 to 16, wherein part (i) comprises the use of a donor-nucleic-acid-sequence-encoding episomal replicon comprising a first backbone sequence, and part (ii) comprises the use of a donor-nucleic-acid-sequence-encoding episomal replicon comprising a second backbone sequence, whereinthe first backbone sequence comprises a first marker or set of markers, the first excision site and the second excision site; andthe second backbone sequence comprises a second marker or set of markers, the first excision site within said second backbone sequence, and the second excision site within said second backbone sequence; whereinthe first marker or set of markers is different from the second marker or set of markers.

18. A method for constructing an episomal replicon comprising the steps of:a) providing a donor episomal replicon, said replicon comprising:a backbone, said backbone comprising a first homology region HRn which is specific for an integration step n, and a second, universal, homology region uHR, a first excision site positioned adjacent to HRn and a second excision site positioned adjacent to uHR;a donor nucleic acid DNAn;a double selection cassette, comprising positive and negative selection markers;b) providing a host cell comprising an assembly episomal replicon comprising a double selection cassette comprising positive and negative selection markers, flanked by HRn and uHR, the double selection cassette in the assembly replicon comprising different markers to the selection cassette in the donor replicon;-70-c) providing helper protein(s) capable of supporting nucleic acid recombination in said host cell;c) providing an endonuclease having a recognition sequence which is not present in the host cell system apart from the first and second excision sites and which recognises the first and second excision sites;e) inducing excision of said donor nucleic acid sequence DNAn by the endonuclease in the host cell; andf) incubating to allow recombination between the excised donor nucleic acid and said assembly replicon to form a second assembly replicon, which comprises the nucleic acid DNAn.

19. The method according to claim 18, wherein the endonuclease is a meganuclease, optionally a LAGLID ADG homing endocuclease (LHE) which can optionally be selected from the group consisting of I-Scel, I-Crel, I-Dmol and I-Anil.

20. The method according to claim 19, wherein the meganuclease is I-Scel.

21. The method according to any one of claims 18 to 20, wherein each terminus of the excised nucleic acid comprises nucleic acid sequence derived from the backbone sequence.

22. The method according to claim 21 , wherein the excised donor nucleic acid comprises 6 or fewer base pairs of nucleic acid sequence derived from the backbone sequence at each terminus.

23. The method according to any one of claims 18 to 22, wherein the episomal replicon is a bacterial artificial chromosome.

24. The method according to any one of claims 18 to 23, wherein the episomal replicon is delivered to the host cell by conjugative transfer.

25. The method according to claim 24, wherein the episomal replicon is comprised in a donor host cell, and the assembly replicon is comprised in a recipient host cell; the donor replicon is transferred to-71-the recipient host cell by conjugative transfer; and the donor host cell comprises a non-transferrable F’ plasmid.

26. The method according to any one of claims 17 to 25, wherein the host cell is a prokaryotic cell.

27. The method according to claim 26, wherein the prokaryotic cell is Escherichia coli.

28. The method of any one of claims 17 to 27, wherein the donor nucleic acid DNAn comprises a homology region HRn+1, and the method further comprises the steps of introducing into the host cell a further donor episomal replicon comprising a second donor nucleic acid DNAn+1, inducing excision of said donor nucleic acid sequence DNAn+1 by the endonuclease in the host cell; and incubating to allow recombination between the excised donor nucleic acid DNAn+1 and said second assembly replicon to form a third assembly replicon, which comprises the nucleic acid DNAn and nucleic acid DNAn+1.

29. The method of claim 28, iteratively repeated.

30. A method according to any one of claims 14 to 17, wherein the episomal replicon of the steps of claims 1 to 14 is constructed according to any one of claims 18 to 29.

31. The method according to any one of claims 1 to 30, wherein the host cell is lacking competent recA and / or recO.

32. The method according to claim 31, wherein the host cell lacks recA (ArecA).