Methods for editing a nucleic acid sequence

JP2024542108A5Pending Publication Date: 2025-11-10UNITED KINGDOM RESEARCH AND INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024526648
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-03
Filing Date
2022-11-03
Publication Date
2025-11-10

AI Technical Summary

Technical Problem

Current methods for introducing synthetic DNA into the E. coli genome are complex, time-consuming, and limited by the need for unique spacers at each locus, complicating scalability and efficiency.

Method used

A method using episomal replicons with universal spacers and RNA-guided DNA endonucleases like CRISPR-Cas9 to excise and integrate donor nucleic acids into the genome, allowing for rapid and scarless incorporation of synthetic DNA.

Benefits of technology

Enables efficient and scalable introduction of synthetic DNA into the E. coli genome, reducing the complexity and time required for large-scale genome modification and synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In one aspect, the present invention relates to a method for introducing a desired sequence into a target nucleic acid. The present invention also relates to a method for assembling nucleic acid sequences, comprising repeating the method for introducing a desired sequence into a target nucleic acid, as well as to the assembly of replicons encoding increased amounts of heterologous nucleic acids.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] In one aspect, the present invention relates to a method for introducing a desired sequence into a target nucleic acid. The present invention also relates to a method for assembling a nucleic acid sequence comprising repeating the method for introducing a desired sequence into a target nucleic acid. [Background technology]

[0002] Strategies for replacing genomic DNA with synthetic DNA (Non-Patent Documents 1-12) allow genome modification and provide the basis for a powerful technique for creating entirely synthetic genomes that cannot be created by editing methodologies alone. Genome synthesis has been used to create synthetic genomes in two organisms: mycoplasma (1 Mb), which has been used to investigate genome minimization (Non-Patent Document 13), and Escherichia coli (E. coli) (4 Mb), which has been used to create a recoded organism (Non-Patent Document 14). Work in E. coli has removed over 18,000 synonymous codons to create a genetic code-compacted organism that uses only 61 codons to code for canonical amino acids. The recoded E. coli with a compressed genetic code provides the basis for creating virus-resistant cells and for the reassignment of sense codons for the incorporation of non-canonical amino acids and the synthesis of encoded non-canonical polymers (Non-Patent Document 15). Attempts to synthesize other genomes, and strategies to recode, minimize, or rearrange genomes, are currently underway (Non-Patent Documents 9, 10, 16-18).

[0003] Genome synthesis in E. coli was based on Replicon Excision Enhanced Recombination (REXER) (Non-Patent Document 5). In REXER (Figure 1a), a bacterial artificial chromosome (BAC) carrying an insert consisting of the desired synthetic DNA flanked by double selection cassettes (consisting of a positive selection marker and a negative selection marker) and a 50-152 bp region of homology to the genome is transformed into cells carrying a helper plasmid encoding CRISPR / Cas9 and the lambda Red recombination machinery. One clone carrying the correct BAC is isolated and made competent, and then a spacer array is introduced into the competent cells to activate Cas9-mediated excision of the insert carrying the synthetic DNA from the BAC. The spacer sequence is designed such that the spacer RNA base pairs with a unique sequence within the homology region of the insert, precisely cleaving the link between the insert and the BAC backbone (Figure 1b). The correctly excised inserts are then used by the Lambda Red recombination machinery to insert the synthetic DNA into the genome in place of the corresponding genomic DNA. Separate double selection cassettes in the genomic and synthetic DNA are used to select for the desired integrations (Figure 1a). Starting from cells with appropriately marked genomes, it takes 4 days to obtain clonal colonies on agar plates after REXER.

[0004] One-step REXER has been used to replace up to 136 kb of genome with synthetic recoded DNA. The genome resulting from one-step REXER provides the template for the next round of REXER, and repeated rounds of REXER (GENESIS, Genome interchange stepwise synthesis) allow larger sections of the E. coli genome to be replaced with synthetic DNA. Thirty-eight REXER steps, each of which requires the design, synthesis, cloning, and validation of custom-made spacer pairs, were used to replace the entire E. coli genome (across seven strains) with synthetic recoded DNA. The recoded DNA was then assembled into a strain by conjugation to create the recoded organism. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Santos, CN, Regitsky, DD & Yoshikuni, Y. Implementation of stable and complex biological systems through recombinase-assisted genome engineering. Nat Commun 4, 2503, doi:10.1038 / ncomms3503 (2013) [Non-Patent Document 2] Santos, CN & Yoshikuni, Y. Engineering complex biological systems in bacteria through recombinase-assisted genome engineering. Nat Protoc 9, 1320-1336, doi:10.1038 / nprot.2014.084 (2014) [Non-Patent Document 3] Krishnakumar, R. et al. Simultaneous non-contiguous deletions using large synthetic DNA and site-specific recombinases. Nucleic Acids Res 42, e111, doi:10.1093 / nar / gku509 (2014).

Outdoor Tools 4

Direct Environment 5

Outdoor Configuration6

Direct Environment 7

Non-licensed Document 13

Non-licensed Document 14

Non-licensed Document 15

Non-licensed Document 16

Non-licensed Document 17

Outdoor Tools 18

Outdoor Tools 19

Outdoor Tools20

Direct Environment21

Outdoor Tools22

[0006] Strategies to further simplify and expedite the introduction of synthetic DNA into the E. coli genome will be key for future large-scale genome engineering and synthesis efforts. [Means for solving the problem]

[0007] In one aspect of the invention, there is provided a method for introducing a desired sequence into a target nucleic acid, comprising the steps of: a) providing a host cell, the host cell comprises an episomal replicon; the episomal replicon comprises a backbone sequence and a donor nucleic acid sequence; The donor nucleic acid sequence comprises, in order: 5'-homologous recombination sequence 1-desired sequence-homologous recombination sequence 2-3'; the backbone sequence comprises a first excision site located adjacent to homologous recombination sequence 1 and a second excision site located adjacent to homologous recombination sequence 2; the host cell further comprises a target nucleic acid, b) providing helper proteins capable of assisting recombination of nucleic acids in said host cell; c) providing an RNA-guided DNA endonuclease; d) providing a first RNA molecule comprising a sequence specific to the first excision site and a second RNA molecule comprising a sequence specific to the second excision site, the first and second RNA molecules being involved in directing an RNA-guided DNA endonuclease during excision; e) directing excision of the donor nucleic acid sequence by an RNA-guided DNA endonuclease; and f) incubating to allow recombination to occur between the excised donor nucleic acid and the target nucleic acid. The method includes:

[0008] The RNA-guided DNA endonuclease can be a CRISPR-Cas nuclease, the first RNA molecule can include a spacer specific to a first excision site, and the second RNA molecule can include a spacer specific to a second excision site. The CRISPR-Cas nuclease can be Cas9. The first RNA molecule and / or the second RNA molecule can be encoded by an episomal replicon.

[0009] In one embodiment, each end of the excised nucleic acid comprises a nucleic acid sequence derived from the backbone sequence. The excised donor nucleic acid may comprise up to 6 or 5 base pairs of nucleic acid sequence derived from the backbone sequence at each end.

[0010] In one embodiment, the episomal replicon is a bacterial artificial chromosome. The episomal replicon can be delivered to the host cell by conjugal transfer.

[0011] The target nucleic acid can be the genome of a host cell. The host cell can be a prokaryotic cell, such as Escherichia coli.

[0012] In one aspect of the invention, there is provided a method for assembling a nucleic acid sequence, comprising the steps of: (i) performing steps of a method of the invention for introducing a desired sequence into a target nucleic acid to introduce a first donor nucleic acid sequence into a first target nucleic acid, thereby generating a second target nucleic acid; and (ii) performing steps of the method of the invention for introducing a desired sequence into a target nucleic acid to introduce a second donor nucleic acid sequence into the second target nucleic acid, thereby generating a third target nucleic acid; Part (i) and part (ii) may be repeated.

[0013] In one embodiment, the sequence of the first RNA molecule for part (i) is identical in each repeat and / or the sequence of the second RNA molecule for part (i) is identical in each repeat, the sequence of the first RNA molecule for part (ii) is identical in each repeat and / or the sequence of the second RNA molecule for part (ii) is identical in each repeat.

[0014] In one embodiment, the method comprises: (iii) performing steps of the method of the invention for introducing a desired sequence into a target nucleic acid to introduce a third donor nucleic acid sequence into a third target nucleic acid, thereby generating a fourth target nucleic acid; Repeating parts (i), (ii), and (iii). Further comprising: The sequence of the first RNA molecule for part (iii) is identical in each repeat, and / or the sequence of the second RNA molecule for part (iii) is identical in each repeat.

[0015] In certain embodiments, part (i) comprises the use of an episomal replicon encoding a donor nucleic acid sequence comprising a first backbone sequence, and part (ii) comprises the use of an episomal replicon encoding a donor nucleic acid sequence comprising a second backbone sequence; a first scaffold sequence comprising a first marker or set of markers and encoding a first RNA molecule specific for a first excision site within said first scaffold sequence, and encoding a second RNA molecule specific for a second excision site within said first scaffold sequence; a second scaffold sequence comprising a second marker or set of markers and encoding a first RNA molecule specific for a first excision site within said second scaffold sequence and encoding a second RNA molecule specific for a second excision site within said second scaffold sequence; The first marker or set of markers is different from the second marker or set of markers.

[0016] In a further aspect of the invention there is provided a method for constructing an episomal replicon comprising the steps of: a) providing a donor episomal replicon, said replicon comprising: a scaffold comprising a universal spacer sequence; a first homology region HRn specific for integration step n, and a second universal homology region uHR; a first excision site located adjacent to HRn and a second excision site located adjacent to uHR; A donor nucleic acid DNAn located between HRn and uHR, and Double selection cassette containing a positive selection marker and a negative selection marker The steps include b) providing a host cell comprising an assembly episomal replicon comprising a double selection cassette comprising a positive selection marker and a negative selection marker flanked by HRn and uHR, wherein the double selection cassette in the assembly replicon comprises a different marker than the selection cassette in the donor replicon; c) providing helper proteins capable of assisting recombination of nucleic acids in said host cell; d) providing an RNA-guided DNA endonuclease; e) providing a first RNA molecule comprising a sequence specific for a first excision site and a second RNA molecule comprising a sequence specific for a second excision site, the first and second RNA molecules being involved in directing an RNA-guided DNA endonuclease during excision; f) inducing excision of the donor nucleic acid sequence DNAn by an RNA-guided DNA endonuclease in a host cell; and g) incubating to allow recombination between the excised donor nucleic acid and the assembly replicon to form a second assembly replicon comprising the nucleic acid DNAn; The method includes:

[0017] The RNA-guided DNA endonuclease can be a CRISPR-Cas nuclease, the first RNA molecule can include a spacer specific for a first excision site, and the second RNA molecule can include a spacer specific for a second excision site. The CRISPR-Cas nuclease can be Cas9. The first RNA molecule and / or the second RNA molecule can be encoded by an episomal replicon.

[0018] In one embodiment, each end of the excised nucleic acid comprises a nucleic acid sequence derived from the backbone sequence. The excised donor nucleic acid may comprise 12, 10, 8, 6, 4, or 2 base pairs of nucleic acid sequence derived from the backbone sequence at each end. Preferably, the excised donor nucleic acid comprises no more than 6 or 5 base pairs of nucleic acid sequence derived from the backbone sequence at each end.

[0019] In one embodiment, the episomal replicon is a bacterial artificial chromosome. The episomal replicon can be delivered to the host cell by conjugal transfer.

[0020] In a preferred embodiment of the present invention, the donor episomal replicon is contained in the donor host cell and the cell assembly replicon is contained in the recipient host cell. At the conjugation between the donor host cell and the recipient host cell, the donor episomal replicon can be advantageously transferred to the recipient host cell. The donor host cell preferably contains a non-transmissible F' plasmid, so that the F' plasmid is not transferred to the recipient host cell. Preferably, the F' plasmid is non-transmissible due to deletion of oriT. The donor episome can be transferable and carry oriT.

[0021] Selection of host cells can be performed as described, except utilizing the positive and negative selection markers present in the recombinant donor and the assembly replicon DNA.

[0022] The target nucleic acid may be the genome of a host cell. The host cell may be a prokaryotic cell, such as E. coli.

[0023] In one embodiment, the donor nucleic acid comprises a homology region HRn+1, and the method further comprises a further step (h) of introducing into the host cell a further donor episomal replicon comprising donor nucleic acid DNAn+1, and inducing excision of said donor nucleic acid sequence DNAn+1 in the host cell by an RNA-guided DNA endonuclease, and incubating to allow recombination between the excised donor nucleic acid DNAn+1 and said second assembly replicon to form a third assembly replicon comprising nucleic acid DNAn and nucleic acid DNAn+1.

[0024] The method steps may be repeated iteratively.

[0025] The replicons provided in this aspect of the invention may be used in the above-described aspects describing methods for assembling nucleic acid sequences.

[0026] In an advantageous embodiment of all aspects of the invention, the host cell lacks competent recA and / or recO. Preferably, the host cell lacks recA (ΔrecA). [Brief description of the drawings]

[0027] [Figure 1]Figure 1 shows the steps in REXER-mediated integration of approximately 100 kb of synthetic DNA into the E. coli genome using homology region (HR)-specific spacers, and the relationship between cleavage directed by HR-specific spacers and universal spacers. a: REXER allows integration of over 100 kb of synthetic DNA (pink) into the genome via replacement of genomic DNA (as shown herein) or by insertion into the genome. Bacterial artificial chromosomes (BACs) carrying the desired synthetic DNA are electroporated into competent cells with appropriately marked genomes, which also carry a helper plasmid encoding the Cas9 protein and lambda Red recombination components. The cells are then induced with arabinose to express the helper plasmid genes and are made electrocompetent again. The HR-specific spacer array (either plasmid-based as shown or as linear DNA) is then electroporated into cells, resulting in CRISPR / Cas9-mediated in vivo excision of the synthetic DNA from the BAC, flanked by the double selection cassette and the HR to the genome. The lambda Red recombination machinery then uses the HR to direct integration of the excised DNA into the genome. Triangles indicate Cas9 cleavage sites in the HR (grey boxes) flanking the synthetic DNA. +1, blue, kanR; -1, yellow, rpsL; +2, green, cat; -2, pink, sacB; +3, dark blue, tetR; -3, purple, pheS*; +4, orange, ampR. b: So far, we have designed spacer RNAs specific for each HR flanking the locus of recombination. The BAC sequences flanking the synthetic DNA insert have a constant PAM sequence (black boxes). Targeting Cas9 with HR-specific spacer RNAs allows precise excision of synthetic DNA at the ends of HR1 and HR2. c: Universal spacer RNA targets Cas9 to constant sequences in the BAC backbone.To create a universal cleavage site for any BAC, independent of the REXER locus, Cas9 is directed to the PAM sequence immediately adjacent to the HR, which adds an additional 6 bp of non-homologous sequence to both ends of the excised DNA fragment. [Figure 2-1] ~ [Figure 2-3]Figure 1 shows that REXER in a universal spacer results in scarless genomic integration of large synthetic DNA fragments. a: BAC backbone used for total synthesis of E. coli genome. Two BACs, called odd BAC and even BAC, have different positive and negative selection cassettes, which repeat REXER. We designed a universal 1 spacer for all odd BACs and a universal 2 spacer for all even BACs (Table 3 and SEQ ID NOs: 24 and 25). One universal spacer RNA (blue) targets a sequence within the BAC backbone from 5' to the insert. This sequence is common to both backbones. A universal spacer RNA targets the BAC backbone from 3' to the insert. These 3' sequences are different in the two backbones, therefore two different spacer sequences (yellow and red, respectively) were designed. Cleavage sites are indicated by colored triangles, PAM sequences are indicated by black boxes, selection cassettes are indicated by colored arrows, and synthetic DNA is indicated in pink. b: Genotyping validation of 5' and 3' genomic integration sites after REXER using universal spacer at five loci. At the 5' locus, successful replacement by REXER removes a double selection cassette from the genome, while another double selection cassette is inserted at the 3' locus. Eleven post-REXER clones were genotyped for each experiment. Triangles indicate the expected PCR product size at each locus before (white) and after (black) REXER. c: Sequence validation of the ends of integration sites after REXER using universal spacer RNA. The excised synthetic DNA flanked by HR and 6 bp non-homologous sequences (tilted sites) is shown above the sequence expected for scarless integration. Five post-REXER clones were sequenced for each experiment. We did not observe non-homologous end integration in any of the clones, nor did any point mutations appear. The colored triangles indicate the Cas9 cleavage sites according to the respective spacer sequences listed in panel a.d: Summary of the recoding status of REXER using HR-specific spacers and universal spacers, respectively. We performed REXER to replace 95.6 kb of wild-type genomic DNA with synthetic DNA (100k24). The synthetic DNA is mostly homologous to the corresponding genomic DNA, except that 410 codons were replaced with synonymous codons (Non-Patent Document 14). Cas9 cleavage was initiated with HR-specific spacer RNA or universal spacer RNA. Ten post-REXER clones were fully sequenced by NGS in each experiment (total of 20 from two independent replicates). The graph summarizing the recoding status shows the average frequency at which each recoded codon was integrated across the locus. Overall, REXER with universal spacer RNA yielded eight clones in which all 410 codons were completely replaced, whereas REXER with HR-specific spacer RNA yielded four completely recoded clones in this region (recoding status of individual clones in Figure 5). [Diagram 3]Conjugation coupled with programmed excision for enhanced recombination (CONEXER). The combination of episomal transfer, excision of synthetic DNA at universal spacers, and homology-guided recombination provides a rapid, simple, and standardized method for large-scale genome modification and genome synthesis. a: A BAC carrying a universal spacer array (gray bar) and oriT sequence (red arrow) is transferred from a donor cell to a recipient cell via conjugal transfer over 1 h using a non-transmissible F' plasmid. The recipient cell has a helper plasmid and a properly marked genome (as in REXER). Recipient cells that have obtained the BAC are then selected, and donor cells are eliminated by selection for both +3 and +1 (1.5 h with arabinose to induce Cas9 and lambda red components, followed by 2.5 h with glucose to stop further expression). Replacement of genomic DNA with synthetic DNA is then selected for by further selection for loss of -2. Selection for loss of -3 ensures loss of the BAC backbone. Selectable markers are: +1, blue, kanR; -1, yellow, rpsL; +2, green, cat; -2, pink, sacB; +3, dark blue, tetR; -3, purple, pheS*. b: Summary of the recoding status of 84 clones obtained with CONEXER using 100k24 (similar to Fig. 2d). [Figure 4] FIG. 1 shows a genome map. [Figure 5-1] ~ [Figure 5-4] Figure 1 shows the recoding status of REXER using HR-specific spacers and universal spacers, respectively. The recoding status of individual clones is shown. [Figure 6-1] ~ [Figure 6-2]Figure 1 shows BAC Stepwise Insertion Synthesis (BASIS). BAC Stepwise Insertion Synthesis (BASIS) for the assembly of large DNA repeats in BAC. a: The donor BAC has HRn and uHR homologous to the recipient BAC. The BAC backbone has oriT and universal spacer. uHR is a universal homology region at all steps of insertion. HRn is specific to the nth step of insertion. The BAC insert has HRn+1, which serves as HRn at the (n+1)th step. The BAC has the double selection cassettes -3, +3 shown. The assembly BAC has different double selection cassettes -1, +1 shown, flanked by HRn and uHR. The DNA insert is excised from the donor BAC and inserted into the assembly BAC of the recipient cell. The green triangle indicates the cut site for Cas9 excision. Note that HRn is described as HR1 in this specification. In the example shown, the selectable markers are +1, blue, kanR; -1, yellow, rpsL; +3, purple, hygroR; -2, orange, PheS; +4, petrol, GentamycinR. b: BASIS workflow. A donor BAC is delivered by conjugation to a recipient cell carrying an assembly BAC and expressing Cas9 and lambda Red components. As shown in (a), an insert is excised from the donor BAC and inserted into the recipient BAC. Repetition of this process using alternating marker sets allows for the insertion of n DNA fragments into the assembly BAC. c: Three BACs encoding segments of the CFTR gene were assembled in yeast. The complete CFTR gene sequence was reconstructed iteratively via BASIS and verified by next generation sequencing (NGS). d: Human BACs covering the indicated regions of chromosome 21 were used as substrates for subsequent assembly. Intermediate and final assembly products were verified by NGS. The final 503 Kbp BAC contained a 495 Kbp human DNA insert.This BAC can serve as a substrate for further repeated insertions, thus allowing much longer lengths of human DNA to be assembled into episomes. [Figure 7] Figure 9 shows that deletion of recA from the host genome increases the frequency of fully recoded clones in CONEXER-mediated genome replacement. a: Screening of KO strains for increased frequency of fully recoded clones in CONEXER-mediated replacement of genome fragment 100k24 reveals that deletion of recA and recO improves intact integration of synthetic DNA. b: When recA is deleted, the frequency of fully recoded clones after CONEXER-mediated genome replacement is increased for all tested fragments (100k24-100k28). Colony forming units (CFU) recorded in these experiments are shown in Figure 9. [Figure 8]Figure 1. Sequential genome synthesis. a: Conjugally delivered episomal excised synthetic DNA recombines by replacing a large piece (100k24) of the genome (-2 / +2 at LS23) that is appropriately marked. Selection ensures that only cells that have lost -2 / +2 (at LS23) and integrated -1 / +1 (at LS24) survive. Clones resulting from selection were pooled and subjected to the next round of CONEXER-mediated genome replacement (100k25). Selection ensures that only cells that have lost -1 / +1 (at LS24) and integrated -2 / +2 (at LS25) survive. This process was repeated three more times (100k26, 100k27, and 100k28) until a population of cells with -1 / +1 (at LS28) was obtained, for a total of five times. A subset of these cells is expected to have synthetic DNA integrated sequentially across the entire 500Kbp region. Selectable markers are +1, blue, kanR; -1, yellow, rpsL; +2, green, cat; -2, pink, sacB. b: Summary of the recoding status of 182 clones obtained from the sequential genome synthesis of 100k24 to 100k28. Nineteen of the 182 sequenced clones were completely recoded across the entire 500Kbp genomic section. [Figure 9] Figure 1 shows colony forming units for CONEXER-mediated genome replacement in different genome backgrounds. Screening of KO strains with CONEXER-mediated replacement of genome section 100k24. Colony forming units (CFU) obtained in each experiment are shown. Deletion of recA results in a reduction of CFU. [Figure 10]Figure 5. Progressive synthesis of a 500Kbp E. coli genome. Conjugation-enhanced replacement of 100Kbp genomic sections with synthetic DNA allows for progressive genome synthesis (Figure 5). The episomal excised synthetic DNA is recombined by replacing a large section (100k24) of the appropriately marked genome (-2 / +2 at LS23). Selection ensures that only cells that have lost -2 / +2 (at LS23) and integrated -1 / +1 (at LS24) survive. Clones resulting from selection are pooled and subjected to the next round of CONEXER-mediated genome replacement (100k25). Selection ensures that only cells that have lost -1 / +1 (at LS24) and integrated -2 / +2 (at LS25) survive. This process was repeated three more times (100k26, 100k27, and 100k28) until a population of cells with -1 / +1 (at LS28) was obtained. A subset of these cells is expected to have continuously integrated synthetic DNA over the entire 500 Kbp region. Selectable markers are +1, blue, kanR; -1, yellow, rpsL; +2, green, cat; -2, pink, sacB. Each step of integration of over 100 Kbp of synthetic DNA by CONEXER requires 2 days. Continuous synthesis of 500 Kbp can therefore be achieved in 10 days. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] Methods disclosed in the prior art, such as REXER, require a new set of homology region (HR)-specific spacers to be cloned for each locus being targeted. These spacers can be difficult to clone, and cloning of spacers can be expensive and time-consuming. For example, a recent E. coli genome synthesis required cloning of 78 unique spacers. Each new set of spacers must be designed to avoid undesired cleavage of the target nucleic acid. In addition, changes in the spacer sequence can affect the excision efficiency, which may be the cause of differences in the efficiency of REXER at different loci. The need for HR-specific spacer RNAs complicates the workflow and may limit the scalability of REXER.

[0029] We hypothesized that if sequences within the episomal replicon scaffold, rather than the insert, could be used to direct the excision of the insert from the episomal replicon, then the same spacer pair, i.e., the "universal spacer," could be used to perform REXER at any target locus with a given episomal replicon scaffold. This would greatly simplify the introduction of synthetic DNA into target nucleic acids, e.g., the E. coli genome, enabling a rapid method for large-scale genome modification and whole genome synthesis. However, previous methods use spacers for Cas9 that cannot direct precise cleavage of the junction between the episomal replicon scaffold and the insert unless these spacers are specifically bound within the insert. Indeed, the placement of spacers that bind within the scaffold and minimize the distance of the cleavage site from the junction between the scaffold and the insert results in the excision of an insert that is flanked by six base pairs of the scaffold at each end (Figure 1c). These 6 bp sequences are not commonly homologous to regions of the target nucleic acid adjacent to the region complementary to the insert. Therefore, it was unclear whether a derivative of REXER using a universal spacer would function and whether it would result in unwanted insertions or deletions.

[0030] As shown herein, we provide universal spacers and demonstrate that these spacers can be used for scarless integration of synthetic DNA into target nucleic acids. Furthermore, we develop a rapid protocol for replacing genomic DNA with synthetic DNA that can introduce 100 kb of synthetic DNA into the genome of a cell in one day. This approach is based on the REXER approach disclosed in (Wang et al. (Defining synonymous codon compression schemes by genome recoding, Nature 2016 November 03; 539(7627):59-64, doi:10.1038 / nature20124) and WO 2018 / 020248, each of which is incorporated by reference. In the examples, we develop universal spacers for two BAC scaffolds used in REXER-based whole genome synthesis.

[0031] These experiments reveal that the presence of short non-homologous ends on the excised synthetic DNA does not prevent recombination and integration efficiency, and scarless integration can still be achieved. Without being bound to a particular theory, the inventors suggest that the non-homologous ends of DNA within the BAC can be removed by an exonuclease prior to recombination, or by a flap endonuclease such as EcoIX during recombination, similar to the mechanism described for eukaryotic FEN1. Thus, the inventors have discovered that the cut site does not need to be exactly at the junction between the donor nucleic acid and the backbone of the episomal replicon.

[0032] This finding is important because recognition of the sequences flanking both sides of the actual cleavage site is a requirement for many RNA-guided DNA endonucleases. Therefore, in order to cleave at the exact junction, it is necessary for the RNA-guided DNA endonucleases to recognize parts of the backbone of the episomal replicon and parts of the donor nucleic acid. The inventors have surprisingly found that the excised donor nucleic acid can tolerate regions of the backbone sequence without affecting recombination, which allows the "excision site", i.e., the combination of the actual cleavage site and the required flanking sequence, to be entirely located within the backbone sequence. Thus, the complexity of the process is reduced because the RNA-guided DNA endonucleases do not need to recognize any of the donor nucleic acid sequences that change between rounds of the method.

[0033] Thus, one embodiment of the present invention provides a method for introducing a desired sequence into a target nucleic acid, comprising the steps of: a) providing a host cell, the host cell comprises an episomal replicon; the episomal replicon comprises a backbone sequence and a donor nucleic acid sequence; The donor nucleic acid sequence comprises, in order: 5'-homologous recombination sequence 1-desired sequence-homologous recombination sequence 2-3'; the backbone sequence comprises a first excision site located adjacent to homologous recombination sequence 1 and a second excision site located adjacent to homologous recombination sequence 2; the host cell further comprises a target nucleic acid, b) providing helper proteins capable of assisting recombination of nucleic acids in said host cell; c) providing an RNA-guided DNA endonuclease; d) providing a first RNA molecule comprising a sequence specific to the first excision site and a second RNA molecule comprising a sequence specific to the second excision site, the first and second RNA molecules being involved in directing an RNA-guided DNA endonuclease during excision. e) directing excision of the donor nucleic acid sequence by an RNA-guided DNA endonuclease; and f) Incubating to allow recombination between the excised donor nucleic acid and the target nucleic acid. The method includes:

[0034] Steps a), b), c), and d) do not have to be performed in order or as separate steps. For example, an RNA-guided DNA endonuclease may be provided before helper proteins that can assist recombination of nucleic acids in said host cell are provided. The endonuclease and associated RNA can be provided separately, at different times, and in any order. However, all components required for excision, e.g., the endonuclease and associated RNA, are provided before induction of excision. In addition, all helper proteins that can assist recombination of nucleic acids are provided before incubation to allow recombination to occur. Steps a), b), c), and d) may be performed simultaneously.

[0035] Episomal replicon comprises a backbone sequence and a donor nucleic acid sequence, and the backbone sequence does not overlap with the donor nucleic acid sequence.Therefore, a constant backbone sequence can be used to introduce multiple different donor nucleic acid sequences.In the embodiment in which multiple donor nucleic acids are introduced into the target by repeating the method of the present invention, two or more types of backbone sequences can be used.For example, backbones containing different markers, such as selection markers, can be used, so that the successful introduction of desired sequences can be identified at each step.

[0036] The first and second excision sites allow for cleavage of the episomal replicon to excise the donor nucleic acid sequence. The excision site provides all the nucleic acid sequences necessary for recognition and cleavage by the endonuclease, and therefore no part of homologous recombination sequence 1 or homologous recombination sequence 2 is recognized to effect cleavage. This is in contrast to the prior art (Figure 1b), where homologous recombination sequence 1 and homologous recombination sequence 2 provide the sequences that are recognized to allow cleavage.

[0037] The excision sites are adjacent to the homologous recombination sequences of the donor nucleic acid. Preferably, each excision site is contiguous with a homologous recombination sequence. In such embodiments, there are no base pairs between the sequence required for excision and the homologous recombination sequence. In other embodiments, an intervening sequence of 1, 2, 3, or 4 base pairs may be tolerated.

[0038] Any backbone sequence between the site where the episomal replicon is cut and the homologous recombination sequence is present as part of the excised nucleic acid. Thus, the excised nucleic acid can include a first portion of the backbone sequence, the donor nucleic acid, and a second portion of the backbone sequence. Thus, in one embodiment, each end of the excised nucleic acid includes a nucleic acid sequence derived from the backbone sequence.

[0039] The first and second portions of the backbone sequence can be 10, 9, 8, 7, 6, or 5 base pairs or less in length. In certain embodiments, the first and / or second portions of the backbone sequence are 6 base pairs in length. This embodiment is particularly relevant when cleavage is performed by Cas9, which can cleave 3 base pairs upstream of a 3 base pair PAM.

[0040] The method of the present invention includes providing a first RNA molecule that contributes to directing the RNA-guided DNA endonuclease to recognize a first excision site, and a second RNA molecule that contributes to directing the RNA-guided DNA endonuclease to recognize a second excision site. The first and second RNA molecules are specific for a region of the episomal replicon entirely contained within the backbone sequence. The first and second RNA molecules do not recognize sequences within the donor nucleic acid sequence or within the homologous recombination sequence. The first and second RNA molecules may be of the same sequence or of different sequences.

[0041] An example of an RNA-guided DNA endonuclease is CRISPR-Cas nuclease. In such an embodiment, the first RNA molecule may include a spacer specific to the first excision site, and the second RNA molecule may include a spacer specific to the second excision site. As discussed herein, the spacers of the first and second RNA molecules are specific to a region of the episomal replicon that is entirely contained within the backbone sequence. The region of the episomal replicon to which the spacer is specific can be called a protospacer. Thus, the first excision site includes the entire protospacer cognate to the spacer of the first RNA molecule, and the second excision site includes the entire protospacer cognate to the spacer of the second RNA molecule. The spacers of the first and second RNA molecules do not recognize sequences in the donor nucleic acid sequence or in the homologous recombination sequence. The spacers of the first and second RNA molecules may be of the same sequence or different sequences.

[0042] Thus, in embodiments in which cleavage is performed by a CRISPR-Cas nuclease, the excision site comprises a protospacer cognate to the spacer RNA. The excision site also comprises a PAM.

[0043] In certain embodiments, the first and second RNA molecules are encoded by episomal replicon and are part of the backbone sequence. In REXER, the RNA molecule encodes a spacer specific to the homologous region of the donor nucleic acid sequence, but the spacer must not be encoded by the backbone of the episomal replicon, since the spacer sequence must vary depending on the donor nucleic acid sequence to be introduced. The present invention is not so limited, and therefore the method can be short, since the RNA molecule for directing excision can be encoded by the episomal replicon that also contains the donor nucleic acid sequence. This is particularly advantageous for the repeat method in which two or more types of backbone sequences are used, since the relevant spacers can be directly encoded by each backbone.

[0044] In some embodiments, the helper protein capable of assisting recombination of nucleic acids and / or at least one endonuclease capable of cleaving the first and / or second excision sites is encoded on an episomal replicon. In some embodiments, the helper protein capable of assisting recombination of nucleic acids and / or at least one endonuclease capable of cleaving the first and / or second excision sites is encoded on a separate episomal replicon, such as a plasmid. This separate episomal replicon may be known as a helper episomal replicon or a helper plasmid.

[0045] In certain embodiments, a method comprising: 1) providing a host cell, the host cell comprises an episomal replicon, e.g., a BAC; the episomal replicon comprises a backbone sequence and a donor nucleic acid sequence; The donor nucleic acid sequence comprises, in order: 5'-homologous recombination sequence 1-desired sequence-homologous recombination sequence 2-3'; the backbone sequence comprises a first excision site located adjacent to homologous recombination sequence 1 and a second excision site located adjacent to homologous recombination sequence 2, the first excision site comprises a first protospacer, and the second excision site comprises a second protospacer; the host cell further comprises a target nucleic acid, e.g., a genome of the host cell; 2) providing helper proteins capable of assisting recombination of nucleic acids in the host cell; 3) providing a CRISPR-Cas nuclease, such as a Cas9 nuclease; 4) providing a first RNA molecule comprising a spacer specific to the first protospacer and a second RNA molecule comprising a spacer specific to the second protospacer; 5) inducing excision of the donor nucleic acid sequence by a CRISPR-Cas nuclease; and 6) Incubating to allow recombination to occur between the excised donor nucleic acid and the target nucleic acid. A method is provided that includes:

[0046] Steps 1), 2), 3), and 4) do not have to be performed in order, nor do they have to be performed as separate steps. For example, CRISPR-Cas nuclease may be provided before helper proteins that can assist recombination of nucleic acids in said host cell are provided. In addition, CRISPR-Cas nuclease and the first and second RNA molecules can be provided separately, at different times, and in any order. However, all components required for excision are provided before induction of excision. In addition, all helper proteins that can assist recombination of nucleic acids are provided before incubation to allow recombination to occur. Steps 1), 2), 3), and 4) may be performed simultaneously.

[0047] The episomal replicon can be provided to the host cell by conjugation. Thus, the method can further comprise delivering the episomal replicon to the host cell by conjugative transfer. The episomal replicon, e.g., BAC, can be included in the donor cell for transfer. The episomal replicon can include an origin of transfer (oriT). The donor cell can include a non-transmittable F' plasmid.

[0048] The methods of the invention are particularly suitable for iterative assembly of large synthetic nucleic acid sequences, for example for the construction of artificial genomes. Thus, in one aspect of the invention, there is provided a method of assembling a nucleic acid sequence, comprising: (i) performing the steps of any method of the invention to introduce a first donor nucleic acid sequence into a first target nucleic acid to generate a second target nucleic acid; and (ii) performing the steps of any method of the invention to introduce a second donor nucleic acid sequence into the second target nucleic acid to generate a third target nucleic acid; A method is provided that includes:

[0049] Parts (i) and (ii) may be repeated multiple times, allowing for the introduction of a first donor nucleic acid, a second donor nucleic acid, a third donor nucleic acid, and potentially further donor nucleic acids. When this technique is repeated, the product of one round of the method of the invention may serve as a target nucleic acid sequence for the next round of nucleic acid introduction.

[0050] The first RNA molecule may be of the same sequence during each iteration of the method of the invention, and / or the second RNA molecule may be of the same sequence during each iteration of the method of the invention. Alternatively, the first RNA molecule pair and the second RNA molecule pair may be used in an alternating manner during the iterations of the invention, such that the first pair is used for every odd iteration and the second pair is used for every even iteration. The first RNA molecule pair may be of the same sequence as the second RNA molecule pair. The first RNA molecule pair may include one RNA molecule that is identical to the RNA molecule in the second pair and one RNA molecule that is different in sequence from the RNA molecule in the second pair. Each of the first RNA molecule pairs may be different in sequence from each of the RNA molecules in the second pair.

[0051] In other embodiments, additional RNA molecule pairs, such as a third pair, can be used as part of the repeating pattern.Thus, the method can further comprise repeating parts (i), (ii) and (iii), where part (iii) comprises carrying out any method step of the present invention to introduce a third donor nucleic acid sequence into a third target nucleic acid, thereby generating a fourth target nucleic acid, and part (iii) comprises using a third RNA molecule pair.This pattern can be extended as necessary.

[0052] The backbone sequence of the episomal replicon may be different between part (i) and between part (ii). For example, the episomal replicon may include one or more markers to allow identification or selection of successful transfer of the desired sequence. To allow rounds of nucleic acid transfer including identification or selection, the episomal replicon of part (i) may include a first marker or set of markers and the episomal replicon of part (ii) may include a second marker or set of markers. In embodiments where parts (i) and (ii) are repeated, this may mean that the first marker or set of markers is used for all odd selections and the second marker or set of markers is used for all even selections. In other embodiments, an additional marker, for example a third marker or set of markers, may be used as part of the repeating pattern. To allow selection, the marker or markers for each round of nucleic acid transfer must be different from the marker or markers used in the previous round.

[0053] The backbone sequence of the episomal replicon during part (i) can be a first backbone sequence comprising a first marker or marker set and can encode a first pair of RNA molecules as described herein. The backbone sequence of the episomal replicon during part (ii) can be a second backbone sequence comprising a second marker or marker set and can encode a second pair of RNA molecules as described herein. This pattern can be maintained during the repeats, such that the first backbone sequence is present during all odd repeats and the second backbone sequence is present during all even repeats. The repeat pattern can also include additional backbone sequences, such as a third backbone sequence, comprising additional markers or marker sets and encoding additional pairs of RNA molecules. The first marker or marker set and the second marker or marker set are different from each other, such that each successful nucleic acid transfer can be selected during a round of recombination. The RNA molecules encoded by each backbone sequence allow for cleavage of the encoding backbone.

[0054] In a further aspect, the principles established for CONEXER can be extended to achieve scarless assembly and cloning through repeated insertion of megabases of DNA into an episome in E. coli. The present invention thus relates to an assembly episome replicon for repeated insertion and assembly of DNA therein (FIG. 6a). In an embodiment, this replicon has an approximately 50 bp sequence (HR1) that is homologous to one end of the next insertion sequence. This is immediately followed by positive and negative selection cassettes, and a universal homology region (uHR) that is complementary to the other end of the sequence to be inserted. Longer homologies, as long as sequences of several tens of kb, can also be utilized.

[0055] The present invention also provides a donor episomal replicon with a CONEXER backbone, with a universal spacer and oriT (Figure 6a). In the donor replicon, HR1 is within one end of the next DNA sequence to be inserted into the recipient BAC. This DNA sequence is followed by different positive and negative selection cassettes and universal homology regions. Each assembly step (Figure 6a, b) proceeds by joining the donor replicon to a recipient cell carrying the assembly replicon, nuclease-mediated excision of the sequence from HR1 to uHR from the donor replicon, and recombination-mediated insertion of this sequence into the assembly replicon. Selection for loss of the negative selection marker on the assembly replicon and acquisition of the positive marker from the sequence excised from the donor replicon selects for cells carrying assembly replicons with correct insertion. Cells carrying the new assembly replicon provide the input for the next insertion step. This approach has been named stepwise insertion synthesis of BAC (BASIS).

[0056] Thus, in a further aspect of the invention there is provided a method for constructing an episomal replicon comprising multiple assembly steps, wherein step n is a) providing a donor episomal replicon, said replicon comprising: A backbone comprising a universal spacer sequence, a first homology region HRn specific for integration step n, and a second universal homology region uHR, a first excision site located adjacent to HRn, and a second excision site located adjacent to uHR, a donor nucleic acid DNAn containing a homology region HRn+1 specific for assembly step n+1; Double selection cassette containing a positive selection marker and a negative selection marker The steps include b) providing a host cell comprising an assembly episomal replicon comprising a double selection cassette comprising a positive selection marker and a negative selection marker flanked by HRn and uHR, wherein the double selection cassette in the assembly replicon comprises a different marker than the selection cassette in the donor replicon; c) providing helper proteins capable of assisting recombination of nucleic acids in said host cell; c) providing an RNA-guided DNA endonuclease; d) providing a first RNA molecule comprising a sequence specific for a first excision site and a second RNA molecule comprising a sequence specific for a second excision site, the first and second RNA molecules being involved in directing an RNA-guided DNA endonuclease during excision; e) inducing excision of the donor nucleic acid sequence DNAn by an RNA-guided DNA endonuclease in a host cell; and f) incubating to allow recombination between the excised donor nucleic acid and the assembly replicon to form a second assembly replicon comprising the nucleic acid DNAn; The method includes:

[0057] The assembly replicon having the donor nucleic acid can then be used as the assembly replicon in a second step, in which a second donor replicon comprising homology regions HRn+1 and uHR and a second donor nucleic acid DNAn+1 is utilized to introduce a second donor nucleic acid into the assembly replicon generated in the first step.

[0058] Alternating sets of positive and negative selection markers allow an infinite number of steps to be repeated to assemble an episomal replicon of any desired size. Preferably, the number of steps performed is at least two, and may be up to 100, up to 50, up to 25, and most preferably about 10, 9, 8, 7, 6, or 5.

[0059] The size of the assembled replicon is preferably between 1 and 100 Mb, more preferably between 2 and 50 Mb, 3 and 25 Mb, 4 and 15 Mb, or 5 and 10 Mb.

[0060] Thus, in one embodiment, the invention comprises carrying out the method of the above aspect of the invention further comprising the steps of introducing into a host cell a further donor episomal replicon comprising a donor nucleic acid DNAn+1, and inducing in the host cell excision of said donor nucleic acid sequence DNAn by an RNA-guided DNA endonuclease, and incubating to allow recombination between the excised donor nucleic acid DNAn+1 and said assembly replicon to form a second assembly replicon comprising nucleic acid DNAn and nucleic acid DNAn+1.

[0061] The steps of this embodiment of the invention may be repeated to insert donor nucleic acid DNA n+2, n+3, n+4, etc. into the assembled episomal replicon.

[0062] BASIS can be used to generate episomes or other DNA vectors or segments useful for continuous genome synthesis (CGS). BASIS can also be used to direct continuous genome synthesis (CGS). 、It can be applied continuously without a sequencing step for the continuous production of artificial DNA, whether episomal, genomic, bacterial, or other genes and DNA. We first demonstrated the assembly of the 208Kbp human cystic fibrosis transmembrane conductance regulator gene by BASIS. The CFTR gene was assembled with a donor BACS carrying a gene fragment of approximately 70Kbp in three steps of BASIS.

[0063] BASIS can be used to assemble large sections of human genomic DNA, including exon, intron, and intergenic regions, into a single episome.

[0064] It is desirable to remove the sequencing step to verify each clone to allow rapid iteration of CONEXER by directly using an unsequenced pool of clones derived from one CONEXER run as input for the next CONEXER run. Therefore, we investigated factors that would substantially increase the fraction of clones in which genomic DNA has been completely replaced by synthetic DNA in one step of CONEXER.

[0065] Both recA and recO have been identified as factors that increase the fraction of clones with completely synthetic sequences (Figure 7a, Figure 9). Deletion of recA (ΔrecA) increased the fraction of completely synthetic DNA from 20% to approximately 80% in 100k24 (Figure 7a). Similar dramatic increases were seen in several other 100 Kbp regions, supporting the generality of these findings (Figure 7b).

[0066] Sequential genome synthesis can be performed by directly using the output from one round of CONEXER as input for the next round of CONEXER, without identifying individual fully encoded clones by sequencing.

[0067] Thus, the host cell used in the present invention is advantageously devoid of competent recA and / or recO. Preferably, the host cell is devoid of recA (ΔrecA).

[0068] Episomal replicons Embodiments of the present invention include episomal replicons that include a donor nucleic acid sequence. The donor nucleic acid can be DNA.

[0069] The term "episome" has its ordinary meaning in the art, e.g., any accessory extra-chromosomal replicating genetic element that can exist autonomously or that can be integrated with a chromosome.

[0070] An episomal replicon is an episomal nucleic acid that has its own origin of replication that is functional within the host cell.

[0071] An episomal replicon may be a plasmid. Plasmid refers to a small circular nucleic acid (generally DNA, most commonly double-stranded DNA) molecule. Plasmids in a cell are physically separate from any chromosomal nucleic acid, such as DNA, and can replicate independently. With respect to plasmids, "small" means that they are typically not more than 10 kb in size. Suitably, plasmids useful in the present invention have the following genetic elements: an origin of replication cognate to the host cell, and at least one selection marker.

[0072] The episomal replicon may be a BAC. A BAC may contain the following genetic elements: an origin of replication cognate to the host cell, and at least one selectable marker.

[0073] BACs and plasmids differ from each other in their origins of replication. BACs have a special origin of replication, which typically makes the BAC one copy in each cell and also helps the BAC to maintain a larger size (up to several hundred kb). Plasmids have a plasmid origin of replication, which typically makes the plasmid multiple copies in each cell (ranging from a few copies to several hundred copies per cell), typically up to a size of approximately 10 kb.

[0074] The episomal replicon may be a yeast artificial chromosome (YAC). A YAC may contain the following genetic elements: an origin of replication cognate to the host cell, and at least one selection marker.

[0075] Multiple origins of replication active for the same single nucleic acid in the same cell are generally undesirable. This is particularly true, for example, when the donor nucleic acid is a multicopy episomal nucleic acid, such as a plasmid, and in this scenario, it is clearly undesirable to integrate the plasmid origin of replication into (for example) the BAC or the host genome. Thus, suitably, the excised linear donor nucleic acid does not include an origin of replication. Suitably, the target nucleic acid sequence includes an origin of replication.

[0076] Suitably, the origin of replication on the episomal replicon containing the donor sequence must be compatible with the host, for example all prokaryotic hosts. Suitably, the origin of replication on the episomal replicon containing the target and on the episomal replicon containing the donor sequence must be compatible with the host, for example all prokaryotic hosts.

[0077] Suitably, the episomal replicon with the donor sequence has a prokaryotic origin of replication. Suitably, the replicon with the target sequence has a prokaryotic origin of replication. Suitably, the replicon with the target sequence is an episomal replicon and has a prokaryotic origin of replication. Suitably, the host cell is a prokaryotic host cell. Suitably, the synthetic genome is a synthetic prokaryotic genome.

[0078] target nucleic acid The target nucleic acid can be any suitable for the introduction of a donor nucleic acid sequence. In particular, the target nucleic acid can be a DNA molecule suitable for the introduction of a donor DNA molecule.

[0079] The target nucleic acid may comprise homologous recombination sequence 1 and homologous recombination sequence 2. The target nucleic acid may comprise one or more selection markers, which may be flanked by homologous recombination sequences. For example, the target nucleic acid may comprise a negative selection marker.

[0080] The target nucleic acid may be a nucleic acid region that also has its own replication origin that can function in a host cell. The target nucleic acid may be a plasmid. The target nucleic acid may be a BAC. The target nucleic acid may be a YAC. In certain embodiments, the target nucleic acid is the genome of a host cell.

[0081] Where the invention is applied to a genome, suitably the genome is a non-human genome, suitably a non-mammalian genome. Suitably the genome is a prokaryotic genome, suitably a bacterial genome. In particular, the genome may be the E. coli genome.

[0082] In certain embodiments, the episomal replicon comprising the donor DNA molecule is a BAC and the target nucleic acid is the genome of an E. coli cell.

[0083] Homologous recombination In theory, any nucleotide sequence can be selected as a site for homologous recombination.

[0084] The nucleotide sequence for homologous recombination can be unique. For example, the nucleotide sequence for homologous recombination can be unique within the target sequence for recombining the donor sequence therein. In other embodiments, the homologous recombination sequence 1 and / or the homologous recombination sequence 2 can be unique within the target sequence for recombining the donor sequence therein.

[0085] Alternatively, the homologous recombination sequence 1 and / or the homologous recombination sequence 2 may not be unique within the target sequence for recombining the donor sequence therein. In such an embodiment, selection may be used to identify successful introduction of the desired sequence into the desired site. For example, off-target integration does not result in removal or destruction of the negative selection marker, or off-target integration does not repair the double-strand break induced at the introduction site.

[0086] Suitably, the sequences for homologous recombination are unique.

[0087] Suitably, the sequence for homologous recombination is at least 30 nucleotides long. A homologous recombination sequence as short as 30 nucleotides may result in low efficiency. Therefore, for high efficiency, suitably, the homologous recombination sequence is at least 40 nucleotides long, suitably at least 50 nucleotides, suitably 50-100 nucleotides, most suitably 50-65 nucleotides long.

[0088] The sequence for homologous recombination is selected according to the target sequence and introduced into the donor sequence, so that the homologous recombination sequence 1 (HR1) and the homologous recombination sequence 2 (HR2) on the donor sequence show 100% sequence identity to the HR1 and HR2 on the target sequence.

[0089] The use of lambda Red recombination allows short nucleotide sequences to be used for homologous recombination, as outlined above. Other recombination assisting systems may be used. For example, the RecBCD system may be used. When the RecBCD system is used, suitably the step of "providing helper proteins capable of assisting recombination of nucleic acids in said host cell" comprises inducing or enabling expression of the RecBCD system in the host cell.

[0090] When using the RecBCD system or other recombination-assisted systems, one of skill in the art will pay attention to the requirements of these systems for sequences selected for homologous recombination. For example, the RecBCD system may require longer homologous recombination sequences, such as 3-10 kb in length.

[0091] More specifically, RecBCD is a natural E. coli recombination system consisting of three components RecB, RecC, and RecD. The three subunits form an ATP-dependent helicase / nuclease complex that is essential for both homologous recombination and double-strand break repair during the transduction and conjugation processes in E. coli. Studies in vivo inducing double-strand breaks in E. coli DNA have shown that double-strand break repair (DSBR) can proceed via one of two recombination pathways. Both pathways require RecBCD and RecA, but one depends on the resolvase enzyme RuvABC, while the other does not depend on the resolvase enzyme RuvABC, relying instead on RecG. The recB and recD genes form an operon, while recC is located in close proximity but has its own promoter. The three gene products form a heterotrimer, also known as exonuclease V. In case any further guidance is needed, details can be found in the publicly available EcoCyc database, e.g., for the E. coli strain K12-MG-1655, called “RecBCD” (Keseler et al. (2013), “EcoCyc: fusing model organism databases with systems biology”, Nucleic Acids Research 41: D605-12).

[0092] Thus, to support recombination according to this embodiment, at least RecBCD should be expressed in the host cell.

[0093] RecA may also be required. More suitably, therefore, at least RecBCD and RecA should be expressed in the host cell to support recombination according to this embodiment. Most suitably, RecBCD and RecA should be expressed in the host cell to support recombination according to this embodiment.

[0094] Another alternative to the Lambda Red system is the RecET system. RecE and RecT are phage-derived E. coli genes. RecE mimics Lambda Red alpha and RecT mimics Lambda Red beta (Muyrers, JP, Zhang, Y., Buchholz, F. & Stewart, AF RecE / RecT and Redalpha / Redbeta initiate double-stranded break repair by specifically interacting with their respective partners. Genes Dev. 14, 1971-1982 (2000)). The RecET combination works relatively better than the Lambda Red alpha / beta combination. Lambda Red alpha and beta are the actual components that carry out the recombination, while Lambda Red gamma is an inhibitor of the RecBCD system.

[0095] Suitably, recombination assistance is provided via the Lambda Red system, for example from the pRed / ET plasmid commercially available from Gene Bridges ("Quick & Easy E. coli Gene Deletion Kit", Gene Bridges GmbH, Im Neuenheimer Feld 584, 69120 Heidelberg, Germany).

[0096] This system in this configuration was first described in Datsenko et al 2000 (Datsenko KA & Wanner, BL One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc. Natl. Acad. Sci. USA 97, 6640-6645 (2000)), which is hereby incorporated by reference, particularly for details of the Lambda Red system.

[0097] The inventors teach that the pRed / ET plasmid is based on the pKD46 plasmid in Datsenko et al 2000 (as judged by sequence identity) and therefore that the pKD46 plasmid can be used as a template to perform PCR for the construction of the Lambda Red system.

[0098] When the helper proteins capable of assisting recombination of nucleic acids include lambda red protein, suitably the following proteins are expressed in the host cell: [Table 1] TIFF2024542108000003.tif253170TIFF2024542108000004.tif23170

[0099] Homologous recombination sequence To select homologous recombination sequences, the following steps can be used: - Select 50-100 nucleotides at the desired position within the sequence of the nucleic acid to be modified (target nucleic acid), such as the bacterial genome or a plasmid backbone. - Conducting a BLAST search of the selected sequence against the target nucleic acid. - consider the E-value of the selected sequence compared to the closest match in the BLAST search. Typically, it is 10 -20Any E value greater than 0.01 relative to an undesirable target site elsewhere in the target nucleic acid will be too high. If this is found, then an alternative homologous recombination sequence is suitably selected.

[0100] Suitably, standard BLAST tools are used to calculate the E-values ​​of homologous recombination (HR) sequences. One such online tool is at http: / / biocyc.org / ECOLI / blast.html. Suitably, the focus is on how unique a given HR sequence is as judged by the E-value. Suitably, affinity does not need to be considered / calculated. In principle, all sequences that can work in conventional recombination will work better in the present invention.

[0101] More specifically, if HR sequences can function in conventional recombination, they will function better in the present invention. Suitably, HR sequences for the present invention are selected according to the precise principles and requirements for conventional recombination using the lambda Red system. For example, we typically design HRs of 50-70 bp in length and 10 -20 A blast is performed against the E. coli genome for an expected value of less than 10. (E value is a measure of how unique a given sequence is; the smaller the E value, the more unique the sequence. Any suitable tool for the calculation can be used, for example, standard BLAST tools for calculating E values ​​for HR sequences. One such online tool is at http: / / biocyc.org / ECOLI / blast.html). -20 It is anticipated that E values ​​smaller than this will be unnecessary, although it will be appreciated that sequences with smaller values ​​will still be useful in the present invention.

[0102] The E value is a measure of how unique a given sequence is. Because conventional recombination relies solely on the specificity of homologous regions, this is -20In principle, the method of the present invention can tolerate less stringent E values ​​(e.g., less stringent homology regions), since the method can increase the specificity of a locus not only by the specificity of the homology region, but also by the simultaneous loss of a negative selection marker and gain of a positive selection marker. However, in practice, it is very easy to create homology regions with stringent E values, and therefore, suitably, a cutoff of 10 -20 An E-value cutoff of 0.01 is used.

[0103] Selectable Marker Any of the methods of the invention can include the further step of selecting recombinants in which the donor nucleic acid has been integrated into the target nucleic acid, which step is performed after the induction step to cause recombination.

[0104] The desired sequence may include a positive selectable marker, including any marker that allows for the identification or selection of cells that contain the marker.

[0105] The target nucleic acid may comprise, in order: 5'-homologous recombination sequence 1-negative selectable marker-homologous recombination sequence 2-3'.

[0106] Selection of recombinants in which the donor nucleic acid has been integrated into the target nucleic acid may comprise selection for the gain of a positive selectable marker of the donor nucleic acid and the loss of a negative selectable marker of the target nucleic acid. Suitably, selection for the gain of a positive selectable marker of the donor nucleic acid and the loss of a negative selectable marker of the target nucleic acid is performed simultaneously. In other embodiments, the step of selecting recombinants comprises sequential selection of a positive marker and a negative marker, or sequential selection of a negative marker and a positive marker.

[0107] The desired sequence may contain both positive and negative selectable markers.

[0108] The method of the present invention comprises the steps of: The method may further comprise the step of inducing at least one double-stranded break in a target nucleic acid sequence, said double-stranded break being between said homologous recombination sequence 1 and said homologous recombination sequence 2.

[0109] Suitably, at least two double stranded breaks are induced in the target nucleic acid sequence, each said double stranded break being between said homologous recombination sequence 1 and said homologous recombination sequence 2.

[0110] The episomal replicon may comprise a negative selectable marker independent of the donor nucleic acid sequence. Suitably, the method comprises the further step of selecting for loss of the episomal replicon by selecting for loss of said negative selectable marker independent of the donor nucleic acid sequence.

[0111] Some methods of the present invention include a combinatorial selection approach involving the loss of a positive marker and a negative marker. The use of this "double selection" scheme actually also aids in site specificity. For example, if a recombination event occurs at an inappropriate site, a positive selectable marker may be gained. However, by using simultaneous selection of a positive marker and a loss of a negative marker, even if a nucleic acid is integrated into a target nucleic acid at an inappropriate site (thereby conferring a positive marker), such molecules will still not be selected, because if they are recombined at an inappropriate site, they will not simultaneously result in the loss of a negative marker. Thus, while being a useful selection in itself, it actually adds the technical benefit of aiding in site specificity by selecting not only for the gain of donor sequences, but also for the simultaneous loss of sequences to be removed / replaced.

[0112] Examples of suitable selectable markers are shown in the table below (Table 2). Additionally, any suitable antibiotic marker may be used. Examples of such antibiotic markers include TetR, AmpR, HyR, and ErmR, which allow for selection schemes including tetracycline resistance, ampicillin resistance, hygromycin resistance, or erythromycin resistance, respectively. [Table 2] TIFF2024542108000006.tif234170

[0113] Thus, in one embodiment, the negative selectable marker is selected from the group consisting of sacB (sucrose sensitivity) and rpsL (S12 ribosomal protein-streptomycin sensitivity), in some embodiments, the positive selectable marker is selected from the group consisting of CmR (chloramphenicol resistance) and KanR (kanamycin resistance).

[0114] Excision / introduction of double-strand breaks The methods of the invention include a mechanism for excising / introducing a double stranded break. Suitably, excision is performed to generate a linear donor nucleic acid.

[0115] In certain embodiments, the system is a CRISPR / Cas9 system, although other systems that provide this function are known.

[0116] For example, there are three published papers on alternative RNA-guided endonucleases to the original Streptococcus pyogenes CRISPR / Cas9 (Ran, FA et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature 520, 186-191 (2015); Zetsche, B. et al. Cpfl is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell 163, 759-771 (2015); Lee, CM, Cradick, TJ & Bao, G. The Neisseria meningitidis CRISPR-Cas9 System Enables Specific Genome Editing in Mammalian Cells. Mal. Ther. 24, 645-654 (2016)). All of these can be used to guide in vivo excision in the present invention. These references are specifically and expressly incorporated by reference herein for their teachings on alternative systems for the introduction of double-strand breaks / excisions used herein.

[0117] CRISPR / Cas9 sequences The CRISPR / Cas9 system is described in Jiang et al 2013 (Jiang, W., Cox, D., Zhang, F., Bikard, D. & Marraffini, LA RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013)).

[0118] In general, guide RNA refers to a single fusion RNA between tracrRNA and spacer RNA.Suitably, as discussed herein, a combination of a certain tracrRNA and a different spacer RNA is used in the present invention.These tracrRNA-spacer RNA combinations may be replaced with multiple different guide RNAs.

[0119] In the art, guide RNA refers only to the fusion of tracrRNA and spacer RNA as one RNA, and does not refer to the double RNA complex of tracrRNA and spacer RNA.

[0120] PAM stands for protospacer adjacent motif. It is typically a three nucleotide motif. A typical guide RNA is 30 nucleotides long. A guide RNA typically includes a 27 nucleotide target sequence and a three nucleotide PAM sequence.

[0121] Suitably, the present invention may use a separate tracrRNA / spacer RNA CRISPR setup, identical to Jiang et al. 2013. Alternatively, a single guide RNA CRISPR setup may be used, e.g., as known in the art (see Le Cong et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013)).

[0122] To assist excision when using CRISPR / Cas9, suitably helper proteins capable of assisting nucleic acid excision include, at a minimum, Cas9 (e.g. see below) and RNAseIII (e.g. rnc, EcoCyc accession ID EG10857) along with associated RNA (spacer RNA guide, tracrRNA (see below)).

[0123] A typical sequence is shown below. Cas9 (SEQ ID NO: 10) tracrRNA (SEQ ID NO: 11) AAAAAAGTTTAAATTAAATCCATAATGATTTGATGATTTCAATAATAGTTTTAATGACCTCCGAAATTAGTTTAATATGCTTTAATTTTTCTTTTTCAAAATATCTCTTCAAAAAATATTACCCAATACTTAATAATAAATAGATTATAACACAAAATTCTTTTAAAAAGTAGTTTATTTTGTTATCATTCTATAGTATTAAGTATTGTTTTATGGCTGATAAATTTCTTT GAATTTCTCCTTGATTATTTGTTATAAAAGTTATAAAATAATCTTGTTGGAACCATTCAAAACAGCATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTTGATACTTCTATTCTACTCTGACTGCAAACCAAAAAAACAAGCGCTTTCAAAAACGCTTGTTTTATCATTTTTAGGGAAATTAATCTCTTAATCCTTTT

[0124] The actual sequences of Cas9 and tracrRNA typically remain constant throughout all experiments. Spacer RNA sequences vary depending on the exact CRISPR / Cas9 cleavage site. As discussed herein, the methods of the present invention allow for the use of universal spacers that remain constant in the repeats of the present invention. For example, a spacer can be constant in all odd repeats of the present invention, and another set of spacers can be constant in all even repeats of the present invention.

[0125] Suitably, Cas9, tracrRNA and spacer RNA are provided together to the cell in which excision occurs (i.e. the host cell). Suitably, all three of these elements are essential for efficient excision.

[0126] In one embodiment, tracrRNA is constitutively expressed in the host cell. Cas9 is induced with helper proteins (e.g., lambda red alpha / beta / gamma) that can assist in recombination of nucleic acids. Spacer RNA can be provided to the host cell by transforming the cell with a small plasmid that expresses the spacer RNA. Alternatively, in a preferred embodiment, the spacer RNA is encoded by an episomal replicon that contains the donor nucleic acid sequence. Excision occurs when all three components are present in the cell.

[0127] In another embodiment, the nucleic acid (e.g. DNA) sequences for expressing Cas9, tracrRNA and spacer RNA may be provided together to a cell, but the actual expression of (some of) the three components may be suppressed (not induced / silent). At the appropriate time, expression may be induced, thus inducing excision.

[0128] In one embodiment, to cause excision, tracrRNA is constantly (constitutively) expressed, expression of Cas9 is induced, and a spacer RNA is provided last.

[0129] Induction of expression is well within the scope of the ability of a person skilled in the art. For example, the desired sequence (e.g., Cas9) is placed under the control of an inducible promoter. The activity of the promoter is induced as needed. For example, the well-known arabinose (pAra) promoter, which is induced in the presence of arabinose, can be used. Similarly, a person skilled in the art may select a constitutive promoter as needed from a vast array of well-known promoters suitable for constitutive expression.

[0130] As is well known for the operation of the CRISPR system, the sequence of the spacer RNA will vary for each target site. Selection of an appropriate spacer RNA is well within the realm of one of skill in the art. [Table 3] TIFF2024542108000008.tif215170 [Table 4]

[0131] The combination of tracrRNA and spacer RNA can be provided individually or as one guide RNA that is a fusion of tracrRNA and spacer RNA.

[0132] It should be noted that different elements of the CRISPR system require different motifs. In Cas9, the PAM is NGG. More specifically, alternative implementations of the CRISPR system may be used depending on the operator's choice, for example implementations that result in alternative PAMs. More specifically, it has been demonstrated that the CRISPR / Cas9 system of Streptococcus pyogenes, which naturally recognizes NGG as the PAM, can be modified to recognize altered PAMs (Kleinstiver, BP et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature (2015). doi:10.1038 / nature14592). The three alternative RNA-guided endonuclease systems described in Ran 2015, Zetsche 2015, and Lee 2016 (see above) naturally have different PAMs.

[0133] Those skilled in the art will know if alternative components of CRISPR system are used in the present invention, and then should use corresponding alternative cognate PAM sequence.This is well within the scope of the art.In case any further guidance is needed, the following table shows alternative elements of CRISPR system together with their PAM sequence. [Table 5] TIFF2024542108000011.tif250170TIFF2024542108000012.tif44170

[0134] Manipulation of the PAM site for control of resection In practicing the invention, the PAM on the target sequence can be compared to the PAM on the donor nucleic acid (e.g., sDNA) to be introduced into the target, and if the PAM sequence matches the donor nucleic acid (e.g., sDNA) and the target nucleic acid (e.g., genomic DNA), it can be mutated, if necessary, to avoid double excision problems (e.g., excision that accidentally includes a homologous recombination sequence). This can be easily accomplished by one of skill in the art when arranging the elements in the order taught herein.

[0135] The homology regions flanking the donor nucleic acid (e.g. synthetic DNA) on the episomal replicon may be further flanked by AvrII sites (CCTAGG). AGG or CCT (depending on the orientation) correspond to the NGG PAM sequence required by the CRISPR / Cas9 system of Streptococcus pyogenes, while the complementary CCT or AGG constitute the last three nucleotides of the protospacer. Any substitution of either the last three nucleotides of the protospacer and / or the G of the NGG PAM will disable CRISPR / Cas9 recognition and / or cleavage.

[0136] Some embodiments include a step of introducing a double-stranded break in the target nucleic acid into which the donor nucleic acid is to be recombined. In such embodiments, care must be taken in the selection of the PAM to be used on the sequence present in the nucleic acid to be modified (target nucleic acid), since the sequence on the donor nucleic acid (e.g. the DNA to be introduced into the target nucleic acid) will not match the PAM on the target nucleic acid. If they match, the excision step of the method of the invention will also introduce a double-stranded break at an inappropriate position in the target nucleic acid. Therefore, suitably, the PAM on the target nucleic acid (e.g. the genome or plasmid or BAC into which the donor DNA is to be introduced) should be compared with the PAM on the episomal replicon carrying the donor nucleic acid. If the PAM sequences are found to match, these sequences are mutated on the target nucleic acid to be modified to avoid possible problems. This is well within the realm of the skilled artisan. More particularly, in these embodiments, the two cut sites on the episomal replicon are distinguished from the corresponding ends of the homologous region on the target nucleic acid in the same manner as disclosed above. The two additional cleavage sites inside the homologous region on the target nucleic acid must be identified by looking for NGG motifs that define the boundaries of the homologous region on the target nucleic acid. The NGG PAMs of the two additional cleavage sites inside the homologous region on the target nucleic acid must also not be present at the corresponding ends of the homologous region of the episomal replicon with the donor nucleic acid to avoid "double excision". This can be achieved very easily because the sequence for insertion is naturally different from the cleavage site on the target nucleic acid (e.g. genome). This should be carefully sequenced if the donor nucleic acid (e.g. synthetic DNA) has a similar sequence to the target nucleic acid (e.g. wild type genomic DNA). This is achieved by replacing the corresponding NGG in the donor nucleic acid (e.g. synthetic DNA) and / or by changing the last three nucleotides in the other protospacer immediately following the NGG. In this way, we mark the cleavage site only to the target nucleic acid (e.g. genome) location.

[0137] In one embodiment, it may be desirable to induce cleavage of the target nucleic acid to aid in the selection of recombinants.In this embodiment, suitably there are three cleavage sites, two on the episomal replicon to excise the donor nucleic acid, and one on the target nucleic acid to aid in selection.Therefore, suitably, the target nucleic acid comprises in the order 5'-homologous recombination sequence 1-cleavage site-homologous recombination sequence 2-3'.

[0138] Suitably, the target nucleic acid comprises a) 5'-homologous recombination sequence 1-cleavage site-homologous recombination sequence 2-3'; b) 5'-homologous recombination sequence 1-positive selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2; c) 5'-homologous recombination sequence 1-negative selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2; d) 5'-homologous recombination sequence 1-positive selectable marker-negative selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2; e) 5'-homologous recombination sequence 1-negative selectable marker-positive selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2 Includes, in order.

[0139] When the invention is applied in multiple rounds, the donor nucleic acid of the first round may participate in / become part of the target nucleic acid of the next round. Thus, suitably, the desired sequence is a) 5'-homologous recombination sequence 1-cleavage site-homologous recombination sequence 2-3' b) 5'-homologous recombination sequence 1-positive selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2; c) 5'-homologous recombination sequence 1-negative selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2; d) 5'-homologous recombination sequence 1-positive selectable marker-negative selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2; e) 5'-homologous recombination sequence 1-negative selectable marker-positive selectable marker-homologous recombination sequence 2-3', further comprising a cleavage site between the homologous recombination sequence 1 and the homologous recombination sequence 2 In turn,

[0140] Suitably, the cleavage site on the target nucleic acid or desired sequence is different to the excision site on the episomal replicon / donor nucleic acid.

[0141] The cleavage site may be between the positive / negative selectable markers or within the positive / negative selectable markers. Suitably, the target nucleic acid comprises two such cleavage sites. Suitably, the cleavage site is adjacent to one of the homologous recombination sequences. Suitably, the two cleavage sites comprise a first cleavage site adjacent to the homologous recombination sequence 1 and a second cleavage site adjacent to the homologous recombination sequence 2.

[0142] Applicable The present invention involves the introduction of a desired sequence into a target nucleic acid. "Introduction of a desired sequence" as used herein means incorporating a desired sequence into a target nucleic acid so that the resulting nucleic acid sequence contains the desired sequence. This can also be described as the incorporation of a desired sequence into a target nucleic acid.

[0143] The introduction of a desired sequence into a target nucleic acid may include replacing a portion of the target nucleic acid sequence with the desired sequence. Thus, after replacement, the resulting nucleic acid sequence contains only the desired sequence and a portion of the original sequence of the target nucleic acid. For example, in an embodiment where a genome is the target nucleic acid sequence, a portion of the genome sequence may be replaced by the desired sequence introduced during the repetition of the method of the present invention. The desired sequence may replace a region in the target nucleic acid that is smaller than the desired sequence, the same size as the desired sequence, or larger than the desired sequence.

[0144] Introduction of a desired sequence can include insertion of the desired sequence into the target nucleic acid. As used herein, "insertion" means that the desired sequence is introduced into a site within the target nucleic acid such that the resulting nucleic acid sequence has all of the original sequence of the target nucleic acid and also contains the inserted desired sequence.

[0145] The methods of the invention can include multiple steps, whereby a donor nucleic acid, for example encoding a selection marker, is introduced into the target nucleic acid, which is then replaced by a desired sequence.

[0146] For example, the methods of the invention can be used to insert a desired sequence into a genome or other target nucleic acid such that all of the original sequence of the genome remains in addition to the newly inserted desired sequence. In another embodiment, the methods of the invention can be used to replace a portion of a genome or other target nucleic acid with a desired sequence such that the overall size of the genome remains unchanged. In yet another embodiment, the methods of the invention can be used to replace a portion of a genome or other target nucleic acid with a desired sequence that is longer than the region being replaced such that the resulting nucleic acid has only a portion of the original genomic sequence and the overall size of the genome is increased.

[0147] The present invention is useful in constructing plasmids.The present invention is useful in modifying a host genome.The present invention is useful in constructing artificial chromosomes such as BACs.

[0148] The present invention has particular application in the generation of large sized nucleic acid constructs. The present invention has particular application in the generation of libraries with high diversity. In this regard, current transformation techniques allow for approximately 10 8 Transformation efficiencies of 10 10 Transformation efficiency of 10 or more is very difficult and / or uncertain. According to the present invention, a first half library can be generated and transformed into a first host cell (population of host cells). This first half library is then transformed with a nucleic acid encoding a second half library. These two half libraries are then combined in vivo using recombination according to the present invention, and transformed into a 10 or more host cell (population of host cells). 10 A library with a diversity of 10 5 Advantageously, transformation efficiencies of 100% were obtained using the plasmid p53.

[0149] host cell Suitably the host cell is a prokaryotic cell. Suitably the host cell is a bacterial cell.

[0150] In one embodiment, the host cell is in vitro, i.e. in a laboratory. In one embodiment of the invention, the method is an in vitro method. In some embodiments, the method is not performed in vivo. Suitably, the host cell is not part of a living human or animal body. Suitably, the host cell is selected from one of the host cells used in the examples below.

[0151] The host cell can be any Gram-negative bacterium. The host cell can be E. coli. The host cell can be any E. coli strain (such as MG1655 or BL21), or a cell derived therefrom.

[0152] MG1655 is considered to be the wild type strain of E. coli. The GenBank ID for the genome sequence of this strain is U00096 (as of the filing date, U00096.3). BL21 is widely available commercially.

[0153] A host organism such as E. coli may be selected or modified to inhibit native repair mechanisms to ensure that double strand repair is absent or highly unlikely. For example, the RecBCD system may be mutated or inhibited, provided that a suitable helper protein capable of assisting nucleic acid recombination in the host cell, such as the lambda Red protein described herein or other suitable recombination auxiliary protein, is present in place of RecBCD. For example, in one embodiment, RecBCD may be inhibited, since it may interfere with the lambda Red component, reducing the efficiency of recombination performed by the lambda Red component using double stranded DNA with short homology regions (e.g., approximately 50 bp) (which are degraded by the RecBCD system). However, if long homology regions (e.g., approximately 3-5 kb) are used, RecBCD may be an alternative recombination auxiliary protein to replace the lambda Red component as a recombination auxiliary protein.

[0154] Optional Additional Steps or Features In one embodiment, the invention may involve a first recombination step carried out by conventional techniques, which has the advantage of allowing the introduction of a reverse selectable marker into the target site.

[0155] The invention may include a final step of final recombination, which may be performed by the methods disclosed herein or by conventional recombination, for example, this may be advantageous in removing selectable markers that have served their purpose and are no longer necessary for further iterations of the methods disclosed herein.

[0156] In some embodiments, the cyclic methods disclosed herein are initiated and continued without a first conventional homologous recombination event.

[0157] In one embodiment, an excision mechanism, such as CRISPR / Cas9, is utilized to cut at the site to be replaced by the recombination event, which creates selection pressure against the cut (unrecombined) target nucleic acid, i.e., negative selection due to the double stranded break. In another embodiment, this negative selection due to the double stranded break within the target sequence is used to improve selection in an embodiment with three double stranded breaks (two double stranded breaks for the excision of the donor nucleic acid, and one double stranded break between the HR1 and HR2 sequences on the target nucleic acid, for a total of three DS breaks / cuts).

[0158] In certain embodiments, there is provided a method for introducing a desired sequence into a target nucleic acid, comprising the steps of: 1) providing an E. coli host cell comprising a genome including a target nucleic acid; 2) delivering the BAC to a host cell by conjugative transfer, the BAC comprises a backbone sequence and a donor nucleic acid sequence; The donor nucleic acid sequence comprises, in order: 5'-homologous recombination sequence 1-desired sequence-homologous recombination sequence 2-3'; the backbone sequence comprises a first excision site located adjacent to homologous recombination sequence 1 and a second excision site located adjacent to homologous recombination sequence 2, the first excision site comprises a first protospacer, and the second excision site comprises a second protospacer; the backbone sequence encodes a first RNA molecule comprising a spacer specific to the first protospacer and a second RNA molecule comprising a spacer specific to the second protospacer; 3) providing a lambda red protein capable of assisting in recombination of nucleic acids in said host cell; 4) providing a Cas9 nuclease; 5) inducing excision of the donor nucleic acid sequence by a Cas9 nuclease; 6) incubating to allow recombination to occur between the excised donor nucleic acid and the target nucleic acid; and 7) selecting recombinants in which the donor nucleic acid has been integrated into the target nucleic acid The method includes: Any of the alternatives of the features of the method discussed herein can be used in place of the corresponding features of the method. For example, the BAC can be an episomal replicon, the target nucleic acid can be any target nucleic acid, the lambda red protein can be used in place of any suitable for the purpose, etc.

[0159] As discussed herein, steps 1)-4) may be performed in any order or simultaneously.

[0160] In another embodiment, the method is for assembly of nucleic acid sequences and includes performing steps 1)-7) to introduce a first donor nucleic acid sequence into a first target nucleic acid, thereby creating a second target nucleic acid, and then performing steps 1)-7) to introduce a second donor nucleic acid sequence into a second target nucleic acid, thereby creating a third target nucleic acid. This process may be repeated, with the product of each repeat being the target for the next repeat. A BAC containing the first backbone sequence may be used for each odd repeat, and a BAC containing the second backbone sequence may be used for each even repeat. The first backbone sequence and the second backbone sequence may encode different spacers and may include different selection markers.

[0161] All of the features described in this specification (including any accompanying claims, abstract, and drawings) and / or all of the steps of any method or process disclosed may be combined in any combination with any of the above aspects, except combinations where at least some of such features and / or steps are mutually exclusive.

[0162] For a better understanding of the present invention and to illustrate how embodiments of the invention may be made, reference will now be made to examples which are not intended to be limiting of the invention in any way. [Example]

[0163] method Strains and plasmids used in this study We used a genome-reduced, streptomycin-resistant E. coli (Mds42, Scarab Genomics) as the starting strain for REXER and CONEXER. The same strain was used for the assembly of the human CFTR gene by BASIS. Large sections of the human genome were assembled with a ΔrecA mutant of the same strain. All cloning procedures were performed in E. coli DH10b. We performed yeast assemblies in strain BY4741.

[0164] In this study, we used the following BACs: 100k13, 100k22, 100k24, 100k25, 100k26, 100k27, 100k28, and 100k37a (Non-Patent Document 14). Each BAC has approximately 100 kb of synthetic DNA with a defined synonymous codon compression scheme in which two serine codons (TCG and TCA) and a stop codon (TAG) are replaced via defined recoding rules (TCG to AGC, TCA to AGT, and TAG to TAA).

[0165] The present inventors have used the following positive and negative selection markers in REXER and CONEXER: sacB (conferring sucrose sensitivity), cat (chloramphenicol resistance), rpsL (streptomycin sensitivity), and kan. R (kanamycin resistance), phe ST251a_a294G (pheS * ) (4-chlorophenylalanine (4-CP) sensitivity), and Hyg R (hygromycin resistance) was used (Non-Patent Document 14). In REXER, the present inventors used the upstream end of the integration site (locus 0 We used a streptomycin-resistant E. coli strain (Mds42) with a reduced genome that carries a genomic double selection cassette at 100k13 and 100k37a: rpsL-kan at REXER; RIn REXER, sacB-cat in 100k22, 100k24, and 100k28. In CONEXER, we located the upstream end of the integration site in the recipient cell (locus 0 We used streptomycin-resistant E. coli (WT or ΔrecA, as indicated) strains with reduced genomes carrying a genomic double selection cassette in the 100k25 and 100k27: rpsL-kan in CONEXER; R , sacB-cat in CONEXER at 100k24, 100k26, and 100k28.

[0166] We used the helper plasmid pKW20 (Wang, K. et al. Nature 539, 59-64, doi:10.1038 / nature20124 (2016)) to enable excision and recombination in REXER and CONEXER. pKW20 constitutively expresses tracrRNA, as well as cas9 and lambda Red components under the control of an arabinose-inducible promoter. In addition, we created a derivative plasmid without cas9 to enable lambda Red recombination without expression of Cas9, which was utilized for modification of the BAC for CONEXER (see below). This was done by PCR amplification of the remaining parts of pKW20 followed by NEBuilder HiFi DNA assembly.

[0167] The BACs for the assembly of a large region of human chromosome 21 (Figure 3C) were based on the BACs of the 32k-human BAC library (BACPACK genomics). These were applied to CONEXER using lambda Red recombination as described in Construction of BACs for CONEXER.

[0168] For the deletion of host genes, we used plasmids with spacer sequences and pKW20. Spacer plasmids were constructed by restriction ligation into the pMB1 plasmid backbone with guide-encoding ssDNA oligonucleotides. All spacer sequences are shown in Tables 3a and 3b.

[0169] Construction of spacer arrays In this study, we perform genomic integration of synthetic DNA from BACs of two different designs (labeled even and odd numbers, respectively), which require a set of universal spacers (universal 1 and universal 2). The spacer for the CONEXER BAC, derived from a human BAC library, is based on a third design (universal 3). Note that to further simplify the method, a series of BACs can be designed such that one universal spacer RNA excises both the 5' and 3' in all BACs.

[0170] All spacer RNAs for REXER are expressed from the plasmid pKW3_MB1amp_tracr_spacer (Table 15), which carries the ampicillin resistance marker, tracrRNA, and spacer array. We constructed each array from overlapping oligonucleotides via two rounds of PCR, and prepared the backbone by restriction digestion of pKW3 with AccI and EcoRI (Non-Patent Document 14). We combined the backbone and each array by NEBuilder HiFi DNA assembly, and then verified by Sanger sequencing. All spacer and oligonucleotide sequences are shown in Tables 5 to 15.

[0171] Building a BAC for CONEXER We modified the even BAC for CONEXER by incorporating the sequence of the origin of transfer (oriT) to enable conjugal transfer, and the universal spacer array (universal2) into the BAC backbone (Table 5). To this end, we inserted the oriT sequence and the spacer array sequence into the selection marker amp R We amplified each sequence by PCR. The plasmid pKW3_MB1amp_tracr_universal2 was R pKW3_MB1amp_tracr_universal2 served as a template for the spacer array, pRK24 (addgene no. 51950) served as a template for oriT, and pKW3_MB1amp_tracr_universal2 served as a template for the spacer array. We concatenated the PCR products of two successive PCRs to construct the pheS * The final amplicon with primers and BAC backbone that generate a 50 bp homology region to R A -oriT-universal2 cassette was generated. We used the helper plasmid pLF118 without cas9 to initiate lambda red recombination and selected for integration of the cassette into the BAC with ampicillin. Complete integration of the cassette was first verified by Sanger sequencing, and the successfully modified BAC 100k24 was further verified by next generation sequencing (NGS) to confirm the integrity of the entire synthetic DNA insert. All oligonucleotide sequences are shown in Tables 5-14.

[0172] Odd-numbered BACs can be modified in a similar manner in CONEXER. The corresponding universal spacer array, universal1, was amplified from the pKW3_Mb1amp_tracr_universal1 plasmid described above. The corresponding oligonucleotide sequences are listed in Table 8.

[0173] The odd and even CONEXER BACs provide a simple and rapid basis for integrating synthetic DNA at any point in the E. coli genome using the CONEXER protocol. For this purpose, BAC backbones may be directly amplified using the described BACs as templates for S. cerevisae-mediated assembly of the BACs with other synthetic DNA (14, 5). Robertson, WE et al. Nat Protoc 16, 2345-2380, doi:10.1038 / s41596-020-00464-3 (2021).

[0174] BACs from human libraries were applied to CONEXER by the integration of oriT sequences, universal spacer arrays, universal homology regions, and double selection cassettes. For this purpose, we cloned plasmids carrying all components in the correct orientation via Gibson assembly. These plasmids served as templates for PCR, in which we amplify the entire sequence to be integrated into the BAC as one linear piece of DNA. We used the helper plasmid pLF118 without cas9 to initiate lambda Red recombination and selected for integration of the cassette into the BAC with the appropriate antibiotic (hygromycin or kanamycin, depending on the type of double selection cassette used). The complete integration of the cassette was first verified by genotyping the ligation at both ends of the cassette. Successfully modified BACs were further verified by next-generation sequencing (NGS) to confirm the integrity of the entire synthetic DNA insert.

[0175] The BAC for assembly of the CFTR gene was assembled from fragments in yeast (Robertson, WE et al. Nature Protocols 16, doi:10.1038 / s41596-020-00464-3 (2021)). Fragments were generated via PCR amplification. CONEXER BAC 100k25 served as a template for amplification of the BAC backbone fragment (with origin of replication, universal homology region, oriT, and universal spacer array). Genomic DNA purified from hTERT RPE-1 cells served as a template for PCR amplification of the fragments of the CFTR gene that we used for assembly.

[0176] REXER The inventors performed REXER (Non-Patent Documents 14, 5). Robertson, WE et al. Nat Protoc 16, 2345-2380, doi:10.1038 / s41596-020-00464-3 (2021). Starting with streptomycin-resistant E. coli cells with a reduced genome carrying the helper plasmid pKW20_CDFtet_pAraRedCas9_tracrRNA and a genomic double selection cassette, the inventors transformed the cells with the relevant BAC and plated them on LB agar with selection for the helper plasmid (5 μg / mL tetracycline), selection for the BAC (either 18 μg / mL chloramphenicol or 50 μg / mL kanamycin), and suppression of Cas9 and lambda red expression (2% glucose). We inoculated isolated colonies into LB medium with 5 μg / mL tetracycline and antibiotic selection for the BAC and incubated the cultures overnight at 37° C. with shaking. To induce and make the cells competent, we diluted the overnight cultures 1:50 into medium with 5 μg / mL tetracycline and antibiotic selection for the BAC. When the cells reached OD 600Once it reached ≈0.2 (usually after 2 h), we induced expression of lambda red and Cas9 by adding arabinose to a final concentration of 0.5% (w / v) and continued incubation with shaking for an additional hour at 37° C. We harvested the cells and made them electrocompetent (Fredens, et al., 2019; Robertson et al., 2021).

[0177] For genomic integration of synthetic DNA by REXER, we transformed electrocompetent induced cells with 2 μg of plasmid pKW3_MB1amp_tracr_spacer encoding the spacer RNA. After 1 h recovery in 4 mL SOB medium with shaking at 37° C., we transferred the culture to 50 mL LB medium with 5 μg / mL tetracycline, selection for the spacer RNA (100 μg / mL ampicillin), antibiotic selection for the BAC, and continued incubation for 3 h at 37° C. with shaking. We transferred the culture to 50 mL LB medium with 5 μg / mL tetracycline, antibiotic selection for the BAC, and agents selecting for the negative markers on the genome and on the BAC backbone (200 μg / mL streptomycin for rpsL, 7.5% sucrose for sacB, and / or pheS). * After overnight incubation at 37°C, we picked 10-11 colonies and dissolved them in 30 μL of water. We identified each clone by colony PCR, which contains the upstream genomic double selection cassette (locus 0 ) and downstream double selection (locus 1 ) were evaluated for genomic integration. We further verified the first five clones by Sanger sequencing of colony PCR products. All oligonucleotide sequences are shown in Tables 5-14.

[0178] CONEXER CONEXER requires the preparation of conjugation between competent donor cells and recipient cells. The donor cells harbor the non-transferable conjugative plasmid pJF146 (accession no. MK809154.1, (Fredens, et al., 2019)) and a BAC carrying synthetic DNA for integration, an oriT sequence, and a universal spacer array. The orientation of oriT ensures that the spacer array enters the recipient cell last to minimize the risk of partial excision due to premature initiation of Cas9 cleavage in the recipient cell. The recipient cells harbor a locus that marks the upstream end of the integration site. 0 and the helper plasmid pKW20 for inducible expression of Cas9 and lambda Red components. Odd and even BACs can be alternated for replacement of adjacent genomic regions in repeated CONEXER steps, essentially in an alternating selection strategy as described for REXER and GENESIS (Non-Patent Documents 14, 5).

[0179] Here, we present a 100 kb synthetic DNA insert and a rpsL-kan gene on a BAC backbone. R (or sacB-cat) followed by pheS * (or rpsL), and a donor strain carrying an even (or odd) number of BACs carrying the locus 0 The genome sacB-cat (or rpsL-kan RWe describe CONEXER with recipient strains carrying the pJF146 (50 μg / mL apramycin) and BAC (50 μg / mL kanamycin or 20 μg / mL chloramphenicol) selection cassettes. We grew donor strains overnight in 25 mL LB medium with selection for pJF146 (50 μg / mL apramycin) and BAC (50 μg / mL kanamycin or 20 μg / mL chloramphenicol) until saturation. We grew recipient strains overnight in 25 mL LB medium with selection for helper plasmid (5 μg / mL tetracycline), genomic double selection cassette (20 μg / mL chloramphenicol or 50 μg / mL kanamycin), and suppression of Cas9 and lambda red expression (2% glucose) until saturation. We harvested cells from each culture by centrifugation and washed the pellets three times in 1 mL LB medium. After the final wash, we resuspended the pellets in 800 μl LB. We mixed 160 μl of recipient with 640 μl of donor, spotted the mixture onto an LB agar plate, and once the spot was dry, incubated the plate for 1 hour at 37° C. After conjugation, we washed the cells off the plate and transferred everything to 250 mL pre-warmed LB medium with selection for recipient cells carrying the helper plasmid (5 μg / mL tetracycline) and the BAC (50 μg / mL kanamycin or 20 μg / mL chloramphenicol) and induced expression of Cas9 and lambda red (0.5% L-arabinose). After incubating at 37° C. for 1.5 hours with shaking, we harvested the cells by centrifugation and immediately transferred all to 250 ml pre-warmed LB with 50 μg / mL kanamycin (or 20 μg / mL chloramphenicol), 5 μg / mL tetracycline, and 2% glucose to terminate recombination by suppressing expression of Cas9 and lambda Red. After incubating at 37° C. for an additional 2.5 hours with shaking, we spun the cultures by centrifugation and resuspended the pellet in 2 mL of filtered Milli-Q water. The cell suspension was incubated for 1.5 hours with selection for helper plasmid (5 μg / mL tetracycline), locus 1Selection for integration of the double selection cassette at the locus (50 μg / ml kanamycin or 20 μg / mL chloramphenicol) 0 Selection for loss of the double selection cassette in 0.5% sucrose or 200 μg / ml streptomycin, and selection for loss of the BAC backbone in 0.5% sucrose or 200 μg / ml streptomycin [in this case, the selection marker on the backbone is at the locus 0 Serial dilutions were spread onto LB agar plates containing sucrose (equivalent to the selection marker 0.01, and therefore not added in addition). For selection plates without sucrose, we added 2% glucose to suppress Cas9 and lambda red expression.

[0180] In experiments in the ΔrecA host (separate from the initial screen), we grew cells for 2-8 hours in 250 ml pre-warmed LB with 50 μg / mL kanamycin (or 20 μg / mL chloramphenicol), 5 μg / mL tetracycline, and 2% glucose to grow the BAC-receiving cells prior to 1.5 hour induction of Cas9 and lambda Red expression, which increased the number of successful recombinants from the CONEXER experiments.

[0181] For assembly of human DNA in episomes, some BACs are inserted into the pheS region after the human DNA insert. * -HygR double selection cassette and rpsL on the backbone were used. * To select for the maintenance of the -HygR double selection cassette, we used 200 μg / mL hygromycin. * To select for loss of the -HygR double selection cassette, we added 2.5 mM 4-CP. To select for loss of the BAC backbone, we used 200 μg / ml streptomycin.

[0182] After overnight incubation at 37°C, we picked 16-24 colonies and dissolved them in 30 μL of water. We then identified each clone by colony PCR, as described for REXER, for the locus 0 Loss of the sacB-cat cassette and the locus 1 rpsL-kan R The cassette was evaluated for integration. We selected 5 to 16 colonies with verified genotypes for whole genome sequencing by NGS. All oligonucleotide sequences are shown in Tables 5 to 14.

[0183] Next Generation Sequencing (NGS) and Analysis of Sequencing Data BACS and genomic DNA (gDNA) were extracted from overnight cultures of E. coli using the QIAprep Spin Miniprep kit and the DNEasy Blood and Tissue kit (QIAGEN), respectively. Preparation for NGS has been described previously (Robertson, WE et al. Nat Protoc 16, 2345-2380, doi:10.1038 / s41596-020-00464-3 (2021)). For preparation of many genomes, an automated workflow was performed on a Biomek FXp (Beckman Coulter) as follows: E. coli cultures (500 μL) were grown overnight in 1.2 mL 96-well plates, then resuspended in ATL buffer (QIAGEN, 90 μL) and proteinase K (10 μL) and incubated at 56 ° C for 2 hours. AMPure XP (Beckman Coulter, 100 μL) was added to each well and the plate was vortexed (1000 rpm, 6 min). The beads were magnetized (5 min), the supernatant removed, and washed with 70% EtOH (3×400 μL), after which the gDNA was eluted with Buffer AE (QIAGEN, 100 μL). The gDNA was then diluted 1:10 in HO and quantified using a standard curve and a Qubit™ dsDNA HS Assay Kit (Thermofisher) fitted to a connected fluorescent plate reader (Molecular Devices SpectraMAX I3) using a total volume of 100 μL (ex / em: 502 / 532) in a 96-well plate. The data was processed on-board and used for immediate dilution of gDNA to 0.25 ng / μL. Finally, we prepared paired-end sequencing libraries using the Nextera XT DNA Library Preparation kit (Illumina) according to the manufacturer's protocol, except for the reduced volumes: input gDNA (0.2–0.25 ng / μL, 2 μL), TD buffer (3 μL), ATM (2 μL), NT buffer (1.5 μL), index (1 μL), and NPM (3.5 μL).Index sequences were generated from the "Illumina Adapter Sequences" supporting document purchased from Biomers (Nextera DNA Index, p. 16, June 2020) and used at 10 μM. Libraries were then purified using AMPure XP magnetic beads (Beckman Coulter) according to the manufacturer's instructions (7:14 bead:reaction volume ratio), quantified by Qubit (Thermofisher), pooled, and denatured according to the manufacturer's instructions.

[0184] Libraries were paired-end sequenced on a MiSeq (Illumina, Reagent Kit v3 (600 cycles)), Illumina HiSeq2500 (200 cycles), or NextSeq 2000 (Illumina, P2 Reagent Kit v3 (100 / 200 cycles)). Downstream sequencing analysis was performed with a custom Python script as previously detailed (Non-Patent Document 14), Robertson et al., 2021. To generate a recoding landscape across the target genomic region, we used a custom Python script as previously detailed (Non-Patent Document 14), Robertson et al., 2021. The output is the frequency of recoding at each target codon plotted across the genomic region of interest.

[0185] Generation of host factor knockout strains For gene deletion and lambda red recombineering by CRISPR / Cas9-mediated cleavage, we adopted the procedure of Jiang et al. (2013) (Jiang, W., et al. Nat Biotechnol 31, 233-239, doi:10.1038 / nbt.2508 (2013)). We cloned the spacer plasmid with the spacer sequence by restriction ligation into the pMSP43 backbone with guide-encoding ssDNA oligonucleotides. Briefly, we phosphorylated the ssDNA oligonucleotides with T4 PNK (NEB), annealed, and ligated with the pMSP43 backbone. We transformed the resulting plasmid into E. coli DH10b and verified the sequence by Sanger sequencing. Single deletions of host factors were performed in streptomycin-resistant E. coli with a reduced genome that contained a sacB-cat double selection cassette integrated into LS23 carrying the helper plasmid pKW20. We cultured the strain at OD 600=0.2 and then L-arabinose (0.5%) was added to induce Cas9 and lambda red. After 1.5 h of arabinose induction, cells were harvested and made electrocompetent by washing three times with 50 mL of 20% (w / v) glycerol in ice-cold Milli-Q water. For CRISPR / Cas9-mediated cleavage, an additional helper plasmid expressing a target-specific spacer sequence (conferring spectinomycin resistance) was co-electroporated with two stop codons and a repair ssDNA oligonucleotide introducing a frameshift mutation into the target gene. After electroporation, the culture was recovered in 1 ml of SOB for 1 h at 37°C and then plated on selective LB agar plates (75 μg / mL spectinomycin, 20 μg / mL chloramphenicol, and 0.5% L-arabinose for continued Cas9 activity). The next day, we picked colonies from the selection plate and amplified the targeted gene region by colony PCR. The deletion was confirmed by Sanger sequencing. The helper plasmid (pHFXX with specR) of the deletion strain was then removed by repeated passage. The removal was confirmed by phenotyping.

[0186] Sequential genome synthesis For sequential genome synthesis, CONEXER 100k24 was first performed with a genome-reduced streptomycin-resistant E. coli ΔrecA carrying a sacB-cat double selection cassette at LS23. The next day, 40 clones were picked from the selection plate and grown individually. We phenotyped each clone to identify the loss of the sacB-cat cassette at LS23 and the loss of the rpsL-kan cassette at LS24, as described for REXER. RThe cassette was assessed for integration. Clones with the correct phenotype (39) were then pooled in equal proportions to a total volume of 25 mL. This cell pool served as the recipient culture for CONEXER 100k25. 96 clones were picked from the selection plate and expanded individually. Again, we identified each clone by phenotyping as a clone that expressed the rpsL-kan phenotype in LS24. R The clones were assessed for loss of the cassette and integration of the sacB-cat cassette at LS25. Clones with the correct phenotype (72 clones) were then pooled in equal proportions to a total volume of 25 mL. This cell pool served as the recipient culture for CONEXER 100k26. 96 clones were picked from the selection plate and grown individually. Again, we identified each clone by phenotyping for loss of the sacB-cat cassette at LS25 and integration of the rpsL-kan cassette at LS26. R The cassette was assessed for integration. Clones with the correct phenotype (53) were then pooled in equal proportions to a total volume of 25 mL. This cell pool served as the recipient culture for CONEXER 100k27. The next day, 96 clones were picked from the selection plate and grown individually. Again, we identified each clone by phenotyping as a clone that was identical to the rpsL-kan clone in LS26. R The clones were assessed for loss of the cassette and integration of the sacB-cat cassette at LS27. Clones with the correct phenotype (77) were then pooled in equal proportions to a total volume of 25 mL. This cell pool served as the recipient culture for CONEXER 100k28. The next day, 288 clones were picked from the selection plate and grown individually. Again, we identified each clone by phenotyping for loss of the sacB-cat cassette at LS27 and integration of the rpsL-kan cassette at LS28. R The integration of the cassette was assessed. Of all the clones with the correct phenotype (284), 182 were sequenced by NGS.

[0187] To calculate the expected frequency of fully recoded clones in sequential genome synthesis, we multiplied the experimentally determined frequencies of fully recoded clones for each step in CONEXER. [Table 6] JPEG2024542108000014.jpg114170JPEG2024542108000015.jpg74170 [Table 7] JPEG2024542108000017.jpg75170 [Table 8] JPEG2024542108000019.jpg110170 [Table 9] JPEG2024542108000021.jpg91170 [Table 10] JPEG2024542108000023.jpg36170 [Table 11] JPEG2024542108000025.jpg119170 [Table 12] JPEG2024542108000027.jpg117170JPEG2024542108000028.jpg106170 [Table 13] JPEG2024542108000030.jpg119170 [Table 14] JPEG2024542108000032.jpg51170 [Table 15] JPEG2024542108000034.jpg116170JPEG2024542108000035.jpg61170 [Table 16] JPEG2024542108000037.jpg95170 [Example 1]

[0188] A rapid and general method for creating custom synthetic genomes in Escherichia coli overview Whole genome synthesis and large-scale genome engineering promise to provide a powerful approach to understand the function of organisms, to engineer biosynthetic pathways on a large scale, and to create organisms with functions superior to those found in nature. A simple, robust, rapid, and scalable method for replacing genomic DNA with synthetic DNA would make genome synthesis and large-scale genome engineering faster.

[0189] Here, we report an approach that simplifies and speeds up the introduction of over 100 kb of synthetic DNA into the E. coli genome. Our method achieves this using a rapid (one day) protocol that can be repeated to introduce even larger synthetic DNA sequences. Importantly, the method standardizes and integrates all the necessary components, so that the user only needs to clone the desired synthetic DNA into a bacterial artificial chromosome and then follow the standard protocol.

[0190] result As discussed herein, it was unclear whether derivatives of REXER using universal spacers would work, because REXER using universal spacers results in non-homologous sequences being present at the ends of the excised donor nucleic acid, and it was unclear whether this would result in unwanted insertions or deletions.

[0191] To examine the efficiency and fidelity of REXER with universal spacers, we first designed and cloned two spacer RNA pairs (universal 1 and universal 2) (Table 3 and SEQ ID NOs: 24 and 25) that bind to and direct cleavage within each of the BAC scaffolds we used for E. coli genome synthesis via REXER and GENESIS (Figure 2a). Universal 1 and universal 2 target the BAC scaffolds used to genome sections designated by odd and even numbers, respectively (Figure 4).

[0192] We first demonstrated that synthetic DNA is scarlessly integrated into the genome despite non-homologous sequences at both ends resulting from Cas9 excision via universal spacer RNA from the BAC. We tested REXER at multiple loci using recoded synthetic DNAs 100k13, 100k22, 100k24, 100k28, and 100k37 (Non-Patent Document 14) (Figure 4 and Table 3). In REXER, genomic integration of synthetic DNA is selected by simultaneous positive and negative selection at antibiotic resistance markers located at the upstream end of the genomic integration site in the BAC and at the downstream end of the synthetic DNA insert (Figure 1a). Successful recombination is characterized by the 5' end of the integration site (locus 0 ) and loss of the genomic selection cassette at the 3' end (locus 1 ) resulting in integration of the selection cassette at each of the five loci tested. We confirmed successful integration of 11 / 11 post-REXER clones at each of the five loci tested (FIG. 2b).

[0193] We verified the sequences surrounding the homologous regions for recombination in clones after REXER and confirmed that all clones had lost the 6 bp non-homologous sequence (Figure 2c). Thus, the mismatched sequences created by both universal spacer sets were efficiently and reliably removed in REXER experiments at all five loci tested. We conclude that the universal spacer enables scarless integration of large synthetic DNA into the genome.

[0194] Next, we evaluated whether the use of universal spacers in REXER would affect recombination across the entire 100 kb genomic region targeted for replacement with synthetic DNA. REXER between a synthetic DNA insert from a BAC and the corresponding genomic DNA targeted for replacement can generate chimeric sequences resulting from recombination crossovers (Non-Patent Document 5). These crossovers are facilitated by a high degree of homology between the recoded DNA and the wild-type DNA: there is approximately 98.5% sequence identity between the synthetic DNA we used to perform synonymous codon compression across the entire E. coli genome and the corresponding natural genomic DNA sequence (Non-Patent Document 14). We have previously shown that chimeras resulting from REXER can be very useful for identifying and fixing sequences in synthetic DNA that are not tolerated by cells. The crossover frequency we observed with REXER is ideal because it is low enough to obtain at least one out of eight clones in which the synthetic DNA completely replaces the corresponding genomic DNA (presumably the cell tolerates the entire synthetic DNA sequence). At the same time, the crossover frequency is high enough to efficiently identify regions (including individual codon positions) in the synthetic DNA sequence that are not allowed. This is done through sequencing several post-REXER clones or a pool of post-REXER clones and analyzing the frequency of recoding at each codon position (Non-Patent Documents 14, 5), which can be visualized by summarizing the recoding situation based on the sequencing data (Figure 2d). Codon positions that have never been recoded may not allow recoding by the scheme under test.

[0195] We first compared the use of universal spacer RNA with HR-specific spacer RNA in REXER with the same recoded synthetic DNA 100k24 to replace the 96.5 kb E. coli genome. From previous REXER experiments, we knew that all designed synonymous codon replacements were tolerated in this region (Non-Patent Document 14). Although the summary situation is similar, REXER initiated with universal spacer RNA led to twice as many clones with fully integrated synthetic DNA and 40% fully recoded clones after REXER (Figure 2d and Figure 5). Thus, despite the presence of short non-homologous ends in the excised synthetic DNA, we did not observe a higher frequency of crossovers. Surprisingly, the overall efficiency of REXER with universal spacers is at least as good as with HR-specific spacers.

[0196] We conclude that the generation of terminal mismatches on the template for recombination does not prevent the efficiency of recombination and integration. We suggest that non-homologous ends of DNA in BACs may be removed by exonucleases before recombination (Non-Patent Document 19) or by flap endonucleases such as EcoIX during recombination (Non-Patent Document 20), similar to the mechanism described for eukaryotic FEN1 (Non-Patent Document 21, 22).

[0197] REXER requires two successive rounds of competent cell preparation and electroporation, which takes 4 days to obtain clonal colonies on agar plates after REXER, starting from cells with a properly marked genome. To speed up and simplify the introduction of synthetic DNA into the genome, we created BACs in which universal spacer arrays and oriT sequences were integrated into the BAC backbone (Figure 3 and SEQ ID NOs: 26 and 27). Once we mixed these new BACs (with synthetic DNA inserts) and "donor cells" carrying the non-transmissible F' plasmid with the desired "recipient cells", we selected for BAC transfer to the recipient via conjugal transfer. We turned on the expression of the Cas9 protein and lambda Red recombination components from the helper plasmid in the recipient for 1.5 hours with arabinose, and then turned off their expression for 2.5 hours with glucose. The spacers were expressed from the BACs. We selected for recipient cells in which the negative selection marker was lost from the genome and the positive selection marker was acquired from the BAC. Using this one-day universal protocol, we introduced synthetic DNA (fully recoded fragment 24, 96 kb) in place of the corresponding genomic DNA in recipient cells. In 19% of the clones, the synthetic DNA completely replaced the corresponding genomic sequence (Figure 3b). We named our short-term approach CONEXER (CONjugation coupled with programmed EXcision for Enhanced Recombination). [Example 2]

[0198] Assembly of Mbp-sized human DNA in episomes The assembly of large DNA in episomes provides a fundamental technology for genome construction. Entirely synthetic mycoplasma genomes have been assembled in yeast prior to transfer to mycoplasma. To synthesize Gbp-sized genomes of plants and animals with chromosomes of tens to hundreds of Mbp in length requires a technology to assemble Mbp-sized DNA that can be used to replace DNA in chromosomes in a reasonable number of steps.

[0199] In particular, human DNA used to sequence the essentially complete human genome was primarily captured in BACs in E. coli. The repetitive nature of many human, animal, and crop genomic sequences makes E. coli an attractive host for assembling DNA to build synthetic Gbp-sized genomes. We hypothesized that the principles we established for CONEXER could be extended to achieve scarless assembly and cloning via repeated insertion of megabase-sized DNA into episomes in E. coli.

[0200] We designed an assembly BAC into which DNA can be repeatedly inserted and assembled (Figure 6a). This BAC has approximately 50 bp of sequence (HR1) that is homologous to one end of the next sequence to be inserted. This is immediately followed by positive and negative selection cassettes and a universal homology region (uHR) that is complementary to the other end of the sequence to be inserted. We also designed a donor BAC with a CONEXER backbone with a universal spacer and oriT (Figure 6a). In the donor BAC, HR1 is within one end of the next DNA sequence to be inserted into the recipient BAC. This DNA sequence is followed by different positive and negative selection cassettes and a universal homology region. Each step of the assembly (Figure 6a, Figure 6b) proceeds by mating the donor BAC to a recipient cell carrying the assembly BAC, Cas9-mediated excision of the sequence from HR1 to the universal homology region from the donor BAC, and lambda red-mediated insertion of this sequence into the assembly BAC. Selection for loss of the negative selection marker on the assembly BAC and gain of the positive marker from the sequence excised from the donor BAC selects for cells carrying the assembly BAC with correct insertion. Cells carrying the new assembly BAC provide input for the next insertion step. We have named our approach stepwise insertion synthesis of BAC (BASIS).

[0201] We first demonstrated that assembly of the 208 Kbp human cystic fibrosis transmembrane conductance regulator gene by BASIS. We assembled the CFTR gene in three steps by BASIS in two donor BACS and one recipient episome, each with a gene fragment of approximately 70 Kbp. We verified each intermediate step assembly and the final assembly by next-generation sequencing (Figure 6c).

[0202] Next, we demonstrated that BASIS can be used to assemble large sections of human genomic DNA, including exon, intron, and intergenic regions, into one episome. We started with a library of human BACs that were used to essentially completely sequence the human genome. Each of these human BACs contains approximately 100 Kbp of human DNA, and there is considerable overlap between the human BAC sequences.

[0203] We used one-step lambda Red recombination to convert members of a human BAC library covering a region of chromosome 21 into donor BACs for BASIS. This step introduced positive and negative selection cassettes, uHR, oriT, and a universal spacer. We performed three steps of BASIS to assemble a 503 Kbp episome with 495 Kbp of human DNA. We identified correctly assembled clones by sequencing and used these clones as input for the next step of BASIS, with the final assembly also verified by sequencing (Figure 6d). [Example 3]

[0204] Minimize Crossings In our E. coli genome synthesis, genome sequencing was performed after each step of REXER to identify one correct clone that could be used as input for the next round of REXER. Since only approximately 20% of the clones from each step replaced all 100Kbp of genomic DNA with synthetic DNA, it was necessary to identify correct intermediate clones to use for progression. Thus, without identification of correct clones at each step, the frequency of completely recoded clones obtained in five steps was 3×10 -4Therefore, sequencing of tens of thousands of clones is required to identify one clone with the correct sequence. Sequencing after each step of REXER was therefore required to complete the synthesis, which significantly delayed the synthesis and increased its cost.

[0205] We envisioned repeating CONEXER by directly using the unsequenced pool of clones derived from one CONEXER as input for the next CONEXER. To this end, we set out to identify factors that would substantially increase the fraction of clones in which genomic DNA has been completely replaced by synthetic DNA in one step of CONEXER.

[0206] We identified 20 factors involved in DNA repair, replication, and recombination to test their contribution to CONEXER. We deleted each of these factors in E. coli and performed CONEXER at 100k 24 in the resulting deletion strains. These experiments identified recA and recO as factors that increase the fraction of clones with fully synthetic sequences (Figure 6a, Figure 9). Deletion of recA (ΔrecA) increased the fraction of fully synthetic DNA from 20% to approximately 80% at 100k 24 (Figure 7a). We observed similar dramatic increases in several other 100Kbp regions, clearly demonstrating the generality of our observations (Figure 7b). [Example 5]

[0207] Rapid and sequential genome synthesis Encouraged by the increased complete replacement of genomic DNA with synthetic DNA in one step of CONEXER that we observed, we investigated whether the output of one round of CONEXER could be used directly as input for the next round of CONEXER without identifying individual fully recoded clones by sequencing.

[0208] We first performed CONEXER to replace the E. coli genome with synthetic recoded DNA between LS23 and LS24 in ΔrecA E. coli carrying a +2 / -2 selection cassette at landing site (LS) 23 of its genome (Figure 8a, Figure 10). We selected for the loss of a negative 7 marker from the genome (in the +2 / -2 cassette at LS23) and the gain of a positive marker associated with the synthetic DNA (in the +1 / -1 cassette at LS24). We picked clones from the selection plates and grew them overnight. In parallel with the overnight growth, we stamped the clones onto selective agar to ensure that the clones had the correct post-CONEXER phenotype. The next day, clones with the correct phenotype, in which the genomic DNA between LS24 and LS25 had been replaced with synthetic DNA, were pooled and used as direct input for the next round of CONEXER (100k25) (Figure 8a). We selected for the loss of a negative marker from the genome (in the +1 / -1 cassette at LS24) and the gain of a positive marker associated with the synthetic DNA (in the +2 / -2 cassette at LS25). We took clones from the previous round of CONEXER, pooled them, and used the resulting pool as input for the next round of CONEXER (100k26), essentially as described. We performed five rounds of CONEXER to replace a 0.5 Mbp section of the E. coli genome between LS23 and LS28 with synthetic DNA (Figure 8a). The entire process took 10 days. Sequencing revealed that after five rounds of CONEXER, 10 percent of clones (19 of 182) were fully recoded across the targeted 0.5 Mbp genomic region (Figure 5b). We conclude that we have developed a method for rapid and continuous genome synthesis (GCS).

[0209] Consideration We have realized a one-step, one-day, universal protocol for introducing at least 100 Kbp of synthetic DNA into the E. coli genome. We have identified host factor knockouts that minimize crossover between the host genome and the synthetic DNA, allowing for continuous genome synthesis. We have demonstrated continuous genome synthesis to construct 0.5 Mbp E. coli genome sections from BACs in 10 days. Because this method is parallelizable, it will be possible to construct synthetic DNA covering the genome in about 10 days in 7-8 strains. By combining this advantage with the rapid and accurate method we obtained to assemble 0.5 Mbp synthetic recoded sections in one strain (Non-Patent Documents 14, 12), we predict that our advantages will reduce the timescale for synthesizing E. coli genomes from years to weeks. Furthermore, we predict that our approach will enable the parallel construction of many genomes, making it possible to generate genomic libraries for large-scale testing of genome-level hypotheses and discovering new cellular functions.

[0210] By extending the principles we established for E. coli genome synthesis, we have achieved scarless assembly of episomes carrying large human genomic regions. Although we have illustrated this principle through the assembly of natural sequences from human BACs (Osoegawa, K. et al. Genome Res 11, 483-496, doi:10.1101 / gr.169601 (2001)), this approach can also be used to assemble synthetic DNA fragments. Furthermore, a number of methods for multiplex editing in E. coli have been developed to edit large regions of assembled human DNA much more rapidly than in human, animal, or plant cells (Wang, HH et al. Nature 460, 894-898, doi:10.1038 / nature08187 (2009); Tong, Y., et al. Nat Commun 12, 5206, doi:10.1038 / s41467-021-25541-3 (2021); Jiang, W., et al., Nat Biotechnol 31, 233-239, doi:10.1038 / nbt.2508 (2013); Farzadfard, F. & Lu, TK Science 346, 1256272, doi:10.1126 / science.1256272 (2014)) may be combined with the assembly.These methods have been used to transfer large episomal DNA into animal cells (Waters, VL Nat Genet 29, 375-376, doi:10.1038 / ng779 (2001); Litzkas, P., Jha, KK & Ozer, HL Mol Cell Biol 4, 2549-2552, doi:10.1128 / mcb.4.11.2549-2552.1984 (1984)) and for repeated recombination of synthetic DNA into animal chromosomes (Martella, A., et al. ACS Synth Biol 6, 1380-1392, doi:10.1021 / acssynbio.7b00016 (2017); Lee, EC et al. Nat Biotechnol 32, 356-363, doi:10.1038 / nbt.2825 (2014); Macdonald, LE et al. Proc Natl Acad Sci USA 111, 5147-5152, doi:10.1073 / pnas.1323896111 (2014)).

[0211] Overall, the ability to rapidly assemble large DNA in episomes and the development of processive genome synthesis methods provide the basis for rapid and scalable genome synthesis.

[0212] (References) 1. Santos, CN, Regitsky, DD & Yoshikuni, Y. Implementation of stable and complex biological systems through recombinase-assisted genome engineering. Nat Commun 4, 2503, doi:10.1038 / ncomms3503 (2013). 2. Santos, C. N. & Yoshikuni, Y. Engineering complex biological systems in bacteria through recombinase-assisted genome engineering. Nat Protoc 9, 1320-1336, doi:10.1038 / nprot.2014.084 (2014). 3. Krishnakumar, R. et al. Simultaneous non-contiguous deletions using large synthetic DNA and site-specific recombinases. Nucleic Acids Res 42, e111, doi:10.1093 / nar / gku509 (2014). 4. Wang, G. et al. CRAGE enables rapid activation of biosynthetic gene clusters in undomesticated bacteria. Nat Microbiol 4, 2498-2510, doi:10.1038 / s41564-019-0573-8 (2019). 5. Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59-64, doi:10.1038 / nature20124 (2016). 6. Itaya, M., Tsuge, K., Koizumi, M. & Fujita, K. Combining two genomes in one cell: stable cloning of the Synechocystis PCC6803 genome in the Bacillus subtilis 168 genome. Proc Natl Acad Sci U S A 102, 15971-15976, doi:10.1073 / pnas.0503868102 (2005). 7. Lau, Y. H. et al. Large-scale recoding of a bacterial genome by iterative recombineering of synthetic DNA. Nucleic Acids Res 45, 6971-6980, doi:10.1093 / nar / gkx415 (2017). 8. Lartigue, C. et al. Creating bacterial strains from genomes that have been cloned and engineered in yeast. Science 325, 1693-1696, doi:10.1126 / science.1173759 (2009). 9. Ostrov, N. et al. Design, synthesis, and testing toward a 57-codon genome. Science 353, 819-822, doi:10.1126 / science.aaf3639 (2016). 10. Dymond, J. S. et al. Synthetic chromosome arms function in yeast and generate phenotypic diversity by design. Nature 477, 471-476, doi:10.1038 / nature10403 (2011). 11. Gibson, D. G. et al. Creation of a bacterial cell controlled by a chemically synthesized genome. Science 329, 52-56, doi:10.1126 / science.1190719 (2010). 12. Wang, K., de la Torre, D., Robertson, W. E. & Chin, J. W. Programmed chromosome fission and fusion enable precise large-scale genome rearrangement and assembly. Science 365, 922-926, doi:10.1126 / science.aay0737 (2019). 13. Hutchison, C. A., 3rd et al. Design and synthesis of a minimal bacterial genome. Science 351, aad6253, doi:10.1126 / science.aad6253 (2016). 14. Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514-518, doi:10.1038 / s41586-019-1192-5 (2019). 15. de la Torre, D. & Chin, J. W. Reprogramming the genetic code. Nat Rev Genet 22, 169-184, doi:10.1038 / s41576-020-00307-7 (2021). 16. Mercy, G. et al. 3D organization of synthetic and scrambled chromosomes. Science 355, doi:10.1126 / science.aaf4597 (2017). 17. Richardson, S. M. et al. Design of a synthetic yeast genome. Science 355, 1040-1044, doi:10.1126 / science.aaf4557 (2017). 18. Venetz, J. E. et al. Chemical synthesis rewriting of a bacterial genome to achieve design flexibility and biological functionality. Proc Natl Acad Sci U S A 116, 8070-8079, doi:10.1073 / pnas.1818259116 (2019). 19. Lovett, S. T. The DNA Damage Response. Bacterial Stress Responses, 341 2nd Edition, 205-228 (2011). 20. Anstey-Gilbert, C. S. et al. The structure of Escherichia coli ExoIX--implications for DNA binding and catalysis in flap endonucleases. Nucleic Acids Res 41, 8357-8367, doi:10.1093 / nar / gkt591 (2013). 21. Liu, Y., Kao, H. I. & Bambara, R. A. Flap endonuclease 1: a central component of DNA metabolism. Annu Rev Biochem 73, 589-615, doi:10.1146 / annurev.biochem.73.012803.092453 (2004). 22. Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157, doi:10.1038 / s41586-019-1711-4 (2019).

Claims

1. 1. A method for introducing a desired sequence into a target nucleic acid, comprising: a) providing a host cell, the host cell comprises an episomal replicon; the episomal replicon comprises a backbone sequence and a donor nucleic acid sequence; the donor nucleic acid sequence comprises, in order: 5'-homologous recombination sequence 1-desired sequence-homologous recombination sequence 2-3'; the backbone sequence comprises a first excision site located adjacent to homologous recombination sequence 1 and a second excision site located adjacent to homologous recombination sequence 2; the host cell further comprises a target nucleic acid, b) providing helper proteins capable of assisting recombination of nucleic acids in said host cell; c) providing an RNA-guided DNA endonuclease; d) providing a first RNA molecule comprising a sequence specific to the first excision site and a second RNA molecule comprising a sequence specific to the second excision site, wherein the first and second RNA molecules are involved in directing a DNA endonuclease guided by the RNA during excision; e) directing excision of the donor nucleic acid sequence by a DNA endonuclease guided by the RNA; and f) incubating to allow recombination between the excised donor nucleic acid and the target nucleic acid The method comprising:

2. 2. The method of claim 1, wherein the RNA-guided DNA endonuclease is a CRISPR-Cas nuclease, the first RNA molecule comprises a spacer specific for a first excision site, and the second RNA molecule comprises a spacer specific for a second excision site.

3. The method of claim 2, wherein the CRISPR-Cas nuclease is Cas9.

4. The method of any one of claims 1 to 3, wherein the first RNA molecule and / or the second RNA molecule is encoded by an episomal replicon.

5. The method of any one of claims 1 to 3, wherein each end of the excised nucleic acid comprises a nucleic acid sequence derived from the backbone sequence.

6. 6. The method of claim 5, wherein the excised donor nucleic acid comprises, at each end, no more than 6 or 5 base pairs of nucleic acid sequence derived from the backbone sequence.

7. The method according to any one of claims 1 to 3, wherein the episomal replicon is a bacterial artificial chromosome.

8. The method of any one of claims 1 to 3, wherein the episomal replicon is delivered to the host cell by conjugal transfer.

9. The method according to any one of claims 1 to 3, wherein the target nucleic acid is the genome of a host cell.

10. The method according to any one of claims 1 to 3, wherein the host cell is a prokaryotic cell.

11. The method according to any one of claims 1 to 3, wherein the prokaryotic cell is Escherichia coli.

12. 1. A method for assembling a nucleic acid sequence, comprising: (i) performing the steps of claim 1 to introduce a first donor nucleic acid sequence into a first target nucleic acid, thereby generating a second target nucleic acid; and (ii) performing the steps of claim 1 to introduce a second donor nucleic acid sequence into the second target nucleic acid, thereby generating a third target nucleic acid. The method comprising:

13. 13. The method of claim 12, wherein parts (i) and (ii) are repeated.

14. the sequence of the first RNA molecule for part (i) is identical in each repeat, and / or the sequence of the second RNA molecule for part (i) is identical in each repeat, 14. The method of claim 13, wherein the sequence of the first RNA molecule for part (ii) is identical in each repeat, and / or the sequence of the second RNA molecule for part (ii) is identical in each repeat.

15. (iii) performing the steps of claim 1 to obtain a third donor nucleic acid sequence; into a third target nucleic acid, thereby creating a fourth target nucleic acid; repeating parts (i), (ii), and (iii). further comprising the sequence of the first RNA molecule for part (iii) is identical in each repeat, and / or wherein the sequence of the second RNA molecule for part (iii) is identical in each repeat.

15. The method according to any one of 2 to 14.

16. wherein part (i) comprises the use of an episomal replicon encoding a donor nucleic acid sequence comprising a first backbone sequence, and part (ii) comprises the use of an episomal replicon encoding a donor nucleic acid sequence comprising a second backbone sequence; the first scaffold sequence comprises a first marker or set of markers and encodes a first RNA molecule specific for a first excision site within the first scaffold sequence, and encodes a second RNA molecule specific for a second excision site within the first scaffold sequence; the second scaffold sequence comprises a second marker or set of markers and encodes a first RNA molecule specific for a first excision site within the second scaffold sequence, and encodes a second RNA molecule specific for a second excision site within the second scaffold sequence; The method of any of claims 12 to 14, wherein the first marker or set of markers is different from the second marker or set of markers.

17. 1. A method for constructing an episomal replicon, comprising: a) providing a donor episomal replicon, said replicon comprising: a backbone comprising a universal spacer sequence, a first homology region HRn specific to integration step n, and a second universal homology region uHR, a first excision site located adjacent to HRn and a second excision site located adjacent to uHR; donor nucleic acid DNAn, Double selection cassette containing a positive and a negative selection marker The steps comprising: b) providing a host cell comprising an assembly episomal replicon comprising a double selection cassette comprising a positive selection marker and a negative selection marker flanked by HRn and uHR, wherein the double selection cassette in the assembly replicon comprises a different marker than the selection cassette in the donor replicon; c) providing helper proteins capable of assisting recombination of nucleic acids in said host cell; c) providing an RNA-guided DNA endonuclease; d) providing a first RNA molecule comprising a sequence specific to the first excision site and a second RNA molecule comprising a sequence specific to the second excision site, wherein the first and second RNA molecules are involved in directing a DNA endonuclease guided by the RNA during excision; e) inducing excision of the donor nucleic acid sequence DNAn by a DNA endonuclease guided by the RNA in the host cell; and f) incubating to allow recombination between the excised donor nucleic acid and the assembly replicon to form a second assembly replicon containing the nucleic acid DNAn; The method comprising:

18. 18. The method of claim 17, wherein the RNA-guided DNA endonuclease is a CRISPR-Cas nuclease, the first RNA molecule comprises a spacer specific for a first excision site, and the second RNA molecule comprises a spacer specific for a second excision site.

19. The method of claim 18, wherein the CRISPR-Cas nuclease is Cas9.

20. The method of any one of claims 17 to 19, wherein the first RNA molecule and / or the second RNA molecule is encoded by a donor episomal replicon.

21. The method of any one of claims 17 to 19, wherein each end of the excised nucleic acid comprises a nucleic acid sequence derived from the backbone sequence.

22. 22. The method of claim 21 , wherein the excised donor nucleic acid comprises no more than 6 or 5 base pairs of nucleic acid sequence derived from the backbone sequence at each end.

23. The method according to any one of claims 17 to 19, wherein the episomal replicon is a bacterial artificial chromosome.

24. The method of any one of claims 17 to 19, wherein the episomal replicon is delivered to the host cell by conjugal transfer.

25. 25. The method of claim 24, wherein an episomal replicon is contained in a donor host cell and an assembly replicon is contained in a recipient host cell, the donor replicon is delivered to the recipient host cell by conjugal transfer, and the donor host cell contains a non-transmissible F' plasmid.

26. The method according to any one of claims 17 to 19, wherein the host cell is a prokaryotic cell.

27. The method according to any one of claims 17 to 19, wherein the prokaryotic cell is Escherichia coli.

28. 20. The method of any one of claims 17 to 19, wherein the donor nucleic acid DNAn comprises a homology region HRn+1, and the method further comprises the steps of introducing into a host cell an additional donor episomal replicon comprising a second donor nucleic acid DNAn+1, inducing excision of the donor nucleic acid sequence DNAn+1 in the host cell by an RNA-guided DNA endonuclease, and incubating to allow recombination between the excised donor nucleic acid DNAn+1 and the second assembly replicon to occur, forming a third assembly replicon comprising nucleic acid DNAn and nucleic acid DNAn+1.

29. 29. The method of claim 28, wherein the method is repeated repeatedly.

30. The method of claim 12, wherein the episomal replicon of the step of claim 1 is constructed according to the method of claim 17.

31. The method of claim 1 , wherein the host cell lacks competent recA and / or recO.

32. 32. The method of claim 31 , wherein the host cell is deficient in recA (ΔrecA).