In vivo DNA assembly and analysis

JP2024509194A5Active Publication Date: 2026-02-17THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023553587
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-03-05
Filing Date
2022-03-04
Publication Date
2026-02-17
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

Current methods for oligonucleotide assembly are expensive, time-consuming, and limited by the size and composition of DNA elements that can be combined, with challenges in identifying and isolating unique DNA sequences in complex mixtures due to complex purification requirements and low sample recovery rates.

Method used

A method for assembling DNA elements in vivo using donor plasmids and recipient oligonucleotides through conjugation and homologous recombination, involving sequential transfer and recombination of DNA fragments with endonuclease sites and homologous recombination regions to form DNA assemblies.

Benefits of technology

Enables efficient, high-throughput assembly of DNA elements up to 500,000 nucleotides in length, overcoming limitations of existing methods by reducing the need for multiple enzymes and improving sample recovery and sequencing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000084_0000
    Figure 00000084_0000
  • Figure 00000084_0001
    Figure 00000084_0001
  • Figure 00000084_0002
    Figure 00000084_0002
Patent Text Reader

Abstract

Among other things, methods and compositions are provided herein for assembling oligonucleotide fragments in vivo. The methods avoid the need for inefficient cloning methods and expensive enzymes. The methods may further be used to assemble long fragments of DNA. The methods may be used to generate variant and combinatorial libraries, and may be used to track biological processes. Among other things, methods are provided herein for in vivo DNA barcoding of oligonucleotide sequences. The methods provided herein are intended to generate unique barcode-oligonucleotide fusion sequences, for example, for identifying and isolating oligonucleotide sequences from a mixture. Thus, methods are also provided for identifying oligonucleotides from a mixture of oligonucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims the benefit of priority under 35 USC Section 119(e) to U.S. Provisional Application No. 63 / 157,497, filed March 5, 2021, and U.S. Provisional Application No. 63 / 157,498, filed March 5, 2021, the entire contents of each of which are incorporated herein by reference in their entirety.

[0002] The contents of the sequence listing text file named "41243-570001WO_Sequence_Listing_ST25.txt", created on February 15, 2022 and having a size of 24,576 bytes, are incorporated by reference in their entirety into this specification.

[0003] This invention was made with Government support under DOE FWP #100582 awarded by the Department of Energy and NIST IAA P18-630-0001 awarded by the National Institute for Standards and Technology. The Government has certain rights in the invention. [Background technology]

[0004] Recent advances in recombinant oligonucleotide technology have stimulated research in traditional biology and biotechnology fields. However, the process of oligonucleotide assembly can be expensive and time-consuming, requiring multiple purification steps and various enzymes. In addition, current molecular biology methods are limited in the size and composition of DNA elements that can be combined. Therefore, methods are needed to assemble DNA elements (e.g., promoters, gene fragments, etc.) together to address these limitations and avoid the need for numerous and expensive enzymes (e.g., ligases, etc.). New methods are needed for efficient, high-throughput, and versatile assembly of DNA fragments on a quantitatively comparable scale. Summary of the Invention [Problem to be solved by the invention]

[0005] Advances in sequencing technology have enabled the identification of long fragments of DNA, but the identification and isolation of unique DNA sequences in complex mixtures remains difficult due to complex purification requirements, low sample recovery, or inefficient sequencing workflows, among other challenges.

[0006] Among other things, solutions to these and other problems in the art are provided herein. [Means for solving the problem]

[0007] Summary of the Invention Among other things, methods and compositions are provided herein for assembling DNA elements in vivo and DNA barcoding oligonucleotide sequences.

[0008] The invention provides a method of assembling multiple DNA elements into an assembled DNA element in a recipient cell, the method comprising the steps of: (a) contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions that (i) transfer a first donor plasmid from the first donor cell to a recipient cell by conjugation, and (ii) allow the first donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination, The plasmid comprises, in sequential order, an optional first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a fragment of a first DNA element (oligo 1), a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2), and an optional third endonuclease site (C3), and the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2, thereby forming a first endonuclease site that is a first endonuclease site (C4) and a second endonuclease site (C5) that is a second endonuclease site (C6) and a third endonuclease site (C7) that is a second endonuclease site (C8) and a fourth endonuclease site (C9) that is a second endonuclease site (C1) and a fourth endonuclease site (C2) that is a second endonuclease site (C3) and a fourth endonuclease site (C4) that is a second endonuclease site (C4) and a fourth endonuclease site (C5) that is a second endonuclease site (C6) and a fourth endonuclease site (C7) that is a second endonuclease site (C8) and a fourth endonuclease site (C9) that is a second endonuclease site (C1) and a fourth endonuclease site (C1) that is a second endonuclease site (C1) and a fourth endonuclease site (C2) that is a second endonuclease site (C2) and a fourth endonuclease site (C3) that is a second endonuclease site (C3) and a fourth endonuclease site (C4) that is a second endonuclease site (C4) and a fourth endonuclease site (C5) that is (b) providing a first recombined recipient oligonucleotide comprising a fragment of the first DNA element following homologous recombination of HR.2 and HR4; (i) transferring the second donor plasmid from the second donor cell to the first recipient cell by conjugation; and (ii) reacting the second donor plasmid with the first recombined recipient oligonucleotide in the recipient cell by homologous recombination to form a second recombined recipient oligonucleotide. contacting a recipient cell containing a bound recipient oligonucleotide, the second donor plasmid containing, in sequential order, an optional fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a fragment of a second DNA element (oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2), and an optional sixth endonuclease site (C6), thereby binding HR5 to HR2.1 and HR6.Following homologous recombination of HR2 and HR4, providing a second recombined recipient oligonucleotide comprising fragments of the first and second DNA elements (oligo1, oligo2), which form a DNA assembly. In an embodiment, HR2.1 and HR2.2 are flanked by non-homologous regions comprising one (C2) or two (C2.1, C2.2) endonuclease sites, and optionally, HR3 and HR4 are flanked by non-homologous regions comprising one (C4) or two (C4.1, C4.2) endonuclease sites. In an embodiment, HR6.1 and HR6.2 are flanked by non-homologous regions comprising one (C7) or two (C7.1, C7.2) endonuclease sites. In an embodiment, the recipient oligonucleotide is present in a recipient cell plasmid or a recipient cell genome. In embodiments, the DNA assembly includes at least a portion of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, a guide RNA (gRNA), or a combination thereof.

[0009] In an embodiment, step (b) is repeated in one or more iterations using a third or subsequent donor cell comprising a third or subsequent donor plasmid comprising a matching HR region and a third or subsequent oligonucleotide (oligo 3, oligo 4, ... oligo N) encoding a fragment of the third or subsequent DNA element, thereby forming a third or subsequent recombined recipient oligonucleotide comprising fragments of the first, second and third or subsequent DNA elements, which together form a DNA assembly.

[0010] In an embodiment, step (a) comprises a plurality of first donor cells each comprising a different first donor plasmid, and step (b) comprises a plurality of second, third or subsequent donor cells each comprising a different second, third or subsequent donor plasmid, optionally each first donor cell being at a position in a first ordered array, and each second, third or subsequent donor cell being at a position in a second, third or subsequent ordered array, and optionally the method generates a combinatorial library comprising a plurality of different assembled DNA elements.

[0011] In an embodiment, the oligonucleotide encoding a first endonuclease targeting the first, third, and / or fourth endonuclease site is present on the first donor plasmid and / or present in the recipient cell. In an embodiment, the oligonucleotide encoding a second endonuclease targeting the second, fifth, and / or sixth endonuclease site is present on the second donor plasmid and / or present in the recipient cell. In an embodiment, the expression of the first and / or second endonuclease is inducible, and the method further comprises inducing the expression of the first and / or second endonuclease. In an embodiment, the first and / or second endonuclease is selected from an RNA-guided endonuclease, a homing endonuclease, a transcription activator-like effector nuclease, and a zinc finger nuclease.

[0012] In an embodiment, the first, second, or subsequent donor plasmid comprises a selectable marker that selects for integration of the first, second, or subsequent oligonucleotide into the recipient oligonucleotide, optionally the selectable marker being present in the non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions. In an embodiment, the recipient oligonucleotide comprises a counter-selectable marker that selects against recipient cells that do not comprise the first, second, third, or subsequent oligonucleotide, optionally the counter-selectable marker being present in the non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions.

[0013] In an embodiment, the donor plasmid comprises an origin of transfer.

[0014] In embodiments, the donor plasmid comprises a conditional origin of replication. In embodiments, the conditional origin of replication is dependent on the presence of the oligonucleotide or on the conditions of cell growth. In embodiments, the donor plasmid or the recipient oligonucleotide comprises an inducible high copy origin of replication.

[0015] In embodiments, the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid greater than 30 kilobases in length.

[0016] In embodiments, the donor plasmid or recipient oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome.

[0017] In embodiments, the donor plasmid or recipient oligonucleotide is a viral vector.

[0018] In embodiments, the donor plasmid contains oligonucleotides that enable conjugation of the plasmid.

[0019] In embodiments, the donor plasmid or recipient cell comprises an oligonucleotide encoding one or more homologous DNA repair genes, and optionally, expression of the one or more homologous DNA repair genes is inducible.

[0020] In embodiments, the donor plasmid or recipient cell comprises an oligonucleotide encoding one or more recombination-mediating genetic engineering genes.

[0021] In embodiments, the donor cells and the recipient cells are independently bacterial cells, and optionally the bacterial cells are E. coli, Vibrio natriegens, or V. cholerae.

[0022] In embodiments, the assembled DNA elements are between 100 nucleotides and 500,000 nucleotides in length.

[0023] In embodiments, the first, second, or subsequent homologous recombination (HR) regions and their corresponding HR regions on the recipient oligonucleotide each comprise from about 20 base pairs to about 500 base pairs, optionally from about 50 to 100 base pairs.

[0024] In embodiments, any of the above methods may further comprise one or more of the steps of lysing the recipient cells, amplifying the assembled DNA elements, isolating the assembled DNA elements, isolating the recipient oligonucleotides, sequencing the assembled DNA elements, and sequencing the recipient oligonucleotides.

[0025] In embodiments, the steps of contacting the first donor cell and the second or subsequent donor cell with the first recipient cell are performed simultaneously, and optionally, only the final donor plasmid contains a selectable marker or each donor plasmid contains a selectable marker that is not present on the recipient oligonucleotide.

[0026] In one embodiment of any of the above methods, the donor plasmid comprising the final DNA element forming part of the assembled DNA element comprises a barcoded homologous recombination (BHR) region that produces recipient cells each comprising the assembled DNA element, a barcoded BHR region, and a recombined recipient oligonucleotide comprising an additional HR, the method further comprising the steps of (i) constructing or obtaining an array of barcoded donor cells, each comprising a barcoded donor plasmid comprising a HR homologous to the BHR, a unique barcode oligonucleotide, and a second HR homologous to the additional HR of the recombined recipient oligonucleotide; (ii) (a) transferring the barcoded donor plasmid from the barcoded donor cell to a recipient cell by conjugation; and (b) contacting the array of barcoded donor cells with the array of recipient cells under conditions that allow the barcoded donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination, thereby producing an array of recipient cells comprising the barcoded assemblies.

[0027] In one embodiment of any of the above methods, each donor plasmid comprises an additional pair of unique endonuclease sites CX, CY flanking the barcode homologous recombination (BHR) region, and the method further comprises contacting an array of recipient cells each comprising a DNA assembly with an array of barcode donor cells comprising barcode donor plasmids each comprising a pair of HR regions homologous to the BHR flanking a unique barcode oligonucleotide to produce an array of recipient cells comprising the barcoded assemblies.

[0028] In one embodiment of any one of the above methods, the method further comprises contacting a reset donor cell comprising the reset donor plasmid with a recipient cell comprising the recombined recipient oligonucleotide, the reset donor plasmid comprising, in sequential order, a homologous recombination region (HRt) homologous to the end sequence of the DNA assembly, a reset endonuclease site, a selectable marker, a reset endonuclease site, a homologous recombination region (HRX), and an origin of transfer, and the recombined recipient oligonucleotide comprising, in sequential order, a reset endonuclease site, a DNA assembly, a homologous recombination region (HRXa) homologous to HRX, and a reset endonuclease site, thereby providing a reset plasmid comprising an origin of transfer and a DNA assembly following homologous recombination between HRt and the end sequence of the DNA assembly and between HRX and HRXa. In an embodiment, the reset plasmid is in the donor cell. In an embodiment, the reset plasmid comprises a restricted origin of replication that functions in both the donor cell and the recipient cell. In an embodiment, the reset donor plasmid is constructed by a method comprising the steps of introducing an oligonucleotide insert HRt-C1-CM-C2-HRX, or a library of such oligonucleotide inserts, comprising homologous recombination regions HRt, HRX flanked by two endonuclease sites (C1, C2) and a counter-selectable marker (CM), allowing an endonuclease to cleave the endonuclease sites, and using homologous recombination to introduce the counter-selectable marker at the cleavage site.

[0029] In an embodiment of any of the above methods, the recipient oligonucleotide comprises a mobile genetic element capable of transferring the DNA assembly to other cell types, including yeast cells, plant cells, mammalian cells, or other bacterial cells.

[0030] In an embodiment of any of the above methods, the method comprises constructing a DNA library utilizing two or more recipient oligonucleotides that have compatible homologous recombination regions.

[0031] In any of the above methods, the oligonucleotides of the donor plasmid include a first linker oligonucleotide that is homologous to the terminal sequence of the first DNA assembly and a second linker oligonucleotide that is homologous to the second oligonucleotide. In an embodiment, the linker oligonucleotide further includes a fragment of an additional DNA element that is not homologous to the first DNA assembly or the second DNA oligonucleotide. In an embodiment, the method is used to assemble a mutagenesis library, to combine genetic regions such as genes, promoters, terminators, and regulatory regions from different species, to construct and / or combine genetic regulatory pathways, to construct a combinatorial gRNA library, or to assemble an array of bacteria containing plasmids for screening assays.

[0032] In any embodiment of the above method, prior to steps (a) and (b), first and second oligonucleotides comprising fragments of the first and second DNA elements are inserted into first and second donor plasmids.

[0033] The present invention also provides a method for conjugating a barcode to an oligonucleotide, the method comprising: (a) inserting each oligonucleotide of the mixture of oligonucleotides into a donor plasmid, each of which optionally comprises a first endonuclease site (C1), a first homologous recombination region (HR1), a second homologous recombination region (HR2), and optionally a second endonuclease site (C2) in sequential order, with each oligonucleotide being inserted between HR1 and HR2, thereby providing a plurality of donor plasmids comprising the donor oligonucleotide, each donor plasmid comprising a single donor oligonucleotide C1-HR1-oligo-HR2-C2 from the mixture of oligonucleotides; (b) transforming a plurality of cells with the plurality of donor plasmids, whereby each cell comprises the donor plasmid, thereby forming a plurality of donor cells; (c) plating and culturing the plurality of donor cells, each at a unique location on the first ordered array, thereby providing a first ordered array of donor cells; and (d) providing a plurality of recipient cells into a second ordered array. wherein each recipient cell comprises a recipient oligonucleotide comprising, in sequential order, a unique barcode sequence identifying the location of the recipient cell in the second ordered array, a third homologous recombination region (HR3) homologous to HR1, optionally a third endonuclease site (C3), and a fourth homologous recombination region (HR4) homologous to HR2; (e) contacting the first ordered array of donor cells with the second ordered array of recipient cells under conditions to (i) transfer the donor plasmid from the donor cell to the recipient cell at the corresponding location on the array by conjugation, (ii) optionally cleave the first, second, and third endonuclease sites, and (ii) transfer the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby forming a third array of fusion oligonucleotides each comprising the unique barcode sequence and a donor oligonucleotide from the mixture of oligonucleotides; and (f) optionally sequencing the fusion oligonucleotides, therebyThe method includes identifying each oligonucleotide in the array by its barcode sequence. In an embodiment, the recipient oligonucleotide is present in a recipient cell plasmid or in a recipient cell genome. In an embodiment, the donor plasmid includes a selectable marker between HR1 and HR2 that selects for integration of the oligonucleotide into the recipient cell oligonucleotide, and optionally the donor plasmid includes a counter-selectable marker. In an embodiment, the recipient cell oligonucleotide includes a fourth endonuclease site (C4).

[0034] The invention also provides a method of identifying an oligonucleotide from a plurality of oligonucleotides, the method comprising the steps of: (a) providing a plurality of donor cells in a first ordered array, each donor cell comprising a donor plasmid each comprising, in sequential order, an oligonucleotide from the plurality of oligonucleotides, optionally a first endonuclease site (C1), a first homologous recombination region (HR1), a unique barcode sequence, a second homologous recombination region (HR2), and optionally a second endonuclease site (C2), wherein the unique barcode sequence identifies a location of the host cell in the first ordered array; (b) providing a plurality of recipient cells, each recipient cell comprising, in sequential order, an oligonucleotide from the plurality of oligonucleotides, a third homologous recombination region homologous to HR1 (HR3), optionally a third endonuclease site (C3), and a fourth homologous recombination region homologous to HR2 (HR4). (c) plating and culturing a plurality of recipient cells, each at a unique location on the second ordered array, thereby providing a second ordered array of recipient cells, (d) contacting the first ordered array with the second ordered array under conditions that (i) transfer the donor plasmid from the donor cells to the recipient cells at corresponding locations on the array by bacterial conjugation, (ii) cleave the first, second, and third endonuclease sites, and (ii) transfer the barcode sequence from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby forming a third array of fusion oligonucleotides each comprising a unique barcode sequence and an oligonucleotide from the mixture of oligonucleotides, and (e) sequencing the fusion oligonucleotides, thereby identifying each oligonucleotide in the array by its barcode sequence. In an embodiment, the recipient oligonucleotide is present in a recipient cell plasmid or a recipient cell genome.In an embodiment, the donor plasmid comprises a selectable marker between HR1 and HR2 that selects for integration of the barcode sequence into the recipient cell oligonucleotide, and optionally the donor plasmid comprises a counter-selectable marker. In an embodiment, the recipient cell oligonucleotide comprises a fourth endonuclease site (C4). In an embodiment, the first endonuclease site, the second endonuclease site, and the third endonuclease site are the same or different. In an embodiment, the donor plasmid comprises an origin of transfer and / or a conditional origin of replication, optionally the origin of transfer is derived from a mobile element, and further optionally the conditional origin of replication is dependent on the presence of the oligonucleotide or the conditions of cell growth. In an embodiment, the donor plasmid or the recipient plasmid comprises a replicon capable of replicating a plasmid at least 30 kilobases in length, and optionally the replicon is derived from a P1-induced artificial chromosome or a bacterial artificial chromosome. In an embodiment, the donor plasmid or the recipient cell oligonucleotide comprises an inducible high-copy origin of replication. In an embodiment, the donor plasmid or the recipient cell oligonucleotide comprises a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome. In an embodiment, the donor plasmid or the recipient cell oligonucleotide comprises a viral vector. In an embodiment, the endonuclease site is cleaved by one or more endonucleases encoded by one or more oligonucleotides in the recipient cell and / or encoded in the donor plasmid, optionally, the one or more endonucleases are homing endonucleases or RNA-guided DNA endonucleases, and further optionally, the endonuclease is HO. In an embodiment, the donor cell or the recipient cell comprises an oligonucleotide that (i) allows conjugation of the plasmid, (ii) encodes one or more homologous DNA repair genes, or (iii) encodes one or more recombination-mediated genetic engineering genes.In embodiments, the donor cell, recipient cell, or recombinant recipient cell is transferred to a location on a third ordered array, a fourth ordered array, or a subsequent ordered array. In embodiments, the donor cell and the recipient cell are independently bacterial cells, and optionally the bacterial cells are E. coli, Vibrio natriegens, or V. cholerae. In embodiments, the barcode sequence is about 4 nucleotides to about 100 nucleotides in length, and optionally the barcode sequence is about 30 nucleotides in length. In an embodiment, the mixture of oligonucleotides is the product of a DNA synthesis or assembly technique selected from chemical coupling, template-independent enzymatic synthesis using polymerase-nucleotide conjugates, polymerase chain assembly (polymerase cycling assembly), Gibson assembly (Chew-back, anneal and repair), ligase chain reaction / ligase cycling reaction, Phi29 polymerase, rolling circle, loop-mediated isothermal (LAMP), strand displacement (SDA), helicase-dependent (HAD), recombinase polymerase (RPA), nucleic acid sequence-based amplification (NASBA), Golden Gate cloning, MoClo cloning, BioBricks or assembled BioBricks, thermodynamically balanced back-to-front synthesis, DNA cloning, ligation-independent cloning, ligation by selection cloning, recombinational engineering, yeast assembly, PCR, capture with molecular inversion probes or LASSO probes, DropSynth, and enzymatic DNA synthesis. In an embodiment, the mixture of oligonucleotides is the product of pooled mutagenesis techniques selected from polymerase chain reaction techniques including error-prone PCR, PCR with degenerate oligos, and regular PCR, chemical or photomutagenesis, in vitro synthesis with libraries of oligos, in vivo editing, e.g., MAGE, MAGESTIC, CRISPR, prime editing, retron editing, and base modification with CRISPR, TALEN, and zinc finger nucleases.In an embodiment, the mixture of oligonucleotides comprises at least a fragment of genomic DNA, cDNA, organelle DNA, or natural plasmid DNA. In an embodiment, the mixture of oligonucleotides comprises captured or amplified DNA from gDNA, cDNA, or organelle DNA, e.g., from a balanced cDNA library, PCR products, e.g., multiplex PCR products, molecular inversion probes, including LASSO probes, capture by annealing or subtractive hybridization, co-transformation and homologous recombination, rolling circle amplification, or LAMP. In an embodiment, the mixture of oligonucleotides comprises captured or amplified DNA from a plasmid or a plasmid library, e.g., an open reading frame (ORF) library, a promoter library, a terminator library, an intron library, a BAC library, a PAC library, a lentiviral library, a gRNA library, a PCR product, a restriction digestion product, or a GATEWAY shuttling product. In an embodiment, the oligonucleotides of the mixture of oligonucleotides are integrated into the donor plasmid by a method comprising co-transformation and recombination, transformation and recombination, or conjugation and recombination. In embodiments, the method of co-transformation and recombination engineering includes constructing a linear or circular donor plasmid containing a selectable marker and two homologous recombination regions that are homologous to sequences at the ends of the oligonucleotides in the mixture, respectively, co-transforming cells with the donor plasmid and the oligonucleotides, inducing homologous recombination, and selecting for the selectable marker, optionally the method is performed with a library or pool of donor plasmids and / or oligonucleotides.In an embodiment, the method of transformation and recombination engineering includes constructing a linear or circular donor plasmid containing a selectable marker and two homologous recombination regions, each homologous to sequences at the ends of the oligonucleotides in the mixture, where the oligonucleotides are present on a plasmid in a host cell, transforming the host cell with the donor plasmid, inducing homologous recombination, and selecting for the selectable marker. In an embodiment, the method of conjugation and recombination engineering includes the steps of constructing a linear or circular donor plasmid containing two optional endonuclease sites and a counter-selectable marker (-1) flanked by two homologous recombination (HR) regions, where the donor plasmid is present in a donor cell containing a disabled F plasmid that can induce conjugation but may not conjugate, and the oligonucleotides of the mixture are present on a plasmid in a recipient cell, each flanked by HR regions homologous to the HR regions of the donor plasmid and flanked by at least one selectable marker (+1) that selects for recombination of each oligonucleotide into the donor plasmid, providing a homologous recombinase and optionally one or more endonucleases in the recipient cell or encoded by the donor plasmid, (i) transferring the donor plasmid from the donor cell to the recipient cell by bacterial conjugation, and (ii) contacting the donor cell with the recipient cell under conditions that recombine the donor and recipient plasmids by homologous recombination, and selecting for cells that contain the selectable marker but do not contain the counter-selectable marker. In an embodiment, the method is performed using a library of donor and / or recipient plasmids.In an embodiment, the oligonucleotides comprise a library such as an ORF library, a promoter library, a terminator library, an intron library, a BAC library, a PAC library, a lentiviral library, a gRNA library, a gDNA library, a cDNA library, a protein domain library, a promoter library, a terminator library, a library of regulatory elements, a library of structural elements, or a library of DNA variants derived from DNA mutagenesis. In an embodiment, the mixture of oligonucleotides comprises an array of cells comprising a plasmid library, such as a gRNA library, a gDNA library, a cDNA library, an open reading frame (ORF) library, a protein domain library, a promoter library, a terminator library, a library of regulatory elements, a library of structural elements, or a library of DNA variants derived from DNA mutagenesis. In an embodiment, the mixture of oligonucleotides comprises an array of cells comprising fragments of DNA elements for use in the method of any one of claims 1 to 35. [Brief description of the drawings]

[0035] [Figure 1]Schematic diagram of an embodiment of a method for DNA assembly described herein showing donor plasmids and recipient oligonucleotide elements as shaded boxes. The diagram shows three "rounds" of "DNA stitching" in each of which a new oligonucleotide containing a fragment of a DNA element is added to a recipient oligonucleotide. Throughout the diagram, the fragments of DNA elements are variously referred to as Input DNA1, Input DNA2, etc. before recombination, DNA1, DNA2, etc. after recombination, or Oligo1, Oligo2, etc., or may be more commonly referred to as "DNA blocks." Boxes designated C1, C2, etc. refer to endonuclease sites, boxes designated HR1, HR2, etc. refer to regions of homologous recombination, and boxes designated Oligo1, Oligo2, etc. refer to oligonucleotides containing fragments of DNA elements. Boxes designated with numbers and plus or minus signs refer to selectable (+) and counterselectable (-) markers. Not all elements shown in the schematic diagrams are required in every embodiment of the methods described herein; for example, the markers and endonuclease sites designated C2.1, C2.2, C7.1, C7.2 are optional. [Figure 2A] Panel A is a schematic map of an exemplary recipient oligonucleotide (in the form of a plasmid) used in the methods described herein. [Figure 2B] Panel B shows two schematic maps of exemplary donor and helper plasmids. [Diagram 3] CRISPR / Cas9 enhances the efficiency of DNA assembly, also referred to herein as "suture." Number of colonies per 6 x 106 cells on selection plates. Two donor plasmids, only one of which expresses a functional gRNA, were transformed into each of three recipient strains: (1) BW28705 (without λ-red, without Cas9), (2) BW28705 / pML300 (with λ-red, without Cas9), and (3) BW28705 / pSL359 (with Cas9 and λ-red). [Figure 4]Schematic map of exemplary plasmids used in the in vivo DNA assembly or "suture" method described herein. In the donor cell, the conjugation-competent helper plasmid may contain genes for plasmid transfer (Tra operon). To immobilize the conjugation plasmid itself, the origin of transfer (oriT) is replaced with a selectable marker (+6). The donor plasmid contains a swapping cassette (+1 / -1 or +2 / -2), two homology regions (H2 and H3), four endonuclease cleavage sites (two circles labeled 1 and two circles labeled 2), a backbone selectable marker (+4), a conditional origin of replication (R6K) depending on the allele in the donor genome (pir1-116), an oriT sequence, and a gRNA expression cassette (gRNA1 or gRNA2). [Diagram 5] Schematic map of exemplary plasmids used in the in vivo suture method described herein. In the recipient cell, the helper plasmid contains a rhamnose-inducible red operon (PrhaBAD-red), an arabinose-inducible Cas9 (ParaBAD-Cas9), an E. coli RecA gene to enhance homologous recombination, a backbone selectable marker (+5), and a curable origin of replication (pSC101 oriTS). The recipient plasmid contains two endonuclease cleavage sites (two circles labeled 1), a swapping cassette (+2 / -2), two regions of homology (H2 and H3), and an origin of replication (ColE1). +1: HygR, +2: NsrR, -1: SacB, -2: PheS, +3: GmR, +4: KanR, +5: SpR, +6: TcR. [Figure 6]Schematic of an exemplary method of in vivo stitching. A donor plasmid carrying a DNA fragment (up or down striped rectangle) is introduced into the donor plasmid and donor cell. The donor plasmid is conjugated to a recipient cell and the DNA fragment is transferred from the donor plasmid to the recipient plasmid. The plasmid is cleaved using CRISPR / Cas9, which is induced by arabinose. Guide RNAs (gRNA1 or gRNA2, which alternate between assembly rounds) on the donor plasmid specify the recognition sequences for cleavage ("1" and "2" circles, which alternate between assembly rounds). Homology regions on both the synthesized oligos and the plasmid backbone (H1 and H3 in round 1) promote rhamnose-induced recombination to seamlessly stitch the oligos together for gene assembly. Alternating selectable (+1 and +2) and counterselectable (-1 and -2) markers on the donor plasmid allow repeatable DNA transfer with a maximum gene length theoretically set by the maximum allowable plasmid size. R6K and ColE1 are origins of replication. +3 and +4 are selectable markers used to maintain the plasmid. [Figure 7A] Example of DNA assembly. Panel A shows three donor plasmids (triads), each carrying a portion of mEGFP, that were sequentially conjugated and assembled into a recipient plasmid. [Figure 7B] Example of DNA assembly. Panel B shows the fluorescence of colonies from the negative control, positive control, and in vivo suture products after three rounds of assembly in liquid. The colonies represent independent mating and recombination events and are 100% fluorescent. [Figure 7C] Examples of DNA assemblies. Panel C shows aligned assemblies of mEGFP in 96-position and 384-position formats. All colonies appear fluorescent. [Figure 7D] Example of DNA assembly. Panel D shows the percentage of fluorescent colonies after a final round of liquid assembly using a third mEGFP fragment of different length that is homologous to the second mEGFP fragment. [Figure 7E] Example of DNA assembly. Panel E shows representative restriction digests of colonies containing the various plasmids scraped from agar during assembly. After selection of recombinants or curing of the helper plasmid, no expected products from the non-recombinant recipient plasmid can be observed (arrow point). [Figure 7F] Example of DNA assembly. Panel F shows a schematic of the analysis of Sanger sequencing results of 96 colonies after assembly of mEGFP. Sequencing products are derived from colony PCR of a pipette tip contacted with each colony. One colony contained an intermediate product (product of the first round assembly) and one contained a stitching error (large deletion). [Figure 7G] Example of DNA assembly. Panel G shows colony fluorescence from in vivo suture products after five rounds of assembly for two fluorescent genes, mPapaya and sfGFP, and four recombinase genes. The colonies may represent independent mating and recombination events. All mPapaya and sfGFP colonies are fluorescent. [Figure 7H] Example of DNA assembly. Panel H is a Sanger sequencing trace file of the in vivo ligation product of mPapaya after five rounds of assembly. Alignment with the expected sequence shows that the assembly is 100% accurate and pure. [Figure 7I]Examples of DNA assemblies. Panel I shows the results of assembly of three ~3 kb fragments for a total assembly length of ~9 kb. Recipient plasmids at various stages of assembly were digested with restriction enzymes and the stitching products were separated from the vector backbone. The digested products were then subjected to agarose gel electrophoresis to examine the size of the stitching products (lanes 1-3). The gel bands corresponding to the stitching products are marked with arrows. The linearized vector backbone without the stitching products is shown in lanes 5-6. The selectable and counterselectable markers in the swapping cassettes differed between assembly rounds, with the swapping cassettes in the first and third assembly rounds being ~1.5 kb longer than the original recipient plasmid or the swapping cassette in the second assembly round. [Figure 8A] Schematic diagrams of exemplary donor and recipient plasmids at the start of the first round of DNA stitching. Panel A shows a schematic diagram of both example plasmids, with the donor plasmid containing the first oligonucleotide (1). [Figure 8B] Schematic diagram of exemplary donor and recipient plasmids at the start of the first round of DNA stitching. Panel B shows the shape legend used to illustrate sequences corresponding to the exemplary genome, positive selectable marker, negative selectable marker, origin of transfer (oriT), gRNA expression unit (gRNA), positional barcode, homology for recombination domain (H), inducible lambda red operon (λred), inducible I-SceI endonuclease, plasmid, inducible endonuclease (Cas9), gRNA target site, I-SceI target site, conjugative Tra operon, deleted oriT (oriΔ::TcR), temperature sensitive origin (pSC101 ori), conditional origin of replication (R6K), and recipient origin of replication (ColE1). The same shape legend in Figure 8B is used in Figures 9-35. [Figure 9]Schematic diagram of an exemplary starting plasmid for use in the methods described herein: a donor plasmid containing the first oligo(1), a recipient plasmid, and a plasmid containing oriT, which mediates conjugation of the donor plasmid into a recipient cell. [Figure 10] Schematic diagram of subsequent steps in an exemplary method of DNA stitching using donor and recipient plasmids shown in Figure 9. gRNA1 guides Cas9 in the recipient cell to generate site-specific double-stranded breaks on the donor and recipient plasmids (indicated by downward arrows). [Figure 11] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-10, where the darkly shaded sequence elements are used as homology regions for lambda Red-mediated homologous recombination. [Figure 12] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-11. Homologous recombination between plasmids showing where and in what orientation sequences from the donor plasmid are inserted into the recipient plasmid with the aid of the λ-red system. [Figure 13] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-12. A fragment containing a first oligonucleotide is integrated into a recipient plasmid as shown. [Figure 14] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-13. Plasmids are selected for acquisition of the +2 positive selectable marker. [Figure 15] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9 to 14. The plasmid is counter-selected for loss of the previous counter-selectable marker. [Figure 16] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-15. Plasmids are further selected for retention of the +3 positive selectable marker on the original recipient backbone. [Figure 17]Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-16. A second donor plasmid containing a second oligonucleotide is ready to be assembled into the previous ligation product (new recipient plasmid) containing the first oligonucleotide. [Figure 18] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-17. oriT directs the conjugation of the donor plasmid. [Figure 19] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-18. Expression of a second gRNA guides Cas9 to generate a double-stranded break at the site indicated by the down arrow. [Figure 20] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-19. Highlighted regions (darker shaded regions) are homology regions for recombination. [Figure 21] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-20. A fragment containing a first oligonucleotide is integrated into a recipient plasmid as shown. [Figure 22] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-21. A second oligonucleotide is assembled adjacent to the 3' end of the first oligo in the recipient plasmid to generate a new recipient plasmid. [Diagram 23] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9 to 22. Plasmids containing the first and second oligonucleotides are selected for acquisition of the +1 positive selectable marker. [Figure 24] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9 to 23. Plasmids containing the first and second oligonucleotides are selected for loss of the -2 counter selectable marker. [Diagram 25]Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9 to 24. Plasmids containing the first and second oligonucleotides are also selected for retention of the backbone selectable marker. [Figure 26] 9-25 Schematic of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-25. Schematic of a donor plasmid containing oligonucleotide 3 and a recipient plasmid containing oligonucleotides 1 and 2 to initiate a third round of DNA stitching. [Figure 27] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9 to 26. The oriT plasmid initiates conjugation. [Figure 28] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-27. gRNA1 guides Cas9 in the recipient cell to generate site-specific double-stranded breaks on the donor and recipient plasmids (indicated by downward arrows). [Figure 29] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-28, where the highlighted sequences are used as homology regions for Lambda Red-mediated homologous recombination. [Diagram 30] 9-29, showing where and in what orientation sequences from the donor plasmid are inserted into the recipient plasmid after homologous recombination. [Diagram 31] Schematic diagram of a subsequent step in the exemplary method of DNA stitching shown in Figures 9-30. A third oligonucleotide is assembled adjacent to the 3' end of the second oligonucleotide in the recipient plasmid to generate a new recipient plasmid. [Diagram 32] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-31. As with round 1 DNA stitching, plasmids are selected for acquisition of the +2 positive selectable marker. [Diagram 33]Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9 to 32. The plasmid is counter-selected for loss of the -1 counter-selectable marker. [Diagram 34] Schematic diagram of subsequent steps in the exemplary method of DNA stitching shown in Figures 9-33. Plasmids are selected for retention of the backbone selectable marker. [Figure 35A] Panel A is a schematic diagram of steps in the methods illustrated in Figures 9-34, illustrating an embodiment in which subsequent oligonucleotides can be incorporated in a similar fashion (up to the upper limit of total sequence length) alternating between the two processes of joining, double-stranded break processing, assembly, and selectable / counterselectable / backbone selection. [Figure 35B] Panel B is a schematic diagram of an exemplary in vivo assembly in which two DNA element fragments (shown in the figure as Input DNA1, Input DNA2 before recombination and DNA1, DNA2 after recombination, and may be referred to herein as Oligo1, Oligo2, etc., or more generally as "DNA blocks") are added in a single round of conjugation. In each well, a recipient cell is conjugated first by a first donor cell containing a first donor plasmid and second by a second donor cell containing a second donor plasmid. The selectable marker introduced by the second donor plasmid and the counter-selectable marker on the recipient plasmid and optionally on the first donor plasmid allow for the selection of recombination assembly products that contain oligonucleotides introduced by both donor plasmids. [Figure 36A] Panel A is a schematic diagram of an exemplary method for in vivo DNA analysis described herein, in this example, where the indexing barcode is located on the recipient plasmid. [Figure 36B]Panel B is a schematic diagram of another exemplary method for in vivo DNA analysis described herein. In this example, indexing barcodes are located on a donor plasmid and are added to the in vivo DNA assembly products using homologous regions at the ends of the assembly products. In this example, the endonuclease sites used are the same as those used for in vivo DNA assembly. [Figure 36C] Panel C is a schematic diagram of another exemplary method for in vivo DNA analysis described herein. In this example, indexing barcodes are located on a donor plasmid and are added to the in vivo DNA assembly product. In this example, the endonuclease target site (C) and homology regions (frames flanking C) are different from those used for in vivo DNA assembly. This example allows for DNA analysis at multiple steps during assembly. [Figure 36D]Panel D is a schematic of a method involving plasmid resetting, which transfers DNA assembly from a recipient plasmid to a donor plasmid to allow for further rounds of assembly with larger DNA blocks. In part A of panel D, a reset donor plasmid in a donor cell with homology to the start of DNA assembly on the recipient plasmid is spliced ​​into a recipient cell. A site-specific endonuclease cleaves at the "D" endonuclease target site in both the reset donor plasmid and the recipient plasmid. Homologous recombination in the regions flanking the "D" endonuclease target site transfers the DNA assembly cassette from the recipient plasmid to the reset donor plasmid. The reset donor plasmid is purified from the recipient cell and transformed into a new donor cell where it can be utilized for further rounds of assembly. In part B of panel D, a schematic shows the workflow for assembly of long DNA constructs. Using four rounds of DNA stitching, small DNA blocks can be assembled into large DNA blocks. The large blocks are transferred into the donor plasmid and donor cells using the reset donor plasmid. The larger blocks can then be assembled into even larger blocks by further rounds of stitching. [Figure 37]Schematic maps of exemplary plasmids for use in in vivo DNA analysis are shown. In donor cells, the conjugative helper plasmid contains genes for plasmid transfer (Tra operon). To immobilize the helper plasmid itself, the origin of transfer (oriT) is replaced by a selectable marker (+6). The donor plasmid contains a swapping cassette (+ and -), two homology regions (H1 and H4), two sites for targeted plasmid cleavage (oval), a backbone selectable marker (+4), a conditional origin of replication (R6K) that depends on the allele in the donor genome (pir1-116), and the oriT sequence. In recipient cells, the helper plasmid contains a lac-inducible red operon (Plac-red), the E. coli RecA gene to enhance homologous recombination, a backbone selectable marker (+5), and a curable temperature-sensitive origin of replication (pSC101 oriTS). The recipient plasmid contains two endonuclease cleavage sites (two ovals), a negative selectable marker (-3), and two homology regions (H1 and H4). In addition to the two plasmids, the recipient cells also have an integrated arabinose-inducible endonuclease I-SceI (ParaBAD-I-SceI) to generate DNA breaks on the target plasmid. +: HygR or NsrR, -: SacB or PheS, +3: GmR, -3: relE, +4: KanR, +5: SpR, +6: TcR. [Figure 38] Schematic maps of exemplary plasmids for use in in vivo DNA analysis. Selectable markers: HygR, KanR, GmR, SpR. Counterselectable markers: SacB, relE. [Figure 39A] Schematic maps of exemplary donor and recipient plasmids used in DNA syntax analysis are shown. Panel A shows a schematic of the plasmid where the donor plasmid contains the first oligonucleotide. [Figure 39B]Schematic maps of exemplary donor and recipient plasmids used for DNA parsing are shown. Panel B shows a legend of shapes used to illustrate sequences corresponding to genome, positive selectable marker, negative selectable marker, origin of transfer (oriT), gRNA expression unit (gRNA), positional barcode, homology for recombination domain (H), inducible lambda red operon (λred), inducible I-SceI endonuclease, plasmid, inducible endonuclease (Cas9), gRNA target site, I-SceI target site, conjugative Tra operon, deleted oriT (oriΔ::TcR), temperature sensitive origin (pSC101 ori), conditional origin of replication (R6K), and recipient of origin of replication (ColE1). The legend in Panel B also applies to Figures 40-46. [Diagram 40] FIG. 1 shows a schematic diagram of donor and recipient plasmids for use in the example methods described herein. [Diagram 41] 41 is an image showing the second step in the example method of FIG. 40. gRNA1 guides Cas9 in the recipient cell to generate a site-specific double-stranded break (indicated by the downward arrow) on the donor and recipient plasmids. SceI is I-SceI, a homing endonuclease. [Diagram 42] 40 and 41 show images of steps in the example method, in which the H1 and H4 sequences are used as homology regions for lambda red-mediated homologous recombination. [Diagram 43] 40-42. Image showing steps in the method of the embodiment of Figures 40-42. Homologous recombination showing where and in what orientation sequences from the donor plasmid are inserted into the recipient plasmid. [Diagram 44] 44 is an image showing steps in the method of the example of Figures 40 to 43. Plasmids are selected for acquisition of a + positive selectable marker. [Diagram 45] Images of steps in the method of the example of Figures 40 to 44 are shown. The plasmid is counter-selected for loss of the previous counter-selectable marker. [Figure 46] 46 is an image showing steps in the method of the example of Figures 40 to 45. Plasmids are further selected for retention of the +3 positive selectable marker on the original recipient backbone. [Figure 47A] Panel A shows the results of an experiment to determine the ability of in vivo DNA analysis to correctly identify sequences at each location on a plate of aligned donor cells containing a DNA barcode unique to each location. Each aligned barcoded donor was mated with two or three barcoded recipient plates, and recombinant cell colonies containing donor and recipient barcodes, respectively, were selected on agar pads. Recombinant cells from the plates were pooled and the dual barcodes were sequenced on an Illumina platform. Sequencing data was used to determine the percentage of aligned barcoded donors that could be correctly indexed when mated to one, two, or three separate barcoded recipient arrays (recovery). No barcoded donors were misassigned to incorrect locations. [Figure 47B] Panel B shows the results of an experiment to index and sequence verify a pool of 100 244-base oligonucleotides ordered from IDT as an oPool. The pool of oligonucleotides was assembled into a donor plasmid, which was then transformed into donor cells. Donor cells were randomly arrayed into 384-well plates at an expected frequency of less than one cell per well. Donor cells were mated to barcoded recipient cell arrays, and the recombinant oligonucleotide-barcoded recipient plasmids were sequenced using an Oxford Nanopore sequencer. The results of the analysis of two 384-well plates are shown. The shading indicates whether the well had input DNA that was sequenced, whether the sequence was 100% identical to one of the 244-base sequences in the oPool, and whether the well was pure (i.e., only one 244-base sequence was detectable in the well). The positions marked as "100% matched pure wells" are typically used for downstream DNA assembly. [Figure 47C]Panel C is a histogram showing the distribution of errors between the consensus sequence determined by Oxford Nanopore sequencing and the predicted DNA sequence in oPool that is closest in sequence to the consensus, using the experimental data in Panel B. The majority of wells contain an oligonucleotide identical to one of the sequences in the oPool. [Figure 47D] Panel D is a histogram showing the distribution of counts of independent clones recovered for each oligonucleotide that could be indexed using the experimental data in Panel B. [Figure 48] Schematic of a DNA assembly workflow for constructing a directed combinatorial library derived from a set of input oligonucleotides. A pool of input DNA from multiple sources is assembled into donor plasmids and parsed into an ordered array. The ordered arrays are rearranged to user-defined locations on multiple donor plates. The donor plates are sequentially mated to recipient plates and the desired construct is assembled. Input oligonucleotides can be used in multiple assemblies by rearranging donor cells containing the oligonucleotides to multiple locations on the donor plates. [Figure 49] Schematic diagram of branched DNA assembly. A partial DNA assembly can be extended by multiple DNA blocks if homologous regions exist. If no homologous regions exist, a "DNA linker" must first be added to the partial DNA assembly. The DNA linker contains a homologous portion between the end of the partial DNA assembly and the beginning of the DNA block to be subsequently joined. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0036] 1. Definitions and Related Embodiments Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with general definitions for many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed., 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.

[0037] The use of singular indefinite or definite articles (e.g., "a," "an," "the," etc.) in this disclosure and in the claims that follow follows the conventional approach in patents to mean "at least one," unless it is clear from the context that the term is intended to specifically mean one and only one in that particular instance. Similarly, the term "comprising" is open-ended and does not exclude additional items, features, components, etc. All references identified herein are expressly incorporated herein by reference in their entirety, unless otherwise indicated.

[0038] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes cases where the event or circumstance occurs and cases where it does not occur.

[0039] The term "about" when used before, for example, temperature, time, amount, concentration, and similar numerical designations, including ranges, indicates approximations that may vary by (+) or (-) 10%, 5%, 1%, or any subrange or subvalue therebetween. Preferably, the term "about" means that the value may vary by ±10%.

[0040] As used herein, the term "comprising" is intended to mean that the compositions and methods include the recited elements but do not exclude others. "Consisting essentially of," when used to define compositions and methods, is intended to mean excluding other elements having some essential characteristic to the combination for the stated purpose. Thus, a composition consisting essentially of the elements defined herein does not exclude other materials or steps that do not materially affect the basic and novel characteristics of the claimed invention. "Consisting of" is intended to mean excluding more than trace amounts of other components and substantial method steps. Embodiments defined by each of these transitional terms are within the scope of this disclosure.

[0041] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleic acid sequence," and "polynucleotide" are used interchangeably and are intended to include, but are not limited to, polymeric forms of nucleotides covalently linked together, whether deoxyribonucleotides or ribonucleotides, which may be of various lengths, or analogs, derivatives, or modifications thereof. Different polynucleotides may have different three-dimensional structures and may perform various functions, known or unknown. Non-limiting examples of polynucleotides include genes, gene fragments, exons, introns, intergenic DNA (including but not limited to heterochromatic DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, sgRNA, guide RNA, tracrRNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of a sequence, isolated RNA of a sequence, PCR products, nucleic acid probes, and primers. Polynucleotides useful in the methods of the present disclosure may include naturally occurring nucleic acid sequences and variants thereof, artificial nucleic acid sequences, or combinations of such sequences.

[0042] "Nucleic acid" refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and polymers thereof in single-, double-, or multiple-stranded form, or their complements, or nucleosides (e.g., deoxyribonucleosides or ribonucleosides). In embodiments, "nucleic acid" does not include nucleosides. The terms "polynucleotide," "oligonucleotide," "oligo," and the like, refer to a linear sequence of nucleotides in the usual and customary sense. The term "nucleoside" refers to a glycosylamine containing a nucleobase and a five-carbon sugar (ribose or deoxyribose) in the usual and customary sense. Non-limiting examples of nucleosides include cytidine, uridine, adenosine, guanosine, thymidine, and inosine. The term "nucleotide" refers to a single unit, or monomer, of a polynucleotide in the usual and customary sense. A nucleotide may be a ribonucleotide, a deoxyribonucleotide, or a modified version thereof. Examples of polynucleotides contemplated herein include single- and double-stranded DNA, single- and double-stranded RNA, and hybrid molecules having a mixture of single- and double-stranded DNA and RNA. Examples of nucleic acids, e.g., polynucleotides, contemplated herein include any type of RNA, e.g., mRNA, siRNA, miRNA, and guide RNA, as well as any type of DNA, genomic DNA, plasmid DNA, and small circular DNA, and any fragments thereof. The term "duplex" in the context of polynucleotides means double-stranded in the usual and customary sense. Nucleic acids may be linear or branched. For example, a nucleic acid may be a straight chain of nucleotides, or the nucleic acid may be branched, e.g., such that the nucleic acid comprises one or more arms or branches of nucleotides. Optionally, branched nucleic acids are repeatedly branched to form dendrimers or other higher order structures.

[0043] For example, a nucleic acid, including a nucleic acid having a phosphothioate backbone, may contain one or more reactive moieties. As used herein, the term "reactive moiety" includes any group that can react with another molecule, such as a nucleic acid or a polypeptide, through covalent, non-covalent, or other interactions. As an example, a nucleic acid may contain an amino acid reactive moiety that reacts with an amino acid on a protein or polypeptide through covalent, non-covalent, or other interactions.

[0044] This term also encompasses nucleic acids that contain known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, and have similar binding properties as the reference nucleic acid, and are metabolized in a similar manner as the reference nucleotide. Examples of such analogs include, but are not limited to, phosphodiester derivatives, including, for example, phosphoramidates, phosphorodiamidates, phosphorothioates (also known as phosphothioates with double bond sulfur replacing the oxygen of the phosphate), phosphorodithioates, phosphonocarboxylic acids, phosphonocarboxylates, phosphonoacetic acids, phosphonoformic acids, methylphosphonates, boron phosphonates, or O-methyl phosphoramidite linkages (see Eckstein, OLIGONUCLEOTIDES AND ANALOGUES: A PRACTICAL APPROACH, Oxford University Press), and modifications of nucleotide bases, such as 5-methylcytidine or pseudouridine, and peptide nucleic acid backbones and linkages. Other analog nucleic acids include those with cationic backbones, non-ionic backbones, modified sugars, and non-ribose backbones (e.g., phosphorodiamidate morpholino oligos or locked nucleic acids (LNAs) known in the art), including those described in U.S. Patent Nos. 5,235,033 and 5,034,506, and ASC Symposium Series 580, CARBOHYDRATE MODIFICATIONS IN ANTISENSE RESEARCH, Sanghui & Cook, Chapters 6 and 7. Nucleic acids containing one or more carbocyclic sugars are also included within one definition of nucleic acid. Modifications of the ribose-phosphate backbone may be made for a variety of reasons, such as to increase the stability and half-life of such molecules in physiological environments or as probes on biochips. Mixtures of naturally occurring nucleic acids and analogs can be made, or alternatively, mixtures of different nucleic acid analogs and mixtures of naturally occurring nucleic acids and analogs may be made. In embodiments, the linkages between nucleotides in the DNA are phosphodiester, phosphodiester derivatives, or a combination of both.

[0045] "Barcode" refers to one or more nucleotide sequences used to identify a cell or cells with which the barcode is associated. A barcode can be 3-1000 or more nucleotides in length, preferably 3-250 nucleotides in length, more preferably 4-40 nucleotides in length, including any length within these ranges, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length. A barcode is "unique" if it is present (statistically) in about one cell in a population of cells. The barcode-containing cells can then be expanded to produce a plurality of clonal cells, whereby each cell in the plurality of cells contains the same barcode. For example, a "plurality of barcoded cells, where each barcoded cell contains a single unique barcode" can refer to a population of cells that (statistically) contains a single cell that contains a given barcode or unique combination of barcodes. Alternatively, it can refer to a population of cells that includes multiple clonal cell populations, where each cell in each clonal population contains the same barcode, but cells in different clonal populations contain different barcodes.

[0046] As used herein, the term "complement" refers to a nucleotide (e.g., RNA or DNA) or sequence of nucleotides that can base pair with a complementary nucleotide or sequence of nucleotides. As described herein and generally known in the art, the complementary (matching) nucleotide of adenosine is thymidine, and the complementary (matching) nucleotide of guanosine is cytosine. That is, a complement can include a sequence of nucleotides that base pair with the corresponding complementary nucleotides of a second nucleic acid sequence. The nucleotides of the complement can partially or completely match with the nucleotides of the second nucleic acid sequence. If the nucleotides of the complement completely match with each nucleotide of the second nucleic acid sequence, the complement will base pair with each nucleotide of the second nucleic acid sequence. If the nucleotides of the complement partially match with the nucleotides of the second nucleic acid sequence, only some of the nucleotides of the complement will base pair with the nucleotides of the second nucleic acid sequence.

[0047] As described herein, sequence complementarity can be partial, where only some of the nucleic acids are matched by base pairing, or complete, where all of the nucleic acids are matched by base pairing, i.e., two sequences that are complementary to each other can have a specified percentage of nucleotides that are identical (i.e., about 60% identity over the specified region, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity).

[0048] As used herein, the term "gene" is used according to its plain and ordinary meaning to refer to a segment of DNA involved in the production of a protein. It includes regions preceding and following the coding region (leader and trailer), as well as intervening sequences (introns) between individual coding segments (exons). Leaders, trailers, and introns contain regulatory elements required during transcription and translation of a gene. Furthermore, a "protein gene product" is a protein expressed from a particular gene.

[0049] The term "expression vector" refers to a nucleic acid molecule that encodes a gene and / or regulatory elements necessary for expression of a gene. Expression of a gene from a vector, which may be in the form of a plasmid, can occur in cis or trans. If a gene is expressed in cis, the gene and the regulatory elements are encoded by the same plasmid. Expression in trans refers to when the gene and the regulatory elements are encoded by separate plasmids.

[0050] As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. A vector may be in the form of a "plasmid", which in this context refers to a linear or circular double stranded DNA loop into which additional DNA segments may be ligated. Another type of vector is a viral vector, in which additional DNA segments may be ligated into the viral genome. Certain vectors (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in a host cell into which they are introduced. Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell, and thereby are replicated along with the host genome. In addition, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors". In general, expression vectors useful in recombinant DNA techniques are often in the form of plasmids. As used herein, "plasmid" and "vector" may be used interchangeably, with the plasmid being the most commonly used form of vector. However, the present invention is intended to include such other forms of expression vectors that serve equivalent functions, such as viral vectors (e.g., replication-defective retroviruses, adenoviruses, and adeno-associated viruses). Furthermore, some viral vectors can target specific cell types, either specifically or non-specifically. Replication-incompetent or replication-defective viral vectors refer to viral vectors that can infect their target cells and deliver their viral payload, but cannot continue down the typical cytolytic pathway that leads to cell lysis and death.

[0051] According to the methods described herein, the oligonucleotide, plasmid, or vector may comprise at least one selectable marker. The selectable marker for use in the methods described herein may be any suitable selectable marker. In embodiments, the selectable marker includes, but is not limited to, HygR, NsrR, ZeoR, TetA, CmR, SpR, GmR, mFabI, TmR, neoR, or kanR. In embodiments, the selectable marker is HygR. In embodiments, the selectable marker is NsrR. In embodiments, the selectable marker is ZeoR. In embodiments, the selectable marker is TetA. In embodiments, the selectable marker is CmR. In embodiments, the selectable marker is SpR. In embodiments, the selectable marker is GmR. In embodiments, the selectable marker is mFabI. In embodiments, the selectable marker is TmR. In embodiments, the selectable marker is neoR. In embodiments, the selectable marker is kanR.

[0052] According to the methods described herein, the oligonucleotide, plasmid, or vector may include at least one counterselectable marker, such as a marker that selects for integration of a second or subsequent oligonucleotide into a recombined recipient oligonucleotide in the methods of assembling a DNA element described herein. The counterselectable marker for use in the methods described herein may be any suitable counterselectable marker. In embodiments, the counterselectable marker includes, but is not limited to, PheS, SacB rpsL, tolC, galK, ccdB, tetA, thyA, lacY, gata-1, URA3, relE, mqsR, chpB, vhaV, or tse2. In embodiments, the counterselectable marker is PheS. In embodiments, the counterselectable marker is SacB. In embodiments, the counterselectable marker is rpsL. In embodiments, the counterselectable marker is tolC. In embodiments, the counterselectable marker is galK. In embodiments, the counterselectable marker is ccdB. In an embodiment, the counter selectable marker is ccdB. In an embodiment, the counter selectable marker is tetA. In an embodiment, the counter selectable marker is thyA. In an embodiment, the counter selectable marker is lacY. In an embodiment, the counter selectable marker is gata-1. In an embodiment, the counter selectable marker is URA3. In an embodiment, the counter selectable marker is relE. In an embodiment, the counter selectable marker is mqsR. In an embodiment, the counter selectable marker is chpB. In an embodiment, the counter selectable marker is vhaV. In an embodiment, the counter selectable marker is tse2.

[0053] The terms "transfection", "transduction", "transfect" or "transduce" can be used interchangeably and are defined as the process of introducing a nucleic acid molecule and / or a protein into a cell. The nucleic acid may be introduced into the cell using non-viral or viral methods. The nucleic acid molecule may be a sequence encoding a complete protein or a functional portion thereof. Typically, a nucleic acid vector containing elements necessary for the expression of the protein (e.g., promoter, transcription start site, etc.). Non-viral methods of transfection include any suitable method that does not use viral DNA or viral particles as a delivery system to introduce the nucleic acid molecule into the cell. Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, heat shock transfection, magnetifection, and electroporation. For viral-based methods, any useful viral vector can be used in the methods described herein. Examples of viral vectors include, but are not limited to, retroviral, adenoviral, lentiviral, and adeno-associated viral vectors. In some embodiments, the nucleic acid molecule is introduced into the cell using a retroviral vector, according to standard procedures well known in the art. The term "transfection" or "transduction" also means introducing a protein into a cell from the external environment. Typically, protein transduction or transfection relies on the attachment of a peptide or protein that can cross the cell membrane to the protein of interest. See, for example, Ford et al. (2001) Gene Therapy 8:1-4 and Prochiantz (2007) Nat. Methods 4:119-20.

[0054] As used herein, the term "promoter" refers to a region of DNA that initiates transcription of a particular gene. Promoters are typically located near the transcription start site of a gene, upstream of the gene, and on the same strand of DNA (i.e., 5' on the sense strand). Promoters may be, for example, about 100 to about 1000 base pairs in length.

[0055] The "position" of a nucleotide base is designated by a number that sequentially identifies each amino acid (or nucleotide base) in a reference sequence based on its position relative to the 5' end. Due to deletions, insertions, truncations, fusions, etc., which must be considered when determining optimal alignment, the number of an amino acid residue in a test sequence, determined by simple counting from the 5' end, will generally not necessarily be the same as the number of the corresponding position in the reference sequence. For example, if a variant has a deletion relative to the aligned reference sequence, there will be no nucleotide base in the variant that corresponds to the position in the reference sequence at the site of the deletion. If there is an insertion in the aligned reference sequence, the insertion will not correspond to a numbered nucleotide position in the reference sequence. In the case of a truncation or fusion, there may be a stretch of nucleotides in the reference sequence or in the aligned sequence that does not correspond to any nucleotide in the corresponding sequence.

[0056] The terms "numbered with respect to" or "corresponding to," when used in reference to the numbering of a given polynucleotide sequence, refer to the numbering of residues in a specified reference sequence when comparing the given polynucleotide sequence to a reference sequence.

[0057] As used herein, the term "virus" or "viral particle" is used according to its plain and ordinary meaning within the context of viral transduction. Transduction with viral vectors can be used to insert genes into or modify genes in mammalian cells.

[0058] As used herein, the terms "genetic modification", "genetic modification", "gene editing", "genetic editing", "genome editing", "genome manipulation", and the like refer to a type of genetic manipulation in which DNA is inserted, deleted, modified, or replaced at one or more specific locations in the genome of a cell. One key step in gene editing is to create a double-stranded break at a specific point in a gene or genome. Examples of gene editing tools, such as nucleases, that accomplish this step include, but are not limited to, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, and clustered regularly interspaced short palindromic repeats system (CRISPR / Cas).

[0059] As used herein, "DNA element" refers to any DNA sequence that can be transferred between cells, for example, between a donor cell and a recipient cell. That is, DNA elements include, but are not limited to, genes, promoters, enhancers, terminators, introns, intergenic regions, barcodes, or gRNAs. The DNA element can be a gene, a promoter, enhancer, terminator, introns, intergenic regions, barcodes, or a fragment of a gRNA. The DNA element can be a gene, a promoter, enhancer, terminator, introns, intergenic regions, barcodes, gRNA, and a combination of a gene, a promoter, enhancer, terminator, introns, intergenic regions, barcodes, and a fragment of a gRNA. In an embodiment, the DNA element is in a donor plasmid. In another embodiment, the DNA element is transferred to or is in a recipient oligonucleotide. In another embodiment, the DNA element is transferred from a recipient oligonucleotide to a reset donor plasmid.

[0060] As used herein, the term "gene editing reagent" refers to components necessary for gene editing tools and may include enzymes, riboproteins, solutions, cofactors, etc. For example, gene editing reagents include one or more components necessary for gene editing with zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, and clustered regularly interspaced short palindromic repeats system (CRISPR / Cas).

[0061] As used herein, the term "endonuclease" refers to an enzyme or component of an endonuclease system (e.g., any component of CRISPR, including gRNA) that has endonuclease cleavage catalytic activity for cleavage of a polynucleotide. For example, an endonuclease or its components can cleave the phosphodiester bond of an oligonucleotide or polynucleotide. The endonuclease cleaves at a phosphodiester bond that is at least 4 bp in length within or near its recognition site sequence. Types of endonucleases include, but are not limited to, restriction enzymes, AP endonucleases, T7 endonuclease, T4 endonuclease, Bal 31 endonuclease, endonuclease I, micrococcal nuclease, endonuclease II, Neurospora endonuclease, S1 endonuclease, P1-nuclease, Mung bean nuclease I, DNase I, RNA-guided DNA endonucleases (e.g., CRISPR including any CRISPR component, e.g., Cas proteins, gRNA, etc.), homozygous switching endonucleases, TALENs, zinc finger nucleases, and EndoR.

[0062] "Cleavage" refers to the destruction of the covalent backbone of a DNA molecule. Cleavage can be initiated by various methods, including but not limited to enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand breaks are possible, and double-strand breaks can result from two different single-strand break events. DNA cleavage can result in the production of either blunt ends or sticky ends. In some embodiments, a complex comprising a guide RNA and a site-specific modification enzyme is used for targeted double-strand DNA cleavage.

[0063] As used herein, the term "CRISPR" or "clustered regularly interspaced short palindromic repeats" is used according to its plain ordinary meaning to refer to a genetic element that bacteria use as a kind of adaptive immunity to protect against viruses. CRISPR comprises short sequences that originate from viral genomes and are integrated into bacterial genomes. Cas (CRISPR associated protein) processes these sequences and cuts the matching viral DNA sequences. That is, the CRISPR sequence acts as a guide for Cas, which recognizes and cuts DNA that is at least partially complementary to the CRISPR sequence. By introducing a plasmid containing the Cas gene and a specifically constructed CRISPR into a eukaryotic cell, the eukaryotic cell genome can be cut at any desired location.

[0064] As used herein, the term "Cas9" or "CRISPR associated protein 9" is used according to its plain ordinary meaning to refer to an enzyme that uses a CRISPR sequence as a guide to recognize and cleave a specific strand of DNA that is at least partially complementary to the CRISPR sequence. The Cas9 enzyme, together with the CRISPR sequence, forms the basis of a technology known as CRISPR-Cas9 that can be used to edit genes in living organisms. This editing process has a wide range of applications, including basic biology research, development of biotechnology products, and treatment of disease.

[0065] As referred to herein, "CRISPR associated protein 9", "Cas9", "Csn1", or "Cas9 protein" includes any recombinant or naturally occurring form of Cas9 endonuclease or variant or homologue thereof that maintains the activity of the Cas9 endonuclease enzyme (e.g., within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% activity compared to Cas9). In embodiments, the variant or homologue has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a 50, 100, 150, or 200 contiguous amino acid portion) compared to a naturally occurring Cas9 protein. In embodiments, the Cas9 protein is substantially identical to a protein identified by UniProt reference number Q99ZW2, or a variant or homologue having substantial identity thereto. In embodiments, the Cas9 protein has at least 75% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 80% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 85% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 90% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 95% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2.

[0066] As referred to herein, "CRISPR-associated endonuclease Cas12a," "Cas12a," "Cas12," or "Cas12 protein" includes any recombinant or naturally occurring form of Cas12 endonuclease or variants or homologs thereof that maintain the activity of the Cas12 endonuclease enzyme (e.g., within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the activity compared to Cas12). In embodiments, the variant or homolog has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a 50, 100, 150, or 200 contiguous amino acid portion) compared to the naturally occurring Cas12 protein. In embodiments, the Cas12 protein is substantially identical to the protein identified by UniProt reference number A0Q7Q2, or a variant or homolog having substantial identity thereto.

[0067] As referred to herein, "CRISPR-associated endoribonuclease Cas13a," "Cas13a," "Cas13," or "Cas13 protein" includes any recombinant or naturally occurring form of Cas13 endoribonuclease or a variant or homolog thereof that maintains the activity of the Cas13 endoribonuclease enzyme (e.g., within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the activity compared to Cas13). In embodiments, the variant or homolog has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a 50, 100, 150, or 200 contiguous amino acid portion) compared to the naturally occurring Cas13 protein. In embodiments, the Cas13 protein is substantially identical to the protein identified by UniProt reference number P0DPB8, or a variant or homolog having substantial identity thereto.

[0068] As used herein, "TALEN" or "transcription activator-like effector nuclease" refers to a restriction enzyme generated by combining a DNA-binding domain (e.g., a TAL effector DNA-binding domain) with a nuclease (e.g., FokI). TALENs typically contain naturally occurring DNA-binding domains that contain multiple modules, referred to as TALs or TALEs. That is, TALs contain variable di-residues that confer DNA-binding specificity.

[0069] As provided herein, "guide RNA" or "gRNA" refers to an RNA sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. For example, the gRNA can direct Cas to a target polynucleotide. In an embodiment, the gRNA comprises a crRNA and a tracrRNA. For example, the gRNA may comprise a crRNA and a tracrRNA hybridized by base pairing. That is, in an embodiment, the two RNAs are encoded as two RNA molecules by the crRNA and the tracrRNA separately, which can then form an RNA / RNA complex by complementary base pairing between the crRNA and the tracrRNA. In an embodiment, the degree of complementarity between a guide RNA sequence and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more when optimally aligned using a suitable alignment algorithm. In embodiments, the degree of complementarity between a guide RNA sequence and its corresponding target sequence is at least about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% when optimally aligned using a suitable alignment algorithm.

[0070] Non-limiting examples of CRISPR enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm 2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. In an embodiment, the CRISPR enzyme is a Cas9 enzyme. In an embodiment, the Cas9 enzyme is a Cas9 from S. pneumoniae, S. pyogenes, or S. thermophilus, or mutants derived therefrom in these organisms. In an embodiment, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In embodiments, the CRISPR enzyme directs one or two strand cleavage at the location of the target sequence, hi embodiments, the CRISPR enzyme lacks DNA strand cleavage activity.

[0071] As used herein, a "zinc finger" is a polypeptide structural motif that folds around a bound zinc cation. In embodiments, a zinc finger polypeptide has the form X-Cys-X 2-4 -Cys-X 12 -His-X 3-5 -His-X4, where X is any amino acid, e.g., X 2-4 represents an oligopeptide having a length of 2 to 4 amino acids.) That is, the term "zinc finger nuclease" used herein means a nuclease that includes a zinc finger motif and a domain capable of inducing cleavage in a target DNA.

[0072] The term "homologous recombination" refers to a type of genetic recombination in which information is exchanged between two similar or identical nucleic acid sequences, which may be referred to herein as "regions of homology." In some embodiments of the methods described herein, the region of homology may include, for example, two areas of homology that may optionally be flanked by non-homologous regions. In embodiments of the methods described herein, the E. coli RecA gene may be used to enhance homologous recombination. "RecA" refers to the bacterial homolog of a family of universal 38 kD homologous DNA repair proteins that mediate ATP-dependent homologous recombination in bacteria. In embodiments, the donor or recipient cells of the methods described herein contain one or more homologous DNA repair genes, such as oligonucleotides encoding RecA. In embodiments, expression of the homologous DNA repair genes is inducible. In embodiments, the homologous DNA repair gene is RecA. In embodiments, the homologous DNA repair genes are the recombinant engineered genes Redα, Redβ, and Redγ. Non-limiting examples of methods of homologous recombination and gene editing using various nuclease systems can be found, for example, in U.S. Patent No. 8,945,839, International PCT Publication No. 2013 / 163394, and U.S. Patent Publication Nos. 2016 / 0060657, 2012 / 0192298A1, and 2007 / 0042462. These and other known methods for homologous recombination can be used in combination with the methods described herein.

[0073] As used herein, the term "transfection" is used according to its plain ordinary meaning to mean the process of deliberately introducing naked or purified nucleic acid into eukaryotic cells. In illustration, "transfection" may refer to other methods and cell types, although other terms are often preferred. For example, the term "transformation" is typically used to describe non-viral DNA transfer in non-animal eukaryotic cells, including bacteria and plant cells. In animal cells, transfection is the preferred term. For example, the term "transduction" is often used to describe viral-mediated gene transfer into eukaryotic cells.

[0074] The terms "bacterial conjugation" and "bacterial mating" are interchangeable and refer to the mode of genetic exchange between bacteria. Typically, bacterial conjugation involves only a portion of the genome of one of the cells (the donor) and the complete genome of its partner (the recipient cell). That is, gene transfer in bacterial conjugation is typically partial. In an embodiment, bacterial conjugation is the transfer of non-genomic bacterial DNA from a donor cell to a recipient cell. In an illustrative example, bacterial conjugation occurs through a plasmid. In an illustrative example, bacterial conjugation occurs through exogenous DNA in the bacteria. In an embodiment, a donor cell and a recipient cell come into contact for bacterial conjugation to occur. In an embodiment, the donor cell and the recipient cell include a connecting bridge (e.g., a pilus) for bacterial conjugation to occur.

[0075] In an embodiment, the recipient cell or donor cell comprises an oligonucleotide that enables plasmid conjugation. In an embodiment, the oligonucleotide that enables plasmid conjugation is in the genome of the donor cell. In an embodiment, the oligonucleotide that enables plasmid conjugation is in a helper plasmid. In an embodiment, the oligonucleotide that enables plasmid conjugation is the Tra operon. In an embodiment, the oligonucleotides enabling conjugation of the plasmids are: IncF1 Tra (traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traR, traS, traT), IncP Tra operon: (trbA, trbB, trbC, trbD, trbE, trbF, trbG, trbH, trbI, trbJ, trbK, trbL, traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO), IncI1 The tra operon: (traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traS, traT, traU, traV, traW, traY), pTiC58 tra genes: (traA, traF, traB, traC, traG, traD, traR, traI), and pIJ101: clt, korB. In an embodiment, the oligonucleotide enabling conjugation of the plasmid is IncF1 Tra (traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traR, traS, traT).In an embodiment, the oligonucleotides enabling conjugation of the plasmid are the IncP Tra operon: (trbA, trbB, trbC, trbD, trbE, trbF, trbG, trbH, trbI, trbJ, trbK, trbL, traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO). In an embodiment, the oligonucleotides enabling conjugation of the plasmid are the IncI1 tra operon: (traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traS, traT, traU, traV, traW, traY). In an embodiment, the oligonucleotides enabling conjugation of the plasmid are the pTiC58 tra genes: (traA, traF, traB, traC, traG, traD, traR, traI). In an embodiment, the oligonucleotides enabling conjugation of the plasmid are pIJ101:clt, korB.

[0076] As used herein, "donor cell" refers to a cell (e.g., a bacterial cell) that transfers genetic material to another cell (e.g., a bacterial cell, a plant cell, etc.). A cell that receives the transferred genetic material is referred to herein as a "recipient cell."

[0077] The term "donor plasmid" as used herein refers to DNA from a donor cell that contains an oligonucleotide sequence (e.g., donor DNA, an oligonucleotide containing a DNA element) that is transferred from a donor cell (e.g., a bacterial cell) to a recipient cell (e.g., a bacterial cell, a yeast cell, a plant cell, etc.). Typically, the donor plasmid is a circular double-stranded DNA that is separated from the genomic DNA. That is, the term "recipient plasmid" refers to DNA from a recipient cell that receives the donor DNA. In an embodiment, the DNA from the donor plasmid is received by DNA other than the DNA from the recipient plasmid. That is, in an embodiment, the donor DNA may be integrated into the genomic DNA.

[0078] In an embodiment, the donor plasmid comprises an origin of transfer. In an embodiment, the origin of transfer is from a mobile element. In an embodiment, the mobile element is a plasmid. In an embodiment, the plasmid is an IncFI plasmid, an IncPα plasmid, an IncI1 plasmid, pTiC58 from Agrobacterium tumefaciens, a pAD1 plasmid, an Inc18 plasmid, or an IncH plasmid. In an embodiment, the plasmid is an IncFI plasmid. In an embodiment, the plasmid is an IncPα plasmid. In an embodiment, the plasmid is an IncI1 plasmid. In an embodiment, the plasmid is pTiC58 from Agrobacterium tumefaciens. In an embodiment, the plasmid is a pAD1 plasmid. In an embodiment, the plasmid is an Inc18 plasmid. In an embodiment, the plasmid is an IncH plasmid. Plasmids may be prepared as described in Ippen-Ihler, KA and Minkley, EG, Jr., (1986) The conjugation system of F, the fertility factor of Escherichia coli. Ann. Rev. Genet. 20:593-624; Guiney, DG and Lanka, E., (1989), Conjugative transfer of IncP plasmids, in: Promiscuous Plasmids of Gram-negative Bacteria (ed. C. M. Thomas), Academic Press, London, pp. 27-56; Catherine E. D. Rees, David E. Bradley, Brian M. Wilkins, (1987) Organization and regulation of the conjugation genes of IncI1 plasmid ColIb-P9. Plasmid. 18:223-236; von Bodman, SB, McCutchan, JE, Farrand, SK.(1989)Characterization of conjugal transfer functions of Agrobacterium tumefaciens Ti plasmid pTiC58.J.Bacteriol.171(10):5281~5289;Clewell DB, Weaver KE.(1989)Sex pheromones and plasmid transfer in Enterococcus faecalis.Plasmid.21(3):pp.175~84;Kohler V,Vaishampayan A,Grohmann E.(2018)Broad-host-range Inc18 plasmids:Occurrence,spread and transfer mechanisms.Plasmid.99:pp.11~21;Andreas Schluter,Patrice Nordmann,Remy A.Bonnin,Yves Millemann,Felix G. Eikmeyer, Daniel Wibberg, Alfred, Puhler, Laurent Poirel.(2014)IncH-Type Plasmid Harboring bla. CTX-M-15 ,bla DHA-1 , and qnrB4 Genes Recovered from Animal Isolates. Antimicrobial Agents and Chemotherapy 58(7):3768-3773. The entire contents of these references are incorporated herein by reference in their entirety for all purposes.

[0079] In an embodiment, the origin of transfer is derived from a mobile element. In an embodiment, the mobile element is derived from a conjugative transposon. In an embodiment, the conjugative transposon is Tn916 from Enterococcus faecalis or CTnDOT from Bacteroides. In an embodiment, the conjugative transposon is Tn916 from Enterococcus faecalis. In an embodiment, the conjugative transposon is CTnDOT from Bacteroides. In an embodiment, the mobile element is derived from an integrated conjugative element. In an embodiment, the mobile element is SXT from Vibrio cholerae or R391 from Providencia rettgeri. In an embodiment, the mobile element is SXT from Vibrio cholerae. In an embodiment, the mobile element is R391 from Providencia rettgeri.The element is described in the following references: Rice LB (1998). Tn916 family conjugative transposons and dissemination of antimicrobial resistance determinants. Antimicrobial agents and chemotherapy, 42(8), pp. 1871-1877; Cheng Q, Paszkiet BJ, Shoemaker NB, Gardner JF, Salyers AA. (2000) Integration and excision of a Bacteroides conjugative transposon, CTnDOT. J Bacteriol. 182(14): pp. 4035-43; Bianca Hochhut and Matthew K. Waldor. (1999) Site-specific integration of the conjugal Vibrio cholerae SXT element into prfC. Mol. Microbiology. 32(1): pp. 99-110; Boltner D, MacMahon C, Pembroke JT, Strike P, Osborn AM.R391: a conjugative integrating mosaic comprised of phage, plasmid, and transposon elements. J Bacteriol. 2002;184(18):5158-5169, each of which is incorporated herein by reference in its entirety.

[0080] In an embodiment, the donor plasmid comprises a conditional origin of replication. In an embodiment, the conditional replicons are R6K-pir, RSF1010 oriV-RepA / B / C, ColE2 P9-RepA, RP4 oriV-trfA, pPS10 oriV-RepA, pSC101 ori-RepC. TS., RK2 oriV, bacteriophage P1 ori, plasmid pSC101 origin of replication, bacteriophage lambda ori, pBR322 plasmid, pSU739 plasmid, or pSU300 plasmid. In an embodiment, the conditional replicon is R6K-pir. In an embodiment, the conditional replicon is RSF1010 oriV-RepA / B / C. In an embodiment, the conditional replicon is ColE2 P9-RepA. In an embodiment, the conditional replicon is RP4 oriV-trfA. In an embodiment, the conditional replicon is pPS10 oriV-RepA. In an embodiment, the conditional replicon is pSC101 ori-RepC. TSIn an embodiment, the conditional replicon is RK2 oriV. In an embodiment, the conditional replicon is bacteriophage P1 ori. In an embodiment, the conditional replicon is the plasmid pSC101 origin of replication. In an embodiment, the conditional replicon is bacteriophage lambda ori. In an embodiment, the conditional replicon is a pBR322 plasmid. In an embodiment, the conditional replicon is a pSU739 plasmid. In an embodiment, the conditional replicon is a pSU300 plasmid. Plasmids are referenced in: Metcalf WW, Jiang W, Daniels LL, Kim SK, Haldimann A, Wanner BL. (1996) Conditionally replicative and conjugative plasmids carrying lacZ alpha for cloning, mutagenesis, and allele replacement in bacteria. Plasmid. 35(1): pp. 1-13; Scherzinger E, Bagdasarian MM, Scholz P, Lurz R, Ruckert B, Bagdasarian M. (1984) Replication of the broad host range plasmid RSF1010: requirement for three plasmid-encoded proteins. Proc Natl Acad Sci USA. 81(3): pp. 654-8; ColE2-P9: Yagura M, Nishio SY, Kurozumi H, Wang CF, Itoh T. (2006) Anatomy of the replication origin of plasmid ColE2-P9. J Bacteriol.188(3):999~1010; Ayres EK, Thomson VJ, Merino G, Balderes D, Figurski DH. Precise deletions in large bacterial genomes by vulector-mediated excision(VEX).(1993) The trfA gene of promiscuous plasmid RK2 is essential for replication in several gram - negative hosts. J Mol Biol. 5;230(1):174 - 85; Maestro B, Sanz JM, Diaz - Orejas R, Fernandez - Tresguerres E. (2003) Modulation of pPS10 host range by plasmid - encoded RepA initiator protein. J Bacteriol. 185(4):1367 - 75; Hashimoto - Gotoh, T., & Sekiguchi, M. (1977). Mutations of temperature sensitivity in R plasmid pSC101. Journal of bacteriology, 131(2), 405 - 412; Ayres EK, Thomson VJ, Merino G, Balderes D, Figurski DH. Precise deletions in large bacterial genomes by vulector - mediated excision (VEX). (1993) The trfA gene of promiscuous plasmid RK2 is essential for replication in several gram - negative hosts. J Mol Biol. 5;230(1):174 - 85 Stenzel TT, Patel P, Bastia D. (1987) The integration host factor of Escherichia coli binds to bent DNA at the origin of replication of the plasmid pSC101. Cell. 5;49(5):709 - 17; Sugiura S, Ohkubo S, Yamaguchi K.(1993)Minimal essential origin of plasmid pSC101 replication:requirement of a region downstream of iterons.J Bacteriol.175(18):5993~6001;Pal SK,Mason RJ,Chattoraj DK.(1986)P1 plasmid replication.Role of initiator titration in copy number control.J Mol Biol.20;192(2):275~85;LeBowitz JH, McMacken R.(1984)The bacteriophage lambda O and P protein initiators promote the replication of single-stranded DNA.Nucleic Acids Res.12(7):3069~3088;Grindley ND,Kelley WS.(1976)Effects of different alleles of the E.coli K12 pol A gene on the replication of non-transferring plasmids.Mol Gen Genet. 2;143(3):311-8; Francia, MV, & Garcia Lobo, JM (1996). Gene integration in the Escherichia coli chromosome mediated by Tn21 integrase (Int21). Journal of bacteriology, 178(3), 894-898; Mendiola MV, de la Cruz F. (1989) Specificity of insertion of IS91, an insertion sequence present in alpha-haemolysin plasmids of Escherichia coli. Mol Microbiol. 3(7):979-84. The references are incorporated herein in their entirety.

[0081] In an embodiment, the conditional origin of replication is dependent on the presence of an oligonucleotide. In an embodiment, the oligonucleotide is pir1, pir1-116, repA / repB / repC (RSF1010 replicon), repA (ColE2-P9 replicon), trfA (RP4 replicon), RepA (pSP10 replicon), RepC TS (pSC101 replicon), or combinations thereof. In embodiments, the oligonucleotide encodes pir1. In embodiments, the oligonucleotide encodes pir1-116. In embodiments, the oligonucleotide encodes repA / repB / repC (RSF1010 replicon). In embodiments, the oligonucleotide encodes repA (ColE2-P9 replicon). In embodiments, the oligonucleotide encodes trfA (RP4 replicon). In embodiments, the oligonucleotide encodes RepA (pSP10 replicon). In embodiments, the oligonucleotide encodes RepC TS (pSC101 replicon).

[0082] In an embodiment, the conditional origin of replication is dependent on the conditions of cell growth. In an embodiment, the condition is temperature.

[0083] For the methods provided herein, in embodiments, the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid from 20 or 30 kilobases in length.

[0084] For the methods provided herein, in embodiments, the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid greater than 30 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid between about 30 kilobases and about 500 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 30 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 50 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 70 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 90 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 100 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 120 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 140 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 160 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 180 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid about 200 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 220 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 240 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 260 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 280 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 300 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 400 kilobases in length. In embodiments, the replicon is capable of replicating a plasmid that is about 500 kilobases in length. The length may be any value or subrange within the indicated ranges, including the endpoints.

[0085] In an embodiment, the replicon is derived from a P1-derived artificial chromosome or a bacterial artificial chromosome. In an embodiment, the replicon is derived from a P1-derived artificial chromosome. In an embodiment, the replicon is derived from a bacterial artificial chromosome. In an embodiment, the donor plasmid or the recipient oligonucleotide comprises an inducible high-copy origin of replication. In an embodiment, the donor plasmid comprises an inducible high-copy origin of replication. In an embodiment, the recipient oligonucleotide comprises an inducible high-copy origin of replication.

[0086] In an embodiment, the donor plasmid or the recipient oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome. In an embodiment, the donor plasmid is a yeast artificial chromosome (YAC). In an embodiment, the donor plasmid is a mammalian artificial chromosome (MAC). In an embodiment, the donor plasmid is a human artificial chromosome (HAC). In an embodiment, the donor plasmid is a plant artificial chromosome. In an embodiment, the recipient oligonucleotide is a yeast artificial chromosome (YAC). In an embodiment, the recipient oligonucleotide is a mammalian artificial chromosome (MAC). In an embodiment, the recipient oligonucleotide is a human artificial chromosome (HAC). In an embodiment, the recipient oligonucleotide is a plant artificial chromosome.

[0087] In embodiments, the donor plasmid or the recipient oligonucleotide comprises a conjugation competent vector, which may be a viral vector. In embodiments, the donor plasmid is a viral vector. In embodiments, the recipient oligonucleotide is a viral vector. In embodiments, the viral vector is a retrovirus. In embodiments, the viral vector is a lentivirus. In embodiments, the viral vector is an adenovirus. In embodiments, the viral vector is an adeno-associated virus. In embodiments, the viral vector is a tobacco mosaic virus. In embodiments, the viral vector is a baculovirus. In embodiments, the viral vector is a herpes simplex virus. In embodiments, the viral vector is a poxvirus. In embodiments, the viral vector is a gamma retrovirus. In embodiments, the viral vector is a Sendai virus.

[0088] As used herein, the term "control" or "control experiment" is used according to its plain and ordinary meaning to mean an experiment in which the experimental subject or agent is treated as in a parallel experiment except that the experimental procedure, agent, or variable is omitted. In some instances, a control is used as a standard of comparison in evaluating the effect of an experiment.

[0089] A "control" sample or value refers to a sample that serves as a reference, usually a known reference, for comparison with a test sample. For example, a test sample can be taken from a test condition, e.g., in the presence of a test compound, and compared to a sample taken from a known condition, e.g., in the absence of a test compound (negative control) or in the presence of a known compound (positive control). A control may represent an average value collected from several tests or results. Those skilled in the art will recognize that controls can be designed for the evaluation of any number of parameters. For example, controls can be designed to compare therapeutic benefits based on pharmacological data (e.g., half-life) or therapeutic measures (e.g., comparison of side effects). Those skilled in the art will understand which controls are valuable in a given situation and can analyze data based on comparison with the control value. Controls are also valuable for determining the significance of data. For example, if the values ​​for a given parameter vary widely in the controls, the variation in the test sample will not be considered significant.

[0090] As used herein, the term "contacting" is used according to its clear and ordinary meaning to mean a process that allows at least two distinct species (e.g., compounds including biomolecules or cells) to come into sufficient proximity to react, interact, or physically contact. However, it should be recognized that the resulting reaction product may be produced directly from the reaction between the added reagents, or from an intermediate from one or more of the added reagents that may be produced in the reaction mixture.

[0091] The term "expression" includes any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion. Expression can be detected using conventional techniques for detecting proteins (e.g., ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, etc.).

[0092] The term "recombinant," when used with reference to, for example, a cell, or a nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein, or vector has been modified by the introduction of a heterologous nucleic acid or protein, or the alteration of a naturally occurring nucleic acid or protein, or that the cell is derived from a cell so modified. That is, for example, a recombinant cell expresses genes that are not found in the native (non-recombinant) form of the cell, or expresses naturally occurring genes that are otherwise aberrantly expressed, under-expressed, or not expressed at all. Transgenic cells and plants are cells or plants that express heterologous genes or coding sequences, typically as a result of recombinant methods.

[0093] As used herein, the term "origin of transfer" or "oriT" refers to a short sequence (up to 500 bp) necessary for the transfer of DNA from a bacterial host and recipient during bacterial conjugation.

[0094] As used herein, a "curable origin of replication" refers to an origin of replication that does not replicate when cells are grown in the presence of certain chemicals or environmental conditions. Under these conditions, a plasmid containing a curable origin of replication is lost from the cell. For example, pSC101 ori TS does not function at high temperatures and is lost.

[0095] As used herein, the term "mobile element" is a type of genetic material that can be moved around within a genome or transferred between genomes, even between species.

[0096] As used herein, the term "conjugative transposon" refers to an integrated DNA element that cleaves itself to form a covalently closed circular intermediate that can be reintegrated within the same cell or transferred via conjugation to a recipient cell.

[0097] As used herein, the term "integration of conjugative elements" refers to a group of self-propagating genetic elements integrated into a chromosome.

[0098] As used herein, the term "P1-derived artificial chromosome" refers to a DNA construct that originates from the P1 bacteriophage.

[0099] As used herein, "bacterial artificial chromosome" refers to an engineered DNA sequence used to clone DNA sequences into bacteria.

[0100] The term "recombination-mediated genetic engineering gene" or "recombinant gene" refers to a gene that aids in the creation of genetic modifications in a DNA sequence. In an embodiment, the recombination-mediated genetic engineering gene allows for the in vivo construction of constructs in cells (e.g., bacterial cells) without in vitro genetic engineering techniques. In an embodiment, the recombination-mediated genetic engineering gene allows for genetic modifications to occur without the introduction of enzymes, including ligases and restriction enzymes. For example, the gene can participate in the natural process of bacterial homologous recombination without traditional molecular biology techniques known in the art. In an embodiment, the recombining gene is a lambda Red gene. The recombining genes are Redα, Redβ, and Redγ. For example, the gene can induce a high rate of homologous recombination in bacteria. In the donor or recipient cells, expression of one or more recombination genes is inducible. In an embodiment, the donor or recipient cells contain oligonucleotides encoding one or more recombination-mediated genetic engineering genes. In an embodiment, the oligonucleotides encoding one or more recombination-mediated genetic engineering genes are in a donor cell plasmid. In an embodiment, the recombination-mediated genetically engineered gene is inducible. In an embodiment, the recombination-mediated genetically engineered gene is Redα, Redβ, and Redγ. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in the recipient cell genome. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in a helper plasmid. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in a recipient oligonucleotide, which may be in the form of a plasmid. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in the recipient cell genome.

[0101] As used herein, the term "inducible high copy origin of replication" refers to a plasmid or vector that contains multiple origins of replication (e.g., 150-200 copies in the E. coli plasmid pUC) that can be induced by environmental conditions such as a change in temperature.

[0102] As used herein, the term "helper plasmid" refers to a plasmid that contains genes or other DNA elements necessary for a bacterium to perform a specified function. A helper plasmid may contain endonucleases, which are elements for transferring foreign DNA into the genome, transferring the plasmid to another cell, or performing homologous recombination. In an embodiment, the helper plasmid is an IncF1 plasmid, an IncPα plasmid, an IncI1 plasmid, pTiC58 from Agrobacterium tumefaciens, a cAD1 plasmid, an Inc18 plasmid, pIJ101 from Streptomyces, or an IncH plasmid. In an embodiment, the helper plasmid is an IncF1 plasmid. In an embodiment, the helper plasmid is an IncPα plasmid. In an embodiment, the helper plasmid is an IncI1 plasmid. In an embodiment, the helper plasmid is pTiC58 from Agrobacterium tumefaciens. In an embodiment, the helper plasmid is a cAD1 plasmid. In an embodiment, the helper plasmid is an Inc18 plasmid. In an embodiment, the helper plasmid is pIJ101 from Streptomyces. In an embodiment, the helper plasmid is an IncH plasmid. In an embodiment, the helper plasmid lacks a functional origin of transfer. In an embodiment, the helper plasmid contains a selectable marker that selects for retention of the helper plasmid in the donor cell.

[0103] As used herein, the term "homing endonuclease" refers to an endonuclease encoded as an autonomous gene in an intron sequence, as a fusion with a host protein, or as a self-splicing protein. Homing endonucleases catalyze the hydrolysis of DNA at longer recognition sites compared to group II restriction enzymes. Examples of homing endonucleases include, but are not limited to, LAGLIDAG, GIY-YIG, His-Cys box, HNH, PD-(D / E)xK, and Vsr-like / EDxHD. In embodiments, the homing endonucleases include I-ScaI, PI-SceI, I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-Po rI, I-PpoI, PI-PspI, I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I -Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I, or I-Vdi141I. In an embodiment, the homing endonuclease is I-ScaI. In an embodiment, the homing endonuclease is PI-SceI. In an embodiment, the homing endonuclease is I-AniI. In an embodiment, the homing endonuclease is I-CeuI. In an embodiment, the homing endonuclease is I-ChuI. In an embodiment, the homing endonuclease is I-CpaI. In an embodiment, the homing endonuclease is I-CpaII. In an embodiment, the homing endonuclease is I-CreI. In an embodiment, the homing endonuclease is I-DmoI. In an embodiment, the homing endonuclease is H-DreI. In an embodiment, the homing endonuclease is I-HmuI. In an embodiment, the homing endonuclease is I-HmuII. In an embodiment, the homing endonuclease is I-LlaI.In an embodiment, the homing endonuclease is I-MsoI. In an embodiment, the homing endonuclease is PI-PfuI. In an embodiment, the homing endonuclease is PI-PkoII. In an embodiment, the homing endonuclease is I-PorI. In an embodiment, the homing endonuclease is I-PpoI. In an embodiment, the homing endonuclease is PI-PspI. In an embodiment, the homing endonuclease is I-SceI. In an embodiment, the homing endonuclease is I-SceII. In an embodiment, the homing endonuclease is I-SceIII. In an embodiment, the homing endonuclease is I-SceIV. In an embodiment, the homing endonuclease is I-SceV. In an embodiment, the homing endonuclease is I-SceVI. In an embodiment, the homing endonuclease is I-SceVII. In an embodiment, the homing endonuclease is I-Ssp6803I. In an embodiment, the homing endonuclease is I-TevI. In an embodiment, the homing endonuclease is I-TevII. In an embodiment, the homing endonuclease is I-TevIII. In an embodiment, the homing endonuclease is PI-TliI. In an embodiment, the homing endonuclease is PI-TliII. In an embodiment, the homing endonuclease is I-Tsp061I or I-Vdi141I. In an embodiment, the homing endonuclease is I-Vdi141I.

[0104] As used herein, the term "RNA-guided DNA endonuclease" refers to any DNA endonuclease that is guided to a target DNA sequence by a helper or guide RNA molecule. Examples of RNA-guided DNA endonucleases include, but are not limited to, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2 , Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4, and all their variants and homologues.

[0105] As used herein, the term "HO" or "homotypic switching endonuclease" refers to a zinc finger nuclease in Saccharomyces cerevisiae that is involved in the initiation of mating type interconversion.

[0106] As used herein, a "cell" refers to a cell that performs metabolic or other functions sufficient to preserve or replicate its genomic DNA. Cells can be identified by methods known in the art, including, for example, the presence of an intact membrane, staining with a particular dye, the ability to produce progeny, or, in the case of gametes, the ability to combine with a second gamete to produce viable progeny. Cells can include prokaryotic and eukaryotic cells. Prokaryotic cells include, but are not limited to, bacteria. Eukaryotic cells include, but are not limited to, yeast cells and cells derived from plants and animals, such as mammalian, insect (e.g., spodoptera), and human cells. Cells can be useful if they are naturally non-adherent or if they are treated to prevent them from adhering to surfaces, for example, by trypsinization.

[0107] As used herein, "donor cell" refers to a cell (e.g., a bacterial cell) that transfers genetic material to another cell (e.g., a bacterial cell, a plant cell, etc.). A cell that receives the transferred genetic material is referred to herein as a "recipient cell."

[0108] The term "donor plasmid" as used herein refers to DNA from a donor cell that contains an oligonucleotide sequence (e.g., donor DNA, an oligonucleotide containing a DNA element or fragment thereof) that is transferred from the donor cell (e.g., a bacterial cell) to a recipient cell (e.g., a bacterial cell, a yeast cell, a plant cell, etc.). Typically, the donor plasmid is a circular double-stranded DNA that is separated from the genomic DNA. That is, the term "recipient oligonucleotide" can refer to the plasmid DNA in the recipient cell that receives the donor DNA (e.g., by homologous recombination of the donor DNA into the recipient), or the term "recipient oligonucleotide" can refer to any oligonucleotide in the recipient cell that receives the donor DNA, e.g., the recipient cell genomic DNA. In an embodiment, the DNA from the donor plasmid is received by DNA other than the DNA from the recipient plasmid. That is, in an embodiment, the donor DNA may be integrated into the genomic DNA.

[0109] The term "isolated," as applied to a nucleic acid or protein, means that the nucleic acid or protein is essentially free from other cellular components with which it is naturally associated. For example, it may be in a homogeneous state, in a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified.

[0110] It is understood that the examples and embodiments described herein are for illustrative purposes only, and that various modifications or changes in light thereof will be suggested to those skilled in the art and are to be included within the spirit and scope of this application and the scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.

[0111] In an embodiment of the method described herein, the donor cell or the recipient cell comprises an oligonucleotide encoding one or more homologous DNA repair genes. In an embodiment, the oligonucleotide encoding one or more homologous DNA repair genes is in the first, second, or subsequent donor plasmid. In an embodiment, the homologous DNA repair gene expression is inducible. In an embodiment, the homologous DNA repair gene is RecA.

[0112] In embodiments of the methods described herein, the donor or recipient cells comprise oligonucleotides encoding one or more recombination-mediated genetic manipulation genes. In embodiments, the oligonucleotides encoding one or more recombination-mediated genetic manipulation genes are in a donor cell plasmid.

[0113] In an embodiment of the method described herein, the recombination-mediated genetically engineered gene is inducible. In an embodiment, the recombination-mediated genetically engineered gene is Redα, Redβ, and Redγ. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in the recipient cell genome. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in a helper plasmid. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in a recipient oligonucleotide, which may be in the form of a plasmid. In an embodiment, the oligonucleotides encoding one or more homologous DNA repair genes are in the recipient cell genome.

[0114] In embodiments of the methods described herein, the donor cells, recipient cells, or recombinant recipient cells may be in an ordered array, or in a first or second ordered array, hi embodiments, the donor cells, recipient cells, or recombinant recipient cells may be transferred to a position on a third ordered array, a fourth ordered array, or a subsequent ordered array.

[0115] In embodiments of the methods described herein, the donor cell and the recipient cell are bacterial cells. In embodiments, the recipient cell is not a bacterial cell. In embodiments, the recipient cell is a plant cell. In embodiments, the recipient cell is a yeast cell. In embodiments, the recipient cell is a mammalian cell.

[0116] II. Methods for Assembling DNA Elements The invention provides a method of assembling multiple DNA elements into an assembled DNA element in a recipient cell, the method comprising the steps of: (a) contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions that (i) transfer a first donor plasmid from the first donor cell to a recipient cell by conjugation, and (ii) allow the first donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination, The plasmid comprises, in sequential order, an optional first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a fragment of a first DNA element (oligo 1), a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2), and an optional third endonuclease site (C3), and the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2, thereby forming a first endonuclease site that is a first endonuclease site (C4) and a second endonuclease site (C5) that is a second endonuclease site (C6) and a third endonuclease site (C7) that is a second endonuclease site (C8) and a fourth endonuclease site (C9) that is a second endonuclease site (C1) and a fourth endonuclease site (C2) that is a second endonuclease site (C3) and a fourth endonuclease site (C4) that is a second endonuclease site (C4) and a fourth endonuclease site (C5) that is a second endonuclease site (C6) and a fourth endonuclease site (C7) that is a second endonuclease site (C8) and a fourth endonuclease site (C9) that is a second endonuclease site (C1) and a fourth endonuclease site (C1) that is a second endonuclease site (C1) and a fourth endonuclease site (C2) that is a second endonuclease site (C2) and a fourth endonuclease site (C3) that is a second endonuclease site (C3) and a fourth endonuclease site (C4) that is a second endonuclease site (C4) and a fourth endonuclease site (C5) that is (b) providing a first recombined recipient oligonucleotide comprising a fragment of the first DNA element following homologous recombination of HR.2 and HR4; (i) transferring the second donor plasmid from the second donor cell to the first recipient cell by conjugation; and (ii) reacting the second donor plasmid with the first recombined recipient oligonucleotide in the recipient cell by homologous recombination to form a second recombined recipient oligonucleotide. contacting a recipient cell containing a bound recipient oligonucleotide, the second donor plasmid containing, in sequential order, an optional fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a fragment of a second DNA element (oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2), and an optional sixth endonuclease site (C6), thereby binding HR5 to HR2.1 and HR6.Following homologous recombination of HR2 and HR4, providing a second recombined recipient oligonucleotide comprising fragments of the first and second DNA elements (oligo1, oligo2), which form a DNA assembly. In an embodiment, HR2.1 and HR2.2 are flanked by non-homologous regions comprising one (C2) or two (C2.1, C2.2) endonuclease sites, and optionally, HR3 and HR4 are flanked by non-homologous regions comprising one (C4) or two (C4.1, C4.2) endonuclease sites. In an embodiment, HR6.1 and HR6.2 are flanked by non-homologous regions comprising one (C7) or two (C7.1, C7.2) endonuclease sites. In an embodiment, the recipient oligonucleotide is present in a recipient cell plasmid or a recipient cell genome. In an embodiment, the DNA assembly comprises at least a portion of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, a guide RNA (gRNA), or a combination thereof. In an embodiment, step (b) is repeated in one or more iterations with a third or subsequent donor cell comprising a third or subsequent donor plasmid comprising a matching HR region and a third or subsequent oligonucleotide (oligo 3, oligo 4, ... oligo N) encoding a fragment of the third or subsequent DNA element, thereby forming a third or subsequent recombined recipient oligonucleotide comprising the first, second, and third or subsequent DNA element fragments, which together form the DNA assembly. In an embodiment, step (a) comprises a plurality of first donor cells each comprising a different first donor plasmid, and step (b) comprises a plurality of second, third or subsequent donor cells each comprising a different second, third or subsequent donor plasmid, optionally each first donor cell being at a position in a first ordered array, and each second, third or subsequent donor cell being at a position in a second, third or subsequent ordered array, and optionally the method generates a combinatorial library comprising a plurality of different assembled DNA elements.

[0117] In an embodiment, the donor plasmid comprising the final DNA element forming part of the assembled DNA element comprises a barcoded homologous recombination (BHR) region to produce recipient cells each comprising the assembled DNA element, a barcoded BHR region, and a recombined recipient oligonucleotide comprising an additional HR, and the method further comprises the steps of (i) constructing or obtaining an array of barcoded donor cells each comprising a barcoded donor plasmid comprising a HR homologous to the BHR, a unique barcode oligonucleotide, and a second HR homologous to the additional HR of the recombined recipient oligonucleotide; (ii) (a) transferring the barcoded donor plasmid from the barcoded donor cell to a recipient cell by conjugation; and (b) contacting the array of barcoded donor cells with the array of recipient cells under conditions that allow the barcoded donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination, thereby producing an array of recipient cells comprising the barcoded assemblies.

[0118] In embodiments, each donor plasmid contains an additional pair of unique endonuclease sites CX, CY flanking the barcode homologous recombination (BHR) region, and the method further comprises contacting the array of recipient cells, each containing a DNA assembly, with an array of barcode donor cells containing barcode donor plasmids, each containing a pair of HR regions homologous to the BHR flanking a unique barcode oligonucleotide, to produce an array of recipient cells containing barcoded assemblies.

[0119] In an embodiment, the DNA assembly method may further comprise contacting a reset donor cell comprising a reset donor plasmid with a recipient cell comprising a recombined recipient oligonucleotide, the reset donor plasmid comprising, in sequential order, a homologous recombination region (HRt) homologous to the end sequence of the DNA assembly, a reset endonuclease site, a selectable marker, a reset endonuclease site, a homologous recombination region (HRX), and an origin of transfer, and the recombined recipient oligonucleotide comprising, in sequential order, a reset endonuclease site, a DNA assembly, a homologous recombination region (HRXa) homologous to HRX, and a reset endonuclease site, thereby providing a reset plasmid comprising an origin of transfer and a DNA assembly following homologous recombination between HRt and the end sequence of the DNA assembly and between HRX and HRXa. In an embodiment, the reset plasmid is in the donor cell. In an embodiment, the reset plasmid comprises a restricted origin of replication that functions in both the donor cell and the recipient cell. In an embodiment, the reset donor plasmid is constructed by a method comprising the steps of introducing an oligonucleotide insert HRt-C1-CM-C2-HRX, or a library of such oligonucleotide inserts, comprising homologous recombination regions HRt, HRX flanked by two endonuclease sites (C1, C2) and a counter-selectable marker (CM), allowing an endonuclease to cleave the endonuclease sites, and introducing the counter-selectable marker at the cleavage site by homologous recombination.

[0120] The present invention also provides a method of conjugating a barcode to an oligonucleotide, the method comprising the steps of: (a) inserting each oligonucleotide of the mixture of oligonucleotides into a donor plasmid, each of which optionally comprises a first endonuclease site (C1), a first homologous recombination region (HR1), a second homologous recombination region (HR2), and optionally a second endonuclease site (C2) in sequential order, wherein each oligonucleotide is inserted between HR1 and HR2, thereby providing a plurality of donor plasmids comprising the donor oligonucleotide, each donor plasmid C1-HR1-oligo-HR2-C2 comprising a single donor oligonucleotide from the mixture of oligonucleotides; (b) transforming a plurality of cells with the plurality of donor plasmids, whereby each cell comprises the donor plasmid, thereby forming a plurality of donor cells; (c) plating and culturing the plurality of donor cells, each at a unique location on the first ordered array, thereby providing a first ordered array of donor cells; (d) providing a plurality of recipient cells in a second ordered array, each recipient cell comprising a recipient oligonucleotide comprising, in sequential order, a unique barcode sequence identifying the location of the recipient cell in the second ordered array, a third homologous recombination region (HR3) homologous to HR1, optionally a third endonuclease site (C3), and a fourth homologous recombination region (HR4) homologous to HR2; (e) contacting the first ordered array of donor cells with the second ordered array of recipient cells under conditions to (i) transfer the donor plasmid from the donor cell to the recipient cell at the corresponding location on the array by conjugation, (ii) optionally cleave the first, second, and third endonuclease sites, and (ii) transfer the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby forming a third array of fusion oligonucleotides each comprising the unique barcode sequence and a donor oligonucleotide from the mixture of oligonucleotides;and (f) optionally sequencing the fusion oligonucleotides, thereby identifying each oligonucleotide in the array by its barcode sequence. In an embodiment, the recipient oligonucleotide is present in a recipient cell plasmid or a recipient cell genome. In an embodiment, the donor plasmid includes a selectable marker between HR1 and HR2 that selects for integration of the oligonucleotide into the recipient cell oligonucleotide, and optionally the donor plasmid includes a counter-selectable marker. In an embodiment, the recipient cell oligonucleotide includes a fourth endonuclease site (C4).

[0121] In a further aspect, a method of assembling DNA elements is provided, the method comprising the steps of: (a) providing a first host cell comprising a first donor plasmid comprising, in sequential order: (i) a first endonuclease target site, (ii) a first homologous recombination region, (iii) a first oligonucleotide optionally comprising a fragment of the first DNA element, (iv) a second homologous recombination region, (v) a second endonuclease target site, and (vi) a third endonuclease target site; (b) providing a recipient cell comprising a recipient oligonucleotide comprising: (i) a third homologous recombination region homologous to the first homologous recombination region, (ii) a fourth endonuclease target site, and (iii) a fourth homologous recombination region homologous to the second homologous recombination region; and (c) (i) assembling the first donor plasmid by bacterial conjugation. The method includes the steps of: (i) transferring a donor plasmid from a first host cell to a recipient cell; (ii) directing a first endonuclease to at least one of the first endonuclease target site, the third endonuclease target site, or the fourth endonuclease target site, thereby generating a double-stranded break in the first donor plasmid and the recipient oligonucleotide; and (iii) contacting the first host cell with the recipient cell under conditions that allow the first donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination via the first and second homologous recombination regions and the third and fourth corresponding homologous recombination regions, thereby forming a recombined recipient oligonucleotide.The method further comprises the steps of (d) providing a second host cell comprising a second donor plasmid comprising, in sequential order, (i) a fifth endonuclease target site, (ii) a fifth homologous recombination region homologous to the second region of homology, (iii) a second oligonucleotide comprising a fragment of a second DNA element, (iv) a sixth homologous recombination region homologous to the fourth region of homology, and (v) a sixth endonuclease target site; (e) (i) transferring the second donor plasmid from the second host cell to a recipient cell by bacterial conjugation, (ii) expressing a second endonuclease, and (iii) expressing the second endonuclease at the second endonuclease target site, the fifth endonuclease target site, and the fifth oligonucleotide comprising a fragment of a second DNA element. and (iv) contacting a second host cell with a recipient cell containing the recombined recipient oligonucleotide under conditions to direct the second endonuclease target site to the fifth and sixth endonuclease target site, thereby generating a double-stranded break, and (v) recombining the second donor plasmid and the recombined recipient oligonucleotide in the recipient cell by homologous recombination via the second and fourth homologous recombination sites corresponding to the fifth and sixth homologous recombination sites, thereby forming a second recombined recipient oligonucleotide comprising an assembled DNA element. In an embodiment, a different portion of the fourth homologous region is homologous to the sixth homologous recombination region compared to a portion of the fourth homologous region that is homologous to the second homologous recombination region.

[0122] In an embodiment, step (a) comprises a plurality of first host cells, each cell comprising a unique first oligonucleotide. In an embodiment, step (d) comprises a plurality of second host cells, each cell comprising a unique second oligonucleotide. In an embodiment, each first host cell comprises a unique plasmid. In an embodiment, each second host cell comprises a unique plasmid. In an embodiment, each first host cell is at a location in a first ordered array. In an embodiment, the plurality of first host cells are at a location in the first ordered array, thereby forming a pool of first host cells in the first ordered array. In an embodiment, each second host cell is at a location in a second ordered array. In an embodiment, the plurality of second host cells are at a location in a second ordered array, thereby forming a pool of second host cells at each location in the second ordered array.

[0123] In an embodiment, the first donor cell is in a first ordered array, the second donor cell is in a second ordered array, and one or more subsequent donor cells are in one or more subsequent arrays. That is, in an embodiment, the method provided herein generates a variant library that includes a plurality of different assembled DNA elements. In an embodiment, the variant library is generated by 1) independently producing each variant at a location in the first, second, or subsequent array using a first host cell, a second host cell, or a subsequent host cell, or 2) generating a pool of variants at a location in the first, second, or subsequent array using a plurality of first host cells, a second host cell, or a subsequent host cell. In an embodiment, the variant library is generated by independently producing each variant at a location in the first, second, or subsequent array using a first host cell, a second host cell, or a subsequent host cell. In embodiments, the variant library is generated by generating a variant pool using a plurality of first host cells, second host cells, or subsequent host cells at locations in the first, second, or subsequent arrays. For example, for the methods provided herein, including embodiments thereof, the first DNA element and / or the second DNA element can be a DNA barcode or a plurality of DNA barcodes. In embodiments, the method generates a repeatable barcoded platform. For example, in embodiments where the first DNA element and / or the second DNA element is a DNA barcode or a plurality of DNA barcodes, the method can be used for tracking cell lineage.

[0124] For example, the first DNA element and / or the DNA element may be a gRNA or multiple gRNAs, i.e., in an embodiment, the method comprises generating a combinatorial gRNA library.

[0125] In embodiments, the first endonuclease targets the first endonuclease target site. In embodiments, the first endonuclease targets the third endonuclease target site. In embodiments, the first endonuclease targets the fourth endonuclease target site. In embodiments, the second endonuclease targets the second endonuclease target site. In embodiments, the second endonuclease targets the fifth endonuclease target site. In embodiments, the second endonuclease targets the sixth endonuclease target site.

[0126] In embodiments, the DNA element is a gene. In embodiments, the DNA element is a promoter. In embodiments, the DNA element is an enhancer. In embodiments, the DNA element is a terminator. In embodiments, the DNA element is an intron. In embodiments, the DNA element is an intergenic region. In embodiments, the DNA element is a barcode. In embodiments, the DNA element is a translation start site. In embodiments, the DNA element is a gRNA. In embodiments, the DNA element is a fragment of any of the above.

[0127] In an embodiment, the recipient oligonucleotide is in a recipient plasmid. In an embodiment, the recipient oligonucleotide is in a recipient cell genome.

[0128] In an embodiment, the second donor plasmid further comprises a seventh homologous recombination region and a seventh endonuclease target site between components (d)iii) and (d)iv). In an embodiment, the first endonuclease targets the seventh endonuclease site.

[0129] In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the recipient cell comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the recipient cell genome comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the recipient plasmid comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the recipient helper plasmid comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the donor plasmid comprises an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the recipient genome comprises an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA and the recipient plasmid comprises an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA and the recipient helper plasmid contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease.

[0130] In an embodiment, the recipient cell comprises an inducible gRNA. In an embodiment, the donor cell comprises an oligonucleotide encoding a DNA endonuclease guided by an RNA. In an embodiment, the recipient cell comprises an oligonucleotide encoding a DNA endonuclease guided by an RNA. In an embodiment, the expression of the DNA guided by an RNA is constitutive. In an embodiment, the expression of the DNA guided by an RNA is inducible.

[0131] In an embodiment, the RNA-guided DNA endonuclease is Cas9. In an embodiment, the RNA-guided DNA endonuclease is Cas10. In an embodiment, the RNA-guided DNA endonuclease is Cpf1. In an embodiment, the RNA-guided DNA endonuclease is C2c1. In an embodiment, the RNA-guided DNA endonuclease is C2c2. In an embodiment, the RNA-guided DNA endonuclease is C2c3. In an embodiment, the RNA-guided DNA endonuclease is Cas12c1. In an embodiment, the RNA-guided DNA endonuclease is Cas12a. In an embodiment, the RNA-guided DNA endonuclease is Cas12b. In an embodiment, the RNA-guided DNA endonuclease is Cas12c2. In an embodiment, the RNA-guided DNA endonuclease is Cas12g. In an embodiment, the RNA-guided DNA endonuclease is Cas12e. In an embodiment, the RNA-guided DNA endonuclease is Cas12i1. In an embodiment, the RNA-guided DNA endonuclease is haCas12i2.

[0132] Regarding the methods provided herein, in an embodiment, the oligonucleotide encoding the first endonuclease is in a donor plasmid. In an embodiment, the oligonucleotide encoding the first endonuclease is a recipient oligonucleotide. In an embodiment, the oligonucleotide encoding the first endonuclease is in a recipient cell helper plasmid. In an embodiment, the oligonucleotide encoding the first endonuclease is in a recipient genome. In an embodiment, the expression of the first endonuclease is inducible. In an embodiment, the oligonucleotide encoding the second endonuclease is in a donor plasmid. In an embodiment, the oligonucleotide encoding the second endonuclease is in a recipient oligonucleotide. In an embodiment, the oligonucleotide encoding the second endonuclease is in a recipient cell helper plasmid.

[0133] In an embodiment, the first endonuclease and / or the second endonuclease is a homing endonuclease. In embodiments, the homing endonuclease is I-ScaI, PI-SceI, I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI. , I-PpoI, PI-PspI, I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-S sp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I, or I-Vdi141I. In an embodiment, the homing endonuclease is I-ScaI. In an embodiment, the homing endonuclease is PI-SceI. In an embodiment, the homing endonuclease is I-AniI. In an embodiment, the homing endonuclease is I-CeuI. In an embodiment, the homing endonuclease is I-ChuI. In an embodiment, the homing endonuclease is I-CpaI. In an embodiment, the homing endonuclease is I-CpaII. In an embodiment, the homing endonuclease is I-CreI. In an embodiment, the homing endonuclease is I-DmoI. In an embodiment, the homing endonuclease is H-DreI. In an embodiment, the homing endonuclease is I-HmuI. In an embodiment, the homing endonuclease is I-HmuII. In an embodiment, the homing endonuclease is I-LlaI. In an embodiment, the homing endonuclease is I-MsoI. In an embodiment, the homing endonuclease is PI-PfuI. In an embodiment, the homing endonuclease is PI-PkoII. In an embodiment, the homing endonuclease is I-PorI. In an embodiment, the homing endonuclease is I-PpoI.In an embodiment, the homing endonuclease is PI-PspI. In an embodiment, the homing endonuclease is I-SceI. In an embodiment, the homing endonuclease is I-SceII. In an embodiment, the homing endonuclease is I-SceIII. In an embodiment, the homing endonuclease is I-SceIV. In an embodiment, the homing endonuclease is I-SceV. In an embodiment, the homing endonuclease is I-SceVI. In an embodiment, the homing endonuclease is I-SceVII. In an embodiment, the homing endonuclease is I-Ssp6803I. In an embodiment, the homing endonuclease is I-TevI. In an embodiment, the homing endonuclease is I-TevII. In an embodiment, the homing endonuclease is I-TevIII. In an embodiment, the homing endonuclease is PI-TliI. In an embodiment, the homing endonuclease is PI-TliII. In an embodiment, the homing endonuclease is I-Tsp061I, or I-Vdi141I. In an embodiment, the homing endonuclease is I-Vdi141I.

[0134] In embodiments, the endonuclease is a transcription activator-like effector nuclease. In embodiments, the endonuclease is a zinc finger nuclease.

[0135] In the methods provided herein, in embodiments, steps (d) and (e) are repeated one or more times, thereby forming one or more subsequent assembled DNA elements. In embodiments, the first, second, or subsequent donor plasmid comprises a selectable marker that selects for integration of the first oligonucleotide, the second oligonucleotide, or subsequent oligonucleotides into the recipient cell oligonucleotide. In embodiments, the first donor plasmid comprises a selectable marker that selects for integration of the first oligonucleotide into the recipient cell oligonucleotide. In embodiments, the second donor plasmid comprises a selectable marker that selects for integration of the second oligonucleotide into the recipient cell oligonucleotide. In embodiments, the subsequent donor plasmid comprises a selectable marker that selects for integration of the subsequent oligonucleotide into the recipient cell oligonucleotide.

[0136] In an embodiment, the first donor plasmid contains a selectable marker that selects for integration of the first oligonucleotide into the recipient oligonucleotide. In an embodiment, the selectable marker is between components (a)(v) and (a)(iv). In an embodiment, the second donor plasmid contains a selectable marker that selects for integration of the second oligonucleotide into the recipient oligonucleotide.

[0137] In embodiments, the recipient cell oligonucleotide comprises a counterselectable marker that selects for integration of a first oligonucleotide into the recipient oligonucleotide. In embodiments, the recombined recipient cell oligonucleotide comprises a counterselectable marker that selects for integration of a second or subsequent oligonucleotide into the recombined recipient cell oligonucleotide. Counterselectable markers for use in the methods described herein are described above.

[0138] In an embodiment, the assembled DNA element, also referred to as DNA assembly, is sequenced. In an embodiment, the recipient oligonucleotide is sequenced. In an embodiment, the recombined recipient oligonucleotide is sequenced. In an embodiment, the recombined recipient oligonucleotide is a plasmid, which is linearized, ligated to a sequencing adaptor, and sequenced. In an embodiment, the assembled DNA element is amplified by PCR and sequenced. In an embodiment, (a) the recipient cell is lysed, (b) the oligonucleotide is digested with an endonuclease or endonucleases, (c) the assembled DNA element is isolated, and (d) the assembled DNA element or assembled genes are ligated to a sequencing adaptor and sequenced. In an embodiment, the assembled DNA element is isolated. In an embodiment, the recombined recipient oligonucleotide is isolated.

[0139] In embodiments, the assembled DNA elements are from about 100 nucleotides to about 500,000 nucleotides in length. The length may be any value or subrange within the indicated ranges, inclusive of the endpoints.

[0140] In embodiments, the assembled DNA elements are about 100 nucleotides, about 1000 nucleotides, about 10,000 nucleotides, about 20,000 nucleotides, about 40,000 nucleotides, about 60,000 nucleotides, about 80,000 nucleotides, about 100,000 nucleotides, about 120,000 nucleotides, about 140,000 nucleotides, about 160,000 nucleotides, about 180,000 nucleotides, about 20,000 nucleotides, about 240,000 nucleotides, about 260,000 nucleotides, about 280,000 nucleotides, about 300,000 nucleotides, 320,000 nucleotides, about 340,000 nucleotides, about 360,000 nucleotides, about 380,000 nucleotides, about 400,000 nucleotides, about 420,000 nucleotides, about 440,000 nucleotides, about 460,000 nucleotides, about 480,000 nucleotides, or about 500,000 nucleotides. The length may be any value or subrange within the indicated range, including the endpoints.

[0141] In embodiments, the first, second, or subsequent homology regions and corresponding first, second, or subsequent homology regions are from about 20 base pairs to about 500 base pairs in length, and the length may be any value or subrange within the indicated ranges, including the endpoints.

[0142] In embodiments, the first, second, or subsequent homology regions and corresponding first, second, or subsequent homology regions are about 20 base pairs, 40 base pairs, 60 base pairs, 80 base pairs, 100 base pairs, 120 base pairs, 140 base pairs, 160 base pairs, 180 base pairs, 200 base pairs, 220 base pairs, 240 base pairs, 260 base pairs, 280 base pairs, 300 base pairs, 320 base pairs, 340 base pairs, 360 base pairs, 380 base pairs, 400 base pairs, 420 base pairs, 440 base pairs, 460 base pairs, 480 base pairs, or 500 base pairs in length. In embodiments, the first, second, or subsequent homology regions and corresponding first, second, or subsequent homology regions are about 50 base pairs in length. The length may be any value or subrange within the indicated ranges, inclusive of the endpoints.

[0143] III. Method of Analysis In one aspect, a method for identifying an oligonucleotide from a mixture of oligonucleotides is provided, the method comprising the steps of: (a) providing a mixture of oligonucleotides; (b) inserting each of the oligonucleotides into a donor plasmid, each of which comprises, in sequential order, i) a first endonuclease cleavage site, ii) a first homologous recombination region, iii) a second homologous recombination region, and iv) a second endonuclease cleavage site, and an oligonucleotide is inserted between the first and second homologous recombination regions, thereby producing a plurality of donor plasmids, each of which comprises a single oligonucleotide from the mixture of oligonucleotides; (c) transforming a plurality of host cells with the plurality of donor plasmids, each of which comprises the donor plasmid, thereby forming a plurality of transformed host cells; and (d) plating and culturing the plurality of transformed host cells on a first ordered array, each of which produces clonal colonies within the first ordered array. (e) providing a plurality of recipient cells in a second ordered array, each recipient cell comprising a recipient oligonucleotide comprising, in sequential order, (i) a unique barcode sequence identifying the location of the recipient cell in the second ordered array, (ii) a corresponding first homologous recombination region to which the first homologous recombination region is homologous, (iii) a third endonuclease cleavage site, and (iv) a corresponding second homologous recombination region to which the second homologous recombination region is homologous, wherein the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are cleavable by an endonuclease; (f) under conditions to (i) transfer the donor plasmid from the clonal colony to the recipient cell by bacterial conjugation, (ii) cleave the first, second, and third endonuclease cleavage sites by an endonuclease, and (ii) transfer the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination.(g) contacting each colony of clones from the first ordered array with a recipient cell at a corresponding site of the second ordered array, thereby producing a fusion sequence comprising the barcode sequence and the oligonucleotide; (g) sequencing the fusion protein, and identifying the sequenced oligonucleotide in the first and / or second ordered array of donor cells and / or recipient cells by identifying the barcode sequence;

[0144] In embodiments, plating and culturing the cells comprises plating and culturing the cells on a surface. For example, the surface may be a solid medium. That is, in embodiments, the clonal colony is a colony of cells on a solid medium. In embodiments, plating and culturing the cells comprises plating and culturing the cells in a liquid medium. For example, a single cell may be plated and cultured in a liquid medium. That is, in embodiments, the clonal colony is a colony of cells in a liquid medium.

[0145] In an embodiment, the recipient oligonucleotide is in a recipient cell plasmid. In an embodiment, the recipient oligonucleotide is in a recipient cell genome. In an embodiment, the donor plasmid comprises a selectable marker between the first and second homologous recombination regions, which selects for integration of the oligonucleotide into the recipient cell oligonucleotide. In an embodiment, the recipient cell oligonucleotide comprises two endonuclease cleavage sites in step (e)(iii). In an embodiment, the method further comprises a counter-selectable marker between the two endonuclease cleavage sites, which selects for integration of the oligonucleotide into the recipient oligonucleotide.

[0146] In one aspect, a method for identifying an oligonucleotide from a mixture of oligonucleotides is provided.The method includes the steps of: (a) providing a plurality of host cells in a first ordered array, each host cell comprising a donor plasmid, each donor plasmid comprising, in sequential order: i) a first endonuclease cleavage site, ii) a first homologous recombination region, iii) a unique barcode sequence, iv) a second homologous recombination region, and v) a second endonuclease cleavage site, wherein the unique barcode sequence identifies a location of the host cell in the first ordered array; (b) providing a plurality of recipient cells, Each recipient cell comprises a recipient oligonucleotide comprising an oligonucleotide from the plurality of oligonucleotides, and each recipient plasmid comprises, in sequential order, i) an oligonucleotide sequence, ii) a corresponding first homologous recombination region to which the first homologous recombination region is homologous, iii) a first endonuclease cleavage site, a second endonuclease cleavage site, and a third endonuclease cleavage site, wherein the third endonuclease cleavage site can be cleaved by an endonuclease, and iv) a second homologous recombination region. (c) plating and culturing a plurality of recipient cells on the second ordered array, each recipient cell producing a clonal colony in the second ordered array; (d) contacting a donor cell from the first ordered array with each colony of clones at a corresponding site in the second ordered array under conditions that (i) transfer the donor plasmid from the donor cell to the clonal colony by bacterial conjugation, (ii) cleave the first, second, and third endonuclease cleavage sites by an endonuclease, and (iii) transfer the barcode sequence from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby producing a fusion sequence comprising the barcode sequence and the oligonucleotide; (e) sequencing the fusion sequence; and (f) identifying the sequenced oligonucleotide in the first and / or second ordered array of recipient cells by identifying the barcode sequence.

[0147] In an embodiment, the recipient oligonucleotide is in a recipient cell plasmid.

[0148] In an embodiment, the recipient oligonucleotide is in the recipient cell genome. In an embodiment, the donor plasmid comprises a selectable marker between the first and second homologous recombination regions that selects for integration of the barcode into the recipient oligonucleotide. In an embodiment, the recipient oligonucleotide comprises two endonuclease cleavage sites in step (b)(iii). In an embodiment, the method further comprises a counter-selectable marker between the two endonuclease cleavage sites that selects for integration of the barcode into the recipient cell oligonucleotide.

[0149] For the methods provided herein, in an embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are the same endonuclease cleavage site. In an embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are different endonuclease cleavage sites. In an embodiment, the endonuclease comprises multiple endonucleases.

[0150] In embodiments, the endonuclease is encoded by an oligonucleotide in the recipient cell. In embodiments, the oligonucleotide is in the recipient cell genome. In embodiments, the oligonucleotide is in the recipient plasmid. In embodiments, the oligonucleotide encoding the endonuclease is in a helper plasmid. In embodiments, the endonuclease is encoded by an oligonucleotide in a donor plasmid. In embodiments, the endonuclease is encoded by an oligonucleotide in a donor plasmid. In embodiments, expression of the endonuclease is inducible.

[0151] In an embodiment, the endonuclease is a transcription activator-like effector nuclease. In an embodiment, the endonuclease is a zinc finger nuclease. A zinc finger nuclease. The endonuclease is HO.

[0152] In an embodiment, the RNA-guided DNA endonuclease is a CRISPR system. In an embodiment, the endonuclease is an RNA-guided DNA endonuclease. In an embodiment, the RNA-guided DNA endonuclease is Cas9. In an embodiment, the RNA-guided DNA endonuclease is Cas10. In an embodiment, the RNA-guided DNA endonuclease is Cpf1. In an embodiment, the RNA-guided DNA endonuclease is C2c1. In an embodiment, the RNA-guided DNA endonuclease is C2c2. In an embodiment, the RNA-guided DNA endonuclease is C2c3. In an embodiment, the RNA-guided DNA endonuclease is Cas12c1. In an embodiment, the RNA-guided DNA endonuclease is Cas12a. In an embodiment, the RNA-guided DNA endonuclease is Cas12b. In an embodiment, the RNA-guided DNA endonuclease is Cas12c2. In an embodiment, the RNA-guided DNA endonuclease is Cas12g. In an embodiment, the RNA-guided DNA endonuclease is Cas12e. In an embodiment, the RNA-guided DNA endonuclease is Cas12i1. In an embodiment, the RNA-guided DNA endonuclease is Cas12i2.

[0153] For the methods provided herein, in an embodiment, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient genome comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient plasmid comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient helper plasmid comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In an embodiment, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient plasmid comprises an oligonucleotide encoding a guided RNA-guided DNA endonuclease.

[0154] In an embodiment, the recipient cell comprises an inducible gRNA. In an embodiment, the donor cell comprises an oligonucleotide encoding a DNA endonuclease guided by an RNA. In an embodiment, the recipient cell comprises an oligonucleotide encoding a DNA endonuclease guided by an RNA. In an embodiment, the expression of the DNA guided by an RNA is constitutive. In an embodiment, the expression of the DNA guided by an RNA is inducible.

[0155] In an embodiment, the method further comprises isolating the donor plasmid. In an embodiment, the method further comprises isolating the recipient cell. In an embodiment, the method further comprises isolating the recombinant recipient plasmid. In an embodiment, the method further comprises isolating the sequenced oligonucleotide. In an embodiment, the method further comprises isolating one or more of the donor cell, the recipient cell, the recombinant recipient cell, the recipient oligonucleotide, or the recombinant recipient oligonucleotide.

[0156] In an embodiment, the method includes combining one or more subsets of colonies. That is, in an embodiment, the method includes combining one or more subsets of colonies and isolating a plurality of donor plasmids from the subsets of colonies. In an embodiment, the method includes combining one or more subsets of colonies and isolating a plurality of recipient plasmids from the subsets of colonies. In an embodiment, the method further includes combining one or more subsets of colonies and isolating a plurality of recombinant recipient plasmids from the subsets of colonies. In an embodiment, the method further includes isolating a plurality of sequenced oligonucleotides.

[0157] In an embodiment, the recipient cell or donor cell comprises an oligonucleotide that enables plasmid conjugation. In an embodiment, the oligonucleotide that enables plasmid conjugation is in the genome of the donor cell. In an embodiment, the oligonucleotide that enables plasmid conjugation is in a helper plasmid. In an embodiment, the oligonucleotide that enables plasmid conjugation is in the Tra operon. Other oligonucleotides that enable plasmid conjugation are described above.

[0158] In embodiments, donor cells, recipient cells, or recombinant recipient cells are transferred to a location on a third ordered array, a fourth ordered array, or a subsequent ordered array.

[0159] In embodiments, the barcode sequence is about 4 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 8 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 12 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 16 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 20 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 24 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 28 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 32 nucleotides to about 50 nucleotides in length. In embodiments, the barcode sequence is about 36 nucleotides to about 50 nucleotides in length. The length of the barcode may be any value or subrange within the ranges provided herein, including the endpoints.

[0160] In an embodiment, the barcode sequence is about 4 nucleotides to about 36 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 32 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 28 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 24 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 20 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 16 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 12 nucleotides in length. In an embodiment, the barcode sequence is about 4 nucleotides to about 8 nucleotides in length. In an embodiment, the barcode sequence is about 4, 8, 12, 16, 20, 24, 28, 32, 36, or 40 nucleotides in length. In an embodiment, the barcode sequence is about 15 nucleotides in length. In an embodiment, the barcode sequence is about 40 nucleotides in length. The length of the barcode can be any value or subrange within the ranges provided herein, inclusive of the endpoints.

[0161] Further embodiments The invention is further described by the following additional embodiments.

[0162] Embodiment 1: A method of assembling a DNA element, comprising the steps of: (1) providing a first host cell comprising a first donor plasmid comprising, in sequential order, a first endonuclease target site, a first homologous recombination region, a first oligonucleotide optionally comprising a fragment of the first DNA element, a second homologous recombination region, a second endonuclease target site, and a third endonuclease target site; (2) providing a first host cell comprising a first donor plasmid comprising, in sequential order, a first endonuclease target site, a first oligonucleotide optionally comprising a fragment of the first DNA element, a second homologous recombination region, a second endonuclease target site, and a third endonuclease target site; providing a recipient cell comprising a recipient oligonucleotide comprising a fourth homologous recombination region homologous to the fourth endonuclease target site; and (3) (i) transferring the first donor plasmid from the first host cell to the recipient cell by bacterial conjugation; (ii) directing a first endonuclease to at least one of the first endonuclease target site, the third endonuclease target site, or the fourth endonuclease target site, thereby generating a double-stranded break in the first donor plasmid and the recipient oligonucleotide; and (iii) transferring the first endonuclease to at least one of the first endonuclease target site, the third endonuclease target site, or the fourth endonuclease target site, thereby generating a double-stranded break in the first donor plasmid and the recipient oligonucleotide. (3) contacting the first host cell with a recipient cell under conditions that allow the first donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination via the first and second homologous recombination regions and the third and fourth corresponding homologous recombination regions, thereby forming a recombined recipient oligonucleotide; (4) contacting the first host cell with a recipient cell under conditions that allow the first donor plasmid and the recipient oligonucleotide to recombine in the recipient cell by homologous recombination via the first and second homologous recombination regions and the third and fourth corresponding homologous recombination regions, thereby forming a recombined recipient oligonucleotide; providing a second host cell comprising a second donor plasmid comprising a sixth homologous recombination region homologous to the fourth region of homology and a sixth endonuclease target site; and (5) (i) transferring the second donor plasmid from the second host cell to the recipient cell by bacterial conjugation; (ii) expressing a second endonuclease; and (iii) directing the second endonuclease to the second endonuclease target site, the fifth endonuclease target site, and / or the sixth endonuclease target site, thereby generating a double-stranded break;(iv) contacting a second host cell with a recipient cell containing the recombined recipient oligonucleotide under conditions in which the second donor plasmid and the recombined recipient oligonucleotide recombine in the recipient cell by homologous recombination via the second and fourth homologous recombination sites corresponding to the fifth and sixth homologous recombination sites, thereby forming a second recombined recipient oligonucleotide comprising an assembled DNA element.

[0163] In a further embodiment, step (a) comprises a plurality of first host cells, each cell comprising a different first oligonucleotide, and / or step (d) comprises a plurality of second host cells, each cell comprising a different second oligonucleotide.

[0164] In a further embodiment, each first host cell is at a position in the first ordered array.

[0165] In a further embodiment, the plurality of first host cells are at positions in the first ordered array, thereby forming a pool of first host cells in the first ordered array.

[0166] In a further embodiment, each second host cell is at a position in a second ordered array.

[0167] In a further embodiment, the plurality of second host cells are at positions in the second ordered array, thereby forming a pool of second host cells at each position in the second ordered array.

[0168] In a further embodiment, the first donor cells are in a first ordered array, the second donor cells are in a second ordered array, and one or more subsequent donor cells are in one or more subsequent arrays.

[0169] In a further embodiment, the method generates a combinatorial library comprising a plurality of different assembled DNA elements.

[0170] In a further embodiment, the first endonuclease targets a first, third, or fourth endonuclease target site.

[0171] In a further embodiment, the second endonuclease targets a second, fifth, or sixth endonuclease target site.

[0172] In further embodiments, the DNA element is a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, or a gRNA.

[0173] In a further embodiment, the recipient oligonucleotide is in a recipient plasmid.

[0174] In a further embodiment, the recipient oligonucleotide is within the recipient cell genome.

[0175] In a further embodiment, the second donor plasmid further comprises a seventh homologous recombination region and a seventh endonuclease target site between components (d)iii) and (d)iv).

[0176] In a further embodiment, the first endonuclease targets a seventh endonuclease site.

[0177] In further embodiments, the first and second endonucleases are independently selected from an RNA-guided DNA endonuclease, a homing endonuclease, a transcription activator-like effector nuclease, and a zinc finger nuclease.

[0178] In a further embodiment, the oligonucleotide encoding the first endonuclease is in the donor cell or the recipient cell.

[0179] In a further embodiment, expression of the first and / or second endonuclease is inducible.

[0180] In a further embodiment, the oligonucleotide encoding the second oligonucleotide is in the donor cell or the recipient cell.

[0181] In a further embodiment, steps (d)-(e) are repeated one or more iterations, thereby forming one or more subsequent assembled DNA elements.

[0182] In a further embodiment, the first, second, or subsequent donor plasmid contains a selectable marker that selects for integration of the first oligonucleotide, the second oligonucleotide, or subsequent oligonucleotide into the recipient cell plasmid.

[0183] In a further embodiment, the first donor plasmid comprises a selectable marker that selects for integration of the first oligonucleotide into the recipient oligonucleotide, optionally the selectable marker being between the second and third endonuclease target sites. In an embodiment, the second donor plasmid comprises a selectable marker that selects for integration of the second oligonucleotide into the recipient oligonucleotide.

[0184] In a further embodiment, the recipient cell oligonucleotide contains a counterselectable marker that selects for integration of the first oligonucleotide into the recipient cell oligonucleotide.

[0185] In a further embodiment, the donor plasmid comprises an origin of transfer, optionally the origin of transfer is from a mobile element.

[0186] In a further embodiment, the donor plasmid comprises a conditional origin of replication, optionally the conditional origin of replication is dependent on the presence of an oligonucleotide or on cell growth conditions.

[0187] In a further embodiment, the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid greater than 30 kilobases in length.

[0188] In a further embodiment, the donor plasmid or recipient oligonucleotide contains an inducible high copy origin of replication.

[0189] In further embodiments, the donor plasmid or recipient oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome. In further embodiments, the donor plasmid or recipient oligonucleotide is a viral vector.

[0190] In a further embodiment, the donor cell comprises oligonucleotides that enable conjugation of the plasmid.

[0191] In a further embodiment, the donor or recipient cells comprise oligonucleotides encoding one or more homologous DNA repair genes, and optionally expression of the homologous DNA repair genes is inducible.

[0192] In a further embodiment, the donor or recipient cells contain oligonucleotides encoding one or more recombination-mediating genetically engineered genes.

[0193] In a further embodiment, the donor cell and the recipient cell are independently bacterial cells.

[0194] In further embodiments, the assembled DNA elements are sequenced, the recipient oligonucleotides are sequenced, and / or the recombined recipient oligonucleotides are sequenced.

[0195] In a further embodiment, the recombinant recipient oligonucleotide is a plasmid, and the plasmid is linearized, ligated to a sequencing adapter, and sequenced.

[0196] In a further embodiment, the assembled DNA elements are amplified by PCR and sequenced, optionally (a) the recipient cells are lysed, (b) the oligonucleotides are digested with an endonuclease or endonucleases, (c) the assembled DNA elements are isolated, and (d) the assembled DNA elements or assembled genes are ligated to sequencing adaptors and sequenced.

[0197] In a further embodiment, the assembled DNA elements or recombined recipient oligonucleotides are isolated.

[0198] In a further embodiment, the assembled DNA fragments are between 100 nucleotides and 500,000 nucleotides in length.

[0199] In a further embodiment, the first, second, or subsequent homology region and the corresponding first, second, or subsequent homology region are from about 20 base pairs to about 500 base pairs in length.

[0200] In a further embodiment, the first, second or subsequent region of homology and the corresponding first, second or subsequent region of homology are about 50 base pairs in length.

[0201] In an embodiment, provided herein is a method of identifying an oligonucleotide from a mixture of oligonucleotides, the method comprising the steps of: (a) providing a mixture of oligonucleotides; (b) inserting each of the oligonucleotides into a donor plasmid, each donor plasmid comprising, in sequential order, (i) a first endonuclease cleavage site, (ii) a first homologous recombination region, (iii) a second homologous recombination region, and (iv) a second endonuclease cleavage site, wherein an oligonucleotide is inserted between the first homologous recombination region and the second homologous recombination region, thereby producing a plurality of donor plasmids, each donor plasmid comprising a single oligonucleotide from the mixture of oligonucleotides; (c) transforming a plurality of host cells with the plurality of donor plasmids, whereby each host cell comprises the donor plasmid, thereby forming a plurality of transformed host cells; (d) plating and culturing the plurality of transformed host cells on the first ordered array. (e) providing a plurality of recipient cells in a second ordered array, each of the recipient cells comprising, in sequential order, (i) a unique barcode sequence that identifies the location of the recipient cell in the second ordered array, (ii) a corresponding first homologous recombination region to which the first homologous recombination region is homologous, (iii) a third endonuclease cleavage site, and (iv) a second endonuclease cleavage site. and (iv) a recipient oligonucleotide comprising a corresponding second homologous recombination region to which the second homologous recombination region is homologous, wherein the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are cleavable by an endonuclease; (f) (i) transferring the donor plasmid from the clonal colony to a recipient cell by bacterial conjugation; and (ii) cleaving the first, second, and third endonuclease cleavage sites by an endonuclease;(ii) contacting each colony of clones from the first ordered array with a recipient cell at a corresponding site of the second ordered array under conditions to transfer the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby producing a fusion sequence comprising the barcode sequence and the oligonucleotide; (g) sequencing the fusion sequence; and (h) identifying the sequenced oligonucleotide in the first and / or second ordered arrays of the donor and / or recipient cells by identifying the barcode sequence.

[0202] In a further embodiment, the recipient oligonucleotide is in a recipient cell plasmid.

[0203] In a further embodiment, the recipient oligonucleotide is within the recipient cell genome.

[0204] In a further embodiment, the method provides that the donor plasmid includes a selectable marker between the first and second homologous recombination regions that selects for integration of the oligonucleotide into the recipient cell.

[0205] In a further embodiment, the methods herein provide that the recipient cell oligonucleotide comprises two endonuclease cleavage sites in step (e)(iii).

[0206] In an embodiment herein, the method further comprises a counter-selectable marker between the two endonuclease cleavage sites, which selects for integration of the oligonucleotide into the recipient oligonucleotide.

[0207] In an embodiment, provided herein is a method of identifying an oligonucleotide from a plurality of oligonucleotides, the method comprising the steps of: (a) providing a plurality of host cells in a first ordered array, each host cell comprising a donor plasmid, each donor plasmid comprising, in sequential order, (i) a first endonuclease cleavage site, (ii) a first homologous recombination region, (iii) a unique barcode sequence, (iv) a second homologous recombination region, and (v) a second endonuclease cleavage site, wherein the unique barcode sequence identifies a location of the host cell in the first ordered array; (b) providing a plurality of recipient cells, each recipient cell comprising a recipient oligonucleotide comprising an oligonucleotide from the plurality of oligonucleotides, each recipient plasmid comprising, in sequential order, i) the oligonucleotide sequence, ii) a corresponding first homologous recombination region to which the first homologous recombination region is homologous, iii) the first endonuclease cleavage site, the second endonuclease cleavage site, and (c) plating and culturing a plurality of recipient cells on the second ordered array, each of the recipient cells producing a clonal colony in the second ordered array; (d) contacting a donor cell from the first ordered array with each of the clonal colonies at a corresponding site in the second ordered array under conditions that (i) transfer the donor plasmid from the donor cell to the clonal colony by bacterial conjugation, (ii) cleave the first, second, and third endonuclease cleavage sites by an endonuclease, and (iii) transfer the barcode sequence from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby producing a fusion sequence comprising the barcode sequence and the oligonucleotide; and (e) sequencing the fusion sequence.and (f) identifying the sequenced oligonucleotides in the first and / or second ordered arrays of donor cells and / or recipient cells by identifying the barcode sequences.

[0208] In an embodiment, the method provides that the recipient oligonucleotide is in a recipient cell plasmid.

[0209] In a further embodiment, the recipient oligonucleotide is within the recipient cell genome.

[0210] In an embodiment, the donor plasmid contains a selectable marker between the first and second homologous recombination regions that selects for integration of the barcode into the recipient oligonucleotide.

[0211] In a further embodiment, the recipient oligonucleotide comprises two endonuclease cleavage sites in step (b)(iii).

[0212] In a further embodiment, the method includes providing a counter-selectable marker between the two endonuclease cleavage sites that selects for integration of the barcode into the recipient cell oligonucleotide.

[0213] In a further embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are the same endonuclease cleavage site.

[0214] In a further embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are different endonuclease cleavage sites.

[0215] In a further embodiment, the endonuclease comprises multiple endonucleases.

[0216] In a further embodiment, the donor plasmid comprises an origin of transfer.

[0217] In a further embodiment, the origin of transport is from a mobile element.

[0218] In a further embodiment, the donor plasmid comprises a conditional origin of replication.

[0219] In a further embodiment, the conditional origin of replication is dependent on the presence of an oligonucleotide.

[0220] In a further embodiment, the conditional origin of replication is dependent on the conditions of cell growth.

[0221] In a further embodiment, the donor or recipient plasmid comprises a replicon capable of replicating a plasmid at least 30 kilobases in length.

[0222] In a further embodiment, the replicon is derived from a P1 derived artificial chromosome or a bacterial artificial chromosome.

[0223] In a further embodiment, the donor plasmid or recipient cell oligonucleotide contains an inducible high copy origin of replication.

[0224] In further embodiments, the donor plasmid or recipient cell oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome.

[0225] In a further embodiment, the donor plasmid or recipient oligonucleotide is a viral vector.

[0226] In a further embodiment, the endonuclease is encoded by an oligonucleotide in the recipient cell.

[0227] In a further embodiment, the endonuclease is encoded by an oligonucleotide in the donor plasmid.

[0228] In a further embodiment, the endonuclease is a homing endonuclease.

[0229] In a further embodiment, the endonuclease is an RNA-guided DNA endonuclease.

[0230] In a further embodiment, the endonuclease is HO.

[0231] In a further embodiment, the method further comprises isolating the donor plasmid.

[0232] In a further embodiment, the method further comprises the step of isolating the recipient plasmid.

[0233] In a further embodiment, the method further comprises the step of isolating the recombinant recipient plasmid.

[0234] In a further embodiment, the method further comprises isolating the sequenced oligonucleotide.

[0235] In a further embodiment, the donor or recipient cells contain oligonucleotides that allow for plasmid conjugation.

[0236] In a further embodiment, the donor or recipient cells comprise oligonucleotides encoding one or more homologous DNA repair genes.

[0237] In a further embodiment, the donor or recipient cells contain oligonucleotides encoding one or more recombination-mediating genetically engineered genes.

[0238] In further embodiments, the donor cells, recipient cells, or recombinant recipient cells are transferred to a location on a third ordered array, a fourth ordered array, or a subsequent ordered array.

[0239] In a further embodiment, the donor cell and the recipient cell are independently bacterial cells.

[0240] In a further embodiment, the barcode sequence is from about 4 nucleotides to about 40 nucleotides in length.

[0241] In a further embodiment, the barcode sequence is about 15 nucleotides in length.

[0242] It is understood that the examples and embodiments described herein are for illustrative purposes only, and that various modifications or variations in light of these will be suggested to those skilled in the art and are to be within the spirit and scope of this application and the scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes. EXAMPLES

[0243] [Example 1] In vivo DNA stitching method Bacterial strains: BUN20 [Δlac-169 rpoS(Am) robA1 creC510 hsdR514 ΔuidA(MluI):pir-116 endA(BT333)recA1 F'(lac+ pro+ ΔoriT:tet)] was used as the donor strain (Li, M. et al., Nat. Genet. 37, 311-319 (2005)). BW23474: [Δlac-169 rpoS(Am) robA1 creC510 hsdR514 ΔuidA(MluI):pir-116 endA(BT333)recA1] was used as the host strain for cloning and propagation of all donor plasmids (Haldimann, A. et al., Proc. Natl. Acad. Sci. 93, 14361 (1996)). BW28705 [lacIQ rrnB3 ΔlacZ4787 hsdR514 Δ(araBAD)567 Δ(rhaBAD)568 galU95 ΔendA9:FRT ΔrecA635:FRT] or RE1133 (Egbert et al., Nucleic Acids Research, vol. 47(6), April 8, 2019, pp. 3244-3256) [cmR::mutS pTet2-gam-bet-exo-dam / tetR::bioA / B ilvG+ dnaG.Q576A lacIQ1 Pcp8-araE ΔaraBAD pConst-araC ΔrecJ ΔxonA Pkm-cymR-Cas9::bioC] were used as recipient strains for in vivo ligation. DH5α and DH10β were used for cloning of recipient plasmids.

[0244] The DNA oligonucleotides used for the first and second step PCR are shown in Tables 1 and 2.

[0245] [Table 1]

[0246] [Table 2]

[0247] Media and Chemicals: Luria-Bertani (LB) broth (1% w / v tryptone, 0.5% w / v yeast extract, 1% w / v NaCl) as a complex medium was routinely used for cloning and propagation of donor and recipient plasmids. Antibiotics were added at the concentrations listed in Table 3 for plasmid maintenance. For LB medium containing hygromycin, 0.5% w / v sodium chloride was used because hygromycin is salt sensitive. araBAD , P rhaBAD , P Tet2 , P km -cymR, and P lacIQ L-arabinose (0.2% w / v), L-rhamnose (0.2% w / v), anhydrotetracycline (100 ng / ml), 4-isopropylbenzoic acid (cuminate, 15 μg / ml), and isopropyl β-d-1 thiogalactopyranoside (IPTG, 500 μM) were used to induce the promoters, respectively. Sucrose agar plates (0.5% w / v yeast extract, 1% w / v tryptone, 6% w / v sucrose, 1.5% agar) and appropriate amounts of antibiotics were used to select for the SacB counterselectable marker. PheS Gly 294 For counterselection of PheS Gly, Cl-Phe agar plates (0.5% w / v yeast extract, 1% w / v NaCl, 0.4% w / v glycerol, 2% w / v agar, 10 mM D,Lp-Cl-Phe) and appropriate amounts of antibiotic were used. 294 For subcloning of recipient plasmids containing fragments, YEG agar plates (0.5% w / v yeast extract, 1% w / v NaCl, 0.4% w / v glucose, 2% w / v agar) and appropriate amounts of antibiotic were used.

[0248] [Table 3]

[0249] The sequences of the plasmids used for in vivo DNA stitching can be found in Table 4.

[0250] [Table 4] TIFF2024509194000005.tif63161

[0251] Construction of the in vivo DNA stitching system: Construction of helper plasmids and host recipient strains. In the related MAGIC cloning system (Li, M. et al., Nat. Genet. 37, 311-319 (2005)), recipient cells contain the helper plasmid pML300, which contains the inducible λ-red and temperature-sensitive origin of replication (pSC101-ori TS ) MAGIC recipient cells also carry a genomically integrated, inducible I-SceI endonuclease allele. A different helper plasmid containing λ-red and Cas9 is required to perform repeatable cleavage and homologous recombination. To construct such a helper plasmid, pML300 was first digested and ligated into the multiple cloning site (MCS) via HindIII and NheI, resulting in pSL270. Using Gibson Assembly, araC-P araBaAD A single DNA fragment containing -Cas9 was constructed and then cloned into pSL270 using the PacI and XhoI restriction sites to create the helper plasmid pSL359. This plasmid expresses the rhamnose-inducible λ-red recombination system (P rhaBAD -red), arabinose-inducible endonuclease (P araBAD -Cas9), pSC101-ori ts, and a spectinomycin resistance marker (SpR). pSL359 was then transformed into BW28705 to create the recipient host cell BW28705 / pSL359. As an alternative recipient strain, RE1133 (Egbert et al., Nucleic Acids Research, vol. 47(6), April 8, 2019, pp. 3244-3256) [cmR::mutS pTet2-gam-bet-exo-dam / tetR::bioA / B ilvG+ dnaG.Q576A lacIQ1 Pcp8-araE ΔaraBAD pConst-araC ΔrecJ ΔxonA Pkm-cymR-Cas9::bioC] was used without the helper plasmid. RE1133 contains the tetracycline-inducible λ-red recombination system (pTet2-gam-bet-exo-dam / tetR::bioA / B) and the cumate-inducible Cas9 endonuclease (Pkm-cymR-Cas9::bioC).

[0252] Construction of swapping cassettes: A swapping cassette is defined as a stretch of DNA on the donor and recipient plasmids that contributes to the swap of DNA. The cassette on the recipient plasmid is replaced by the cassette originally found on the donor plasmid via homologous recombination. To select reproducibly for cassette swapping in vivo, each cassette is engineered to contain both a selectable marker and a counterselectable marker. A selectable marker in the donor cassette and a counterselectable marker in the recipient cassette are required for each round. To implement such a double selection strategy, two different selection cassettes were constructed by standard cloning methods from the following sources: 1) PheS Gly 294(D,Lp-Cl-Phe sensitive) (Kast, P. Gene 138, pp. 109-114 (1994)), 2) SacB (sucrose sensitive) (Pelicic, V. et al., J. Bacteriol. 178, pp. 1197-1199 (1996)), 3) HygR (hygromycin resistance) containing the EM7 bacterial promoter (Gritz, L. et al., Gene 25, pp. 179-188 (1983)), and 4) NsrR (nourseothricin resistance) (Gene 62, pp. 209-217 (1988)). One cassette contained PheS Gly 294 and NsrR, and a second cassette was constructed containing HygR and SacB. Several other selection cassettes were also constructed to carry out some experiments to characterize the in vivo suturing system: 5) ZeoR (zeocin resistance) (Drocourt, D. Nucleic Acids Res. 18, 4009-4009 (1990)), 6) ampR (ampicillin resistance), and 7) CmR (chloramphenicol resistance). In some cases, one cassette contained PheS GlyR and NsrR. 294 and NsrR, and a second cassette was constructed containing HygR and SacB.

[0253] Construction of backbones for donor and recipient vectors: The donor vector was constructed to contain the following key components: kanR, oriT, R6K oriγ, and a constitutive gRNA expression cassette driven by the strong bacterial promoter J23119 (Standage-Beier, K. et al., ACS Synth. Biol. 4, 1217-1225 (2015)). The swapping regions were reconstructed to generate the T1(F)-H1-T2(R)-T2(F)-H3-T1(R) fragment. where T1 (5'-GGGGCCACTAGGGACAGGATtgg-3' (SEQ ID NO:37) and T2 (5'-CAGGCGGGCTCACCTCCGTGtgg-3' (SEQ ID NO:38) are two unique target sequences for CRISPR-Cas9 cleavage, H1 (5'-CGAGGGCTAGAATTACCTACCGGCCTCCACCATGCCTGCG-3' (SEQ ID NO:39) and H3 (5'-GTACGGGCAACCCGAGAAGGCTGAGCCTGGACTCAACGGGTTGCTGGGTGGACTCCAGACTCGGGGCGACGACTCTTCACGCGCAGAGCAAGGGCGTCGAGCGGTCGTGAAAGTCTTAGTACCGCACGTGCCGACTCACTGGGATATTGCCTGGAGCTGTACCGTTCTAGGGGGGGAGGTTGGAGACCTCC TCTTCTCACGACTGGACCCGCGAGGGCCGCGTTGCCGGTTCCCCCAGAGGCTGAAGAACAAGGGCTTACTGTGGGCAGGGGGACGCCCATTCAGCGGCTGGCGCTTT-3' (SEQ ID NO: 40) are the homology sites for homologous recombination, and (F) and (R) indicate whether the DNA fragment is included in the forward or reverse (reverse complement) orientation. A selection cassette (HygR-SacB or NsrR-PheS) is inserted between T2(R) and T2(F) to generate a donor plasmid ready for stitching. In some cases, the swapping region of the donor plasmid was reconstituted to generate a T2(F)-T1(R)-T1(F)-H3-T2(R) fragment in which a selection cassette (HygR-SacB or NsrR-PheS) is inserted between T1(R) and T1(F).Both gRNA target sites are positioned in the appropriate orientation to ensure that the distance between the locus of the double-stranded break and the homologous region is as short as possible. H3 is a 300 bp synthetic DNA fragment that is used as one homologous arm in every round of assembly. H1 is the homologous region that is used in the first round of in vivo stitching and is either incorporated into the donor scaffold or introduced as part of the first oligonucleotide to be stitched. Other homologous regions (H2, H4, H5, etc.) are introduced into the donor plasmid as part of subsequent oligonucleotides and overlap with the homologous region of the previous oligonucleotide in the assembly to allow seamless stitching. The entry recipient vector contains a selectable marker (GmR) and an origin of replication (ColE1). The swapping region was modified to the H1-T1(R)-T1(F)-H3 configuration and a selection cassette (HygR-SacB or NsrR-PheS) was cloned between T1(R) and T1(F).

[0254] Testing the efficiency of endonuclease cleavage and homologous recombination: To test whether the CRISPR / Cas9 system provides precise DNA cleavage and promotes homologous recombination, the stitching procedure was completed in the presence or absence of target gRNA, Cas9, and λred. The recipient plasmid pSL402 was transformed into three different recipient host cells to construct 1) BW28705 / pSL402 (-λred / -Cas9), 2) BW28705 / pML300 / pSL402 (+λred / -Cas9), and 3) BW28705 / pSL359 / pSL402 (+λred / +Cas9). Two different donor plasmids, pSL414 and pSL415, were then transformed into BUN20 to create BUN20 / pSL414 and BUN20 / pSL415, which contain functional and mock gRNA units, respectively. Each of the two donor strains was mated with one of each of the three different types of recipient cells as described above. The cells were then diluted and plated onto selection plates (Cl-Phe+Gm+Cm+0.2% glucose) at 37° C. overnight to recover recombinant clones. Colonies were counted to quantitate recombination events.

[0255] Assembly of the mEGFP gene in liquid: Three fragments of the mEGFP gene were generated by PCR. The first fragment, containing the constitutive promoter pJ23100, the ribosome binding site, and nucleotides 1-251, was cloned into the donor backbone to create pSL485. The second fragment, containing nucleotides 198-517, was cloned into the donor vector to create pSL486. The third fragment, containing nucleotides 454-720 and the rrnB T1 terminator, was cloned into the donor vector to create pSL488. All three donor vectors were transformed into BUN20 and grown overnight on LB+Kan plates at 37°C. The entry recipient vector pSL398 was transformed into BW28705 / pSL359 and grown on LB agar plates containing gentamicin, spectinomycin, and glucose. Donor BUN20 / pSL485 (D1) and recipient BW28705 / pSL359 / pSL398 (R0) clones were grown overnight at 37°C and 30°C, respectively, in the appropriate liquid medium. Cells (1 ml) from both donor and recipient were then spun down, mixed, and resuspended in 1 ml pre-warmed LB+Ara+Rha liquid medium. After approximately 4 hours of incubation at 30°C without shaking, serial dilutions were performed in mating broth and cells were plated on 6% Suc+Carb+Gm+Sp+0.2% glucose to select for recombinants (R1). Colony PCR was performed to confirm the correct R1 clone, which was then incubated overnight at 30°C in LB+Gm+Sp+0.2% glucose. Freshly cultured donor cells (D2) containing pSL486 were then spun down, mixed, and resuspended with R1 in liquid mating medium at 30° C. for approximately 4 hours. Recombinant clones were selected by plating on Cl-Phe+Hyg+Gm+Sp+0.2% glucose. Correct clones (R2), confirmed by colony PCR, were grown overnight at 30° C. in LB+Gm+Sp+0.2% glucose. As in the previous round, D3 (BUN20 / pSL488) and R2 cells were mixed and resuspended in liquid mating medium. Serial dilutions were performed and cells were plated on 6% Suc+Carb+Gm+Sp+0.2% glucose.Plates from each round of assembly were imaged under UV light and the percentage of GFP fluorescent colonies were counted. After each round of assembly, selected clones were picked and plasmids were purified for diagnostic restriction digestion and Sanger sequencing.

[0256] Testing the effect of homology length on suture: To test the effect of homology length on suture accuracy, a series of plasmids were constructed to contain different sizes of homology to the second mEGFP fragment in pSL486: 1) pSL684 (0 bp), 2) pSL685 (10 bp), 3) pSL510 (20 bp), 4) pSL511 (30 bp), 5) pSL512 (40 bp), 6) pSL681 (53 bp). These plasmids and pSL488 (63 bp) were transformed into donor host cells to construct a group of D3 that mated with recipient cells containing R2. Cells from each mating pair were diluted and plated on selection plates (6% Suc+Carb+Gm+Sp+0.2% glucose). Colonies were allowed to recover overnight and examined under UV light to observe fluorescence. The numbers of fluorescent and non-fluorescent colonies were counted and the percentage of correct assembly was calculated.

[0257] Arrayed assembly of mEGFP: All strains were arrayed on agar plates in 384 format. First, BUN20 / pSL1065(D1) and BW28705 / pSL359 / pSL1060(R0) were arrayed and mixed together on pre-warmed mating plates (LB+Ara+Rha) using a SINGER ROTOR HDA and grown at 30°C for approximately 5 hours. The mated cells were then transferred to the first selection plate (Cl-Phe+Hyg+Gm+Sp+0.2% glucose). Recombinant clones (R1) were enriched overnight at 30°C and then transferred to a pre-mating plate (LB+Hyg+Gm+Sp+0.2% glucose) to optimize the growth of the assembled plasmid for the next round. A new overnight array of BUN20 / pSL1062(D2) was then mated with R1 on the mating plate. The mated cells were then transferred to the first selection (LB+Nat+Gm+Sp+0.2% glucose) to select for recombinant clones overnight at 30°C. The selected clones (R2) were then transferred to a pre-mating plate (6% Suc+Nat+Gm+Sp+0.2% glucose). A new overnight array of BUN20 / pSL1066(D3) was then mated with R2 on the mating plate. Following selection on Cl-Phe+Hyg+Gm+Sp+0.2% glucose and then LB+Hyg+Gm+Sp+0.2% glucose, the final assembly product (R3) was selected. Plates from each round of assembly were imaged on a UV transilluminator under UV light to monitor GFP fluorescence. Over the course of the assembly, selected clones were picked and plasmids were purified for diagnostic restriction digestion and Sanger sequencing.

[0258] Aligned assembly of 12 genes using pooled oligonucleotides: 9 distinct serine / tyrosine recombinases and 3 distinct fluorescent molecules (mPapaya, mPlum, sfGFP) were selected for assembly. To generate a list of oligonucleotides required for each gene assembly, a Python script was written that inputs a FASTA file containing the genes to be synthesized and outputs a list of oligonucleotides to be ordered from a commercial supplier (IDT oPool) with user-defined variables including the length of the synthesized oligonucleotides, the minimum homologous overlap length between adjacent oligos, and the maximum homologous overlap length. DNA hairpins and / or repeats may interfere with the homologous recombination machinery and reduce the fidelity of the assembly, although quantitative studies on this effect are unknown. Therefore, for each gene, these regions were identified using the Primer3 Python extension (Untergasser, A. et al., Nucleic Acids Res. 35, W71-W74 (2007)) and a nucleotide distribution uniformity metric. Each nucleotide position is scored, and if a small homology region is likely to contain an interfering element, the homology region is extended using a user-defined threshold. Once the necessary oligonucleotides for each gene assembly are determined, restriction sites (NotI and AscI) are added to each end, which are used to clone the oligonucleotides into a donor vector and round-specific priming sites. The priming sites allow oligonucleotides from a particular round of parallel gene assembly to be amplified together and parsed by an in vivo parsing platform. Round-specific primers are selected from a pre-designed primer list to reduce the possibility of cross-reactivity between primers (reducing the number of undesired PCR products) when used with a large oligonucleotide pool (Kosuri, S. et al., Nat. Biotechnol. 28, pp. 1295-1299 (2010)).Using this Python script, each gene was divided into five approximately 300 bp oligonucleotides with 50-70 bp of homology between successive oligonucleotides. The PCR amplified oligonucleotides were inserted into the donor plasmid using restriction digestion and ligation. For the first round of assembly of each gene, the oligonucleotides were amplified and cloned into the donor backbone pSL1064, which contains a 40 bp initial H1 region that is homologous to a region in the ingress recipient plasmid pSL1060. Oligonucleotides to be added for further odd-numbered (e.g. 3, 5, 7, 9) stitching rounds were cloned into pSL1063, which contains the same elements as pSL1064 except that it lacks the H1 homology region. Oligonucleotides to be added for even-numbered (e.g. 2, 4, 6, 8) stitching rounds were cloned into pSL1071. The PCR amplified oligonucleotides and donor plasmid were digested with AscI and NotI for 4 hours at 37°C. The digested products were then size selected and purified by gel extraction using a Zymoclean Gel DNA Recovery kit. 0.02 pmol of digested donor plasmid and 0.06 pmol of digested oligonucleotide were ligated by mixing with 1 μl of T4 ligase and incubating at 22°C for 1 hour. The ligated donor plasmid was transferred to the BUN20 donor strain using a standard bacterial transformation protocol. 2 μl of the ligation product was added to 50 μl of chemically competent BUN20 donor strain. The mixture was then incubated on ice for 30 minutes, heat shocked at 42°C for 30 seconds, and incubated again on ice for 3 minutes. The cells were resuspended in 950 ml of NEB SOC recovery medium and allowed to recover at 37°C for 1 hour. The cells were then plated on selection plates (LB+Hyg for odd-round donor plasmids and LB+Nat for even-round donor plasmids) and incubated at 37°C overnight. The resulting colonies containing the cloned oligonucleotides were randomly selected and arrayed into 96-well plates.The arrayed oligonucleotide libraries were parsed and sequences were verified using an in vivo DNA parsing system. Each plate was mated with two different recipient barcode plates. For the odd assembly round oligonucleotide plates, the donor backbone (pSL1063 or pSL1064) contained the HygR-SacB cassette. Following mating with the BPS recipient array on LB+Ara+IPTG agar at 37°C for approximately 3 hours, recombinant plasmids were selected on LB+Hyg+Gm+Rha+Ara plates at 37°C overnight. For the even round oligonucleotide arrays containing the NsrR-PheS cassette, following mating with the BPS collection, recombinant plasmids were selected on LB+Nat+Gm+Rha+Ara plates at 37°C overnight. To assemble the 12 genes, the BUN20 donor strain carrying the first round oligonucleotides was first mated with the RE1133 recipient strain carrying the pSL1086 recipient plasmid. 50 μl of each overnight culture was mixed, spun down at 8000 rpm for 1 min, resuspended in 50 μl LB, and incubated at 37° C. for 30 min. The mated cells were then plated on LB+aTC+cumate plates and incubated at 37° C. for 4 h to induce Cas9 and λ-red. To isolate cells carrying the recombinant recipient plasmid with the first round oligonucleotides, the mated cells were streaked onto LB+Hyg+Gm+IPTG agar plates and incubated at 37° C. overnight. The R6Kγ origin on the donor plasmid is pir. +Unrecombined donor plasmids are rapidly removed as they are non-functional in the background of the RE1133 recipient strain. Colonies from the selection plates were further purified by selecting on LB+Hyg+Gm+IPTG+4CP to remove any remaining unrecombined pSL1086 recipient plasmid. To assemble the second round of oligonucleotides, the purified colonies carrying the recombinant recipient plasmid along with the first round of oligonucleotides were then mated with the BUN20 donor strain carrying the second round of oligonucleotides. The same procedure was used as for the first assembly, except that cells carrying the recombinant recipient plasmid were selected on LB+Nat+Gm+IPTG and purified on LB+Nat+Gm+IPTG+6% sucrose. The same assembly procedure was used for all subsequent assembly steps, with LB+Hyg+Gm+IPTG+6% sucrose for selection of odd round assemblies and LB+Nat+Gm+IPTG for even round assemblies. This process was repeated five times until all 12 genes were fully assembled (Figure 7G). The sequence of the assembled product was verified to be the correct sequence using Sanger sequencing (Figure 7H) and an Oxford Nanopore MinION sequencer.

[0259] Assembly of the 9 kb fragment: A 9 kb DNA block from positions 41489-50489 of chromosome II of the BY4741 Saccharomyces cerevisiae strain was assembled. Using the same Python script used for the assembly of the 12 genes, the 9 kb block was divided into three approximately 3 kb DNA blocks with 50-75 bp of homology between successive DNA blocks. These three DNA blocks were PCR amplified using genomic DNA from yeast strain BY4741 as a DNA template. Genomic DNA was extracted using the MasterPure Yeast DNA Purification kit. The first, second, and third DNA blocks were inserted into donor plasmids pSL1064, pSL1063, and pSL1107, respectively, using AscI / NotI restriction digestion and T4 ligation. The resulting ligated products were transformed into the BUN20 donor strain using standard bacterial transformation procedures. After sequence verification of the donor plasmid using Sanger sequencing, the DNA blocks were assembled in the RE1133 / pSL1086 recipient strain. The donor strain carrying the first DNA block was mated with RE1133 / pSL1086 and grown on LB+aTC+cumate at 37°C for 4 hours. The cells carrying the recombinant recipient plasmid were selected on LB+Hyg+Gm+IPTG and further purified on LB+Hyg+Gm+IPTG+6% sucrose. The resulting colonies were mated with the donor strain carrying the second DNA block, selected for the recombinant recipient plasmid on LB+Nat+Gm+IPTG and purified on LB+Nat+Gm+IPTG+4CP. Finally, the resulting colonies were mated with the donor strain carrying the third DNA block, selected for the recombinant recipient plasmid on LB+Hyg+Gm+IPTG and purified on LB+Hyg+Gm+IPTG+6% sucrose. The sequence of the assembled product was then verified to be the expected sequence using an Oxford Nanopore MinION sequencer and gel electrophoresis (Figure 7I).

[0260] Amplicon sequencing to analyze oligonucleotide libraries: To extract recombinant plasmids, cells were scraped from the selection plates and minipreps were performed using the Plasmid Plus Mini Kit (QIAGEN). Plasmid DNA was then quantified and diluted to approximately 1ng / μl. This is approximately 1.5e6 copies per unique barcode pair on a 96-array mating plate. A two-step PCR as described (Levy, SF et al., Nature 519, pp. 181-186 (2015)) was performed with modifications. First, 4-5 cycles of PCR were performed with OneTaq polymerase (New England Biolabs) using the forward (pBPS_fwr) and reverse (pBPS_rev) primers listed in Table 1. Approximately 1ng of recombinant plasmid DNA was amplified in a single 50 μl PCR reaction. The primers for the first step PCR have the following general configuration: pBPS_fwr: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNNXXXXXXttcggttagagcggatgtg (SEQ ID NO: 41) pBPS_rev: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCNNNNNNNNXXXXXXXXXaggtaacccatatgcatggc (SEQ ID NO: 42)

[0261] N in these sequences corresponds to any random nucleotide and is used in downstream analysis to remove asymmetries in counts caused by PCR jackpotting. X corresponds to one of several multiplexing tags (e.g., the multiplexing tags in Table 1 above), allowing different samples to be distinguished when loaded onto the same sequencing flow cell. Examples of multiplexing tags are the underlined sequences in Table 1. Lowercase sequences correspond to priming sites on the recombinant plasmid. Uppercase sequences correspond to Illumina Read 1 or Read 2 sequencing primers. PCR products were purified using NucleoSpin columns (Macherey-Nagel) and eluted in 33 μl of water. A second 23-25 ​​cycle PCR was performed with PrimeStar HS polymerase (Takara) with 33 μl of purified product from the first PCR as template and a total volume of 50 μl per tube. Primers for this reaction were standard Illumina TruSeq dual-index primers (D501-D508 and D701-D712) listed in Table 2. The PCR products were then purified using NucleoSpin columns. Amplicons from each mating plate were uniquely labeled with custom multiplexing tags as well as Illumina's standard indexes. This quadruple-index strategy not only increases the multiplexing capacity of the sequencing library but is also beneficial for downstream analysis for amplicon chimeras. The purified amplicons were pooled and paired-end sequenced on an Illumina MiSeq (2 × 300 bp) with 25% PhiX DNA spike-in. Sequencing reads were integrated into barcodes by using Bartender (Zhao, L. et al., Bioinformatics 34, 739-747 (2017)).

[0262] Design of a repeatable in vivo suturing technique: The in vivo suturing system exploits the bacterial conjugation machinery and lambda-Red homologous recombination. The donor vector contains a conditional origin of replication derived from R6K oriγ, which depends on a functional trans-acting factor π encoded by the gene pir1 or its relaxed copy number controlled version, pir1-116 (Metcalf, W. Gene 138, pp. 1-7 (1994)). This special origin allows the plasmid to be maintained in donor host cells (e.g., BUN20) that have a genomically integrated pir116 allele, but not in recipient cells that lack this allele. Other important features in the donor vector include oriT, a backbone marker (kanR), a constitutive gRNA expression cassette (gRNA T1 or gRNA T2 Within the swapping region are two pairs of unique gRNA target sites, a dual selectable cassette, and a common long homology sequence (300 bp) for each round of stitching.

[0263] The entry recipient vector contains an origin of replication (ColE1 or pLacIQ-p15A), a backbone marker (GmR), and a swapping region, which consists of homology sequences, two gRNA target sites, and a dual selectable cassette. Some recipient host cells harbor a helper plasmid containing the rhamnose-inducible λ-red recombination system and the arabinose-inducible Cas9 endonuclease. A temperature-sensitive mutant derivative of the origin of replication (pSC101-ori TS ) provides a convenient means of curing the helper plasmid at 42°C once assembly is complete (Hashimoto, TJ Biol. 127, pp. 1561-1563 (1976)). Another recipient host cell (RE1133) contains a genetically integrated tetracycline-inducible λ-red recombination system and a cumate-inducible Cas9 endonuclease.

[0264] In each round, after the donor vector is transferred into the recipient cell, both λ-red and Cas9 are induced in the presence of arabinose and rhamnose or cumate and tetracycline depending on the recipient cell. T1 guides Cas9 into both the donor and recipient plasmids to generate double-stranded breaks. DNA cleavage has been shown to greatly stimulate homologous recombination (see below) (Kuzminov, A. Microbiol. Mol. Biol. Rev. 63, 751 (1999)). The donor-derived fragment contains the DNA of interest, two different gRNA target sites for the next round of assembly, a double selectable marker, and an H3 homology region. The DNA for assembly is designed such that the first 50 bp is homologous to the last 50 bp of the assembled sequence present on the recipient plasmid. This 50 bp serves as one homology arm for double crossover homologous recombination, the other is H3. This swapping event results in a recombined recipient plasmid containing new DNA from the donor. Three different selections are performed to ensure the accuracy of the recombination. 1) Selection against a counterselectable marker (PheS or SacB), 2) Selection for a positive selectable marker (HygR or NsrR), and 3) Selection for a marker on the recipient backbone (GmR). Selection for the helper plasmid and P araBAD and P rhaBAD Promoter or P km-cymR and P Tet2 Promoter suppression is also implemented to prevent excessive recombination and unwanted DNA cleavage. Alternating between two different gRNAs and two different dual-selectable cassettes allows repeatable in vivo stitching to efficiently assemble new DNA fragments in a linear fashion, resulting in the desired sequence with the only theoretical limit being the tolerable plasmid size. Liquid handling and multiplexed pinning robots make this platform highly scalable, allowing thousands of parallel gene assemblies per round.

[0265] CRISPR-Cas9 can efficiently stimulate in vivo suture: To test whether the CRISPR / Cas9 system provides precise DNA cleavage and promotes homologous recombination, in vivo suture was completed in the presence or absence of targeting gRNA, Cas9, and λ-red. Recombinant plasmids were recovered in the presence of all three (Figure 3), indicating that CRISPR / Cas9 facilitates double-stranded DNA breaks to enhance the efficiency of λ-red recombination.

[0266] Assembly of a functional fluorescent gene in liquid: To demonstrate the ability to assemble multiple fragments into a functional gene using the in vivo suture system, three pieces of the mEGFP gene were constructed by PCR and cloned into an appropriate donor backbone. The three fragments were then assembled sequentially into the ingress recipient vector. After each round and before the next round, recovered clones were examined by both restriction digestion and Sanger sequencing to verify assembly accuracy. To ensure that the helper plasmid was retained in each round, both colony touch PCR and plasmid extraction were performed. Green fluorescence was observed in all colonies (approximately 300) on the selection plate in the final round of assembly, indicating high fidelity of assembly (Figures 7A-7I).

[0267] Effect of homology length on suture fidelity: To test how in vivo suture fidelity depends on the length of homology between the fragments, seven donor vectors were constructed that contain the third fragment of the mEGFP assembly with different lengths of homology with the second fragment. After conjugation and recombination, cells derived from different conjugation / recombination events were plated on selective medium and the fraction of fluorescent colonies was counted. The results show that in this example, perhaps 40 bp of homology produces an error-free fusion product (Figures 7A-7I).

[0268] Multiplexed assembly of functional fluorescent genes on agar: To demonstrate the ability to assemble functional genes on agar plates, three pieces of the mEGFP gene were constructed by PCR and cloned into the appropriate donor backbone. The three fragments were then assembled sequentially into the ingress recipient vector in either a 96-pin or 384-pin format. After the final round of assembly, green fluorescence was observed at positions 96 / 96 in the 96-pin format and 383 / 384 in the 384-pin format, indicating that the suture fidelity on agar was comparable to that in liquid (Figures 7A-7I). Over the course of the assembly, plasmids were recovered from different colonies (whole colonies were scraped) and examined by restriction digestion to verify the accuracy of the assembly and / or retention of the helper plasmid. Typical digestion patterns of the final assembly products indicated high purity recombinant plasmids with no observable undesired products (e.g., non-recombinant plasmids). To further characterize suture fidelity, 96 positions from the 384-position assembly were sequenced by Sanger sequencing. Sequencing products were derived from colony PCR of pipette tips contacted with each colony. 94 / 96 colonies were found to contain the correct mEGFP sequence. One colony contained an intermediate product (the product of the first round of assembly) and one colony contained a suture error (a large deletion).

[0269] Multiplexed assembly of genes from oligonucleotide pools: To demonstrate the ability to assemble various DNA constructs from oligonucleotide pools, nine distinct serine / tyrosine recombinases and three distinct fluorescent molecules (mPapaya, mPlum, sfGFP) were constructed from a pool of 300 bp oligonucleotides purchased from IDT (oPools). Each gene was assembled sequentially from five oligonucleotides stitched together. The pool of oligonucleotides was integrated into a matched donor plasmid, which transformed donor bacteria, and the bacterial pool was parsed in a sequence-verified ordered array using the method described in Example 2 (Methods for in vivo DNA analysis). The donor cell arrays were spliced ​​sequentially into recipient cells, and each gene was assembled in at least three replicates. Each assembly was determined to contain the correct DNA sequence by Sanger sequencing (Figure 7G).

[0270] Assembly of long DNA: To demonstrate the ability to assemble and produce long assemblies using long DNA blocks, we reconstructed a 9 kb segment of the Saccharomyces cerevisiae genome from three 3 kb blocks. The 3 kb blocks were amplified from genomic DNA and integrated into a matched donor plasmid, which was then transformed into donor bacteria and sequence verified. The donor cells were sequentially mated to recipient cells. To verify the correct assembly products at each assembly step, the recombinant recipient plasmids were purified from the recipient cells, linearized with restriction enzymes, and analyzed by gel electrophoresis (Figure 7H). The assembly products were also verified to be sequence correct by Sanger sequencing.

[0271] [Example 2] Method for in vivo DNA analysis Plasmid sequences: Information regarding the plasmids used in the in vivo DNA analysis can be found in Table 5 and FIG.

[0272] [Table 5]

[0273] Construction of barcoded donor plasmid: The donor vector was constructed using standard cloning methods. It contains 1) KanR (kanamycin resistance), 2) oriT (origin of transfer), 3) R6K oriγ (conditional origin of replication dependent on phage derived pir1 expression), and 4) swapping region, I-SceI-H1-H4-I-SceI construct, where I-SceI is the recognition site for endonuclease SceI, and H1 (5'-ttgccctctcttcattcagggtcatgagaggcacgccattcaaggggagaagtgagatc-3' (SEQ ID NO: 43)) and H4 (5'-aagaacttttctatttctgggtaggcatcatcaggagcagga-3' (SEQ ID NO: 44)) are the homology regions for recombination. A selection cassette (HygR-SacB or NsrR-PheS) was cloned between H1 and H4 in the swapping region of the donor vector to generate donor backbone plasmids for analysis. To insert random barcodes into the donor backbone (pSL438 and pSL439), an oligonucleotide (pXL633) containing a NotI restriction site, a barcode region containing random 15 nucleotides, and a region of homology to both donor backbones was ordered from IDT. PCR of the barcodes was performed using pXL633 paired with pXL585 and approximately 1 ng of pSL438 or pSL439 as template. The resulting PCR product was restriction digested and ligated into the corresponding donor vector via the NotI and XmaI sites. Following the same cloning protocol described above, competent donor cells BUN20 were transformed with the ligation products and barcoded donor clones were selected on LB agar plates containing 50 μg / ml kanamycin (Kan) at 37° C. Transformants were then randomly selected and arrayed to generate two barcoded 96-well donor collections, pSL438_BC and pSL439_BC. To identify the barcode sequences in the arrayed donor collections, the regions containing the barcodes were amplified by colony touch PCR using pXL583 and pXL584 as primers.The amplicons were then purified and Sanger sequenced using pXL583.Barcodes were then extracted and two lists of known donor barcode collections were compiled.

[0274] Construction of barcoded recipient plasmids: Plasmid pSL937, used as a backbone to insert and align random barcodes to generate a barcoded recipient collection, was constructed by standard methods from the following sources: 1) pBR322-derived plasmid backbone / origin of replication, 2) pUC18-mini-Tn7T-Gm 3 3) the homologous sequences H1 and H4, and the two I-SceI recognition sites in the H1-I-SceI-I-SceI-H4 construct; and 4) pSLC-217. 4 Rhamnose-inducible toxin relE (P rhaBAD -relE) was cloned between two SceI sites. Oligonucleotides containing random barcodes were synthesized by IDT and inserted into pSL937 by restriction digestion and ligation.

[0275] To insert the random barcode into the recipient backbone (pSL937), an oligonucleotide (pXL631) was ordered from IDT that contained a XhoI restriction site, a barcode region containing 20 random nucleotides, and a region of homology to pSL937. Barcodes were generated by PCR using pXL631 paired with pXL154 and approximately 1 ng of pSL937 as template. The resulting PCR product was digested and ligated into pSL937 using the MluI and XhoI restriction sites. The ligation reaction was performed overnight at 16°C with a molar ratio of barcode insert to vector of 3:1. The ligation product was then transformed into the spectinomycin resistance helper plasmid pML104. 1Competent BUN21 cells containing the ribosomal DNA were transformed with the ribosomal DNA of interest. Barcoded recipient clones were selected on LB agar plates containing 50 μg / ml spectinomycin (Sp), 20 μg / ml gentamicin (Gm), and 2% glucose at 30°C. Transformants were then randomly selected and arrayed into 96-well plates. The barcode sequence at each position in the arrayed recipient collection was identified by sequencing. A total of 841 barcodes could be unambiguously identified. These barcodes were then re-arrayed into eight new 96-well plates so that a unique barcode was present at each position.

[0276] Arrayed mating: Each barcoded donor plate (two 96 position plates) was mated with each barcoded recipient plate (eight 96 position plates). The donor barcode collection was grown overnight at 37°C on LB+Kan plates. The recipient array was grown overnight at 30°C on LB+Sp+Gm+2% glucose. Agar medium for arrayed mating contained 0.2% arabinose (Ara) and 0.1 mM IPTG and was pre-warmed at 37°C for 1 hour. Both donor and recipient clones were transferred to the mating plate using a SINGER ROTOR HDA pin pad and grown at 37°C for approximately 3 hours. Each recipient plate was mated with two donor barcoded plates (pSL438_BC and pSL439_BC). The mated cells were then transferred onto selective LB plates containing 0.2% arabinose, 0.2% rhamnose (Rha), 25 μg / ml gentamicin, and 50 μg / ml hygromycin (Hyg) (LB+Ara+Rha+Gm+Hyg). Recombinant clones were then selected overnight at 37°C.

[0277] Amplicon sequencing: To extract recombinant plasmids, cells were scraped from the selection plates and miniprepped using the Plasmid Plus Mini Kit (QIAGEN). Plasmid DNA was quantified and diluted to approximately 1 ng / μl, which represents approximately 1.5 × 10 of each unique barcode pair per 96-array plate. 6 copies. A two-step PCR was performed. First, 4-5 cycles of PCR with OneTaq polymerase (New England Biolabs) were performed using the forward (pBPS_fwr) and reverse (pBPS_rev) primers listed in Table 1. Approximately 1 ng of recombinant plasmid DNA was amplified in a single 50 μl PCR reaction. To increase the multiplexing of sequencing samples, unique pairs of primers for the first and second PCR (see Tables 1 and 2) were used to amplify plasmid DNA from specific pairs of mated plates. This allows multiple mated plates to be pooled together in one sequencing library. The cycling conditions for the first step are shown below in Table 6.

[0278] [Table 6]

[0279] The primers for the first step PCR have this general structure:

[0280] pBPS_fwr: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNNXXXXXXttcggttagagcggatgtg (SEQ ID NO: 45) pBPS_rev:GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCNNNNNNNNXXXXXXXXXaggtaacccatatgcatggc (SEQ ID NO: 46)

[0281] The N in these sequences corresponds to any random nucleotide and is used in downstream analysis to remove asymmetries in the counts caused by PCR jackpotting. The X corresponds to one of several multiplexing tags, allowing different samples to be distinguished when loaded onto the same sequencing flow cell. The sequences in lowercase correspond to priming sites on the recombinant plasmid. The sequences in uppercase correspond to Illumina Read 1 or Read 2 sequencing primers. PCR products were purified using NucleoSpin columns (Macherey-Nagel) and eluted in 33 μl of water. A second 23-25 ​​cycle PCR was performed with PrimeStar HS polymerase (Takara) with 33 μl of purified product from the first PCR as template and a total volume of 50 μl per tube. Primers for this reaction were standard Illumina TruSeq dual index primers (D501-D508 and D701-D712) listed in Tables 1 and 2. The cycling conditions for the second step are in Table 7 below.

[0282] [Table 7]

[0283] The PCR products were then purified using NucleoSpin columns. Amplicons from each mating plate were uniquely labeled with a custom primer index (first PCR) as well as a standard Illumina index (second PCR). This quadruple index strategy increases the multiplexing capacity for sequencing. Purified amplicons were pooled and subjected to paired-end sequencing on an Illumina MiSeq, HiSeq, or NextSeq with 25% PhiX genomic DNA spike-in at approximately 800 reads per barcode pair.

[0284] Sequencing analysis: Donor-recipient dual barcoded amplicon sequencing data were analyzed using the following steps with custom Python scripts and bartender. First, Illumina reads were demultiplexed using the Illumina index. Any sequences that did not match exactly in the two Illumina indexes were discarded. Normal expression was extracted from the demultiplexed sequences. "\D * ?(.GGC|T.GC|TG.C|TGG.)\D{4,7}?AA\D{4,7}?TT\D{4,7}?(.CGG|G.GG|GC.G|GCG.)\D * " (donor barcode) and "\D * ?(.ACA|G.CA|GA.A|GAC.)\D{4,7}?AA\D{4,7}?AA\D{4,7}?TT\D{4,7}?(.TCG|C.CG|CT.G|CTC.)\D * " (Recipient barcode) Barcodes were extracted using Illumina. Unique molecular identifiers (UMI, N in pBPS_fwr and pBPS_rev) were also extracted based on their expected location in the Illumina reads. Barcode reads, which contained a mixture of true barcode sequences and sequences containing errors due to PCR or sequencing, were then clustered into consensus sequences using Bartender. Each barcode cluster was then examined for duplicate UMIs (indicating PCR duplicates) using Bartender, and all duplicates were removed to generate a final count for each barcode pair. Duplicate barcodes with fewer than 20 reads, many of which were expected to be PCR chimeras (barcodes fused by PCR amplification), were excluded. The remaining reads were used to resolve the location of each donor barcode from its corresponding recipient barcode.

[0285] Whole plasmid sequencing on the Oxford Nanopore platform: Recombinant plasmids containing positioning barcodes and oligonucleotides were extracted as already described in the amplicon sequencing section. Circular plasmids were linearized with the restriction enzyme PmlI (NEB) at 37°C for 2 hours. Linearized products were size-selected by running on a 1.2% agarose gel and recovered with the Zymoclean Gel DNA Recovery Kit (Zymoresearch). Sequencing libraries for the Oxford Nanopore platform were constructed using the Ligation Sequencing Kit (SQK-LSK110, Nanoporetech). 300ng (~100fmol) of the linearized recombinant plasmid library was end-repaired using NEBNext FFPE Repair Mix and NEBNext Ultra II End repair / dA-tailing Module (NEB). Nanopore sequencing adapters (AMX-F) were ligated with NEBNext Quick T4 DNA Ligase (NEB). 30 ng (~10 fmol) of the library was loaded onto a Flongle flow cell (R9.4.1, Oxford Nanopore) to generate reads for the recombinant plasmids. The flow cell was run for 16 hours using Miniknow sequencer control software (version 21.11.7, Oxford Nanopore).

[0286] Oxford Nanopore sequencing analysis: (2) Sequencing adapters were identified and removed, and files were separated by sample multiplexing barcode using "guppy_barcoder" from Guppy version 6.0.1+652ffd179. (3) Alignments were used to query contaminating sequences (sequences of origins of replication or transfer for donor and / or helper plasmids) using "minimap2" version 2.22-r1110-dirty, and unwanted sequences that aligned with these contaminants were removed. (4) Alignments to the expected backbone sequence of the recombinant recipient plasmid were generated using "minimap2" version 2.22-r1110-dirty, and identical backbone sequences were removed from each read using custom-written Python scripts. (5) Position-specific barcodes were extracted from this sequence using ambiguous regular expressions for the sequences surrounding the barcodes using "itermae" version 0.6.0.1 and aggregated using the message-passing Levenshtein distance approach in "starcode" version 1.4. (6) Using these barcodes, the plasmid backbone-removed sequences were separated into separate files for each demultiplexed sample and aggregated barcode sequences using custom shell / oak scripts. (7) The sequences for each barcode in each sample were used to generate a multiple sequence alignment using "kalign3" version 3.3.1. (8) A single draft consensus sequence was generated from the multiple sequence alignment by a voting process using custom Python scripts. (9) This draft consensus sequence was refined using "racon" version 1.5.0 to update the consensus based on read agreement and sequence quality. (10) This consensus sequence was further refined using “medaka” version 1.5.0 to generate a refined sequence for each localization barcode in each sample. (11) Payload sequences, the intended targets of the realignment project, were extracted from the refined regions using “itermae” version 0.6.0.1 again containing different normal representations.The refined and localized payloads were analyzed by aligning the raw regions with the scaffold removed with the refined regions, aligning the payload extracted from the refined regions with the intended target sequences, and aligning the raw regions with the scaffold removed with all refined regions generated in the dataset. Alignments were performed with "minimap" version 2.22-r1110-dirty or a custom Python script using BioPython PairwiseAlignment functionality. A custom R script was used to identify reads as "on-target", i.e., reads with greater than 90% identity to the refined regions generated for that sample and barcode (i.e., well). A well was classified as "pure" if more than 90% of the raw reads were "on-target" with the refined consensus sequence. Sequences were compared to the length of the refined payload and alignment with the intended target sequences, and were defined as "correct" if the refined payload was completely identical to one of the intended target sequences.

[0287] Construction of a donor plasmid library containing oligonucleotide pools: Plasmid pSL1071, containing an NsrR-PheS cassette, two I-SceI sites, and two homology regions for recombination (H1 and H4), was used as a backbone for inserting the oligonucleotide pools into it. The oligonucleotide pools, containing one 300 bp oligonucleotide, were designed: GCTTATTCGTGCCGTGTTATGGCGCGCCNN···NNGCGGCCGCGGGCACAGCAATCAAAAGTA (SEQ ID NO: 47) I ordered from IDT according to the following: GCTTATTCGTGCCGTGTTAT and GGGCACAGCAATCAAAAGTA (SEQ ID NO: 48) are priming sites for the forward and reverse primers to amplify the oligonucleotide pool, GGCGCGCC (SEQ ID NO: 49) and GCGGCCGC (SEQ ID NO: 50) are recognition sites for the restriction enzymes AscI and NotI, and NN...NN represents a 244 nt sequence randomly selected from the human genome assembly GRCh38. Amplification of the oligonucleotide pool was performed using 7 ng of template DNA and KAPA HiFi polymerase (Roche) with the cycling conditions described in Table 8.

[0288] [Table 8]

[0289] The PCR products were purified using DNA Clean & Concentrator-5 (Zymoresearch). AscI and NotI restriction sites were used to clone the PCR products into the donor plasmid pSL1071. The digestion reaction of the PCR products with pSL1071 was carried out at 37°C for 4 hours. The digested products were then size-selected by running on a 1.2% agarose gel and recovered using a Zymoclean Gel DNA Recovery Kit (Zymoresearch). The ligation reaction was carried out at 16°C for 15 hours using T4 DNA ligase (NEB) with 25 ng of digested vector and 3.8 ng of insert. The ligation products were transformed into BUN20 and joined with an array of barcoded recipient plasmids (above) and the constructs were sequenced at each position in the donor array.

[0290] Results: Barcode array localization. To validate the parsing and localization accuracy, two 96-well plates of known donor barcodes (pSL438_BC and pSL439_BC) were each mated with eight 96-well recipient plates. Data from these 1536 mating events found that the correct location was identified and sequence verified for 93.82% ± 0.34%, 95.59% ± 0.27%, and 96.04% ± 0.21% of donors using one, two, and three events, respectively (Figures 47A-47D). All failures were due to missing sequencing data. Incorrect locations were never identified for donors in the sequencing data. Similar results were found when the location of recipient barcodes was determined from the donor barcodes.

[0291] Results: Analysis of oligonucleotide pools To further validate the accuracy of the parsing, a pool of 100 oligonucleotides was aligned and sequence verified. This pool contained 244 nucleotide sequences randomly selected from the human genome, synthesized by IDT (Integrated DNA Technologies) as "oPool", and inserted into our donor plasmid pSL1071 by ligation. BUN20 transformants carrying these plasmids were pooled and then randomly aligned in a total of 20 384-well plates. These aligned bacterial plates were then mated with an aligned collection of recipient barcode strains (barcode positions known). Recipient cells containing the recombinant oligonucleotide-barcode plasmids were pooled. Plasmids were sequenced by Nanopore sequencing. The sequencing results were used to determine for each well in each plate the consensus sequence of the oligonucleotide, whether the consensus sequence was identical to the expected sequence in the oligonucleotide pool, and whether any other oligonucleotide sequences were present at low frequency (contaminants) (Figure 47G, Figure 47C, Figure 47D). Consensus sequences were generated in 5,101 (66.4%) of the 7,680 wells available across all plates. Of these consensus sequence wells, 2,329 (45.6%) were pure and had a perfect match to the target oligo. These 2,329 perfect match oligos represented 82% of the oligonucleotides predicted to be present in the pool.

Claims

1. 1. A method of assembling a plurality of DNA elements into an assembled DNA element in a recipient cell, comprising: (a) contacting a first donor cell containing a first donor plasmid with a recipient cell containing a recipient oligonucleotide under conditions to transfer the first donor plasmid from the first donor cell to the recipient cell by conjugation, wherein the recipient oligonucleotide is present in the recipient cell plasmid or in the recipient cell genome; a first donor plasmid comprising, in sequential order, a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a fragment of a first DNA element (oligo 1), a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2), and a third endonuclease site (C3), wherein HR2.1 and HR2.2 are flanked by a non-homologous region (NHR) comprising two endonuclease sites (C2.1, C2.2) flanking a selectable marker; the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2, and HR3 and HR4 are flanked by non-homologous regions containing two endonuclease sites (C4.1, C4.2) flanking a selectable marker; (b) subjecting the first donor plasmid and the recipient oligonucleotide to endonucleolytic cleavage with a first endonuclease to cleave the first donor plasmid at C1 and C3 and the recipient oligonucleotide at C4.1 and C4.2, thereby producing a first donor cassette comprising homologous recombination regions HR1 and HR2.2 at each end and a recipient oligonucleotide having compatible recombination regions HR3 and HR4; (c) subjecting the first donor cassette and recipient oligonucleotide to conditions that recombine the first donor cassette into the recipient oligonucleotide by homologous recombination, thereby providing a first recombined recipient oligonucleotide that comprises a fragment of the first DNA element following homologous recombination of HR1 with HR3 and HR2.2 with HR4; (d)(i) contacting a second donor cell containing the second donor plasmid with a recipient cell containing the first recombined recipient oligonucleotide under conditions to transfer the second donor plasmid from the second donor cell to the first recipient cell by conjugation; a second donor plasmid comprising, in sequential order, a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a fragment of a second DNA element (oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2), and a sixth endonuclease site (C6), wherein HR6.1 and HR6.2 are flanked by non-homologous regions (NHR) comprising two endonuclease sites (C7.1, C7.2) flanking a selectable marker; (e) subjecting the second donor plasmid and the recipient oligonucleotide to endonucleolytic cleavage with a second endonuclease to cleave the second donor plasmid at C5 and C6 and the recipient oligonucleotide at C2.1 and C2.2, thereby producing a second donor cassette comprising homologous recombination regions HR5 and HR6.2 at each end and a recipient oligonucleotide with matching recombination regions HR2.1 and HR4; (f) subjecting the second donor cassette and the recombined recipient oligonucleotide to conditions that allow the first donor fragment and the recombined recipient oligonucleotide to recombine in the recipient cell by homologous recombination, thereby providing a second recombined recipient oligonucleotide comprising fragments of the first and second DNA elements (oligo 1, oligo 2), following homologous recombination of HR5 with HR2.1 and HR6.2 with HR4, which form a DNA assembly; each donor plasmid contains an origin of transfer (oriT) and a conditional origin of replication; An oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third, and / or fourth endonuclease site is present on the first donor plasmid and / or is present in the recipient cell; and The method, wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth, and / or sixth endonuclease site is present on a second donor plasmid and / or is present in the recipient cell.

2. 2. The method of claim 1, wherein steps (d) through (f) are repeated one or more iterations using a third or subsequent donor cell containing a third or subsequent donor plasmid comprising a matching HR region and a third or subsequent oligonucleotide encoding a fragment of a third or subsequent DNA element (oligo 3, oligo 4, ... oligo N), thereby forming third or subsequent recombined recipient oligonucleotides comprising fragments of the first, second, and third or subsequent DNA elements, which together form a DNA assembly.

3. 3. The method of claim 1 or 2, wherein step (a) comprises a plurality of first donor cells each containing a different first donor plasmid, and step (b) comprises a plurality of second, third, or subsequent donor cells each containing a different second, third, or subsequent donor plasmid.

4. 4. The method of claim 1, wherein expression of the first and / or second endonuclease is inducible, and further comprising the step of inducing expression of the first and / or second endonuclease.

5. 5. The method of claim 4, wherein the first and / or second endonuclease is selected from an RNA-guided endonuclease, a homing endonuclease, a transcription activator-like effector nuclease, and a zinc finger nuclease.

6. 6. The method of any one of claims 1 to 5, wherein the first, second, or subsequent donor plasmid comprises a selectable marker that selects for integration of the first oligonucleotide, the second oligonucleotide, or the subsequent oligonucleotide into the recipient oligonucleotide.

7. 7. The method of any one of claims 1 to 6, wherein the recipient oligonucleotide comprises a counter-selectable marker that selects against recipient cells that do not contain the first, second, third, or subsequent oligonucleotide.

8. The method of any one of claims 1 to 7, wherein the donor plasmid or the recipient oligonucleotide comprises an inducible high-copy origin of replication.

9. 9. The method of any one of claims 1 to 8, wherein the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid greater than 30 kilobases in length.

10. The method of any one of claims 1 to 9, wherein the donor plasmid or the recipient cell comprises oligonucleotides encoding one or more homologous DNA repair genes.

11. The method of any one of claims 1 to 10, wherein the donor plasmid or the recipient cell comprises oligonucleotides encoding one or more recombination-mediated genetically engineered genes.

12. 12. The method of any one of claims 1 to 11, wherein the first, second or subsequent homologous recombination (HR) region and their corresponding HR regions on the recipient oligonucleotide comprise 20 base pairs to 500 base pairs, or 20 base pairs to 60 base pairs, respectively.

13. The method of any one of claims 1 to 12, comprising constructing a DNA library using two or more recipient oligonucleotides with compatible homologous recombination regions.

14. 14. The method of claim 13, used to combine genetic regions such as genes, promoters, terminators, and control regions from different species, to construct and / or combine gene regulatory pathways, or to assemble arrays of bacteria containing plasmids for screening assays.

15. 15. The method of any one of claims 1 to 14, wherein prior to steps (a) and (b) of claim 1, first and second oligonucleotides comprising fragments of the first and second DNA elements are inserted into first and second donor plasmids.

16. 16. The method of any one of claims 1 to 15, wherein the DNA assembly comprises at least a portion of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, a guide RNA (gRNA), or a combination thereof.

17. The method of any one of claims 1 to 16, wherein the DNA assembly is between 100 nucleotides and 500,000 nucleotides in length.

18. 18. The method of any one of claims 1 to 17, wherein each first donor cell is at a position in a first ordered array and each second, third, or subsequent donor cell is at a position in a second, third, or subsequent ordered array.

19. The method of any one of claims 1 to 18, wherein the method generates a combinatorial library comprising a plurality of different assembled DNA elements.

20. 20. The method of any one of claims 1 to 19, wherein the selectable marker is located in the non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions.

21. 21. The method of any one of claims 1 to 20, wherein the counter-selectable marker is located in the non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions.

22. The method of any one of claims 1 to 21, wherein the conditional origin of replication is dependent on the presence of an oligonucleotide or on cell growth conditions.

23. The method of any one of claims 1 to 22, wherein expression of one or more homologous DNA repair genes is inducible.

24. 24. The method of any one of claims 1 to 23, wherein the method is used to assemble a mutagenesis library or to construct a combinatorial gRNA library.