In vivo DNA assembly and analysis

In vivo DNA assembly through homologous recombination addresses the inefficiencies of traditional methods by enabling efficient and cost-effective assembly of large DNA fragments using donor and recipient plasmids with specific recombination regions and endonuclease sites.

JP7836831B2Active Publication Date: 2026-03-27THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Current oligonucleotide assembly processes are expensive, time-consuming, and limited by the size and composition of DNA elements that can be combined, requiring numerous purification steps and various enzymes.

Method used

A method for assembling DNA elements in vivo through homologous recombination using donor and recipient plasmids, involving sequential transfer and recombination of oligonucleotides with specific homologous recombination regions and endonuclease sites, allowing for efficient, high-throughput, and versatile DNA fragment assembly.

Benefits of technology

Enables efficient assembly of DNA elements up to 500,000 nucleotides in length with high accuracy and versatility, reducing the need for expensive enzymes and purification steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007836831000010
    Figure 0007836831000010
  • Figure 0007836831000011
    Figure 0007836831000011
  • Figure 0007836831000012
    Figure 0007836831000012
Patent Text Reader

Abstract

Among other things, methods and compositions are provided herein for assembling oligonucleotide fragments in vivo. The methods avoid the need for inefficient cloning methods and expensive enzymes. The methods may further be used to assemble long fragments of DNA. The methods may be used to generate variant and combinatorial libraries, and may be used to track biological processes. Among other things, methods are provided herein for in vivo DNA barcoding of oligonucleotide sequences. The methods provided herein are intended to generate unique barcode-oligonucleotide fusion sequences, for example, for identifying and isolating oligonucleotide sequences from a mixture. Thus, methods are also provided for identifying oligonucleotides from a mixture of oligonucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of priority under 35 U.S.C. Section 119(e) of U.S. Provisional Application No. 63 / 157,497 and U.S. Provisional Application No. 63 / 157,498, filed March 5, 2021. The entire contents of each are incorporated herein by reference in their entirety.

[0002] The contents of the sequence listing text file named "41243-570001WO_Sequence_Listing_ST25.txt", created on February 15, 2022, and measuring 24,576 bytes, are incorporated herein by reference in their entirety.

[0003] This invention was made with government assistance from DOE FWP #100582, awarded by the Department of Energy, and NIST IAA P18-630-0001, awarded by the National Institute for Standards and Technology. The government has certain rights to this invention. [Background technology]

[0004] Recent advances in recombinant oligonucleotide technology have stimulated research in traditional biology and biotechnology fields. However, oligonucleotide assembly processes can be expensive and time-consuming, requiring numerous purification steps and various enzymes. Furthermore, current molecular biology methods have limitations in the size and composition of DNA elements that can be combined. Therefore, there is a need for a method to assemble DNA elements (e.g., promoters, gene fragments, etc.) together to address these limitations and avoid the need for numerous and expensive enzymes (e.g., ligases). A new method is needed for efficient, high-throughput, and versatile DNA fragment assembly on a quantitatively comparable scale. [Overview of the project] [Problems that the invention aims to solve]

[0005] Advances in sequencing technology have made it possible to identify long fragments of DNA. However, identifying and isolating specific DNA sequences in complex mixtures remains challenging, among other issues, due to complex purification requirements, low sample recovery rates, or inefficient sequencing workflows.

[0006] In particular, solutions to these and other problems in the present technology are provided herein. [Means for solving the problem]

[0007] (Summary of the invention) In particular, methods and compositions for assembling DNA elements in vivo and for DNA barcoding oligonucleotide sequences are provided herein.

[0008] The present invention provides a method for assembling multiple DNA elements into an assembled DNA element in a recipient cell, the method comprising the steps of (a) (i) transferring a first donor plasmid from a first donor cell to a recipient cell by conjugation, and (ii) contacting a first donor cell containing a first donor plasmid with a recipient cell containing a recipient oligonucleotide under conditions that allow the first donor plasmid and the recipient oligonucleotide to be recombined in the recipient cell by homologous recombination, wherein the first donor plasmid Rasmid comprises, in a sequential order, an optional first endonuclease site (C1), a first homologous recombination region (HR1), a first DNA element fragment (oligo 1) as a first oligonucleotide, a second homologous recombination region (HR2) containing two homologous recombination regions (HR2.1, HR2.2), and an optional third endonuclease site (C3), and the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2, thereby HR1, HR3, and HR2 (b) The steps of providing a first recombined recipient oligonucleotide containing a fragment of the first DNA element, following homologous recombination of .2 and HR4, (i) transferring the second donor plasmid from the second donor cell to the first recipient cell by conjugation, and (ii) recombining the second donor plasmid and the first recombined recipient oligonucleotide in the recipient cell by homologous recombination to form the second recombined recipient oligonucleotide, under the conditions that the second donor cell containing the second donor plasmid and the first recombined recipient oligonucleotide A step of contacting recipient cells containing a bound recipient oligonucleotide, wherein the second donor plasmid comprises, in a sequential order, an optional fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a fragment of a second DNA element (oligo 2), a sixth homologous recombination region (HR6) containing two homologous recombination regions (HR6.1, HR6.2), and an optional sixth endonuclease site (C6), thereby connecting HR5 with HR2.1 and HR6.The method includes a step in which homologous recombination of 2 and HR4 is followed by providing a second recombined recipient oligonucleotide containing fragments of the first and second DNA elements (oligo 1, oligo 2), which together form a DNA assembly. In embodiments, HR2.1 and HR2.2 are adjacent to a non-homologous region containing one (C2) or two (C2.1, C2.2) endonuclease sites, and optionally, HR3 and HR4 are adjacent to a non-homologous region containing one (C4) or two (C4.1, C4.2) endonuclease sites. In embodiments, HR6.1 and HR6.2 are adjacent to a non-homologous region containing one (C7) or two (C7.1, C7.2) endonuclease sites. In embodiments, the recipient oligonucleotides are located in a recipient cell plasmid or recipient cell genome. In embodiments, the DNA assembly includes at least a portion of genes, promoters, enhancers, terminators, introns, intergenetic regions, barcodes, guide RNA (gRNA), or combinations thereof.

[0009] In embodiments, step (b) is repeated one or more times using a third or subsequent donor cell containing a third or subsequent donor plasmid containing a third or subsequent oligonucleotide (oligo 3, oligo 4, ... oligo N) encoding a suitable HR region and a fragment of a third or subsequent DNA element, thereby forming a third or subsequent recombined recipient oligonucleotide containing fragments of the first, second, and third or subsequent DNA elements, which together form a DNA assembly.

[0010] In an embodiment, step (a) includes a plurality of first donor cells each containing a different first donor plasmid, step (b) includes a plurality of second, third, or subsequent donor cells each containing a different second, third, or subsequent donor plasmid, optionally each first donor cell is in a position in a first ordered array, and each second, third, or subsequent donor cell is in a position in a second, third, or subsequent ordered array, and optionally the method generates a combinatorial library comprising a plurality of assembled different DNA elements.

[0011] In an embodiment, an oligonucleotide encoding a first endonuclease targeting a first, third, and / or fourth endonuclease site is present on the first donor plasmid and / or present in the recipient cell. In an embodiment, an oligonucleotide encoding a second endonuclease targeting a second, fifth, and / or sixth endonuclease site is present on the second donor plasmid and / or present in the recipient cell. In an embodiment, the expression of the first and / or second endonuclease is inducible, and the method further includes a step of inducing the expression of the first and / or second endonuclease. In an embodiment, the first and / or second endonuclease is selected from an RNA-guided endonuclease, a homing endonuclease, a transcription activator-like effector nuclease, and a zinc finger nuclease.

[0012] In embodiments, the first, second, or subsequent donor plasmid includes a selectable marker selected to integrate the first oligonucleotide, second oligonucleotide, or subsequent oligonucleotide into the recipient oligonucleotide, and optionally, the selectable marker is present in the non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions. In embodiments, the recipient oligonucleotide includes a counter-selectable marker selected against recipient cells that do not include the first, second, third, or subsequent oligonucleotide, and optionally, the counter-selectable marker is present in the non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions.

[0013] In embodiments, the donor plasmid includes a starting point for transfer.

[0014] In embodiments, the donor plasmid includes a conditional origin of replication. In embodiments, the conditional origin of replication depends on the presence of an oligonucleotide or conditions of cell growth. In embodiments, the donor plasmid or recipient oligonucleotide includes an inducible high-copy origin of replication.

[0015] In embodiments, the donor plasmid or recipient oligonucleotide includes a replicon capable of replicating plasmids longer than 30 kilobases.

[0016] In embodiments, the donor plasmid or recipient oligonucleotide is a yeast artificial chromosome (YAC), mammalian artificial chromosome (MAC), human artificial chromosome (HAC), or plant artificial chromosome.

[0017] In embodiments, the donor plasmid or recipient oligonucleotide is a viral vector.

[0018] In the embodiment, the donor plasmid contains oligonucleotides that enable plasmid conjugation.

[0019] In the embodiment, the donor plasmid or recipient cells contain oligonucleotides encoding one or more homologous DNA repair genes, and the expression of one or more homologous DNA repair genes is optionally inducible.

[0020] In the embodiment, the donor plasmid or recipient cells contain oligonucleotides encoding one or more recombinant-mediated genetically modified genes.

[0021] In this embodiment, the donor cells and recipient cells are independently bacterial cells, and optionally, the bacterial cells are E. coli, Vibrio natriegens, or V. cholerae.

[0022] In the embodiment, the assembled DNA element has a length of 100 to 500,000 nucleotides.

[0023] In the embodiment, the first, second, or subsequent homologous recombination (HR) regions and their corresponding HR regions on the recipient oligonucleotide each contain about 20 to 500 base pairs, and optionally about 50 to 100 base pairs.

[0024] In embodiments, any of the above methods may further include one or more steps of: lysing recipient cells; amplifying assembled DNA elements; isolating assembled DNA elements; isolating recipient oligonucleotides; sequencing assembled DNA elements; and sequencing recipient oligonucleotides.

[0025] In the embodiment, the steps of contacting the first donor cells and the second or subsequent donor cells with the first recipient cells are performed simultaneously, and optionally, only the final donor plasmid contains a selectable marker, or each donor plasmid contains a selectable marker that is not present on the recipient oligonucleotide.

[0026] In any one embodiment of the above method, a donor plasmid containing the last DNA element that forms part of the assembled DNA element contains a BHR region that produces recipient cells each containing the assembled DNA element, a barcode homologous recombination (BHR) region, and a recombined recipient oligonucleotide containing a further HR, the method further comprises (i) constructing or obtaining an array of barcode donor cells each containing a barcode donor plasmid containing a BHR homologous HR, a specific barcode oligonucleotide, and a second HR homologous to a further HR of the recombined recipient oligonucleotide; and (ii) contacting the array of barcode donor cells and the array of recipient cells under conditions that (a) transfer the barcode donor plasmid from the barcode donor cells to recipient cells by conjugation, and (b) recombining the barcode donor plasmid and recipient oligonucleotide in the recipient cells by homologous recombination, thereby producing an array of recipient cells containing a barcoded assembly.

[0027] In any one embodiment of the above method, each donor plasmid comprises a further pair of specific endonuclease sites CX,CY adjacent to a barcode homologous recombination (BHR) region, and the method further comprises the step of contacting an array of recipient cells, each containing a DNA assembly, with an array of barcode donor cells, each containing a barcode donor plasmid, each containing a pair of HR regions homologous to BHR adjacent to a specific barcode oligonucleotide, to produce an array of recipient cells containing barcoded assemblies.

[0028] In any one embodiment of the above method, the method further comprises the step of contacting a reset donor cell containing a reset donor plasmid with a recipient cell containing a recombined recipient oligonucleotide, wherein the reset donor plasmid comprises, in sequential order, a homologous recombination region (HRt) homologous to the terminal sequence of the DNA assembly, a reset endonuclease site, a selectable marker, a reset endonuclease site, a homologous recombination region (HRX), and a transport origin, and the recombined recipient oligonucleotide comprises, in sequential order, a reset endonuclease site, the DNA assembly, a homologous recombination region (HRXa) homologous to the HRX, and a reset endonuclease site, thereby providing a reset plasmid containing a transport origin and the DNA assembly, following homologous recombination between the HRt and the terminal sequence of the DNA assembly and between the HRX and HRXa. In an embodiment, the reset plasmid is located in the donor cell. In an embodiment, the reset plasmid contains a restricted origin of replication that functions in both the donor cell and the recipient cell. In embodiments, the reset donor plasmid is constructed by introducing an oligonucleotide insert HRt-C1-CM-C2-HRX, or a library of such oligonucleotide inserts, which include two endonuclease sites (C1, C2) and homologous recombination regions HRt, HRX adjacent to an antiselectable marker (CM), thereby enabling the endonuclease to cleave the endonuclease sites and introducing the antiselectable marker at the cleavage site using homologous recombination.

[0029] In any embodiment of the above method, the recipient oligonucleotide comprises a mobile genetic element capable of transferring the DNA assembly to other cell types, including yeast cells, plant cells, mammalian cells, or other bacterial cells.

[0030] In any embodiment of the above method, the method includes the step of constructing a DNA library using two or more recipient oligonucleotides having compatible homologous recombination regions.

[0031] In any embodiment of the above method, the donor plasmid oligonucleotide comprises a first linker oligonucleotide homologous to the terminal sequence of the first DNA assembly and a second linker oligonucleotide homologous to the second oligonucleotide. In embodiments, the linker oligonucleotide further comprises fragments of further DNA elements that are not homologous to the first DNA assembly or the second DNA oligonucleotide. In embodiments, the method assembles a mutagenesis library using genes, promoters, terminators, and from different species. control In order to combine genetic regions such as regions, Child control It is used to construct and / or combine pathways, to construct combinatorial gRNA libraries, or to assemble bacterial arrays containing plasmids for screening assays.

[0032] In any embodiment of the above method, prior to steps (a) and (b), first and second oligonucleotides containing fragments of the first and second DNA elements are inserted into the first and second donor plasmids.

[0033] The present invention also provides a method for attaching barcodes to oligonucleotides, and this method is (a) Inserting each oligonucleotide from a mixture of oligonucleotides into a donor plasmid that optionally contains a first endonuclease site (C1), a first homologous recombination site (HR1), a second homologous recombination site (HR2), and optionally a second endonuclease site (C2) in a sequential order, so that each oligonucleotide is inserted between HR1 and HR2, thereby providing multiple donor plasmids containing donor oligonucleotides, each donor plasmid containing a single donor oligonucleotide C1-HR1-oligo-HR2-C2 from the mixture of oligonucleotides; (b) Transforming multiple cells with multiple donor plasmids so that each cell contains a donor plasmid, thereby forming multiple donor cells; (c) Plate seeding and culturing the multiple donor cells at specific positions on a first ordered array, thereby providing a first ordered array of donor cells; (d) Providing multiple recipient cells into a second ordered array. The present invention relates to a recipient oligonucleotide comprising, in a sequential order, a specific barcode sequence that identifies the location of the recipient cell in a second ordered array, a third homologous recombination region (HR3) homologous to HR1, optionally a third endonuclease region (C3), and a fourth homologous recombination region (HR4) homologous to HR2; (e) (i) transferring the donor plasmid from the donor cell to the recipient cell at the corresponding location on the array by conjugation; (ii) optionally cleaving the first, second, and third endonuclease regions; and (ii) contacting the first ordered array of donor cells with the second ordered array of recipient cells under conditions that transfer the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby forming a third array of fusion oligonucleotides, each containing the donor oligonucleotide from the specific barcode sequence and the mixture of oligonucleotides; and (f) optionally sequencing the fusion oligonucleotides, therebyThe process includes identifying each oligonucleotide in the array by its barcode sequence. In embodiments, the recipient oligonucleotide is located in a recipient cell plasmid or recipient cell genome. In embodiments, the donor plasmid contains selectable markers between HR1 and HR2 for the integration of the oligonucleotide into the recipient cell oligonucleotide, and optionally the donor plasmid contains antiselectable markers. In embodiments, the recipient cell oligonucleotide contains a fourth endonuclease site (C4).

[0034] The present invention also provides a method for identifying oligonucleotides from a plurality of oligonucleotides, the method comprising: (a) providing a plurality of donor cells in a first ordered array, wherein each donor cell comprises a donor plasmid containing, optionally, a first endonuclease site (C1), a first homologous recombination region (HR1), a specific barcode sequence, a second homologous recombination region (HR2), and optionally, a second endonuclease site (C2) in a sequential order, the specific barcode sequence identifying the location of the host cell in the first ordered array; and (b) providing a plurality of recipient cells, wherein each recipient cell comprises a recipient plasmid containing, in a sequential order, oligonucleotides from the plurality of oligonucleotides, a third homologous recombination region homologous to HR1 (HR3), optionally, a third endonuclease site (C3), and a fourth homologous recombination region homologous to HR2 (HR4). The method includes (c) plate seeding and culturing multiple recipient cells, each at specific locations on a second ordered array, thereby providing a second ordered array of recipient cells; (d) contacting the first ordered array with the second ordered array under conditions that (i) transfer a donor plasmid from donor cells to recipient cells at corresponding locations on the array by bacterial conjugation; (ii) cleaving first, second, and third endonuclease sites and (ii) transferring barcode sequences from the donor plasmid to recipient cell oligonucleotides by homologous recombination, thereby forming a third array of fusion oligonucleotides, each containing oligonucleotides from a mixture of specific barcode sequences and oligonucleotides; and (e) sequencing the fusion oligonucleotides, thereby identifying each oligonucleotide in the array by its barcode sequence. In embodiments, the recipient oligonucleotides are located in the recipient cell plasmid or the recipient cell genome.In embodiments, the donor plasmid includes a selectable marker between HR1 and HR2 for integration of a barcode sequence into the recipient cell oligonucleotide, and optionally the donor plasmid includes an antiselectable marker. In embodiments, the recipient cell oligonucleotide includes a fourth endonuclease site (C4). In embodiments, the first, second, and third endonuclease sites are identical or different. In embodiments, the donor plasmid includes a transport origin and / or a conditional replication origin, optionally the transport origin originating from a mobile element, and further optionally the conditional replication origin depending on the presence of oligonucleotides or conditions of cell proliferation. In embodiments, the donor plasmid or recipient plasmid includes a replicon capable of replicating a plasmid of at least 30 kilobases in length, and optionally the replicon is derived from a P1-inducible artificial chromosome or a bacterial artificial chromosome. In embodiments, the donor plasmid or recipient cell oligonucleotide includes an inducible high-copy replication origin. In embodiments, the donor plasmid or recipient cell oligonucleotide includes a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome. In embodiments, the donor plasmid or recipient cell oligonucleotide includes a viral vector. In embodiments, the endonuclease site is encoded by one or more oligonucleotides in the recipient cell and / or cleaved by one or more endonucleases encoded in the donor plasmid, optionally, one or more endonucleases are homing endonucleases or RNA-guided DNA endonucleases, and further optionally, the endonucleases are HO. In embodiments, the donor cell or recipient cell includes an oligonucleotide that (i) enables plasmid conjugation, (ii) encodes one or more homologous DNA repair genes, or (iii) encodes one or more recombination-mediated genetic manipulation genes.In embodiments, donor cells, recipient cells, or recombinant recipient cells are transported to a position on a third ordered array, a fourth ordered array, or a subsequent ordered array. In embodiments, the donor cells and recipient cells are independently bacterial cells, and optionally, the bacterial cells are E. coli, Vibrio sodiumgens, or V. cholera. In embodiments, the barcode sequence is approximately 4 to 100 nucleotides in length, and optionally, the barcode sequence is approximately 30 nucleotides in length. In embodiments, oligonucleotide mixtures are products of DNA synthesis or assembly techniques selected from chemical coupling, template-independent enzymatic synthesis using polymerase nucleotide conjugates, polymerase chain assembly (polymerase cycling assembly), Gibson assembly (chew-back, annealing, and repair), ligase chain reaction / ligase cycling reaction, Phi29 polymerase, rolling circle, loop-mediated isothermal (LAMP), strand substitution (SDA), helicase-dependent (HAD), recombinase polymerase (RPA), nucleic acid sequence-based amplification (NASBA), golden gate cloning, MoClo cloning, BioBricks or assembled BioBricks, thermodynamically balanced front-to-back synthesis, DNA cloning, ligation-independent cloning, ligation by selective cloning, recombination operations, yeast assembly, PCR, capture by molecular inversion probes or LASSO probes, DropSynth, and enzymatic DNA synthesis. In embodiments, the oligonucleotide mixture is the product of polymerase chain reaction techniques including error-prone PCR, PCR with degenerate oligonucleotides, and conventional PCR; chemical or photomutagenesis; in vitro synthesis of oligo-edited libraries; in vivo editing, such as MAGE, MAGESTIC, CRISPR, prime editing, retron editing; and pooled mutagenesis techniques selected from base modification by CRISPR, TALEN, and zinc finger nucleases.In embodiments, the oligonucleotide mixture comprises at least a fragment of genomic DNA, cDNA, organelle DNA, or native plasmid DNA. In embodiments, the oligonucleotide mixture comprises captured or amplified DNA derived from gDNA, cDNA, or organelle DNA, such as a balanced cDNA library, PCR products, such as multiplex PCR products, molecular inversion probes including LASSO probes, capture by annealing or subtractive hybridization, cotransformation and homologous recombination, rolling circle amplification, or LAMP. In embodiments, the oligonucleotide mixture comprises captured or amplified DNA derived from plasmids or plasmid libraries, such as open reading frame (ORF) libraries, promoter libraries, terminator libraries, intron libraries, BAC libraries, PAC libraries, lentivirus libraries, gRNA libraries, PCR products, restriction digest products, or GATEWAY shuttle products. In embodiments, the oligonucleotides of the oligonucleotide mixture are integrated into a donor plasmid by methods including cotransformation and recombination, transformation and recombination, or conjugation and recombination. In embodiments, the cotransformation and recombination procedure includes the steps of constructing a linear or circular donor plasmid containing a selectable marker and two homologous recombination regions homologous to the terminal sequences of oligonucleotides in the mixture, respectively; cotransforming cells with the donor plasmid and oligonucleotides; inducing homologous recombination; and selecting a selectable marker, and optionally, the procedure is carried out using a library or pool of donor plasmids and / or oligonucleotides.In embodiments, a transformation and recombination method comprises the steps of constructing a linear or circular donor plasmid containing a selectable marker and two homologous recombination regions homologous to the terminal sequences of oligonucleotides in a mixture, wherein the oligonucleotides are present on the plasmid in a host cell; transforming the host cell with the donor plasmid; inducing homologous recombination; and selecting a selectable marker. In embodiments, a method for conjugation and recombination includes the steps of: constructing a linear or circular donor plasmid containing two optionally selected endonuclease sites and two homologous recombination (HR) regions adjacent to an antiselectable marker (-1), wherein the donor plasmid is present in a donor cell containing a deactivated F plasmid capable of inducing conjugation but not necessarily conjugating; a mixture of oligonucleotides present on a plasmid in a recipient cell, each adjacent to an HR region homologous to the HR region of the donor plasmid and adjacent to at least one selectable marker (+1) to be selected for recombination of each oligonucleotide into the donor plasmid; providing homologous recombinases and one or more endonucleases optionally present in the recipient cell or encoded by the donor plasmid; contacting the donor cell and recipient cell under conditions to transfer the donor plasmid from the donor cell to the recipient cell by bacterial conjugation; and selecting cells containing a selectable marker but not an antiselectable marker. In embodiments, the method is carried out using a library of donor and / or recipient plasmids. In embodiments, oligonucleotides include an ORF library, promoter library, terminator library, intron library, BAC library, PAC library, lentivirus library, gRNA library, gDNA library, cDNA library, protein domain library, promoter library, and terminator library. controlThe library includes libraries of elements, structural elements, or DNA variants derived from DNA mutagenesis. In embodiments, the oligonucleotide mixture includes plasmid libraries, such as gRNA libraries, gDNA libraries, cDNA libraries, open reading frame (ORF) libraries, protein domain libraries, promoter libraries, terminator libraries, control The present invention includes an array of cells comprising a library of elements, a library of structural elements, or a library of DNA variants derived from DNA mutagenesis. In embodiments, the oligonucleotide mixture comprises an array of cells comprising fragments of DNA elements for use in the method according to any one of claims 1 to 35. [Brief explanation of the drawing]

[0035] [Figure 1] A schematic diagram of an embodiment of the DNA assembly method described herein, showing donor plasmid and recipient oligonucleotide elements as shaded frames. The diagram shows three "rounds" of "DNA stitching," in which a new oligonucleotide containing a fragment of DNA element is added to the recipient oligonucleotide. Throughout the diagram, fragments of DNA element are variously represented as input DNA1, input DNA2, etc. before recombination, DNA1, DNA2, etc., or oligo1, oligo2, etc., or more generally referred to as "DNA blocks" after recombination. Frames designated C1, C2, etc., represent endonuclease sites, frames designated HR1, HR2, etc., represent homologous recombination regions, and frames designated oligo1, oligo2, etc., represent oligonucleotides containing fragments of DNA element. Frames designated with numbers and plus or minus signs represent selectable (+) and antiselectable (-) markers. Not all elements shown in the schematic diagrams are necessary in all embodiments of the methods described herein; for example, the markers and endonuclease sites designated as C2.1, C2.2, C7.1, and C7.2 are optional. [Figure 2A] Panel A is a schematic map of exemplary recipient oligonucleotides (in plasmid form) used in the methods described herein. [Figure 2B] Panel B shows two schematic maps of exemplary donor plasmids and helper plasmids. [Figure 3] CRISPR / Cas9 enhances the efficiency of DNA assembly, also referred to herein as “suturing.” Colony count per 6 × 10⁶ cells on a selection plate. Three recipient strains, (1) BW28705 (without λ-red, without Cas9), (2) BW28705 / pML300 (with λ-red, without Cas9), and (3) BW28705 / pSL359 (with Cas9 and λ-red), were transformed with two donor plasmids, one of which expressed a functional gRNA. [Figure 4] A schematic map of an exemplary plasmid used in the in vivo DNA assembly or “suturing” method described herein. In donor cells, a conjugation-capable helper plasmid may contain a gene for plasmid transport (Tra operon). To immobilize the conjugation plasmid itself, the origin of transport (oriT) is replaced with a selectable marker (+6). The donor plasmid includes a swapping cassette (+1 / -1 or +2 / -2), two homology regions (H2 and H3), four endonuclease cleavage sites (two circles labeled 1 and two circles labeled 2), a selectable marker for the skeleton (+4), an allele-dependent conditional origin of replication (R6K) in the donor genome (pir1-116), an oriT sequence, and a gRNA expression cassette (gRNA1 or gRNA2). [Figure 5]A schematic map of exemplary plasmids used in the in vivo suturing method described herein. In recipient cells, the helper plasmid contains a rhamnose-inducible red operon (PrhaBAD-red), an arabinose-inducible Cas9 (ParaBAD-Cas9), the E. coli RecA gene for enhancing homologous recombination, a selectable skeletal marker (+5), and a cureable origin of replication (pSC101 oriTS). The recipient plasmid contains two endonuclease cleavage sites (two circles labeled 1), a swapping cassette (+2 / -2), two homology regions (H2 and H3), and an origin of replication (ColE1). +1: HygR, +2: NsrR, -1: SacB, -2: PheS, +3: GmR, +4: KanR, +5: SpR, +6: TcR. [Figure 6] Schematic diagram of an exemplary method of in vivo suturing. A donor plasmid (rectangle with upward or downward stripes) carrying a DNA fragment is introduced into the donor plasmid and donor cells. The donor plasmid is joined to the recipient cells, and the DNA fragment is transferred from the donor plasmid to the recipient plasmid. The plasmid is cleaved using CRISPR / Cas9, which is induced by arabinose. Guide RNA (gRNA1 or gRNA2, interchangeable between assembly rounds) on the donor plasmid identifies recognition sequences for cleavage (circles labeled "1" and "2," interchangeable between assembly rounds). Homology regions on both the synthesized oligo and the plasmid backbone (H1 and H3 in round 1) facilitate rhamnose-induced recombination, suturing the oligos together seamlessly for gene assembly. Interchangeable selectable (+1 and +2) and anti-selectable (-1 and -2) markers on the donor plasmid allow for the transfer of repeatable DNA with a maximum gene length theoretically set by the maximum allowable plasmid size. R6K and ColE1 are the origins of replication. +3 and +4 are selectable markers used for plasmid maintenance. [Figure 7A]An example of DNA assembly. Panel A shows three donor plasmids (3 stitches) each carrying a portion of mEGFP, sequentially joined and assembled into a recipient plasmid. [Figure 7B] Example of DNA assembly. Panel B shows the fluorescence of colonies from negative control, positive control, and in vivo suture products after three rounds of assembly in liquid. Colonies represent independent conjugation and recombination events and are 100% fluorescent. [Figure 7C] Examples of DNA assemblies. Panel C shows aligned assemblies of mEGFP in 96-position and 384-position formats. All colonies appear to be fluorescent. [Figure 7D] Example of DNA assembly. Panel D shows the percentage of fluorescent colonies after the final round of liquid assembly using a third mEGFP fragment of a different length homologous to a second mEGFP fragment. [Figure 7E] Example of DNA assembly. Panel E shows representative restriction digests of colonies containing various plasmids scraped from agar during assembly. After recombinant selection or helper plasmid curing, the expected product from the non-recombinant recipient plasmid is not observed (indicated by arrows). [Figure 7F] Example of DNA assembly. Panel F shows a schematic diagram of the analysis of Sanger sequencing results for 96 colonies after mEGFP assembly. Sequencing products are derived from colony PCR of pipette tips in contact with each colony. One colony contained an intermediate product (product of the first round of assembly), and one contained a suture error (large deletion). [Figure 7G] Examples of DNA assemblies. Panel G shows the fluorescence of colonies from in vivo suture products after five rounds of assembly for two fluorescent genes, namely mPapaya and sfGFP, and four recombinase genes. Colonies may represent independent conjugation and recombination events. All colonies for mPapaya and sfGFP are fluorescent. [Figure 7H]Example of DNA assembly. Panel H is a trace file of Sanger sequencing of mPapaya's in vivo suture product after 5 rounds of assembly. Alignment with expected sequences indicates that the assembly is 100% accurate and pure. [Figure 7I] Examples of DNA assembly. Panel I shows the results of assembling three approximately 3kb fragments for a total assembly length of approximately 9kb. The recipient plasmid was digested with restriction enzymes at various stages of assembly, and the suture product was separated from the vector backbone. The digested products were then subjected to agarose gel electrophoresis, and the size of the suture product was examined (lanes 1-3). Gel bands corresponding to the suture product are marked with arrows. The linearized vector backbone without the suture product is shown in lanes 5-6. Selectable and anti-selectable markers in the swapping cassette differ between assembly rounds, with the swapping cassette in the first and third assembly rounds being approximately 1.5kb longer than the original recipient plasmid or the swapping cassette in the second assembly round. [Figure 8A] Schematic diagrams of exemplary donor and recipient plasmids at the start of the first round of DNA suturing. Panel A shows schematic diagrams of plasmids for both examples, with the donor plasmid containing the first oligonucleotide (1). [Figure 8B]Schematic diagrams of exemplary donor and recipient plasmids at the start of the first round of DNA suturing. Panel B shows a shape legend used to illustrate sequences corresponding to the example genome, positive selectable marker, negative selectable marker, origin of transport (oriT), gRNA expression unit (gRNA), positional barcode, homology region for recombinant domain (H), inducible lambda red operon (λred), inducible I-SceI endonuclease, plasmid, inducible endonuclease (Cas9), gRNA target site, I-SceI target site, conjugation Tra operon, deleted oriT (oriΔ::TcR), temperature-sensitive origin (pSC101 ori), conditional origin of replication (R6K), and recipient origin of replication (ColE1). The same shape legend in Figure 8B is also used in Figures 9-35. [Figure 9] Schematic diagram of the first exemplary plasmid for use in the methods described herein. A donor plasmid containing the first oligo(1), a recipient plasmid, and a plasmid containing oriT that mediates the conjugation of the donor plasmid to the recipient cell. [Figure 10] Schematic diagram of subsequent steps in an exemplary method of DNA suturing using donor and recipient plasmids shown in Figure 9. gRNA1 guides Cas9 in recipient cells to generate site-specific double-strand breaks on the donor and recipient plasmids (indicated by downward arrows). [Figure 11] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-10. Here, the darkly shaded sequence elements are used as homology regions for lambda red-mediated homologous recombination. [Figure 12] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-11. Homologous recombination between plasmids, showing where and how sequences derived from the donor plasmid are inserted into the recipient plasmid with the assistance of the λ-RED system. [Figure 13]Schematic diagrams of the subsequent steps in the exemplary method of DNA suturing shown in Figures 9-12. The fragment containing the first oligonucleotide is integrated into the recipient plasmid as shown. [Figure 14] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-13. Plasmids are selected for the acquisition of a +2 positive selectable marker. [Figure 15] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-14. Plasmids are antiselected for the loss of previously antiselectable markers. [Figure 16] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-15. Plasmids are further selected for retention of a +3 positive selectable marker on the original recipient skeleton. [Figure 17] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-16. The second donor plasmid, containing the second oligonucleotide, is ready to be assembled into the previous ligation product (new recipient plasmid) containing the first oligonucleotide. [Figure 18] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-17. oriT instructs the conjugation of the donor plasmid. [Figure 19] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-18. The expression of the second gRNA guides Cas9 to generate a double-strand break at the site indicated by the downward arrow. [Figure 20] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-19. The highlighted regions (darkly shaded regions) are homology regions for recombination. [Figure 21] Schematic diagrams of the subsequent steps in the exemplary method of DNA suturing shown in Figures 9-20. The fragment containing the first oligonucleotide is integrated into the recipient plasmid as shown. [Figure 22]Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-21. A second oligonucleotide is assembled adjacent to the 3' end of the first oligo in the recipient plasmid to generate a new recipient plasmid. [Figure 23] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-22. Plasmids containing the first and second oligonucleotides are selected for the acquisition of a +1 positive selectable marker. [Figure 24] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-23. Plasmids containing the first and second oligonucleotides are selected for the loss of a marker that is 2-2 antiselectable. [Figure 25] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-24. Plasmids containing the first and second oligonucleotides are also selected for the retention of selectable markers on the backbone. [Figure 26] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-25. Schematic diagrams of the donor plasmid containing oligonucleotide 3 and the recipient plasmids containing oligonucleotides 1 and 2 for initiating the third round of DNA suturing. [Figure 27] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-26. The oriT plasmid initiates ligation. [Figure 28] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-27. gRNA1 guides Cas9 in recipient cells to generate site-directed double-strand breaks on donor and recipient plasmids (indicated by downward arrows). [Figure 29] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-28. The sequences highlighted here are used as homology regions for lambda red-mediated homologous recombination. [Figure 30]Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-29. These diagrams show where the donor plasmid-derived sequence is inserted into the recipient plasmid and its orientation after homologous recombination. [Figure 31] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-30. The third oligonucleotide is assembled adjacent to the 3' end of the second oligonucleotide in the recipient plasmid to generate a new recipient plasmid. [Figure 32] Schematic diagrams of the subsequent steps in the exemplary DNA suturing method shown in Figures 9-31. Similar to the DNA suturing in Round 1, plasmids are selected for the acquisition of a +2 positive selectable marker. [Figure 33] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-32. Plasmids are antiselected for the loss of a -1 antiselectable marker. [Figure 34] Schematic diagrams of subsequent steps in the exemplary DNA suturing method shown in Figures 9-33. Plasmids are selected for the retention of selectable markers in the skeleton. [Figure 35A] Panel A is a schematic diagram of the steps in the method illustrated in Figures 9-34, illustrating embodiments in which subsequent oligonucleotides can be incorporated (up to an upper limit of the total sequence length) in a similar manner, alternating between two processes: conjugation, double-strand cleavage, assembly, and selectable / anti-selectable / skeleton selection. [Figure 35B]Panel B is a schematic diagram of an exemplary in vivo assembly in which fragments of two DNA elements (shown in the figure as Input DNA1 and Input DNA2 before recombination, and as DNA1 and DNA2 after recombination, and may be referred to herein as Oligo1, Oligo2, etc., or more generally as a “DNA block”) are added in a single conjugate round. In each well, recipient cells are first conjugated by a first donor cell containing a first donor plasmid, and secondly by a second donor cell containing a second donor plasmid. Selectable markers introduced by the second donor plasmid, as well as contraselectable markers on the recipient plasmid and optionally on the first donor plasmid, allow for selection of a recombinant assembly product containing oligonucleotides introduced by both donor plasmids. [Figure 36A] Panel A is a schematic diagram of an exemplary method for in vivo DNA analysis described herein. In this example, the indexing barcode is located on the recipient plasmid. [Figure 36B] Panel B is a schematic diagram of another exemplary method for in vivo DNA analysis as described herein. In this example, the indexing barcode is located on the donor plasmid and is added to the in vivo DNA assembly product using a homologous region at the end of the assembly product. In this example, the endonuclease site used is the same as that used for in vivo DNA assembly. [Figure 36C] Panel C is a schematic diagram of another exemplary method for in vivo DNA analysis described herein. In this example, the indexing barcode is located on the donor plasmid and added to the in vivo DNA assembly product. In this example, the endonuclease target site (C) and homologous region (frame adjacent to C) are different from those used for in vivo DNA assembly. In this example, DNA analysis is possible at multiple steps during assembly. [Figure 36D]Panel D is a schematic diagram of a method involving plasmid reset to move a DNA assembly from a recipient plasmid to a donor plasmid, enabling further rounds of assembly using a larger DNA block. In Part A of Panel D, a reset donor plasmid in a donor cell, homologous to the start of the DNA assembly on the recipient plasmid, is joined into the recipient cell. Site-specific endonucleases cleave at the "D" endonuclease target site in both the reset donor plasmid and the recipient plasmid. Homologous recombination in the region adjacent to the "D" endonuclease target site moves the DNA assembly cassette from the recipient plasmid to the reset donor plasmid. The reset donor plasmid can be purified from the recipient cells and transformed into new donor cells for use in further rounds of assembly. In Part B of Panel D, the schematic diagram shows a workflow for assembling a long DNA construct. Using four rounds of DNA suturing, small DNA blocks can be assembled into a larger DNA block. The larger block is moved to the donor plasmid and donor cells using the reset donor plasmid. The next larger block can be assembled into an even larger block through further suture rounds. [Figure 37]A schematic map of an exemplary plasmid for use in in vivo DNA analysis is shown. In donor cells, the conjugation-capable helper plasmid contains a gene (Tra operon) for plasmid transport. To immobilize the helper plasmid itself, the origin of transport (oriT) is replaced by a selectable marker (+6). The donor plasmid contains a swapping cassette (+ and -), two homology regions (H1 and H4), two sites for targeted plasmid cleavage (ellipse), a selectable marker for the skeleton (+4), an allele-dependent conditional origin of replication (R6K) in the donor genome (pir1-116), and the oriT sequence. In recipient cells, the helper plasmid contains a lac-inducible red operon (Plac-red), the E. coli RecA gene for enhancing homologous recombination, a selectable marker for the skeleton (+5), and a cureable temperature-sensitive origin of replication (pSC101 oriTS). The recipient plasmid contains two endonuclease cleavage sites (two ovals), a negatively selectable marker (-3), and two homology regions (H1 and H4). In addition to the two plasmids, the recipient cell also possesses an integrated arabinose-inducible endonuclease I-SceI (ParaBAD-I-SceI) for generating DNA cleavage on the target plasmid. +: HygR or NsrR, -: SacB or PheS, +3: GmR, -3: relE, +4: KanR, +5: SpR, +6: TcR. [Figure 38] A schematic map of an exemplary plasmid for use in in vivo DNA analysis is shown. Selectable markers: HygR, KanR, GmR, SpR. Conversely selectable markers: SacB, relE. [Figure 39A] This panel shows schematic maps of exemplary donor and recipient plasmids used in DNA syntactic analysis. Panel A schematic shows a plasmid in which the donor plasmid contains the first oligonucleotide. [Figure 39B]Panel B shows schematic maps of exemplary donor and recipient plasmids used in DNA syntactic analysis. Panel B shows a legend of shapes used to illustrate sequences corresponding to the genome, positive selectable markers, negative selectable markers, origin of transport (oriT), gRNA expression unit (gRNA), positional barcode, homology region for recombinant domain (H), inducible lambda red operon (λred), inducible I-SceI endonuclease, plasmid, inducible endonuclease (Cas9), gRNA target site, I-SceI target site, conjugation Tra operon, deleted oriT (oriΔ::TcR), temperature-sensitive origin (pSC101 ori), conditional origin of replication (R6K), and recipient of the origin of replication (ColE1). The legend in Panel B also applies to Figures 40-46. [Figure 40] A schematic diagram of the donor and recipient plasmids for use in the methods of the examples described herein is shown. [Figure 41] Figure 40 shows an image illustrating the second step in the method of the example. gRNA1 guides Cas9 in the recipient cells to generate site-directed double-strand breaks (indicated by downward arrows) on the donor and recipient plasmids. SceI is I-SceI, i.e., a homing endonuclease. [Figure 42] Figures 40 and 41 illustrate the steps in the method of the embodiment. Here, the sequences of H1 and H4 are used as homology regions for lambda red-mediated homologous recombination. [Figure 43] Figures 40-42 are images illustrating the steps in the methods of the examples. Homologous recombination shows where the sequence derived from the donor plasmid is inserted into the recipient plasmid and its orientation. [Figure 44] Figures 40-43 are images illustrating the steps in the method of the example. Plasmids are selected for the acquisition of a +positive selectable marker. [Figure 45] Figures 40-44 illustrate the steps in the method of the example. The plasmid is antiselected for the loss of a previously antiselectable marker. [Figure 46] Figures 40-45 are images illustrating the steps in the method of the example. Plasmids are further selected for retention of a +3 positive selectable marker on the original recipient skeleton. [Figure 47A] Panel A shows experimental results to determine the ability of in vivo DNA analysis to correctly identify the sequence at each location on plates of aligned donor cells containing DNA barcodes specific to each location. Each aligned barcoded donor was crossed with two or three barcoded recipient plates, and recombinant cell colonies containing donor and recipient barcodes, respectively, were selected on agar pads. Recombinant cells from the plates were pooled, and the dual barcodes were sequenced on an Illumina platform. Using the sequencing data, the percentage (recovery rate) of aligned barcoded donors that were correctly indexed when joined to one, two, or three isolated barcoded recipient arrays was determined. No barcoded donors were misassigned to incorrect locations. [Figure 47B] Panel B shows the results of an experiment in which a pool of 100 244-base oligonucleotides ordered from IDT as an oPool was indexed and sequenced. The oligonucleotide pool was integrated into a donor plasmid, which was then used to transform donor cells. The donor cells were randomly sequenced into a 384-well plate at an expected frequency of less than one cell per well. The donor cells were conjugated to a barcoded recipient cell array, and the recombinant oligonucleotide-barcoded recipient plasmid was sequenced using an Oxford Nanopore sequencer. The results of the analysis of two 384-well plates are shown. Shading indicates whether the well contains sequenced input DNA, whether the sequence is 100% identical to one of the 244-base sequences in the oPool, and whether the well is pure (i.e., only one of the 244-base sequences is detected in the well). Positions marked as "100% identical pure wells" are typically used for downstream DNA assembly. [Figure 47C]Panel C is a histogram showing the distribution of errors between the consensus sequence determined by Oxford Nanopore sequencing using experimental data from Panel B and the predicted DNA sequence in the oPool that is closest to the consensus sequence. Most wells contain oligonucleotides identical to one of the sequences in the oPool. [Figure 47D] Panel D is a histogram showing the distribution of the counts of independent clones recovered for each indexed oligonucleotide, using the experimental data from Panel B. [Figure 48] This is a schematic diagram of a DNA assembly workflow for constructing a directed combinatorial library derived from a set of input oligonucleotides. A pool of input DNA from multiple sources is integrated into a donor plasmid and parsed in an ordered array. The ordered array is rearranged to user-defined positions on multiple donor plates. The donor plates are sequentially joined to recipient plates, and the desired construct is assembled. Input oligonucleotides can be used in multiple assemblies by rearranging the donor cells containing those oligonucleotides to multiple positions on the donor plate. [Figure 49] This is a schematic diagram of a branched DNA assembly. A partial DNA assembly can be expanded by multiple DNA blocks if homologous regions exist. If homologous regions do not exist, a "DNA linker" must first be added to the partial DNA assembly. The DNA linker contains homologous regions between the end of the partial DNA assembly and the start of the subsequent DNA block to be joined. [Modes for carrying out the invention]

[0036] 1. Definitions and related embodiments Unless otherwise defined, all technical and scientific terms used herein have meanings commonly understood by those skilled in the art in the field to which this invention pertains. The following references provide general definitions for many of the terms used herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings attributed to them unless otherwise specified.

[0037] The use of singular indefinite or definite articles (e.g., "a," "an," "the," etc.) in this disclosure and the following claims follows the conventional patent approach of meaning "at least one" unless it is evident from the context that in a particular case the term is intended to specifically mean one and only one in that particular case. Similarly, the term "contains" is open-ended and does not exclude additional items, features, components, etc. References identified herein are expressly incorporated herein as a whole by reference unless otherwise indicated.

[0038] "Optional" or "optionally" means that the event or circumstance described thereafter may or may not occur, and that the description includes both cases where the event or circumstance occurs and cases where it does not.

[0039] When the term "approximately" is used before a numerical specification that includes a range, such as temperature, time, quantity, concentration, and similar values, it indicates an approximation that the value may vary by (+) or (-) 10%, 5%, 1%, or any lower range or lower value in between. Preferably, the term "approximately" means that the value may vary by ±10%.

[0040] As used herein, the term “contains” is intended to mean that the composition and method include the elements referenced but not exclude others. “Essentially consisting of” is intended to mean that other elements having some essential character to the combination for the stated purpose are excluded. Thus, a composition consisting essentially of the elements defined herein does not exclude other materials or steps that do not substantially affect the fundamental and novel features of the claimed invention. “Consists of” is intended to mean that elements exceeding trace amounts of other components and substantial method steps are excluded. Embodiments defined by each of these transitional terms are within the scope of this disclosure.

[0041] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid sequence,” and “polynucleotide” are interchangeable and intended to include, but are not limited to, nucleotides in polymeric form covalently linked together, which may be deoxyribonucleotides or ribonucleotides, and may have varying lengths, or analogs, derivatives, or variants thereof. Different polynucleotides may have different three-dimensional structures and may perform a variety of known or unknown functions. Non-exclusive examples of polynucleotides include genes, gene fragments, exons, introns, intergenetic DNA (including, but not limited to, heterochromatin DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, sgRNA, guide RNA, tracrRNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of a sequence, isolated RNA of a sequence, PCR products, nucleic acid probes, and primers. Polynucleotides useful in the methods of this disclosure may include natural nucleic acid sequences and their variants, artificial nucleic acid sequences, or combinations of such sequences.

[0042] "Nucleic acid" means nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and their polymers or complements in single-stranded, double-stranded, or multi-stranded forms, or nucleosides (e.g., deoxyribonucleosides or ribonucleosides). In embodiments, "nucleic acid" does not include nucleosides. The terms "polynucleotide," "oligonucleotide," "oligo," etc., mean linearly arranged nucleotides in the usual and conventional sense. The term "nucleoside" means glycosylamines containing a nucleic acid base and a pentose (ribose or deoxyribose) in the usual and conventional sense. Non-limiting examples of nucleosides include cytidine, uridine, adenosine, guanosine, thymidine, and inosine. The term "nucleotide" means a single unit of a polynucleotide, i.e., a monomer, in the usual and conventional sense. A nucleotide may be a ribonucleotide, a deoxyribonucleotide, or a modified version thereof. Examples of polynucleotides as intended herein include single-stranded and double-stranded DNA, single-stranded and double-stranded RNA, and hybrid molecules having mixtures of single-stranded and double-stranded DNA and RNA. Examples of nucleic acids as intended herein, such as polynucleotides, include any type of RNA, such as mRNA, siRNA, miRNA, and guide RNA, as well as any type of DNA, such as genomic DNA, plasmid DNA, and microcyclic DNA, and any fragments thereof. In the context of polynucleotides, the term “duplex” means double-stranded in its usual and conventional sense. Nucleic acids may be linear or branched. For example, a nucleic acid may be a straight chain of nucleotides, or a nucleic acid may be branched such that the nucleic acid contains one or more arms or branches of nucleotides. Optionally, branched nucleic acids may repeatedly branch to form dendrimers or other higher-order structures.

[0043] For example, a nucleic acid having a phosphothioate skeleton may contain one or more reactive moieties. As used herein, the term “reactive moiety” includes any group that can react with another molecule, such as a nucleic acid or polypeptide, by covalent, non-covalent, or other interactions. For example, a nucleic acid may contain an amino acid-reactive moiety that reacts with amino acids on a protein or polypeptide by covalent, non-covalent, or other interactions.

[0044] This term also encompasses nucleic acids, including known nucleotide analogs or modified skeletal residues or linkages, which are synthetic, naturally occurring, and unnaturally occurring, possess similar binding properties to the reference nucleic acid, and are metabolized in a similar manner to the reference nucleotide. Examples of such analogs include, but are not limited to, phosphoramidates, phosphorodiamidates, phosphorothioates (also known as phosphothioates having a double-bonded sulfur replacing the oxygen of the phosphate), phosphodiester derivatives including phosphorodithioates, phosphonocarboxylic acids, phosphonocarboxylates, phosphonoacetic acids, phosphonoformic acids, methylphosphonates, boron phosphonates, or O-methylphosphoamidite linkages (see Eckstein, Oligonucleotides and Analogues: A Practice Approach, Oxford University Press), as well as nucleotide base variants, such as 5-methylcytidine or pseudouridine, and peptide nucleic acid skeletons and linkages. Other analog nucleic acids include those having positive skeletons, nonionic skeletons, modified sugars, and non-ribose skeletons (e.g., phosphorodiamidate morpholino oligos or locked nucleic acids (LNAs) known in this art), including those described in U.S. Patents 5,235,033 and 5,034,506, and Chapters 6 and 7 of ASC Symposium Series 580, CARBOHYDRATE MODIFICATIONS IN ANTISENSE RESEARCH, edited by Sanghui & Cook. Nucleic acids containing one or more carbocyclic sugars are also included in one definition of nucleic acids. Modifications of the ribose-phosphate skeleton may be made for various reasons, such as to increase the stability and half-life of such molecules in a physiological environment or as probes on a biochip. Mixtures of naturally occurring nucleic acids and analogs may be prepared, or alternatively, mixtures of different nucleic acid analogs and mixtures of naturally occurring nucleic acids and analogs may be prepared. In the embodiment, the linkage between nucleotides in DNA is a phosphodiester, a phosphodiester derivative, or a combination of both.

[0045] A “barcode” means one or more nucleotide sequences used to identify the cell or group of cells to which the barcode is associated. A barcode may be 3 to 1000 or more nucleotides in length, preferably 3 to 250 nucleotides, more preferably 4 to 40 nucleotides, and may include any length within this range, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. A barcode is “unique” if it is (statistically) present on about one cell in a population of cells. Cells containing barcodes can then proliferate to produce multiple cloned cells, so that each cell in the multiple cells contains the same barcode. For example, "multiple barcoded cells, each barcoded cell containing a single unique barcode" could mean a population of cells that (statistically) contain a single cell containing a given barcode or a unique combination of barcodes. Alternatively, this could mean a population of cells containing multiple clonal cell populations, where each cell in each clonal population contains the same barcode, but cells in different clonal populations contain different barcodes.

[0046] As used herein, the term “complementary” means a nucleotide (e.g., RNA or DNA) or sequence of nucleotides that can form base pairs with a complementary nucleotide or sequence of nucleotides. As described herein and as commonly known in the art, the complementary (matching) nucleotide of adenosine is thymidine, and the complementary (matching) nucleotide of guanosine is cytosine. That is, a complementary may include a sequence of nucleotides that form base pairs with the corresponding complementary nucleotide of a second nucleic acid sequence. The nucleotides of the complementary may partially or completely match the nucleotides of the second nucleic acid sequence. If the nucleotides of the complementary may completely match each nucleotide of the second nucleic acid sequence, the complementary will form base pairs with each nucleotide of the second nucleic acid sequence. If the nucleotides of the complementary may partially match the nucleotides of the second nucleic acid sequence, only some of the nucleotides of the complementary will form base pairs with the nucleotides of the second nucleic acid sequence.

[0047] As described herein, sequence complementarity may be partial, in which case only some of the nucleic acids match by base pairing, or it may be complete, in which case all of the nucleic acids match by base pairing. That is, two complementary sequences may have a percentage of identified nucleotides that are identical (i.e., about 60% identity across the identified region, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity).

[0048] As used herein, the term “gene” is used in its obvious and ordinary sense and means a segment of DNA involved in protein production. This includes the regions preceding and following the coding region (leaders and trailers), as well as the sequences interposed between individual coding segments (exons) (introns). Leaders, trailers, and introns are necessary during gene transcription and translation. control It includes elements. Furthermore, a "protein gene product" is a protein expressed from a specific gene.

[0049] The term "expression vector" refers to a gene and / or a vector necessary for gene expression. control This refers to nucleic acid molecules that code for an element. Gene expression from a vector, which may be in plasmid form, can occur in cis or trans. If a gene is expressed in cis, then that gene and control The elements are encoded by the same plasmid. Expression in trans is that of that gene and control This means that the element is encoded by a separate plasmid.

[0050] As used herein, the term “vector” means a nucleic acid molecule capable of transporting another nucleic acid to which it was ligated. A vector may be in the form of a “plasmid,” which in this context means a linear or circular double-stranded DNA loop into which further DNA segments can be ligated. Another type of vector is a viral vector, in which case further DNA segments can be ligated into the viral genome. Certain types of vectors (e.g., bacterial vectors with bacterial origins of replication and episomal mammalian vectors) can self-replicate within the host cell into which they are introduced. Other vectors (e.g., non-episomal mammalian vectors) integrate into the host cell's genome upon introduction into the host cell and thereby replicate together with the host genome. Furthermore, certain types of vectors can direct the expression of genes to which they are operably ligated. Such vectors are referred to herein as “expression vectors.” Generally, expression vectors useful in recombinant DNA techniques are often in the form of plasmids. In this specification, “plasmid” and “vector” can be used interchangeably, and plasmids are the most commonly used form of vector. However, the present invention is intended to include other forms of expression vectors that provide equivalent functionality, such as viral vectors (e.g., replication-deficient retroviruses, adenoviruses, and adeno-associated viruses). Furthermore, some viral vectors can target specific cell types, either specifically or nonspecifically. Non-replicating or replication-deficient viral vectors mean viral vectors that can infect their target cells and deliver their viral payload, but cannot proceed through the typical cytolytic pathway that leads to cell lysis and death.

[0051] According to the methods described herein, oligonucleotides, plasmids, or vectors may contain at least one selectable marker. The selectable marker for use in the methods described herein may be any suitable selectable marker. In embodiments, but not limited to, selectable markers include HygR, NsrR, ZeoR, TetA, CmR, SpR, GmR, mFabI, TmR, neoR, or kanR. In embodiments, the selectable marker is HygR. In embodiments, the selectable marker is NsrR. In embodiments, the selectable marker is ZeoR. In embodiments, the selectable marker is TetA. In embodiments, the selectable marker is CmR. In embodiments, the selectable marker is SpR. In embodiments, the selectable marker is GmR. In embodiments, the selectable marker is mFabI. In embodiments, the selectable marker is TmR. In embodiments, the selectable marker is neoR. In embodiments, the selectable marker is kanR.

[0052] According to the methods described herein, an oligonucleotide, plasmid, or vector may include at least one antiselectable marker, for example, a marker selected for integration into the recipient oligonucleotide to which a second or subsequent oligonucleotide has been recombined in the method of assembling the DNA elements described herein. The antiselectable marker for use in the methods described herein may be any suitable antiselectable marker. In embodiments, but not limited to, antiselectable markers include PheS, SacB rpsL, tolC, galK, ccdB, tetA, thyA, lacY, gata-1, URA3, relE, mqsR, chpB, vhaV, or tse2. In embodiments, the antiselectable marker is PheS. In embodiments, the antiselectable marker is SacB. In embodiments, the antiselectable marker is rpsL. In embodiments, the antiselectable marker is tolC. In embodiments, the antiselectable marker is galK. In embodiments, the antiselectable marker is ccdB. In this embodiment, the anti-selectable marker is ccdB. In this embodiment, the anti-selectable marker is tetA. In this embodiment, the anti-selectable marker is thyA. In this embodiment, the anti-selectable marker is lacY. In this embodiment, the anti-selectable marker is gata-1. In this embodiment, the anti-selectable marker is URA3. In this embodiment, the anti-selectable marker is relE. In this embodiment, the anti-selectable marker is mqsR. In this embodiment, the anti-selectable marker is chpB. In this embodiment, the anti-selectable marker is vhaV. In this embodiment, the anti-selectable marker is tse2.

[0053] The terms “transfection,” “transfer,” “transfect,” or “transfer” are interchangeable and are defined as the process of introducing nucleic acid molecules and / or proteins into cells. Nucleic acids may be introduced into cells using non-viral or viral methods. Nucleic acid molecules may be complete proteins or sequences encoding functional portions thereof. Typically, nucleic acid vectors contain elements necessary for protein expression (e.g., promoters, transcription start sites, etc.). Non-viral methods of transfection include any suitable method that does not use viral DNA or viral particles as a delivery system for introducing nucleic acid molecules into cells. Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, heat shock transfection, magnetiffection, and electroporation. For virus-based methods, any useful viral vector may be used in the methods described herein. Examples of viral vectors include, but are not limited to, retrovirus, adenovirus, lentivirus, and adeno-associated virus vectors. In some embodiments, nucleic acid molecules are introduced into cells using retroviral vectors according to standard procedures well known in this technique. The term “transfection” also refers to the introduction of a protein from the external environment into a cell. Typically, protein transfection relies on the attachment of a peptide or protein that can cross the cell membrane toward the protein of interest. See, for example, Ford et al., (2001) Gene Therapy 8: pp. 1-4 and Prochiantz (2007) Nat. Methods 4: pp. 119-20.

[0054] As used herein, the term “promoter” refers to a region of DNA that initiates the transcription of a particular gene. Promoters are typically located near the transcription start site of a gene, upstream of the gene, and on the same strand of DNA (i.e., 5' on the sense strand). Promoters may be, for example, about 100 to 1000 base pairs long.

[0055] The "position" of a nucleotide base is represented by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the 5' end. Due to deletions, insertions, excisions, fusions, and other factors that must be considered when determining optimal alignment, the numbers of amino acid residues in the test sequence, determined by simply counting from the 5' end, will generally not be the same as the numbers of the corresponding positions in the reference sequence. For example, if a variant has a deletion relative to the aligned reference sequence, there will be no nucleotide base in the variant corresponding to the position in the reference sequence at the site of the deletion. If there is an insertion in the aligned reference sequence, that insertion will not correspond to the position of a numbered nucleotide in the reference sequence. In the case of excisions or fusions, there may be nucleotide extensions in the reference sequence or aligned sequence that do not correspond to any nucleotide in the corresponding sequence.

[0056] The terms "numbered with respect to" or "corresponding to" refer, when used in relation to the numbering of a given polynucleotide sequence, to the numbering of residues in a specific reference sequence when comparing the given polynucleotide sequence with a reference sequence.

[0057] As used herein, the terms “virus” or “viral particle” are used in the context of viral transduction, according to their obvious and ordinary meanings. Viral vector transduction can be used to insert genes into mammalian cells or to modify genes within mammalian cells.

[0058] As used herein, the terms “genetic modification,” “gene alteration,” “gene editing,” “genome editing,” “genome manipulation,” and others refer to a type of genetic manipulation in which DNA is inserted, deleted, modified, or replaced at one or more specific locations in the genome of a cell. One key step in gene editing is to create a double-strand break at a specific point in a gene or genome. Examples of gene editing tools that achieve this step include, but are not limited to, zinc finger nucleases (ZFNs), transcriptional activator-like effector nucleases (TALENs), meganucleases, and clustered, regularly spaced short palindromic repeat systems (CRISPR / Cas).

[0059] As used herein, “DNA element” means any DNA sequence that can be transferred between cells, for example, between a donor cell and a recipient cell. That is, a DNA element includes, but is not limited to, a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, or gRNA. A DNA element may be a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, or gRNA fragment. A DNA element may be a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, gRNA, or a combination of a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, and gRNA fragment. In one embodiment, the DNA element is in a donor plasmid. In another embodiment, the DNA element is moved to or contained within a recipient oligonucleotide. In another embodiment, the DNA element is moved from the recipient oligonucleotide to a reset donor plasmid.

[0060] As used herein, the term "gene editing reagent" means the components necessary for a gene editing tool, and may include enzymes, riboproteins, solutions, cofactors, etc. For example, a gene editing reagent may include zinc finger nucleases (ZFNs), transcriptional activator-like effector nucleases (TALENs), meganucleases, and one or more components necessary for gene editing using clustered, regularly spaced short palindromic repeat systems (CRISPR / Cas).

[0061] As used herein, the term “endonuclease” means an enzyme or component of an endonuclease system having endonucleotide cleavage catalytic activity for cleaving polynucleotides (e.g., any component of CRISPR including gRNA). For example, an endonuclease or its component can cleave a phosphodiester bond of an oligonucleotide or polynucleotide. The endonuclease cleaves a phosphodiester bond in or near its recognition site sequence, extending at least 4 bp in length. Endonucleases include, but are not limited to, restriction enzymes, AP endonucleases, T7 endonucleases, T4 endonucleases, Bal 31 endonucleases, endonucleases I, micrococcal nucleases, endonucleases II, neurospora endonucleases, S1 endonucleases, P1-nucleases, Mung bean nucleases I, DNases I, RNA-guided DNA endonucleases (e.g., CRISPR containing any CRISPR component, e.g., Cas protein, gRNA, etc.), homosynthetic switching endonucleases, TALENs, zinc finger nucleases, and EndoR.

[0062] "Cutting" refers to the breakdown of the covalent backbone of a DNA molecule. Cutting can be initiated by a variety of methods, including but not limited to enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand breaks are possible, and a double-strand break can result from two different single-strand break events. DNA cutting can result in the production of either blunt or adherent ends. In some embodiments, a complex comprising guide RNA and a site-directed modifying enzyme is used for targeted double-strand DNA cutting.

[0063] As used herein, the terms “CRISPR” or “clustered, regularly spaced, short palindromic repeats” are used according to their obvious and ordinary meanings, referring to the genetic elements that bacteria use as a form of acquired immunity to protect themselves against viruses. CRISPR consists of short sequences originating from viral genomes and incorporated into bacterial genomes. Cas (CRISPR-associated protein) processes these sequences to cleave matching viral DNA sequences. That is, CRISPR sequences act as guides for Cas to recognize and cleave DNA that is at least partially complementary to the CRISPR sequences. By introducing a plasmid containing the Cas gene and specifically constructed CRISPR into eukaryotic cells, the eukaryotic cell genome can be cleaved at any desired location.

[0064] As used herein, the terms “Cas9” or “CRISPR-related protein 9” are used in their obvious and ordinary sense, meaning an enzyme that uses a CRISPR sequence as a guide to recognize and cleave a specific strand of DNA that is at least partially complementary to the CRISPR sequence. The Cas9 enzyme, together with the CRISPR sequence, forms the basis of a technology known as CRISPR-Cas9, which can be used to edit genes in living organisms. This editing process has a wide range of applications, including basic biological research, the development of biotechnology products, and the treatment of diseases.

[0065] In this specification, “CRISPR-related protein 9,” “Cas9,” “Csn1,” or “Cas9 protein” includes either a recombinant or naturally occurring form of Cas9 endonuclease or its variant or homolog that maintains the activity of the Cas9 endonuclease enzyme (e.g., activity within 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% compared to Cas9). In embodiments, the variant or homolog has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a continuous portion of 50, 100, 150, or 200 amino acids) compared to a naturally occurring Cas9 protein. In embodiments, the Cas9 protein is substantially identical to the protein identified by UniProt reference number Q99ZW2 or a variant or homolog substantially identical thereto. In one embodiment, the Cas9 protein has at least 75% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In another embodiment, the Cas9 protein has at least 80% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In another embodiment, the Cas9 protein has at least 85% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In another embodiment, the Cas9 protein has at least 90% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2. In yet another embodiment, the Cas9 protein has at least 95% sequence identity with the amino acid sequence of the protein identified by UniProt reference number Q99ZW2.

[0066] In this specification, “CRISPR-related endonuclease Cas12a,” “Cas12a,” “Cas12,” or “Cas12 protein” includes any recombinant or naturally occurring form of Cas12 endonuclease or its variant or homolog that maintains the activity of the Cas12 endonuclease enzyme (e.g., activity within 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% compared to Cas12). In embodiments, the variant or homolog has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a continuous portion of 50, 100, 150, or 200 amino acids) compared to a naturally occurring Cas12 protein. In one embodiment, the Cas12 protein is substantially identical to the protein identified by UniProt reference number A0Q7Q2, or to a variant or homolog having substantial identity therewith.

[0067] In this specification, “CRISPR-related endoribonuclease Cas13a,” “Cas13a,” “Cas13,” or “Cas13 protein” includes any recombinant or naturally occurring form of Cas13 endoribonuclease or its variant or homolog that maintains the activity of the Cas13 endoribonuclease enzyme (e.g., activity within 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% compared to Cas13). In some embodiments, the variant or homolog has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a continuous portion of 50, 100, 150, or 200 amino acids) compared to a naturally occurring Cas13 protein. In one embodiment, the Cas13 protein is substantially identical to the protein identified by UniProt reference number P0DPB8, or a variant or homolog having substantial identity therewith.

[0068] As used herein, "TALEN" or "transcriptional activator-like effector nuclease" refers to restriction enzymes produced by binding a DNA-binding domain (e.g., a TAL effector DNA-binding domain) to a nuclease (e.g., FokI). TALENs typically contain a naturally occurring DNA-binding domain comprising numerous modules, referred to as TAL or TALE. Specifically, TAL contains a variable of two residues that confer DNA-binding specificity.

[0069] The “guide RNA” or “gRNA” provided herein means an RNA sequence that hybridizes with a target sequence and has sufficient complementarity to the target polynucleotide sequence to direct the sequence-specific binding of the CRISPR complex to the target sequence. For example, a gRNA can direct Cas to the target polynucleotide. In embodiments, the gRNA includes crRNA and tracrRNA. For example, the gRNA may include crRNA and tracrRNA hybridized by base pairing. That is, in embodiments, two RNAs are individually encoded as two RNA molecules by crRNA and tracrRNA, which can then form an RNA / RNA complex by complementary base pairing between crRNA and tracrRNA. In embodiments, the degree of complementarity between the guide RNA sequence and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher when optimally aligned using a suitable alignment algorithm. In this embodiment, the degree of complementarity between the guide RNA sequence and its corresponding target sequence is at least approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, and 99% when optimally aligned using a suitable alignment algorithm.

[0070] Non-exclusive examples of CRISPR enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, and Csm 2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologs, or modified versions thereof. In embodiments, the CRISPR enzyme is the Cas9 enzyme. In embodiments, the Cas9 enzyme is Cas9 of S. pneumoniae, S. pyogenes, or S. thermophilus, or mutants derived from these organisms. In embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In one embodiment, the CRISPR enzyme directs the cleavage of one or two strands at the location of the target sequence. In another embodiment, the CRISPR enzyme lacks DNA strand cleavage activity.

[0071] As used herein, "zinc finger" is a polypeptide structural motif folded around a bound zinc cation. In embodiments, the polypeptide of the zinc finger is of form X3-Cys-X 2-4 -Cys-X 12 -His-X 3-5 It has the sequence -His-X4, where X is any amino acid (e.g., X 2-4 This represents an oligopeptide with a length of 2 to 4 amino acids. That is, as used herein, "zinc finger nuclease" means a nuclease that contains a zinc finger motif and a domain capable of inducing cleavage in target DNA.

[0072] The term “homologous recombination” means a type of genetic recombination in which information is exchanged between two similar or identical nucleic acid sequences, which may be referred to herein as “homologous regions.” In some embodiments of the methods described herein, the homologous regions may include two homologous regions, which may, for example, be optionally adjacent to non-homologous regions. In embodiments of the methods described herein, the E. coli RecA gene may be used to enhance homologous recombination. “RecA” means a bacterial homolog of a family of universal 38kD homologous DNA repair proteins that mediate ATP-dependent homologous recombination in bacteria. In embodiments, the donor or recipient cells of the methods described herein contain one or more homologous DNA repair genes, such as oligonucleotides encoding RecA. In embodiments, the expression of the homologous DNA repair gene is inducible. In embodiments, the homologous DNA repair gene is RecA. In embodiments, the homologous DNA repair gene is the recombination gene, Redα, Redβ, and Redγ. Non-limiting examples of homologous recombination and gene editing methods using various nuclease systems can be found, for example, in U.S. Patent No. 8,945,839, International PCT Application Publication No. 2013 / 163394, and U.S. Patent Application Publication No. 2016 / 0060657, No. 2012 / 0192298A1, and No. 2007 / 0042462. These and other known methods for homologous recombination can be used in combination with the methods described herein.

[0073] As used herein, the term “transfection” is used in its obvious and ordinary sense, meaning the process of intentionally introducing naked or purified nucleic acids into eukaryotic cells. In practice, “transfection” may also mean other methods and cell types, but other terms are often preferred. For example, the term “transformation” is typically used to describe nonviral DNA transfer in non-animal eukaryotic cells, including bacterial and plant cells. In animal cells, transfection is the preferred term. For example, the term “transduction” is often used to describe virus-mediated gene transfer into eukaryotic cells.

[0074] The terms "bacterial conjugation" and "bacterial mating" are interchangeable and refer to a mode of gene exchange between bacteria. Typically, bacterial conjugation involves only a portion of one genome from a cell (donor) and the complete genome of its partner (recipient cell). That is, gene transfer in bacterial conjugation is typically partial. In embodiments, bacterial conjugation is the transfer of non-genomic bacterial DNA from a donor cell to a recipient cell. In examples, bacterial conjugation occurs via plasmids. In examples, bacterial conjugation occurs via exogenous DNA in bacteria. In embodiments, donor and recipient cells come into contact for bacterial conjugation to occur. In embodiments, donor and recipient cells include a ligation bridge (e.g., pili) for bacterial conjugation to occur.

[0075] In embodiments, recipient cells or donor cells contain oligonucleotides that enable plasmid conjugation. In embodiments, the oligonucleotides that enable plasmid conjugation are located in the genome of the donor cells. In embodiments, the oligonucleotides that enable plasmid conjugation are located in a helper plasmid. In embodiments, the oligonucleotides that enable plasmid conjugation are the Tra operon. In this embodiment, the oligonucleotides that enable plasmid conjugation are IncF1 Tra (traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traR, traS, traT), IncP Tra operon: (trbA, trbB, trbC, trbD, trbE, trbF, trbG, trbH, trbI, trbJ, trbK, trbL, traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO), IncI1 The tra operon is selected from (traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traS, traT, traU, traV, traW, traY), the pTiC58 tra gene is selected from (traA, traF, traB, traC, traG, traD, traR, traI), and pIJ101 is selected from clt, korB. In this embodiment, the oligonucleotide that enables plasmid conjugation is IncF1 Tra (traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traR, traS, traT).In one embodiment, the oligonucleotides that enable plasmid conjugation are the IncP Tra operon: (trbA, trbB, trbC, trbD, trbE, trbF, trbG, trbH, trbI, trbJ, trbK, trbL, traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO). In another embodiment, the oligonucleotides that enable plasmid conjugation are the IncI1 tra operon: (traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traS, traT, traU, traV, traW, traY). In yet another embodiment, the oligonucleotides that enable plasmid conjugation are the pTiC58 tra gene: (traA, traF, traB, traC, traG, traD, traR, traI). In this embodiment, the oligonucleotide that enables plasmid conjugation is pIJ101:clt,korB.

[0076] As used herein, “donor cell” means a cell that transfers genetic material to another cell (e.g., a bacterial cell, a plant cell, or another). The cell that receives the transferred genetic material is referred to herein as “recipient cell.”

[0077] As used herein, the term "donor plasmid" refers to donor cell-derived DNA containing an oligonucleotide sequence (e.g., donor DNA, an oligonucleotide containing a DNA element) that is transferred from a donor cell (e.g., a bacterial cell) to a recipient cell (e.g., a bacterial cell, yeast cell, plant cell, etc.). Typically, a donor plasmid is a circular double-stranded DNA separated from genomic DNA. That is, the term "recipient plasmid" refers to DNA from a recipient cell that receives the donor DNA. In embodiments, the donor plasmid-derived DNA is received by DNA other than the recipient plasmid-derived DNA. That is, in embodiments, the donor DNA may be incorporated into the genomic DNA.

[0078] In an embodiment, the donor plasmid includes a transport origin. In an embodiment, the transport origin is derived from a mobile element. In an embodiment, the mobile element is a plasmid. In an embodiment, the plasmid is an IncFI plasmid, an IncPα plasmid, an IncI1 plasmid, a pTiC58 plasmid derived from Agrobacterium tumefaciens, a pAD1 plasmid, an Inc18 plasmid, or an IncH plasmid. In an embodiment, the plasmid is an IncFI plasmid. In an embodiment, the plasmid is an IncPα plasmid. In an embodiment, the plasmid is an IncI1 plasmid. In an embodiment, the plasmid is a pTiC58 plasmid derived from Agrobacterium tumefaciens. In an embodiment, the plasmid is a pAD1 plasmid. In an embodiment, the plasmid is an Inc18 plasmid. In an embodiment, the plasmid is an IncH plasmid. Plasmids are described in Ippen-Ihler, KA and Minkley, EG, Jr., (1986) The conjugation system of F, the fertility factor of Escherichia coli. Ann. Rev. Genet. 20: pp. 593-624; Guiney, DG and Lanka, E., (1989) Conjugative transfer of IncP plasmids, in: Promiscuous Plasmids of Gram-negative Bacteria (edited by CM Thomas), Academic Press, London, pp. 27-56; Catherine EDRees, David E. Bradley, Brian M. Wilkins, (1987) Organization and regulation of the conjugation genes of IncI1 plasmid ColIb-P9. Plasmid. 18: pp. 223-236; von Bodman SB, McCutchan JE, Farrand SK.(1989)Characterization of conjugal transfer functions of Agrobacterium tumefaciens Ti plasmid pTiC58.J.Bacteriol.171(10):5281~5289;Clewell DB, Weaver KE.(1989)Sex pheromones and plasmid transfer in Enterococcus faecalis.Plasmid.21(3):pp.175~84;Kohler V,Vaishampayan A,Grohmann E.(2018)Broad-host-range Inc18 plasmids:Occurrence,spread and transfer mechanisms.Plasmid.99:pp.11~21;Andreas Schluter,Patrice Nordmann,Remy A.Bonnin,Yves Millemann,Felix G. Eikmeyer, Daniel Wibberg, Alfred, Puhler, Laurent Poirel.(2014)IncH-Type Plasmid Harboring bla. CTX-M-15 bla DHA-1 This is discussed in more detail in *QNR B4 Genes Recovered from Animal Isolates*, *Antimicrobial Agents and Chemotherapy*, 58(7): pp. 3768-3773. The entire contents of these references are incorporated herein by reference for all purposes.

[0079] In an embodiment, the origin of transport is a mobile element. In an embodiment, the mobile element is a conjugating transposon. In an embodiment, the conjugating transposon is Tn916 derived from Enterococcus faecalis or CTnDOT derived from Bacteroides. In an embodiment, the conjugating transposon is Tn916 derived from Enterococcus faecalis. In an embodiment, the conjugating transposon is CTnDOT derived from Bacteroides. In an embodiment, the mobile element is derived from an integrated conjugating element. In an embodiment, the mobile element is SXT derived from Vibrio cholera or R391 derived from Providencia rettgeri. In an embodiment, the mobile element is SXT derived from Vibrio cholera. In an embodiment, the mobile element is R391 derived from Providencia rettgeri.The elements are described in the following references: Rice LB (1998). Tn916 family conjugative transposons and dissemination of antimicrobial resistance determinants. Antimicrobial agents and chemotherapy, 42(8), pp. 1871-1877; Cheng Q, Paszkiet BJ, Shoemaker NB, Gardner JF, Salyers AA. (2000) Integration and excision of a Bacteroides conjugative transposon, CTnDOT. J Bacteriol. 182(14): pp. 4035-43; Bianca Hochhut and Matthew K. Waldor. (1999) Site-specific integration of the conjugal Vibrio cholerae SXT element into prfC.Mol.Microbiology. 32(1): pp. 99-110; Boltner D, MacMahon C, Pembroke JT, Strike P, Osborn AM.R391: a conjugative integrating mosaic comprised of phage, plasmaid, and transposon elements. J Bacteriol. 2002;184(18):5158-5169, which is described in more detail, and each of these is incorporated herein by reference as a whole.

[0080] In embodiments, the donor plasmid includes a conditional origin of replication. In embodiments, the conditional replicons are R6K-pir, RSF1010 oriV-RepA / B / C, ColE2 P9-RepA, RP4 oriV-trfA, pPS10 oriV-RepA, and pSC101 ori-RepC. TSThe conditional replicon is RK2 oriV, bacteriophage P1 ori, plasmid pSC101 replication origin, bacteriophage lambda ori, pBR322 plasmid, pSU739 plasmid, or pSU300 plasmid. In embodiments, the conditional replicon is R6K-pir. In embodiments, the conditional replicon is RSF1010 oriV-RepA / B / C. In embodiments, the conditional replicon is ColE2 P9-RepA. In embodiments, the conditional replicon is RP4 oriV-trfA. In embodiments, the conditional replicon is pPS10 oriV-RepA. In embodiments, the conditional replicon is pSC101 ori-RepC TSIn an embodiment, the conditional replicon is RK2 oriV. In an embodiment, the conditional replicon is bacteriophage P1 ori. In an embodiment, the conditional replicon is the origin of plasmid pSC101 replication. In an embodiment, the conditional replicon is bacteriophage lambda ori. In an embodiment, the conditional replicon is pBR322 plasmid. In an embodiment, the conditional replicon is pSU739 plasmid. In an embodiment, the conditional replicon is pSU300 plasmid. Plasmids are referenced in: Metcalf WW, Jiang W, Daniels LL, Kim SK, Haldimann A, Wanner BL. (1996) Conditionally replicative and conjugative plasmids carrying lacZ alpha for cloning, mutagenesis, and allele replacement in bacteria. Plasmid. 35(1): pp. 1-13; Scherzinger E, Bagdasarian MM, Scholz P, Lurz R, Ruckert B, Bagdasarian M. (1984) Replication of the broad host range plasmid RSF1010: requirement for three plasmid-encoded proteins. Proc Natl Acad Sci USA. 81(3): pp. 654-8; ColE2-P9: Yagura M, Nishio SY, Kurozumi H, Wang CF, Itoh T. (2006) Anatomy of the replication origin of plasmid ColE2-P9. Bacteriol.188(3):999~1010; Ayres EK, Thomson VJ, Merino G, Balderes D, Figurski DH. Precise deletions in large bacterial genomes by vulector-mediated excision(VEX).(1993) The trfA gene of promiscuous plasmid RK2 is essential for replication in several gram-negative hosts. J Mol Biol. 5; 230(1): 174 - 85; Maestro B, Sanz JM, Diaz-Orejas R, Fernandez-Tresguerres E. (2003) Modulation of pPS10 host range by plasmid-encoded RepA initiator protein. J Bacteriol. 185(4): 1367 - 75; Hashimoto-Gotoh, T., & Sekiguchi, M. (1977). Mutations of temperature sensitivity in R plasmid pSC101. Journal of bacteriology, 131(2), 405 - 412; Ayres EK, Thomson VJ, Merino G, Balderes D, Figurski DH. Precise deletions in large bacterial genomes by vulector-mediated excision (VEX). (1993) The trfA gene of promiscuous plasmid RK2 is essential for replication in several gram-negative hosts. J Mol Biol. 5; 230(1): 174 - 85 Stenzel TT, Patel P, Bastia D. (1987) The integration host factor of Escherichia coli binds to bent DNA at the origin of replication of the plasmid pSC101. Cell. 5; 49(5): 709 - 17; Sugiura S, Ohkubo S, Yamaguchi K.(1993)Minimal essential origin of plasmid pSC101 replication:requirement of a region downstream of iterons.J Bacteriol.175(18):5993~6001;Pal SK,Mason RJ,Chattoraj DK.(1986)P1 plasmid replication.Role of initiator titration in copy number control.J Mol Biol.20;192(2):275~85;LeBowitz JH,McMacken R.(1984)The bacteriophage lambda O and P protein initiators promote the replication of single-stranded DNA.Nucleic Acids Res.12(7):3069~3088;Grindley ND,Kelley WS.(1976)Effects of different alleles of the E.coli K12 pol A gene on the replication of non-transferring plasmids.Mol Gen This information is described in Genet. 2; 143(3): pp. 311-318; Francia, MV, & Garcia Lobo, JM (1996). Gene integration in the Escherichia coli chromosome mediated by Tn21 integrase (Int21). Journal of bacteriology, 178(3), pp. 894-898; and Mendiola MV, de la Cruz F. (1989). Specificity of insertion of IS91, an insertion sequence present in alpha-haemolysin plasmids of Escherichia coli. Mol Microbiol. 3(7): pp. 979-974. The references are incorporated herein by reference.

[0081] In an embodiment, the conditional replication origin depends on the presence of an oligonucleotide. In an embodiment, the oligonucleotide encodes pir1, pir1-116, repA / repB / repC (RSF1010 replicon), repA (ColE2-P9 replicon), trfA (RP4 replicon), RepA (pSP10 replicon), RepC TS (pSC101 replicon), or a combination thereof. In an embodiment, the oligonucleotide encodes pir1. In an embodiment, the oligonucleotide encodes pir1-116. In an embodiment, the oligonucleotide encodes repA / repB / repC (RSF1010 replicon). In an embodiment, the oligonucleotide encodes repA (ColE2-P9 replicon). In an embodiment, the oligonucleotide encodes trfA (RP4 replicon). In an embodiment, the oligonucleotide encodes RepA (pSP10 replicon). In an embodiment, the oligonucleotide encodes RepC TS (pSC101 replicon).

[0082] In an embodiment, the conditional replication origin depends on the conditions of cell growth. In an embodiment, the condition is temperature.

[0083] For the methods provided herein, in an embodiment, the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid from 20 or 30 kilobases in length.

[0084] For the methods provided herein, in embodiments, the donor plasmid or recipient oligonucleotide includes a replicon capable of replicating plasmids longer than 30 kilobases. In embodiments, the replicon can replicate plasmids from about 30 kilobases to about 500 kilobases in length. In embodiments, the replicon can replicate plasmids about 30 kilobases in length. In embodiments, the replicon can replicate plasmids about 50 kilobases in length. In embodiments, the replicon can replicate plasmids about 70 kilobases in length. In embodiments, the replicon can replicate plasmids about 90 kilobases in length. In embodiments, the replicon can replicate plasmids about 100 kilobases in length. In embodiments, the replicon can replicate plasmids about 120 kilobases in length. In embodiments, the replicon can replicate plasmids about 140 kilobases in length. In embodiments, the replicon can replicate plasmids about 160 kilobases in length. In embodiments, the replicon can replicate plasmids about 180 kilobases in length. In embodiments, the replicon can replicate plasmids about 200 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 220 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 240 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 260 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 280 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 300 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 400 kilobases in length. In an embodiment, the replicon can replicate a plasmid approximately 500 kilobases in length. The length may be any value or a subrange within the indicated range, including the endpoint.

[0085] In embodiments, the replicon is derived from a P1-derived artificial chromosome or a bacterial artificial chromosome. In embodiments, the replicon is derived from a P1-derived artificial chromosome. In embodiments, the replicon is derived from a bacterial artificial chromosome. In embodiments, the donor plasmid or recipient oligonucleotide contains an inducible high-copy origin. In embodiments, the donor plasmid contains an inducible high-copy origin. In embodiments, the recipient oligonucleotide contains an inducible high-copy origin.

[0086] In embodiments, the donor plasmid or recipient oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome. In embodiments, the donor plasmid is a yeast artificial chromosome (YAC). In embodiments, the donor plasmid is a mammalian artificial chromosome (MAC). In embodiments, the donor plasmid is a human artificial chromosome (HAC). In embodiments, the donor plasmid is a plant artificial chromosome. In embodiments, the recipient oligonucleotide is a yeast artificial chromosome (YAC). In embodiments, the recipient oligonucleotide is a mammalian artificial chromosome (MAC). In embodiments, the recipient oligonucleotide is a human artificial chromosome (HAC). In embodiments, the recipient oligonucleotide is a plant artificial chromosome.

[0087] In embodiments, the donor plasmid or recipient oligonucleotide comprises a conjugate-capable vector, which may be a viral vector. In embodiments, the donor plasmid is a viral vector. In embodiments, the recipient oligonucleotide is a viral vector. In embodiments, the viral vector is a retrovirus. In embodiments, the viral vector is a lentivirus. In embodiments, the viral vector is an adenovirus. In embodiments, the viral vector is an adeno-associated virus. In embodiments, the viral vector is a tobacco mosaic virus. In embodiments, the viral vector is a baculovirus. In embodiments, the viral vector is a herpes simplex virus. In embodiments, the viral vector is a poxvirus. In embodiments, the viral vector is a gamma retrovirus. In embodiments, the viral vector is a Sendai virus.

[0088] As used herein, the terms “control” or “control experiment” are used in their obvious and ordinary sense, meaning an experiment in which the subject or reagents of the experiment are treated in the same way as in a parallel experiment, except that the experimental procedure, reagents, or variables are omitted. In some examples, the control is used as a standard of comparison in the evaluation of experimental effects.

[0089] A "control" sample or value refers to a sample that serves as a reference, usually a known reference, for comparison with the test sample. For example, a test sample may be taken from test conditions, for example, in the presence of the test compound, and compared to a sample obtained from known conditions, for example, in the absence of the test compound (negative control) or in the presence of a known compound (positive control). A control may also represent the average value collected from several tests or results. Those skilled in the art will recognize that controls can be designed to evaluate any number of parameters. For example, controls can be designed to compare therapeutic benefits based on pharmacological data (e.g., half-life) or therapeutic means (e.g., comparison of side effects). Those skilled in the art will understand which controls are valuable in a given situation and can analyze data based on comparison with control values. Controls are also valuable for determining the significance of data. For example, if the values ​​for a given parameter vary widely in the controls, the variation in the test sample may not be considered significant.

[0090] As used herein, the term “contact” is used in its obvious and ordinary sense, meaning a process that allows at least two obviously different species (e.g., compounds or cells containing biomolecules) to come into close enough proximity to react, interact, or physically come into contact. However, it should be recognized that the reaction products obtained may be produced directly from the reaction between the added reagents, or from intermediates from one or more of the added reagents that may be produced in the reaction mixture.

[0091] The term "expression" includes, but is not limited to, any steps involved in polypeptide production, including transcription, post-transcriptional modification, translation, post-translational modification, and secretion. Expression can be detected using conventional methods for detecting proteins (e.g., ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, etc.).

[0092] The term "recombinant," when used in relation to cells, nucleic acids, proteins, or vectors, indicates that the cell, nucleic acid, protein, or vector has been modified by the introduction of a different nucleic acid or protein, or by alteration of a native nucleic acid or protein, or that the cell has been derived from such a modified cell. That is, for example, recombinant cells express genes not found in the native (non-recombinant) cell form, or express native genes that would otherwise be abnormally expressed, reduced in expression, or not expressed at all. Transgenic cells and plants are typically cells or plants that express a different gene or coding sequence as a result of a recombination method.

[0093] As used herein, the terms “transfer origin” or “oriT” refer to a short sequence (up to 500 bp) necessary for the transfer of DNA from the bacterial host and recipient during bacterial conjugation.

[0094] As used herein, “cureable origin of replication” means an origin of replication that does not replicate when cells proliferate in the presence of certain chemical or environmental conditions. Under these conditions, plasmids containing a cureable origin of replication are lost from the cell. For example, pSC101 ori TS It does not function and is lost at high temperatures.

[0095] As used herein, the term “mobile element” refers to a type of genetic material that can move around within a genome or be transported between genomes, even between species.

[0096] As used herein, the term “conjugating transposon” means an integrated DNA element that can cleave itself to form a covalently circumferential intermediate that can be reassembled within the same cell or transported via conjugation with a recipient cell.

[0097] As used herein, the term “integration of conjugated elements” refers to a group of self-propagating genetic elements integrated into a chromosome.

[0098] As used herein, the term "P1-induced artificial chromosome" means a DNA construct derived from the P1 bacteriophage.

[0099] As used herein, “bacterial artificial chromosome” means an engineered DNA sequence used to clone a DNA sequence into a bacterium.

[0100] The term “recombination-mediated gene” or “recombinant gene” refers to a gene that facilitates the creation of genetic alterations in a DNA sequence. In practice, a recombination-mediated gene enables the in vivo construction of constructs in cells (e.g., bacterial cells) without in vitro genetic engineering techniques. In practice, a recombination-mediated gene enables genetic alteration without the introduction of enzymes, including ligases and restriction enzymes. For example, a gene can participate in the innate process of homologous recombination in bacteria without the use of conventional molecular biology techniques known in this technology. In an embodiment, the gene to be recombined is the lambda red gene. The genes to be recombined are Redα, Redβ, and Redγ. For example, the gene can induce homologous recombination at a high rate in bacteria. In donor or recipient cells, the expression of one or more recombinant genes is inducible. In an embodiment, the donor or recipient cell contains oligonucleotides encoding one or more recombination-mediated genes. In an embodiment, the oligonucleotides encoding one or more recombination-mediated genes are located in a donor cell plasmid. In embodiments, the recombination-mediated gene is inducible. In embodiments, the recombination-mediated gene is Redα, Redβ, and Redγ. In embodiments, one or more oligonucleotides encoding homologous DNA repair genes are located in the recipient cell genome. In embodiments, one or more oligonucleotides encoding homologous DNA repair genes are located in a helper plasmid. In embodiments, one or more oligonucleotides encoding homologous DNA repair genes are located in a recipient oligonucleotide, which may be in plasmid form. In embodiments, one or more oligonucleotides encoding homologous DNA repair genes are located in the recipient cell genome.

[0101] As used herein, the term "inducible high-copy origin of replication" means a plasmid or vector containing a large number of origins of replication (e.g., 150-200 copies in an E. coli plasmid pUC) that can be induced by environmental conditions such as temperature changes.

[0102] As used herein, the term “helper plasmid” refers to a plasmid containing genes or other DNA elements necessary for a bacterium to perform a specific function. A helper plasmid may include an endonuclease, which is an element for transporting foreign DNA into the genome, transporting a plasmid to another cell, or performing homologous recombination. In embodiments, the helper plasmid is IncF1 plasmid, IncPα plasmid, IncI1 plasmid, pTiC58 from Agrobacterium tumefaciens, cAD1 plasmid, Inc18 plasmid, pIJ101 from Streptomyces, or IncH plasmid. In embodiments, the helper plasmid is IncF1 plasmid. In embodiments, the helper plasmid is IncPα plasmid. In embodiments, the helper plasmid is IncI1 plasmid. In embodiments, the helper plasmid is pTiC58 from Agrobacterium tumefaciens. In embodiments, the helper plasmid is cAD1 plasmid. In embodiments, the helper plasmid is the Inc18 plasmid. In embodiments, the helper plasmid is the Streptomyces-derived pIJ101. In embodiments, the helper plasmid is the IncH plasmid. In embodiments, the helper plasmid lacks a functional transport origin. In embodiments, the helper plasmid includes a selectable marker for selecting the retention of the helper plasmid in donor cells.

[0103] As used herein, the term “homing endonuclease” means an endonuclease encoded as an autonomous gene in an intronic sequence, as a fusion with a host protein, or as a self-splicing protein. Homing endonucleases catalyze the hydrolysis of DNA at longer recognition sites compared to group II restriction enzymes. Examples of homing endonucleases include, but are not limited to, LAGLIDAG, GIY-YIG, His-Cys box, HNH, PD-(D / E)xK, and Vsr-like / EDxHD. In embodiments, the homing endonucleases include I-ScaI, PI-SceI, I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-Po rI, I-PpoI, PI-PspI, I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I -Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I, or I-Vdi141I. In an embodiment, the homing endonuclease is I-ScaI. In an embodiment, the homing endonuclease is PI-SceI. In an embodiment, the homing endonuclease is I-AniI. In an embodiment, the homing endonuclease is I-CeuI. In an embodiment, the homing endonuclease is I-ChuI. In an embodiment, the homing endonuclease is I-CpaI. In an embodiment, the homing endonuclease is I-CpaII. In an embodiment, the homing endonuclease is I-CreI. In an embodiment, the homing endonuclease is I-DmoI. In an embodiment, the homing endonuclease is H-DreI. In an embodiment, the homing endonuclease is I-HmuI. In an embodiment, the homing endonuclease is I-HmuII. In an embodiment, the homing endonuclease is I-LlaI.In an embodiment, the homing endonuclease is I-MsoI. In an embodiment, the homing endonuclease is PI-PfuI. In an embodiment, the homing endonuclease is PI-PkoII. In an embodiment, the homing endonuclease is I-PorI. In an embodiment, the homing endonuclease is I-PpoI. In an embodiment, the homing endonuclease is PI-PspI. In an embodiment, the homing endonuclease is I-SceI. In an embodiment, the homing endonuclease is I-SceII. In an embodiment, the homing endonuclease is I-SceIII. In an embodiment, the homing endonuclease is I-SceIV. In an embodiment, the homing endonuclease is I-SceV. In an embodiment, the homing endonuclease is I-SceVI. In an embodiment, the homing endonuclease is I-SceVII. In an embodiment, the homing endonuclease is I-Ssp6803I. In an embodiment, the homing endonuclease is I-TevI. In an embodiment, the homing endonuclease is I-TevII. In an embodiment, the homing endonuclease is I-TevIII. In an embodiment, the homing endonuclease is PI-TliI. In an embodiment, the homing endonuclease is PI-TliII. In an embodiment, the homing endonuclease is I-Tsp061I or I-Vdi141I. In an embodiment, the homing endonuclease is I-Vdi141I.

[0104] As used herein, the term “RNA-guided DNA endonucleases” means any DNA endonucleases that are guided to a target DNA sequence by a helper or guide RNA molecule. Examples of RNA-guided DNA endonucleases include, but are not limited to, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, and Csc2. This includes Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4, as well as all their variants and homologs.

[0105] As used herein, the term "HO" or "homogeneous interconversion endonuclease" refers to the zinc finger nuclease in Saccharomyces cerevisiae that is involved in initiating mating type interconversion.

[0106] As used herein, “cell” means a cell that performs sufficient metabolic or other functions to preserve or replicate its genomic DNA. Cells can be identified by methods known in this art, including, for example, the presence of an intact membrane, staining with a specific dye, the ability to produce offspring, or, in the case of gametes, the ability to combine with a second gamete to produce viable offspring. Cells may include prokaryotic cells and eukaryotic cells. Prokaryotic cells include, but are not limited to, bacteria. Eukaryotic cells include, but are not limited to, yeast cells and cells derived from plants and animals, such as mammalian, insect (e.g., Spodoptera), and human cells. Cells may be useful if they are naturally non-adherent or treated to prevent adhesion to a surface, for example, by trypsin treatment.

[0107] As used herein, “donor cell” means a cell that transfers genetic material to another cell (e.g., a bacterial cell, a plant cell, or another). The cell that receives the transferred genetic material is referred to herein as “recipient cell.”

[0108] As used herein, the term "donor plasmid" means donor cell-derived DNA containing an oligonucleotide sequence (e.g., donor DNA, an oligonucleotide containing a DNA element or fragment thereof) that is transferred from a donor cell (e.g., a bacterial cell) to a recipient cell (e.g., a bacterial cell, yeast cell, plant cell, etc.). Typically, a donor plasmid is a circular double-stranded DNA separated from genomic DNA. That is, the term "recipient oligonucleotide" means plasmid DNA in a recipient cell that receives donor DNA (e.g., by homologous recombination of donor DNA into the recipient), or the term "recipient oligonucleotide" may mean any oligonucleotide in a recipient cell that receives donor DNA, e.g., recipient cell genomic DNA. In embodiments, the donor plasmid-derived DNA is received by DNA other than the recipient plasmid-derived DNA. That is, in embodiments, the donor DNA may be incorporated into genomic DNA.

[0109] The term "isolated," when applied to nucleic acids or proteins, means that the nucleic acid or protein essentially does not contain other cellular components that would naturally accompany it. For example, it may be in a homogeneous state, a dry solution, or an aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The protein, which is the dominant species present in the preparation, is substantially purified.

[0110] The examples and embodiments described herein are for illustrative purposes only, and it will be understood that various modifications or changes in view thereof will be suggested to those skilled in the art and should be included in the spirit and scope of this application and the scope of the attached claims. All publications, patents and patent applications referenced herein are incorporated herein by reference in their entirety for all purposes.

[0111] In embodiments of the methods described herein, donor or recipient cells contain oligonucleotides encoding one or more homologous DNA repair genes. In embodiments, the oligonucleotides encoding one or more homologous DNA repair genes are located in a first, second, or subsequent donor plasmid. In embodiments, homologous DNA repair gene expression is inducible. In embodiments, the homologous DNA repair gene is RecA.

[0112] In embodiments of the methods described herein, the donor or recipient cells contain oligonucleotides encoding one or more recombinant gene-mediated genetic manipulation genes. In embodiments, the oligonucleotides encoding one or more recombinant gene-mediated genetic manipulation genes are located within the donor cell plasmid.

[0113] In embodiments of the methods described herein, the recombinant-mediated gene manipulation gene is inducible. In embodiments, the recombinant-mediated gene manipulation genes are Redα, Redβ, and Redγ. In embodiments, oligonucleotides encoding one or more homologous DNA repair genes are located in the recipient cell genome. In embodiments, oligonucleotides encoding one or more homologous DNA repair genes are located in a helper plasmid. In embodiments, oligonucleotides encoding one or more homologous DNA repair genes are located in a recipient oligonucleotide, which may be in plasmid form. In embodiments, oligonucleotides encoding one or more homologous DNA repair genes are located in the recipient cell genome.

[0114] In embodiments of the methods described herein, donor cells, recipient cells, or recombinant recipient cells may be in an ordered array or in a first or second ordered array. In embodiments, donor cells, recipient cells, or recombinant recipient cells may be transported to a position on a third ordered array, a fourth ordered array, or a subsequent ordered array.

[0115] In embodiments of the methods described herein, the donor cells and recipient cells are bacterial cells. In embodiments, the recipient cells are not bacterial cells. In embodiments, the recipient cells are plant cells. In embodiments, the recipient cells are yeast cells. In embodiments, the recipient cells are mammalian cells.

[0116] II. How to Assemble DNA Elements The present invention provides a method for assembling multiple DNA elements into an assembled DNA element in a recipient cell, the method comprising the steps of (a) (i) transferring a first donor plasmid from a first donor cell to a recipient cell by conjugation, and (ii) contacting a first donor cell containing a first donor plasmid with a recipient cell containing a recipient oligonucleotide under conditions that allow the first donor plasmid and the recipient oligonucleotide to be recombined in the recipient cell by homologous recombination, wherein the first donor plasmid Rasmid comprises, in a sequential order, an optional first endonuclease site (C1), a first homologous recombination region (HR1), a first DNA element fragment (oligo 1) as a first oligonucleotide, a second homologous recombination region (HR2) containing two homologous recombination regions (HR2.1, HR2.2), and an optional third endonuclease site (C3), and the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2, thereby HR1, HR3, and HR2 (b) The steps of providing a first recombined recipient oligonucleotide containing a fragment of the first DNA element, following homologous recombination of .2 and HR4, (i) transferring the second donor plasmid from the second donor cell to the first recipient cell by conjugation, and (ii) recombining the second donor plasmid and the first recombined recipient oligonucleotide in the recipient cell by homologous recombination to form the second recombined recipient oligonucleotide, under the conditions that the second donor cell containing the second donor plasmid and the first recombined recipient oligonucleotide A step of contacting recipient cells containing a bound recipient oligonucleotide, wherein the second donor plasmid comprises, in a sequential order, an optional fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a fragment of a second DNA element (oligo 2), a sixth homologous recombination region (HR6) containing two homologous recombination regions (HR6.1, HR6.2), and an optional sixth endonuclease site (C6), thereby connecting HR5 with HR2.1 and HR6.The method includes a step in which homologous recombination of 2 and HR4 is followed by providing a second recombined recipient oligonucleotide containing fragments of the first and second DNA elements (oligo 1, oligo 2), which together form a DNA assembly. In embodiments, HR2.1 and HR2.2 are adjacent to a non-homologous region containing one (C2) or two (C2.1, C2.2) endonuclease sites, and optionally, HR3 and HR4 are adjacent to a non-homologous region containing one (C4) or two (C4.1, C4.2) endonuclease sites. In embodiments, HR6.1 and HR6.2 are adjacent to a non-homologous region containing one (C7) or two (C7.1, C7.2) endonuclease sites. In embodiments, the recipient oligonucleotides are located in a recipient cell plasmid or recipient cell genome. In embodiments, the DNA assembly includes at least some of a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, guide RNA (gRNA), or a combination thereof. In embodiments, step (b) is repeated one or more times using a third or subsequent donor cell containing a third or subsequent donor plasmid containing a third or subsequent oligonucleotide (oligo 3, oligo 4, ... oligo N) encoding a matching HR region and a fragment of a third or subsequent DNA element, thereby forming a third or subsequent recombined recipient oligonucleotide containing fragments of the first, second, and third or subsequent DNA elements, which together form a DNA assembly. In the embodiment, step (a) comprises a plurality of first donor cells each containing a different first donor plasmid, and step (b) comprises a plurality of second, third, or subsequent donor cells each containing a different second, third, or subsequent donor plasmid, wherein each first donor cell is located in a first ordered array, and each second, third, or subsequent donor cell is located in a second, third, or subsequent ordered array, wherein the method optionally generates a combinatorial library containing a plurality of different assembled DNA elements.

[0117] In an embodiment, a donor plasmid containing the last DNA element that forms part of an assembled DNA element contains a BHR region that produces recipient cells each containing an assembled DNA element, a barcode homologous recombination (BHR) region, and a recombined recipient oligonucleotide containing a further HR, the method further comprising the steps of (i) constructing or obtaining an array of barcode donor cells each containing a barcode donor plasmid containing a BHR homologous HR, a specific barcode oligonucleotide, and a second HR homologous to a further HR of the recombined recipient oligonucleotide; and (ii) contacting the array of barcode donor cells and the array of recipient cells under conditions that (a) transfer the barcode donor plasmid from the barcode donor cells to recipient cells by conjugation, and (b) recombining the barcode donor plasmid and recipient oligonucleotide in the recipient cells by homologous recombination, thereby producing an array of recipient cells containing a barcoded assembly.

[0118] In embodiments, each donor plasmid contains a further pair of specific endonuclease sites CX and CY adjacent to a barcode homologous recombination (BHR) region, and the method further comprises the step of contacting an array of recipient cells, each containing a DNA assembly, with an array of barcode donor cells, each containing a barcode donor plasmid, each containing a pair of BHR homologous HR regions adjacent to a specific barcode oligonucleotide, to produce an array of recipient cells containing barcoded assemblies.

[0119] In embodiments, the DNA assembly method may further include the step of contacting reset donor cells containing a reset donor plasmid with recipient cells containing a recombined recipient oligonucleotide, wherein the reset donor plasmid comprises, in sequential order, a homologous recombination region (HRt) homologous to the terminal sequence of the DNA assembly, a reset endonuclease site, a selectable marker, a reset endonuclease site, a homologous recombination region (HRX), and a transport origin, and the recombined recipient oligonucleotide comprises, in sequential order, a reset endonuclease site, the DNA assembly, a homologous recombination region (HRXa) homologous to the HRX, and a reset endonuclease site, thereby providing a reset plasmid containing a transport origin and the DNA assembly, following homologous recombination between the HRt and the terminal sequence of the DNA assembly and between the HRX and HRXa. In embodiments, the reset plasmid is located in the donor cells. In embodiments, the reset plasmid includes a restricted origin of replication that functions in both the donor and recipient cells. In embodiments, the reset donor plasmid is constructed by introducing an oligonucleotide insert HRt-C1-CM-C2-HRX, or a library of such oligonucleotide inserts, which include two endonuclease sites (C1, C2) and homologous recombination regions HRt, HRX adjacent to an antiselectable marker (CM), thereby enabling the endonuclease to cleave the endonuclease sites and introducing the antiselectable marker at the cleavage site by homologous recombination.

[0120] The present invention also provides a method for affixing barcodes to oligonucleotides, the method comprising the steps of: (a) inserting each oligonucleotide from a mixture of oligonucleotides into a donor plasmid which optionally contains a first endonuclease site (C1), a first homologous recombination site (HR1), a second homologous recombination site (HR2), and optionally a second endonuclease site (C2) in a sequential order, so that each oligonucleotide is inserted between HR1 and HR2, thereby providing a plurality of donor plasmids containing donor oligonucleotides, and each donor plasmid C1-HR1-oligo-HR2-C2 contains a single donor oligonucleotide from the mixture of oligonucleotides; (b) transforming a plurality of cells with a plurality of donor plasmids so that each cell contains a donor plasmid, thereby forming a plurality of donor cells; and (c) plate seeding and culturing the plurality of donor cells at specific positions on a first ordered array, thereby providing a first ordered array of donor cells. (d) A step of providing a plurality of recipient cells into a second ordered array, wherein each recipient cell comprises a recipient oligonucleotide comprising, in a sequential order, a specific barcode sequence that identifies the location of the recipient cell in the second ordered array, a third homologous recombination region homologous to HR1 (HR3), optionally a third endonuclease site (C3), and a fourth homologous recombination region homologous to HR2 (HR4); (e) A step of (i) transferring a donor plasmid from a donor cell to a recipient cell at a corresponding location on the array by conjugation, (ii) optionally cleaving the first, second, and third endonuclease sites, and (ii) contacting a first ordered array of donor cells with a second ordered array of recipient cells under conditions that transfer the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination, thereby forming a third array of fusion oligonucleotides, each comprising a donor oligonucleotide from a mixture of specific barcode sequences and oligonucleotides.(f) The method also includes the step of optionally sequencing the fused oligonucleotides to identify each oligonucleotide in the array by its barcode sequence. In embodiments, the recipient oligonucleotide is located in a recipient cell plasmid or recipient cell genome. In embodiments, the donor plasmid contains selectable markers between HR1 and HR2 for the integration of the oligonucleotide into the recipient cell oligonucleotide, and optionally the donor plasmid contains anti-selectable markers. In embodiments, the recipient cell oligonucleotide contains a fourth endonuclease site (C4).

[0121] In a further embodiment, a method for assembling DNA elements is provided. The method includes the steps of (a) providing a first host cell comprising a first donor plasmid comprising (i) a first endonuclease target site, (ii) a first homologous recombination region, (iii) a first oligonucleotide optionally comprising a fragment of the first DNA element, (iv) a second homologous recombination region, (v) a second endonuclease target site, and (vi) a third endonuclease target site in a sequential order; (b) providing a recipient cell comprising a recipient oligonucleotide comprising (i) a third homologous recombination region homologous to the first homologous recombination region, (ii) a fourth endonuclease target site, and (iii) a fourth homologous recombination region homologous to the second homologous recombination region; and (c) (i) the first donor plasmid by bacterial conjugation. The method includes the steps of: (ii) transferring a donor plasmid from a first host cell to a recipient cell; (ii) directing a first endonuclease to at least one of a first endonuclease target site, a third endonuclease target site, or a fourth endonuclease target site, thereby generating double-strand breaks in the first donor plasmid and the recipient oligonucleotide; and (iii) contacting the first host cell and the recipient cell under conditions that cause the first donor plasmid and the recipient oligonucleotide to be recombined in the recipient cell by homologous recombination via first and second homologous recombination regions and third and fourth corresponding homologous recombination regions, thereby forming a recombined recipient oligonucleotide.The method comprises the steps of (d) providing a second host cell containing a second donor plasmid containing (i) a fifth endonuclease target site, (ii) a fifth homologous recombination region homologous to the second homologous region, (iii) a second oligonucleotide containing a fragment of the second DNA element, (iv) a sixth homologous recombination region homologous to the fourth homologous region, and (v) a sixth endonuclease target site, in a sequential order; and (e) (i) transferring the second donor plasmid from the second host cell to a recipient cell by bacterial conjugation, (ii) expressing the second endonuclease, and (iii) expressing the second endonuclease to the second endonuclease target site, the fifth endonuclease The steps include (iv) contacting a second host cell with a recipient cell containing the recombined recipient oligonucleotide under conditions that the DNA is directed to an endonuclease target site and / or a sixth endonuclease target site, thereby generating a double-strand break, and (iv) recombining the recipient oligonucleotide, which has been recombined with a second donor plasmid in the recipient cell, by homologous recombination via the second and fourth homologous recombination sites corresponding to the fifth and sixth homologous recombination sites, thereby forming a second recombined recipient oligonucleotide containing an assembled DNA element. In embodiments, different portions of the fourth homologous region are homologous to the sixth homologous recombination region, compared to portions of the fourth homologous region that are homologous to the second homologous recombination region.

[0122] In an embodiment, step (a) comprises a plurality of first host cells, each containing a specific first oligonucleotide. In an embodiment, step (d) comprises a plurality of second host cells, each containing a specific second oligonucleotide. In an embodiment, each first host cell contains a specific plasmid. In an embodiment, each second host cell contains a specific plasmid. In an embodiment, each first host cell is located in a position in a first ordered array. In an embodiment, a plurality of first host cells are located in a position in a first ordered array, thereby forming a pool of first host cells in the first ordered array. In an embodiment, each second host cell is located in a position in a second ordered array. In an embodiment, a plurality of second host cells are located in a position in a second ordered array, thereby forming a pool of second host cells at each position in the second ordered array.

[0123] In the embodiment, the first donor cells are in the first ordered array, the second donor cells are in the second ordered array, and one or more subsequent donor cells are in one or more subsequent arrays. That is, in the embodiment, the method provided herein generates a variant library containing multiple different assembled DNA elements. In the embodiment, the variant library is generated by 1) independently constructing each variant using the first host cell, the second host cell, or the subsequent host cell at positions in the first, second, or subsequent arrays, or 2) generating a pool of variants using multiple first host cells, second host cells, or subsequent host cells at positions in the first, second, or subsequent arrays. In the embodiment, the variant library is generated by independently constructing each variant using the first host cell, the second host cell, or the subsequent host cell at positions in the first, second, or subsequent arrays. In embodiments, the variant library is generated by generating a variant pool using a plurality of first host cells, second host cells, or successive host cells at positions in a first, second, or successive array. For example, in the method provided herein, including embodiments thereof, the first DNA element and / or the second DNA element may be a DNA barcode or a plurality of DNA barcodes. In embodiments, the method generates a repeatable barcoding platform. For example, in embodiments where the first DNA element and / or the second DNA element is a DNA barcode or a plurality of DNA barcodes, the method can be used for tracking cell lineage.

[0124] For example, the first DNA element and / or DNA element may be a gRNA or multiple gRNAs. That is, in embodiments, the method includes the generation of a combinatorial gRNA library.

[0125] In an embodiment, the first endonuclease targets the first endonuclease target site. In an embodiment, the first endonuclease targets the third endonuclease target site. In an embodiment, the first endonuclease targets the fourth endonuclease target site. In an embodiment, the second endonuclease targets the second endonuclease target site. In an embodiment, the second endonuclease targets the fifth endonuclease target site. In an embodiment, the second endonuclease targets the sixth endonuclease target site.

[0126] In an embodiment, the DNA element is a gene. In an embodiment, the DNA element is a promoter. In an embodiment, the DNA element is an enhancer. In an embodiment, the DNA element is a terminator. In an embodiment, the DNA element is an intron. In an embodiment, the DNA element is an intergenetic region. In an embodiment, the DNA element is a barcode. In an embodiment, the DNA element is a translation initiation site. In an embodiment, the DNA element is a gRNA. In an embodiment, the DNA element is any of the above fragments.

[0127] In one embodiment, the recipient oligonucleotide is located within the recipient plasmid. In another embodiment, the recipient oligonucleotide is located within the recipient cell genome.

[0128] In embodiments, the second donor plasmid further includes a seventh homologous recombination region and a seventh endonuclease target site between components (d)iii) and (d)iv). In embodiments, the first endonuclease targets the seventh endonuclease site.

[0129] In one embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the recipient cell contains an oligonucleotide encoding an RNA-guided DNA endonuclease. In another embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the recipient cell genome contains an oligonucleotide encoding an RNA-guided DNA endonuclease. In another embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the recipient plasmid contains an oligonucleotide encoding an RNA-guided DNA endonuclease. In yet another embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the recipient helper plasmid contains an oligonucleotide encoding an RNA-guided DNA endonuclease. In yet another embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the donor plasmid contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In yet another embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the recipient genome contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In yet another embodiment, the donor plasmid contains an oligonucleotide encoding gRNA, and the recipient plasmid contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In one embodiment, the donor plasmid comprises an oligonucleotide encoding a gRNA, and the recipient helper plasmid comprises an oligonucleotide encoding an inducible RNA-guided DNA endonuclease.

[0130] In one embodiment, the recipient cells contain inducible gRNA. In another embodiment, the donor cells contain oligonucleotides encoding RNA-guided DNA endonuclease. In yet another embodiment, the recipient cells contain oligonucleotides encoding RNA-guided DNA endonuclease. In yet another embodiment, RNA-guided DNA expression is constitutive. In yet another embodiment, RNA-guided DNA expression is inducible.

[0131] In an embodiment, the RNA-guided DNA endonucleases are Cas9. In an embodiment, the RNA-guided DNA endonucleases are Cas10. In an embodiment, the RNA-guided DNA endonucleases are Cpf1. In an embodiment, the RNA-guided DNA endonucleases are C2c1. In an embodiment, the RNA-guided DNA endonucleases are C2c2. In an embodiment, the RNA-guided DNA endonucleases are C2c3. In an embodiment, the RNA-guided DNA endonucleases are Cas12c1. In an embodiment, the RNA-guided DNA endonucleases are Cas12a. In an embodiment, the RNA-guided DNA endonucleases are Cas12b. In an embodiment, the RNA-guided DNA endonucleases are Cas12c2. In an embodiment, the RNA-guided DNA endonucleases are Cas12g. In an embodiment, the RNA-guided DNA endonucleases are Cas12e. In an embodiment, the RNA-guided DNA endonucleases are Cas12i1. In an embodiment, the RNA-guided DNA endonucleases are Cas12i2.

[0132] In the methods provided herein, in embodiments, the oligonucleotide encoding the first endonuclease is located in a donor plasmid. In embodiments, the oligonucleotide encoding the first endonuclease is a recipient oligonucleotide. In embodiments, the oligonucleotide encoding the first endonuclease is located in a recipient cell helper plasmid. In embodiments, the oligonucleotide encoding the first endonuclease is located in the recipient genome. In embodiments, the expression of the first endonuclease is inducible. In embodiments, the oligonucleotide encoding the second endonuclease is located in a donor plasmid. In embodiments, the oligonucleotide encoding the second endonuclease is located in a recipient oligonucleotide. In embodiments, the oligonucleotide encoding the second endonuclease is located in a recipient cell helper plasmid.

[0133] In an embodiment, the first endonuclease and / or the second endonuclease is a homing endonuclease. In embodiments, the homing endonuclease is I-ScaI, PI-SceI, I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI. , I-PpoI, PI-PspI, I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-S sp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I, or I-Vdi141I. In an embodiment, the homing endonuclease is I-ScaI. In an embodiment, the homing endonuclease is PI-SceI. In an embodiment, the homing endonuclease is I-AniI. In an embodiment, the homing endonuclease is I-CeuI. In an embodiment, the homing endonuclease is I-ChuI. In an embodiment, the homing endonuclease is I-CpaI. In an embodiment, the homing endonuclease is I-CpaII. In an embodiment, the homing endonuclease is I-CreI. In an embodiment, the homing endonuclease is I-DmoI. In an embodiment, the homing endonuclease is H-DreI. In an embodiment, the homing endonuclease is I-HmuI. In an embodiment, the homing endonuclease is I-HmuII. In an embodiment, the homing endonuclease is I-LlaI. In an embodiment, the homing endonuclease is I-MsoI. In an embodiment, the homing endonuclease is PI-PfuI. In an embodiment, the homing endonuclease is PI-PkoII. In an embodiment, the homing endonuclease is I-PorI. In an embodiment, the homing endonuclease is I-PpoI.In an embodiment, the homing endonuclease is PI-PspI. In an embodiment, the homing endonuclease is I-SceI. In an embodiment, the homing endonuclease is I-SceII. In an embodiment, the homing endonuclease is I-SceIII. In an embodiment, the homing endonuclease is I-SceIV. In an embodiment, the homing endonuclease is I-SceV. In an embodiment, the homing endonuclease is I-SceVI. In an embodiment, the homing endonuclease is I-SceVII. In an embodiment, the homing endonuclease is I-Ssp6803I. In an embodiment, the homing endonuclease is I-TevI. In an embodiment, the homing endonuclease is I-TevII. In an embodiment, the homing endonuclease is I-TevIII. In an embodiment, the homing endonuclease is PI-TliI. In one embodiment, the homing endonuclease is PI-TliII. In another embodiment, the homing endonuclease is I-Tsp061I or I-Vdi141I. In yet another embodiment, the homing endonuclease is I-Vdi141I.

[0134] In one embodiment, the endonuclease is a transcription activator-like effector nuclease. In another embodiment, the endonuclease is a zinc finger nuclease.

[0135] In the embodiments of the method provided herein, steps (d) and (e) are repeated one or more times to form one or more subsequent assembled DNA elements. In embodiments, the first, second, or subsequent donor plasmid includes a selectable marker selected for integration of the first oligonucleotide, the second oligonucleotide, or the subsequent oligonucleotide into the recipient cell oligonucleotide. In embodiments, the first donor plasmid includes a selectable marker selected for integration of the first oligonucleotide into the recipient cell oligonucleotide. In embodiments, the second donor plasmid includes a selectable marker selected for integration of the second oligonucleotide into the recipient cell oligonucleotide. In embodiments, the subsequent donor plasmid includes a selectable marker selected for integration of the subsequent oligonucleotide into the recipient cell oligonucleotide.

[0136] In an embodiment, the first donor plasmid includes a selectable marker selected for integration of the first oligonucleotide into the recipient oligonucleotide. In an embodiment, the selectable marker is between components (a)(v) and (a)(iv). In an embodiment, the second donor plasmid includes a selectable marker selected for integration of the second oligonucleotide into the recipient oligonucleotide.

[0137] In embodiments, the recipient cell oligonucleotide includes an anti-selectable marker selected for the integration of a first oligonucleotide into the recipient oligonucleotide. In embodiments, the recombined recipient cell oligonucleotide includes an anti-selectable marker selected for the integration of a second or subsequent oligonucleotide into the recombined recipient cell oligonucleotide. The anti-selectable markers for use in the methods described herein are listed above.

[0138] In one embodiment, assembled DNA elements, also called DNA assemblies, are sequenced. In another embodiment, recipient oligonucleotides are sequenced. In another embodiment, recombinant recipient oligonucleotides are sequenced. In another embodiment, recombinant recipient oligonucleotides are plasmids, which are linearized, ligated to a sequencing adapter, and sequenced. In yet another embodiment, assembled DNA elements are amplified by PCR and sequenced. In yet another embodiment, (a) recipient cells are lysed, (b) oligonucleotides are digested by an endonuclease or a plurality of endonucleases, (c) assembled DNA elements are isolated, and (d) assembled DNA elements or a plurality of assembled genes are ligated to a sequencing adapter and sequenced. In yet another embodiment, assembled DNA elements are isolated. In yet another embodiment, recombinant recipient oligonucleotides are isolated.

[0139] In the embodiment, the assembled DNA element is approximately 100 to 500,000 nucleotides in length. The length may be any value or a subrange within the indicated range, including the endpoint.

[0140] In an embodiment, the assembled DNA elements are approximately 100 nucleotides, 1,000 nucleotides, 10,000 nucleotides, 20,000 nucleotides, 40,000 nucleotides, 60,000 nucleotides, 80,000 nucleotides, 100,000 nucleotides, 120,000 nucleotides, 140,000 nucleotides, 160,000 nucleotides, 180,000 nucleotides, 20,000 nucleotides, and approximately The length is 240,000 nucleotides, approximately 260,000 nucleotides, approximately 280,000 nucleotides, approximately 300,000 nucleotides, 320,000 nucleotides, approximately 340,000 nucleotides, approximately 360,000 nucleotides, approximately 380,000 nucleotides, approximately 400,000 nucleotides, approximately 420,000 nucleotides, approximately 440,000 nucleotides, approximately 460,000 nucleotides, approximately 480,000 nucleotides, or approximately 500,000 nucleotides. The length may be any value or subrange within the indicated range, including the endpoint.

[0141] In the embodiment, the first, second, or subsequent homology regions and the corresponding first, second, or subsequent homology regions are approximately 20 to 500 base pairs in length. The length may be any value or subrange within the indicated range, including the endpoint.

[0142] In an appropriate manner, the first, second, or subsequent homology region and the corresponding first, second, or subsequent homology region are approximately 20 base pairs, 40 base pairs, 60 base pairs, 80 base pairs, 100 base pairs, 120 base pairs, 140 base pairs, 160 base pairs, 180 base pairs, 200 base pairs, 220 base pairs, 240 base pairs, 260 base pairs, 280 base pairs, 300 base pairs, 320 base pairs, 340 base pairs, 360 base pairs, 380 base pairs, 400 base pairs, 420 base pairs, 440 base pairs, 460 base pairs, 480 base pairs, or 500 base pairs. In an appropriate manner, the first, second, or subsequent homology region and the corresponding first, second, or subsequent homology region are approximately 50 base pairs in length. The length may be any value or subrange within the specified range, including the endpoint.

[0143] III. Methods of Analysis In one embodiment, a method for identifying an oligonucleotide from a mixture of oligonucleotides is provided. The method comprises the steps of (a) providing a mixture of oligonucleotides; (b) inserting each oligonucleotide into a donor plasmid, wherein each donor plasmid comprises, in a sequential order, i) a first endonuclease cleavage site, ii) a first homologous recombination region, iii) a second homologous recombination region, and iv) a second endonuclease cleavage site, wherein the oligonucleotide is inserted between the first homologous recombination region and the second homologous recombination region, thereby producing a plurality of donor plasmids, each donor plasmid containing a single oligonucleotide from the mixture of oligonucleotides; (c) transforming a plurality of host cells with the plurality of donor plasmids, thereby forming a plurality of transformed host cells, each host cell containing a donor plasmid; and (d) plate seeding and culturing the plurality of transformed host cells on a first ordered array, wherein each transformed host cell produces a clonal colony in the first ordered array. Step (e) providing a plurality of recipient cells in a second ordered array, wherein each recipient cell comprises a recipient oligonucleotide comprising, in a sequential order (i) a specific barcode sequence that identifies the location of the recipient cell in the second ordered array, (ii) a corresponding first homologous recombination region having a first homologous recombination region homologous thereto, (iii) a third endonuclease cleavage site, and (iv) a corresponding second homologous recombination region having a second homologous recombination region homologous thereto, wherein the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site can be cleaved by an endonuclease, Step (f) under conditions of (i) transferring a donor plasmid from a clonal colony to recipient cells by bacterial conjugation, (ii) cleaving the first, second, and third endonuclease cleavage sites with an endonuclease, and (ii) transferring the oligonucleotide from the donor plasmid to the recipient cell oligonucleotide by homologous recombination,The process includes (g) contacting each colony of clones from a first ordered array with recipient cells at corresponding sites in a second ordered array to produce a fusion sequence containing a barcode sequence and oligonucleotides, (g) sequencing the fusion protein, and (h) identifying the sequenced oligonucleotides in the first and / or second ordered arrays of donor and / or recipient cells by identifying the barcode sequence.

[0144] In embodiments, the step of plate seeding and culturing cells includes the step of plate seeding and culturing cells on a surface. For example, the surface may be a solid medium. That is, in embodiments, the clonal colonies are colonies of cells on a solid medium. In embodiments, the step of plate seeding and culturing cells includes the step of plate seeding and culturing cells in a liquid medium. For example, a single cell can be plate seeded and cultured in a liquid medium. That is, in embodiments, the clonal colonies are colonies of cells in a liquid medium.

[0145] In an embodiment, the recipient oligonucleotide is located within a recipient cell plasmid. In an embodiment, the recipient oligonucleotide is located within the recipient cell genome. In an embodiment, the donor plasmid includes a selectable marker between a first homologous recombination region and a second homologous recombination region, selected for integration of the oligonucleotide into the recipient cell oligonucleotide. In an embodiment, the recipient cell oligonucleotide includes two endonuclease cleavage sites in step (e)(iii). In an embodiment, the method further includes an anti-selectable marker between the two endonuclease cleavage sites, the anti-selectable marker being selected for integration of the oligonucleotide into the recipient oligonucleotide.

[0146] In one embodiment, a method for identifying oligonucleotides from a mixture of oligonucleotides is provided.The method comprises the steps of (a) providing a plurality of host cells in a first ordered array, each host cell comprising a donor plasmid, each donor plasmid comprising, in a sequential order, i) a first endonuclease cleavage site, ii) a first homologous recombination region, iii) a specific barcode sequence, iv) a second homologous recombination region, and v) a second endonuclease cleavage site, wherein the specific barcode sequence identifies the location of the host cell in the first ordered array; and (b) providing a plurality of recipient cells. Each recipient cell contains a recipient oligonucleotide comprising oligonucleotides from multiple oligonucleotides, and each recipient plasmid contains, in a sequential order, i) an oligonucleotide sequence, ii) a corresponding first homologous recombination region homologous to the first homologous recombination region, iii) a first endonuclease cleavage site, a second endonuclease cleavage site, and a third endonuclease cleavage site that can be cleaved by an endonuclease, and iv) a second homologous recombination region (c) plate seeding and culturing a plurality of recipient cells on a second ordered array, wherein each recipient cell produces a colony of clones in the second ordered array, the step of (i) transferring a donor plasmid from the donor cells to the colonies of clones by bacterial conjugation, (ii) cleaving first, second, and third endonuclease cleavage sites by endonuclease, and (iii) transferring a barcode sequence from the donor plasmid to the recipient cell oligonucleotides by homologous recombination, the step of (e) sequencing the fusion sequence, and (f) identifying the sequenced oligonucleotides in the first and / or second ordered arrays of recipient cells by identifying the barcode sequence.

[0147] In this embodiment, the recipient oligonucleotide is located within the recipient cell plasmid.

[0148] In an embodiment, the recipient oligonucleotide is located within the recipient cell genome. In an embodiment, the donor plasmid contains a selectable marker between a first homologous recombination region and a second homologous recombination region, selected for integration of the barcode into the recipient oligonucleotide. In an embodiment, the recipient oligonucleotide contains two endonuclease cleavage sites in step (b)(iii). In an embodiment, the method further includes an anti-selectable marker between the two endonuclease cleavage sites, selected for integration of the barcode into the recipient cell oligonucleotide.

[0149] For the method provided herein, in an embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are the same endonuclease cleavage site. In an embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are different endonuclease cleavage sites. In an embodiment, the endonuclease comprises a number of endonucleases.

[0150] In an embodiment, the endonuclease is encoded by an oligonucleotide in the recipient cell. In an embodiment, the oligonucleotide is located in the recipient cell genome. In an embodiment, the oligonucleotide is located in the recipient plasmid. In an embodiment, the oligonucleotide encoding the endonuclease is located in a helper plasmid. In an embodiment, the endonuclease is encoded by an oligonucleotide in the donor plasmid. In an embodiment, the endonuclease is encoded by an oligonucleotide in the donor plasmid. In an embodiment, the expression of the endonuclease is inducible.

[0151] In a subset, the endonuclease is a transcription activator-like effector nuclease. In a subset, the endonuclease is a zinc finger nuclease. It is a zinc finger nuclease. The endonuclease is HO.

[0152] In an embodiment, the RNA-guided DNA endonucleases are CRISPR systems. In an embodiment, the endonucleases are RNA-guided DNA endonucleases. In an embodiment, the RNA-guided DNA endonucleases are Cas9. In an embodiment, the RNA-guided DNA endonucleases are Cas10. In an embodiment, the RNA-guided DNA endonucleases are Cpf1. In an embodiment, the RNA-guided DNA endonucleases are C2c1. In an embodiment, the RNA-guided DNA endonucleases are C2c2. In an embodiment, the RNA-guided DNA endonucleases are C2c3. In an embodiment, the RNA-guided DNA endonucleases are Cas12c1. In an embodiment, the RNA-guided DNA endonucleases are Cas12a. In an embodiment, the RNA-guided DNA endonucleases are Cas12b. In an embodiment, the RNA-guided DNA endonucleases are Cas12c2. In an embodiment, the RNA-guided DNA endonucleases are Cas12g. In an embodiment, the RNA-guided DNA endonucleases are Cas12e. In an embodiment, the RNA-guided DNA endonucleases are Cas12i1. In an embodiment, the RNA-guided DNA endonucleases are Cas12i2.

[0153] For the methods provided herein, in embodiments, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient genome comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In embodiments, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient plasmid comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In embodiments, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient helper plasmid comprises an oligonucleotide encoding an RNA-guided DNA endonuclease. In embodiments, the donor plasmid comprises an oligonucleotide encoding a guide RNA, and the recipient plasmid comprises an oligonucleotide encoding an inducible RNA-guided DNA endonuclease.

[0154] In one embodiment, the recipient cells contain inducible gRNA. In another embodiment, the donor cells contain oligonucleotides encoding RNA-guided DNA endonuclease. In yet another embodiment, the recipient cells contain oligonucleotides encoding RNA-guided DNA endonuclease. In yet another embodiment, RNA-guided DNA expression is constitutive. In yet another embodiment, RNA-guided DNA expression is inducible.

[0155] In embodiments, the method further includes the step of isolating a donor plasmid. In embodiments, the method further includes the step of isolating recipient cells. In embodiments, the method further includes the step of isolating a recombinant recipient plasmid. In embodiments, the method further includes the step of isolating a sequenced oligonucleotide. In embodiments, the method further includes the step of isolating one or more donor cells, recipient cells, recombinant recipient cells, recipient oligonucleotides, or recombinant recipient oligonucleotides.

[0156] In embodiments, the method includes the step of combining one or more subsets of a colony. That is, in embodiments, the method includes the step of combining one or more subsets of a colony to isolate multiple donor plasmids from the colony subset. In embodiments, the method includes the step of combining one or more subsets of a colony to isolate multiple recipient plasmids from the colony subset. In embodiments, the method further includes the step of combining one or more subsets of a colony to isolate multiple recombinant recipient plasmids from the colony subset. In embodiments, the method further includes the step of isolating multiple sequenced oligonucleotides.

[0157] In embodiments, recipient or donor cells contain oligonucleotides that enable plasmid conjugation. In embodiments, the oligonucleotides that enable plasmid conjugation are located in the genome of the donor cells. In embodiments, the oligonucleotides that enable plasmid conjugation are located in a helper plasmid. In embodiments, the oligonucleotides that enable plasmid conjugation are located in the Tra operon. Other oligonucleotides that enable plasmid conjugation are listed above.

[0158] In the embodiment, donor cells, recipient cells, or recombinant recipient cells are transported to a position on a third ordered array, a fourth ordered array, or a subsequent ordered array.

[0159] In embodiments, the barcode sequence is approximately 4 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 8 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 12 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 16 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 20 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 24 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 28 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 32 to 50 nucleotides in length. In embodiments, the barcode sequence is approximately 36 to 50 nucleotides in length. The length of the barcode may be any value or a subrange within the range provided herein, including the endpoints.

[0160] In embodiments, the barcode sequence is approximately 4 to 36 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 32 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 28 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 24 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 20 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 16 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 12 nucleotides in length. In embodiments, the barcode sequence is approximately 4 to 8 nucleotides in length. In embodiments, the barcode sequence is approximately 4, 8, 12, 16, 20, 24, 28, 32, 36, or 40 nucleotides in length. In embodiments, the barcode sequence is approximately 15 nucleotides in length. In embodiments, the barcode sequence is approximately 40 nucleotides in length. The length of the barcode, including the endpoint, may be any value or a sub-range within the range provided herein.

[0161] Further embodiments The present invention is further described by the following additional embodiments.

[0162] Embodiment 1: A method for assembling a DNA element, comprising: (1) providing a first host cell comprising, in a sequential order, a first donor plasmid comprising a first endonuclease target site, a first homologous recombination region, a first oligonucleotide optionally comprising a fragment of the first DNA element, a second homologous recombination region, a second endonuclease target site, and a third endonuclease target site; and (2) a third homologous recombination region homologous to the first homologous recombination region, a fourth endonuclease target site, and a second homologous recombination region (3)(i) Transferring a first donor plasmid from a first host cell to a recipient cell by bacterial conjugation; (ii) Directing a first endonuclease to at least one of the first endonuclease target site, the third endonuclease target site, or the fourth endonuclease target site, thereby generating double-strand breaks in the first donor plasmid and the recipient oligonucleotide; (iii) Step (4) In a sequential order, a first host cell and a recipient cell are brought into contact under conditions that cause the first donor plasmid and the recipient oligonucleotide to be recombined in the recipient cell by homologous recombination via homologous recombination regions 1 and 2 and the corresponding homologous recombination regions 3 and 4, thereby forming the recombined recipient oligonucleotide, (4) a fifth endonuclease target site, a fifth homologous recombination region homologous to the second homologous region, and a second oligonucleotide encoding a fragment of the second DNA element, (5) (i) transferring the second donor plasmid from the second host cell to a recipient cell by bacterial conjugation; (ii) expressing the second endonuclease; (iii) directing the second endonuclease to the second endonuclease target site, the fifth endonuclease target site, and / or the sixth endonuclease target site, thereby generating double-strand breaks.(iv) A method comprising the steps of contacting a second host cell with a recipient cell containing a recombined recipient oligonucleotide, under conditions that recombine the recipient oligonucleotide recombined with a second donor plasmid in the recipient cell by homologous recombination via the second and fourth homologous recombination sites corresponding to the fifth and sixth homologous recombination sites, thereby forming a second recombined recipient oligonucleotide containing an assembled DNA element.

[0163] In a further embodiment, step (a) comprises a plurality of first host cells, each containing a different first oligonucleotide, and / or step (d) comprises a plurality of second host cells, each containing a different second oligonucleotide.

[0164] In a further embodiment, each first host cell is located in a position within a first ordered array.

[0165] In a further embodiment, a plurality of first host cells are located in a first ordered array, thereby forming a pool of first host cells in the first ordered array.

[0166] In a further embodiment, each second host cell is located in a second ordered array.

[0167] In a further embodiment, a plurality of second host cells are located in a second ordered array, thereby forming a pool of second host cells at each of the locations in the second ordered array.

[0168] In a further embodiment, the first donor cells are in a first ordered array, the second donor cells are in a second ordered array, and one or more subsequent donor cells are in one or more subsequent arrays.

[0169] In a further embodiment, the method generates a combinatorial library containing multiple different assembled DNA elements.

[0170] In further embodiments, the first endonuclease targets the first, third, or fourth endonuclease target site.

[0171] In further embodiments, the second endonuclease targets the second, fifth, or sixth endonuclease target site.

[0172] In further embodiments, the DNA element is a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, or gRNA.

[0173] In a further embodiment, the recipient oligonucleotide is located within the recipient plasmid.

[0174] In a further embodiment, the recipient oligonucleotide is located within the recipient cell genome.

[0175] In a further embodiment, the second donor plasmid further comprises a seventh homologous recombination region and a seventh endonuclease target site between components (d)iii) and (d)iv).

[0176] In a further embodiment, the first endonuclease targets the seventh endonuclease site.

[0177] In further embodiments, the first and second endonucleases are independently selected from RNA-guided DNA endonucleases, homing endonucleases, transcription activator-like effector nucleases, and zinc finger nucleases.

[0178] In a further embodiment, the oligonucleotide encoding the first endonuclease is present in the donor or recipient cell.

[0179] In further embodiments, the expression of the first and / or second endonuclease is inducible.

[0180] In a further embodiment, the oligonucleotide encoding the second oligonucleotide is present in the donor cell or recipient cell.

[0181] In a further embodiment, steps (d) to (e) are repeated one or more times to form one or more successive assembled DNA elements.

[0182] In further embodiments, the first, second, or subsequent donor plasmid includes a selectable marker chosen for the integration of the first, second, or subsequent oligonucleotides into the recipient cell oligonucleotide.

[0183] In a further embodiment, the first donor plasmid includes a selectable marker for the integration of the first oligonucleotide into the recipient oligonucleotide, the optionally selectable marker being located between the second and third endonuclease target sites. In the embodiment, the second donor plasmid includes a selectable marker for the integration of the second oligonucleotide into the recipient oligonucleotide.

[0184] In a further embodiment, the recipient cell oligonucleotide includes an antiselectable marker selected for integration of the first oligonucleotide into the recipient cell oligonucleotide.

[0185] In a further embodiment, the donor plasmid includes a transport origin, which optionally originates from a mobile element.

[0186] In a further embodiment, the donor plasmid includes a conditional origin of replication, which optionally depends on the presence of an oligonucleotide or on conditions for cell proliferation.

[0187] In further embodiments, the donor plasmid or recipient oligonucleotide includes a replicon capable of replicating plasmids longer than 30 kilobases.

[0188] In further embodiments, the donor plasmid or recipient oligonucleotide includes an inducible high-copy origin for replication.

[0189] In further embodiments, the donor plasmid or recipient oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome. In further embodiments, the donor plasmid or recipient oligonucleotide is a viral vector.

[0190] In a further embodiment, the donor cells contain oligonucleotides that enable plasmid conjugation.

[0191] In a further embodiment, the donor or recipient cells contain oligonucleotides encoding one or more homologous DNA repair genes, and the expression of the homologous DNA repair genes is optionally inducible.

[0192] In further embodiments, the donor cells or recipient cells include oligonucleotides encoding one or more recombinant-mediated gene manipulation genes.

[0193] In a further embodiment, the donor cells and recipient cells are independently bacterial cells.

[0194] In a further embodiment, the assembled DNA elements are sequenced, the recipient oligonucleotides are sequenced, and / or the recombinant recipient oligonucleotides are sequenced.

[0195] In a further embodiment, the recombinant recipient oligonucleotide is a plasmid, which is linearized, ligated to a sequencing adapter, and sequenced.

[0196] In a further embodiment, the assembled DNA elements are amplified and sequenced by PCR, and optionally (a) recipient cells are lysed, (b) oligonucleotides are digested by an endonuclease or a plurality of endonucleases, (c) the assembled DNA elements are isolated, and (d) the assembled DNA elements or a plurality of assembled genes are ligated to a sequencing adapter and sequenced.

[0197] In a further embodiment, the assembled DNA element or recombinant recipient oligonucleotide is isolated.

[0198] In a further embodiment, the assembled DNA fragment is 100 to 500,000 nucleotides in length.

[0199] In further embodiments, the first, second, or subsequent homology regions and the corresponding first, second, or subsequent homology regions are approximately 20 to 500 base pairs in length.

[0200] In further embodiments, the first, second, or subsequent homology regions and the corresponding first, second, or subsequent homology regions are approximately 50 base pairs in length.

[0201] In an embodiment, a method for identifying an oligonucleotide from a mixture of oligonucleotides is provided herein, the method comprising: (a) providing a mixture of oligonucleotides; (b) inserting each oligonucleotide into a donor plasmid, wherein each donor plasmid comprises, in a sequential order, (i) a first endonuclease cleavage site, (ii) a first homologous recombination region, (iii) a second homologous recombination region, and (iv) a second endonuclease cleavage site, and the oligonucleotide is inserted between the first homologous recombination region and the second homologous recombination region, thereby producing a plurality of donor plasmids, each donor plasmid containing a single oligonucleotide from the mixture of oligonucleotides; (c) transforming a plurality of host cells with the plurality of donor plasmids, thereby forming a plurality of transformed host cells, each host cell containing a donor plasmid; and (d) plate seeding and culturing the plurality of transformed host cells on a first ordered array. (e) a nourishing step, wherein each transformed host cell produces a colony of clones in a first ordered array; (f) a step of providing a plurality of recipient cells in a second ordered array, wherein each recipient cell is sequentially provided with (i) a specific barcode sequence that identifies the location of the recipient cell in the second ordered array; (ii) a corresponding first homologous recombination region homologous thereto; and (iii) a third endonuclease cleavage site. (iv) comprising a recipient oligonucleotide comprising a corresponding second homologous recombination region wherein the second homologous recombination region is homologous thereto, wherein the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site can be cleaved by an endonuclease, (f) (i) transferring the donor plasmid from a clonal colony to a recipient cell by bacterial conjugation, (ii) cleaving the first, second, and third endonuclease cleavage sites with an endonuclease,(ii) The process includes the steps of (g) contacting each colony of clones from a first ordered array with recipient cells at a corresponding site in a second ordered array under conditions that transfer oligonucleotides from a donor plasmid to recipient cell oligonucleotides by homologous recombination, thereby producing a fusion sequence containing a barcode sequence and oligonucleotides; (h) sequencing the fusion sequence; and (i) identifying the sequenced oligonucleotides in the first and / or second ordered arrays of donor cells and / or recipient cells by identifying the barcode sequence.

[0202] In a further embodiment, the recipient oligonucleotide is located within the recipient cell plasmid.

[0203] In a further embodiment, the recipient oligonucleotide is located within the recipient cell genome.

[0204] In a further embodiment, the method provides that the donor plasmid includes a selectable marker between a first homologous recombination region and a second homologous recombination region, which is selected for the integration of the oligonucleotide into the recipient cell oligonucleotide.

[0205] In a further embodiment, the method herein provides that the recipient cell oligonucleotide comprises two endonuclease cleavage sites in step (e)(iii).

[0206] In embodiments of this specification, the method further comprises an antiselectable marker between two endonuclease cleavage sites, the antiselectable marker being selected for the integration of the oligonucleotide into the recipient oligonucleotide.

[0207] In embodiments, a method for identifying an oligonucleotide from a plurality of oligonucleotides is provided herein, the method comprising: (a) providing a plurality of host cells in a first ordered array, each host cell comprising a donor plasmid, each donor plasmid comprising, in a sequential order, (i) a first endonuclease cleavage site, (ii) a first homologous recombination region, (iii) a specific barcode sequence, (iv) a second homologous recombination region, and (v) a second endonuclease cleavage site, wherein the specific barcode sequence identifies the location of the host cell in the first ordered array; (b) providing a plurality of recipient cells, each recipient cell comprising a recipient oligonucleotide comprising an oligonucleotide from a plurality of oligonucleotides, each recipient plasmid comprising, in a sequential order, i) an oligonucleotide sequence, ii) a corresponding first homologous recombination region homologous to the first homologous recombination region, iii) a first endonuclease cleavage site, a second endonuclease cleavage site, and (i) a donor plasmid is transferred from the donor cells to the colonies of clones at the corresponding sites of the second ordered array under the conditions that (i) the donor plasmid is transferred from the donor cells to the colonies of clones by bacterial conjugation, (ii) the donor plasmid is transferred from the donor plasmid to the oligonucleotides of the recipient cells by endonuclease, (iii) the donor plasmid is transferred from the donor plasmid to the oligonucleotides of the recipient cells by homologous recombination, thereby producing a fusion sequence containing the barcode sequence and the oligonucleotides, (e) sequencing the fusion sequence.(f) The procedure includes the step of identifying sequenced oligonucleotides in a first and / or second ordered array of donor cells and / or recipient cells by identifying the barcode sequence.

[0208] In one embodiment, the method provides that the recipient oligonucleotide is located within a recipient cell plasmid.

[0209] In a further embodiment, the recipient oligonucleotide is located within the recipient cell genome.

[0210] In one embodiment, the donor plasmid includes a selectable marker between a first homologous recombination region and a second homologous recombination region, which is selected for the integration of a barcode into the recipient oligonucleotide.

[0211] In a further embodiment, the recipient oligonucleotide comprises two endonuclease cleavage sites in step (b)(iii).

[0212] In a further embodiment, the method includes the step of providing an anti-selectable marker between two endonuclease cleavage sites to be selected for integration of a barcode into a recipient cell oligonucleotide.

[0213] In a further embodiment, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are the same endonuclease cleavage site.

[0214] In further embodiments, the first endonuclease cleavage site, the second endonuclease cleavage site, and the third endonuclease cleavage site are different endonuclease cleavage sites.

[0215] In a further embodiment, the endonuclease comprises a number of endonucleases.

[0216] In a further embodiment, the donor plasmid includes a transfer origin.

[0217] In a further embodiment, the origin of the transfer is from the movable element.

[0218] In a further embodiment, the donor plasmid includes a conditional origin of replication.

[0219] In a further embodiment, the conditional replication origin depends on the presence of an oligonucleotide.

[0220] In a further embodiment, the conditional origin of replication depends on the conditions for cell proliferation.

[0221] In a further embodiment, the donor plasmid or recipient plasmid includes a replicon capable of replicating a plasmid of at least 30 kilobases in length.

[0222] In a further embodiment, the replicon is derived from a P1-induced artificial chromosome or a bacterial artificial chromosome.

[0223] In a further embodiment, the donor plasmid or recipient cell oligonucleotide includes an inducible high-copy replication origin.

[0224] In further embodiments, the donor plasmid or recipient cell oligonucleotide is a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC), a human artificial chromosome (HAC), or a plant artificial chromosome.

[0225] In a further embodiment, the donor plasmid or recipient oligonucleotide is a viral vector.

[0226] In a further embodiment, the endonuclease is encoded by an oligonucleotide in the recipient cell.

[0227] In a further embodiment, the endonuclease is encoded by an oligonucleotide in the donor plasmid.

[0228] In a further embodiment, the endonuclease is a homing endonuclease.

[0229] In a further embodiment, the endonuclease is an RNA-guided DNA endonuclease.

[0230] In a further embodiment, the endonuclease is HO.

[0231] In a further embodiment, the method further comprises the step of isolating the donor plasmid.

[0232] In a further embodiment, the method further comprises the step of isolating the recipient plasmid.

[0233] In a further embodiment, the method further comprises the step of isolating the recombinant recipient plasmid.

[0234] In a further embodiment, the method further comprises the step of isolating the sequenced oligonucleotide.

[0235] In a further embodiment, the donor cell or recipient cell comprises an oligonucleotide that enables plasmid conjugation.

[0236] In a further embodiment, the donor cell or recipient cell comprises an oligonucleotide encoding one or more homologous DNA repair genes.

[0237] In a further embodiment, the donor cell or recipient cell comprises an oligonucleotide encoding one or more recombinase-mediated genetic manipulation genes.

[0238] In further embodiments, donor cells, recipient cells, or recombinant recipient cells are transported to a position on a third ordered array, a fourth ordered array, or a subsequent ordered array.

[0239] In a further embodiment, the donor cells and recipient cells are independently bacterial cells.

[0240] In a further embodiment, the barcode sequence is approximately 4 to 40 nucleotides in length.

[0241] In a further embodiment, the barcode sequence is approximately 15 nucleotides in length.

[0242] The examples and embodiments described herein are for illustrative purposes only, and it will be understood by those skilled in the art that various modifications or changes taking them into consideration are suggested and should be included in the spirit and scope of this application and the scope of the appended claims. All publications, patents and patent applications cited herein are incorporated herein by reference in their entirety for all purposes. [Examples]

[0243] [Example 1] Method of in vivo DNA suturing The bacterial strain BUN20 [Δlac-169 rpoS(Am) robA1 creC510 hsdR514 ΔuidA(MluI):pir-116 endA(BT333)recA1 F'(lac+ pro+ ΔoriT:tet)] was used as the donor strain (Li, M. et al., Nat. Genet. 37, pp. 311-319 (2005)). BW23474: [Δlac-169 rpoS(Am)robA1 creC510 hsdR514 ΔuidA(MluI):pir-116 endA(BT333)recA1] was used as the host strain for cloning and proliferation of all donor plasmids (Haldimann, A. et al., Proc. Natl. Acad. Sci. 93, pp. 14361 (1996)). BW28705[lacIQ rrnB3 ΔlacZ4787 hsdR514 Δ(araBAD)567 Δ(rhaBAD)568 galU95 ΔendA9:FRT ΔrecA635:FRT] or RE1133 (Egbert et al., Nucleic Acids Research, Vol. 47 (6), April 8, 2019, pp. 3244-3256)[cmR::mutS pTet2-gam-bet-exo-dam / tetR::bioA / B ilvG+ dnaG.Q576A lacIQ1 Pcp8-araE ΔaraBAD pConst-araC ΔrecJ ΔxonA Pkm-cymR-Cas9::bioC] were used as recipient strains for in vivo suturing. DH5α and DH10β were used for cloning recipient plasmids.

[0244] The DNA oligonucleotides used for the first and second steps of PCR are shown in Tables 1 and 2.

[0245] [Table 1]

[0246] [Table 2]

[0247] Media and Chemicals: For the cloning and propagation of plasmids of donors and recipients, Luria-Bertani (LB) broth (1% w / v tryptone, 0.5% w / v yeast extract, 1% w / v NaCl) as a complex medium was routinely used. To maintain the plasmids, antibiotics were added at the concentrations listed in Table 3. For LB medium containing hygromycin, 0.5% w / v sodium chloride was used because hygromycin is sensitive to salts. P araBAD P rhaBAD P Tet2 P km -cymR, and P lacIQ To induce the respective promoters, L-arabinose (0.2% w / v), L-rhamnose (0.2% w / v), anhydrotetracycline (100 ng / ml), 4-isopropylbenzoic acid (cumate, 15 μg / ml), and isopropyl β-d-1-thiogalactopyranoside (IPTG, 500 μM) were used. To select against the SacB counterselectable marker, sucrose agar plates (0.5% w / v yeast extract, 1% w / v tryptone, 6% w / v sucrose, 1.5% agar) and appropriate amounts of antibiotics were used. PheS Gly 294 For the counterselection of PheS Gly, Cl-Phe agar plates (0.5% w / v yeast extract, 1% w / v NaCl, 0.4% w / v glycerol, 2% w / v agar, 10 mM D,L-p-Cl-Phe) and appropriate amounts of antibiotics were used. PheS Gly 294 To subclone the recipient plasmid containing the PheS Gly fragment, YEG agar plates (0.5% w / v yeast extract, 1% w / v NaCl, 0.4% w / v glucose, 2% w / v agar) and appropriate amounts of antibiotics were used.

[0248]

Table 3

[0249] The sequences of the plasmids used for in vivo DNA stitching can be found in Table 4.

[0250] [Table 4] TIFF0007836831000005.tif63161

[0251] Construction of an in vivo DNA suturing system: Construction of helper plasmids and host recipient strains. In the related MAGIC cloning system (Li, M. et al., Nat. Genet. 37, pp. 311-319 (2005)), recipient cells contain the helper plasmid pML300, which contains inducible λ-red and a temperature-sensitive origin of replication (pSC101-ori). TS ) possesses. MAGIC recipient cells also possess a genomically integrated inducible I-SceI endonuclease allele. Different helper plasmids containing λ-red and Cas9 are required to perform repeatable cleavage and homologous recombination. To construct such a helper plasmid, pML300 was first digested and ligated into a multiple cloning site (MCS) via HindIII and NheI, thereby obtaining pSL270. Using Gibson Assembly, araC-P araBaAD -A single DNA fragment containing Cas9 was constructed and then cloned into pSL270 using PacI and XhoI restriction sites to create helper plasmid pSL359. This plasmid is a rhamnose-inducible λ-red recombinant system (P rhaBAD -red), arabinose-induced endonuclease (P araBAD -Cas9), pSC101-ori tsIt contains, and a spectinomycin resistance marker (SpR). Next, BW28705 was transformed with pSL359 to create recipient host cells BW28705 / pSL359. As an alternative recipient strain, RE1133 (Egbert et al., Nucleic Acids Research, vol. 47(6), April 8, 2019, pp. 3244-3256) [cmR::mutS pTet2-gam-bet-exo-dam / tetR::bioA / B ilvG+ dnaG.Q576A lacIQ1 Pcp8-araE ΔaraBAD pConst-araC ΔrecJ ΔxonA Pkm-cymR-Cas9::bioC] was used without a helper plasmid. RE1133 contains a tetracycline-inducible λ-red recombinant system (pTet2-gam-bet-exo-dam / tetR::bioA / B) and a cuminate-inducible Cas9 endonuclease (Pkm-cymR-Cas9::bioC).

[0252] Construction of Swapping Cassettes: A swapping cassette is defined as a stretch of DNA on donor and recipient plasmids that contributes to DNA swapping. The cassette on the recipient plasmid is replaced by the cassette originally found on the donor plasmid via homologous recombination. For repeatable selection for in vivo cassette swapping, each cassette is engineered to contain both a selectable marker and an anti-selectable marker. A selectable marker in the donor cassette and an anti-selectable marker in the recipient cassette are required in each round. To implement such a double-selection strategy, two different selection cassettes were constructed from the following sources using standard cloning methods: 1) PheS Gly 2941) (D,Lp-Cl-Phe susceptible) (Kast, P. Gene 138, pp. 109-114 (1994)), 2) SacB (sucrose susceptible) (Pelicic, V. et al., J. Bacteriol. 178, pp. 1197-1199 (1996)), 3) HygR (hygromycin resistant) containing EM7 bacterial promoter (Gritz, L et al., Gene 25, pp. 179-188 (1983)), 4) NsrR (nucleoslysin resistant) (Gene 62, pp. 209-217 (1988)). One cassette contains PheS Gly 294 The first cassette was constructed to include NsrR, and the second cassette was constructed to include HygR and SacB. Several other optional cassettes were also constructed to conduct several experiments to characterize the in vivo suturing system. 5) ZeoR (zeosin resistance) (Drocourt, D. Nucleic Acids Res. 18, pp. 4009-4009 (1990)), 6) ampR (ampicillin resistance), and 7) CmR (chloramphenicol resistance). In some cases, one cassette was constructed to include PheS Gly 294 The first cassette was constructed to include NsrR, and the second cassette was constructed to include HygR and SacB.

[0253] Construction of the donor and recipient vector skeletons: The donor vector was constructed to include the following key components: kanR, oriT, R6K oriγ, and a constitutive gRNA expression cassette driven by the strong bacterial promoter J23119 (Standage-Beier, K. et al., ACS Synth. Biol. 4, pp. 1217-1225 (2015)). The swapping region was reconstituted to generate the T1(F)-H1-T2(R)-T2(F)-H3-T1(R) fragment. Here, T1(5'-GGGGCCACTAGGGACAGGATtgg-3' (SEQ ID NO: 37) and T2(5'-CAGGCGGGCTCACCTCCGTGtgg-3' (SEQ ID NO: 38) are two specific target sequences for CRISPR-Cas9 cleavage, H1(5'-CGAGGGCTAGAATTACCTACCGGCCTCCACCATGCCTGCG-3' (SEQ ID NO: 39) and H3(5'-GTACGGGCAACCCGAGAAGGCTGAGCCTGGACTCAACGGGTTGCTGGGTGGACTCCAGACTCGGGGCGACGACTCTTCACGCGCAGAGCAAGGGCGTCGAGCGGTCGTGAAAGTCTTAGTACCGCACGTGCCGACTCACTGGGGATATTGCCTGGAGCTGTACCGTTCTAGGGGGGGGAGGTTGGAGACCTCC TCTTCTCACGACTGGACCCGCGAGGGCCGCGTTGCCGGTTCCCCCAGAGGCTGAAGAACAAGGGCTTACTGTGGGCAGGGGGACGCCCATTCAGCGGCTGGCGCTTT-3' (SEQ ID NO: 40) is a homology site for homologous recombination, and (F) and (R) indicate whether the DNA fragment is included in forward or reverse (reverse complement) orientation. A selection cassette (HygR-SacB or NsrR-PheS) is inserted between T2(R) and T2(F) to generate a donor plasmid ready for suturing. In some cases, the swapping region of the donor plasmid was reconstituted to generate a T2(F)-T1(R)-T1(F)-H3-T2(R) fragment in which the selection cassette (HygR-SacB or NsrR-PheS) was inserted between T1(R) and T1(F).To ensure that the distance between the double-strand break locus and the homologous region is as short as possible, both gRNA target sites are positioned with appropriate orientation. H3 is a 300 bp synthetic DNA fragment used as a single homology arm throughout all rounds of assembly. H1 is a homology region used in the first round of in vivo suturing, either incorporated into the donor skeleton or introduced as part of the first oligonucleotide to be sutured. Other homology regions (H2, H4, H5, etc.) are introduced into the donor plasmid as parts of subsequent oligonucleotides, overlapping with homology regions of previous oligonucleotides in the assembly to enable seamless suturing. The entry recipient vector contains a selectable marker (GmR) and origin of replication (ColE1). The swapping region was modified to an H1-T1(R)-T1(F)-H3 configuration, with a selectable cassette (HygR-SacB or NsrR-PheS) cloned between T1(R) and T1(F).

[0254] Testing the efficiency of endonuclease cleavage and homologous recombination: To test whether the CRISPR / Cas9 system provides precise DNA cleavage and promotes homologous recombination, suturing procedures were completed in or without the presence of target gRNA, Cas9, and λred. Three different recipient host cells were transformed with recipient plasmid pSL402 to construct 1) BW28705 / pSL402 (-λred / -Cas9), 2) BW28705 / pML300 / pSL402 (+λred / -Cas9), and 3) BW28705 / pSL359 / pSL402 (+λred / +Cas9). BUN20 was then transformed with two different donor plasmids, pSL414 and pSL415, to create BUN20 / pSL414 and BUN20 / pSL415, respectively, containing functional and mock gRNA units. Each of the two donor cell lines was crossed with one of each of the three different recipient cell types as described above. The cells were then diluted and seeded onto a selection plate (Cl-Phe+Gm+Cm+0.2% glucose) and recombinant clones were harvested overnight at 37°C. The number of colonies was counted to quantify the recombination event.

[0255] Assembly of the mEGFP gene in liquid: Three fragments of the mEGFP gene were generated by PCR. The first fragment, containing the constitutive promoter pJ23100, ribosome binding site, and nucleotides 1-251, was cloned into the donor skeleton to create pSL485. The second fragment, containing nucleotides 198-517, was cloned into the donor vector to create pSL486. The third fragment, containing nucleotides 454-720 and the rrnB T1 terminator, was cloned into the donor vector to create pSL488. BUN20 was transformed with all three donor vectors and grown overnight on LB+Kan plates at 37°C. BW28705 / pSL359 was transformed with the entry recipient vector pSL398 and grown on LB agar plates containing gentamicin, spectinomycin, and glucose. Clones from donor BUN20 / pSL485 (D1) and recipient BW28705 / pSL359 / pSL398 (R0) were grown overnight in appropriate liquid medium at 37°C and 30°C, respectively. Cells (1 ml) from both the donor and recipient were then centrifuged, mixed, and resuspended in 1 ml of pre-warmed LB+Ara+Rha liquid medium. After incubation at 30°C for approximately 4 hours without shaking, serial dilutions were performed in the mating culture medium, and the cells were plate-seeded on 6% Suc+Carb+Gm+Sp+0.2% glucose to select recombinants (R1). Colony PCR was performed to confirm appropriate R1 clones, which were then incubated overnight at 30°C in LB+Gm+Sp+0.2% glucose. Next, newly cultured donor cells (D2) containing pSL486 were centrifuged, mixed, and resuspended with R1 in liquid mating medium at 30°C for approximately 4 hours. Recombinant clones were selected by plate seeding on Cl-Phe+Hyg+Gm+Sp+0.2% glucose. Appropriate clones (R2), confirmed by colony PCR, were grown overnight in LB+Gm+Sp+0.2% glucose at 30°C. As in the previous round, D3 (BUN20 / pSL488) and R2 cells were mixed and resuspended in liquid mating medium. Serial dilutions were performed, and the cells were plate seeded on 6% Suc+Carb+Gm+Sp+0.2% glucose.Plates from each round of assembly were imaged under UV light, and the percentage of GFP-fluorescent colonies was counted. After each round of assembly, selected clones were harvested, and plasmids were purified for diagnostic restriction digestion and Sanger sequencing.

[0256] Testing the effect of homology length on sutures: To test the effect of homology length on suture accuracy, a series of plasmids were constructed containing homology of different sizes to the second mEGFP fragment in pSL486: 1) pSL684 (0 bp), 2) pSL685 (10 bp), 3) pSL510 (20 bp), 4) pSL511 (30 bp), 5) pSL512 (40 bp), 6) pSL681 (53 bp). Donor host cells were transformed with these plasmids and pSL488 (63 bp) to construct a group of D3 cells that were mated with recipient cells containing R2. Cells from each mating pair were diluted and seeded on a selection plate (6% Suc + Carb + Gm + Sp + 0.2% glucose). Colonies were harvested overnight and examined under UV light for fluorescence. The number of fluorescent and non-fluorescent colonies was counted to calculate the percentage of properly assembled specimens.

[0257] Aligned mEGFP assemblies: All strains were aligned on agar plates in 384 format. First, BUN20 / pSL1065(D1) and BW28705 / pSL359 / pSL1060(R0) were aligned and mixed together using SINGER ROTOR HDA on a preheated mating plate (LB+Ara+Rha), and grown at 30°C for approximately 5 hours. The mated cells were then transferred to a first selection plate (Cl-Phe+Hyg+Gm+Sp+0.2% glucose). Recombinant clones (R1) were enriched overnight at 30°C and then transferred to a preliminary mating plate (LB+Hyg+Gm+Sp+0.2% glucose) to optimize the growth of the assembled plasmid for the next round. A fresh overnight array of BUN20 / pSL1062(D2) was then mated with R1 on the mating plate. The mated cells were then transferred to a first selection plate (LB+Nat+Gm+Sp+0.2% glucose) and recombinant clones were selected overnight at 30°C. The selected clone (R2) was then transferred to a preliminary mating plate (6% Suc+Nat+Gm+Sp+0.2% glucose). A fresh overnight array of BUN20 / pSL1066 (D3) was then mated with R2 on the mating plate. Following selection on Cl-Phe+Hyg+Gm+Sp+0.2% glucose and then on LB+Hyg+Gm+Sp+0.2% glucose, the final assembly product (R3) was selected. Plates from each round of assembly were imaged under UV light using a UV transilluminator, and GFP fluorescence was monitored. Throughout the assembly process, selected clones were harvested, and plasmids were purified for diagnostic restriction digestion and Sanger sequencing.

[0258] Aligned assemblies of 12 genes using pooled oligonucleotides: Nine distinguishable serine / tyrosine recombinases and three distinguishable fluorescent molecules (mPapaya, mPlum, sfGFP) were selected for assembly. To generate a list of oligonucleotides required for each gene assembly, a Python script was created that takes a FASTA file containing the genes to be synthesized as input and outputs a list of oligonucleotides to be ordered from a commercial supplier (IDT oPool), including user-defined variables such as the length of the synthesized oligonucleotide, the minimum homologous overlap length between adjacent oligos, and the maximum homologous overlap length. DNA hairpins and / or repeats may interfere with the homologous recombination mechanism and reduce assembly fidelity, but quantitative studies on this effect are unknown. Therefore, for each gene, these regions were identified using the Primer3 Python extension (Untergasser, A. et al., Nucleic Acids Res. 35, pp. W71-W74 (2007)) and nucleotide distribution uniformity metrics. The position of each nucleotide is scored, and if a small homology region may contain interfering elements, the homology region is expanded using a user-defined threshold. Once the oligonucleotides required for each gene assembly are determined, restriction sites (NotI and AscI) are added to each end, which are used to clone the oligonucleotides into donor vectors and round-specific priming sites. The priming sites allow oligonucleotides from specific rounds of parallel gene assemblies to be amplified together and parsed by an in vivo parsing platform. Round-specific primers are selected from a list of primers pre-designed to reduce the possibility of cross-reactivity between primers (reducing the number of undesirable PCR products) when used for a large oligonucleotide pool (Kosuri, S. et al., Nat. Biotechnol. 28, pp. 1295-1299 (2010)).Using this Python script, each gene was split into five approximately 300 bp oligonucleotides with 50–70 bp homology between the subsequent oligonucleotides. The PCR-amplified oligonucleotides were inserted into the donor plasmid using restriction digestion and ligation. For the first round of assembly of each gene, the oligonucleotide was amplified and cloned into the donor skeleton pSL1064. This contains a 40 bp initiation H1 region homologous to the region in the entry recipient plasmid pSL1060. Oligonucleotides to be added to further odd-numbered (e.g., 3, 5, 7, 9) suture rounds were cloned into pSL1063. This contains the same elements as pSL1064 except that it lacks the H1 homology region. Oligonucleotides to be added to even-numbered (e.g., 2, 4, 6, 8) suture rounds were cloned into pSL1071. The PCR-amplified oligonucleotides and donor plasmids were digested by AscI and NotI at 37°C for 4 hours. The digested products were then size-selected and purified by gel extraction using the Zymoclean Gel DNA Recovery Kit. 0.02 pmol of digested donor plasmid and 0.06 pmol of digested oligonucleotide were mixed with 1 μl of T4 ligase and ligated by incubation at 22°C for 1 hour. The ligated donor plasmid was transferred to a BUN20 donor strain using a standard bacterial transformation protocol. 2 μl of the ligation product was added to 50 μl of chemically competent BUN20 donor strain. The mixture was then incubated on ice for 30 minutes, subjected to a heat shock at 42°C for 30 seconds, and incubated again on ice for 3 minutes. The cells were resuspended in 950 ml of NEB SOC recovery medium and collected at 37°C for 1 hour. The cells were then seeded onto selection plates (LB+Hyg for odd-round donor plasmids, LB+Nat for even-round donor plasmids) and incubated overnight at 37°C. Colonies containing cloned oligonucleotides were randomly selected and aligned in a 96-well plate.Aligned oligonucleotide libraries were syntactically analyzed, and their sequences were validated using an in vivo DNA syntactic analysis system. Each plate was crossed with two different recipient barcode plates. For oligonucleotide plates in odd-numbered assembly rounds, the donor skeleton (pSL1063 or pSL1064) contained a HygR-SacB cassette. Recombinant plasmids were selected overnight on LB+Hyg+Gm+Rha+Ara plates at 37°C, following crossing with a BPS recipient array on LB+Ara+IPTG agar at approximately 3 hours. For oligonucleotide arrays in even-numbered rounds containing an NsrR-PheS cassette, recombinant plasmids were selected overnight on LB+Nat+Gm+Rha+Ara plates at 37°C, following crossing with a BPS collection. To assemble 12 genes, the BUN20 donor strain carrying the first round of oligonucleotides was first crossed with the RE1133 recipient strain carrying the pSL1086 recipient plasmid. 50 μl of each overnight culture was mixed, centrifuged at 8000 rpm for 1 minute, resuspended in 50 μl of LB, and incubated at 37°C for 30 minutes. The crossed cells were then seeded onto an LB+aTC+cuminate plate and incubated at 37°C for 4 hours to induce Cas9 and λ-red. To isolate cells carrying the recombinant recipient plasmid along with the first round of oligonucleotides, the crossed cells were smeared onto an LB+Hyg+Gm+IPTG agar plate and incubated overnight at 37°C. The R6Kγ origin on the donor plasmid was pir. +Since the RE1133 recipient strain is non-functional in the background, unrecombined donor plasmids are promptly removed. Colonies from the selection plate were further purified by selecting them on LB+Hyg+Gm+IPTG+4CP and removing any remaining unrecombined pSL1086 recipient plasmids. To assemble the second round of oligonucleotides, the purified colonies carrying the recombinant recipient plasmid along with the first round of oligonucleotides were then crossed with the BUN20 donor strain carrying the second round of oligonucleotides. The same procedure as the first assembly was used, except that the cells carrying the recombinant recipient plasmid were selected on LB+Nat+Gm+IPTG and purified on LB+Nat+Gm+IPTG+6% sucrose. The same assembly procedure was used for all subsequent assembly steps, with LB+Hyg+Gm+IPTG+6% sucrose used for selection of odd-numbered rounds of assembly and LB+Nat+Gm+IPTG used for even-numbered rounds of assembly. This process was repeated five times until all 12 genes were completely assembled (Figure 7G). The sequences of the assembled products were verified for correctness using Sanger sequencing (Figure 7H) and an Oxford Nanopore MinION sequencer.

[0259] 9kb Fragment Assembly: A 9kb DNA block was assembled from position 41489–50489 of chromosome II of the BY4741 Saccharomyces cerevisiae strain. Using the same Python script used for the assembly of 12 genes, the 9kb block was split into three approximately 3kb DNA blocks with 50–75 bp homology between the subsequent DNA blocks. These three DNA blocks were PCR-amplified using genomic DNA from yeast strain BY4741 as a DNA template. Genomic DNA was extracted using the MasterPure Yeast DNA Purification Kit. The first, second, and third DNA blocks were inserted into donor plasmids pSL1064, pSL1063, and pSL1107, respectively, using AscI / NotI restriction digestion and T4 ligation. BUN20 donor strains were transformed with the resulting ligated products using a standard bacterial transformation procedure. After sequence validation of the donor plasmid using Sanger sequencing, the DNA blocks were assembled in the RE1133 / pSL1086 recipient strain. The donor strain carrying the first DNA block was mated with RE1133 / pSL1086 and grown on LB+aTC+cuminate for 4 hours at 37°C. Cells carrying the recombinant recipient plasmid were selected on LB+Hyg+Gm+IPTG and further purified on LB+Hyg+Gm+IPTG+6% sucrose. The resulting colonies were mated with the donor strain carrying the second DNA block, the recombinant recipient plasmid was selected on LB+Nat+Gm+IPTG and purified on LB+Nat+Gm+IPTG+4CP. Finally, the resulting colonies were mated with the donor strain carrying the third DNA block, the recombinant recipient plasmid was selected on LB+Hyg+Gm+IPTG and purified on LB+Hyg+Gm+IPTG+6% sucrose. Next, the sequence of the assembled product was verified to match the expected sequence using the Oxford Nanopore MinION sequencer and gel electrophoresis (Figure 7I).

[0260] Amplicon sequencing for syntactic analysis of oligonucleotide libraries: To extract recombinant plasmids, cells were scraped from selection plates and minipreps were performed using the Plasmid Plus Mini Kit (QIAGEN). Plasmid DNA was then quantified and diluted to approximately 1 ng / μl. This is approximately 1.5e6 copies per characteristic barcode pair on a 96-array mating plate. The two-step PCR described (Levy, SF et al., Nature 519, pp. 181-186 (2015)) was modified and performed. First, 4-5 cycles of PCR were performed using OneTaq polymerase (New England Biolabs) with the forward (pBPS_fwr) and reverse (pBPS_rev) primers listed in Table 1. Approximately 1 ng of recombinant plasmid DNA was amplified in a single 50 μl PCR reaction. The primers for the first step of PCR have the following general configuration. pBPS_fwr: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNNXXXXXXttcggttagagcggatgtg(Sequence ID 41) pBPS_rev: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCNNNNNNNNXXXXXXXXXaggtaacccatatgcatggc(Sequence ID 42)

[0261] In these sequences, N corresponds to any random nucleotide and is used in downstream analysis to eliminate asymmetry in counts induced by PCR jackpotting. X corresponds to one of several multiplexing tags (e.g., the multiplexing tags in Table 1 above) that allow for differentiation when different samples are loaded into the same sequencing flow cell. Examples of multiplexing tags are the underlined sequences in Table 1. Lowercase sequences correspond to priming sites on recombinant plasmids. Uppercase sequences correspond to sequencing primers for Illumina Read 1 or Read 2. PCR products were purified using a NucleoSpin column (Macherey-Nagel) and eluted in 33 μl of water. A second 23–25 cycle PCR was performed using PrimeStar HS polymerase (Takara) with 33 μl of purified product from the first PCR as a template per tube and a total volume of 50 μl. The primers for this reaction were the standard Illumina TruSeq dual-index primers (D501-D508 and D701-D712) listed in Table 2. The PCR products were then purified using a NucleoSpin column. Amplicons from each mating plate were uniquely labeled with custom-made multiplexing tags and Illumina standard indices. This quadruple-indexing strategy not only increases the multiplexing capacity of the sequencing library but is also beneficial for downstream analysis of amplicon chimeras. The purified amplicons were pooled and paired-end sequenced on Illumina MiSeq (2 × 300 bp) with 25% PhiX DNA spike-in. Sequencing reads were aggregated into barcodes using Bartender (Zhao, L. et al., Bioinformatics 34, pp. 739-747 (2017)).

[0262] Design of repeatable in vivo suturing techniques: The in vivo suturing system leverages bacterial conjugation mechanisms and lambda-Red homologous recombination. The donor vector contains a conditional replication origin derived from R6K oriγ, which depends on a functional transacting factor π encoded by the gene pir1 or its relaxed copy number control version, pir1-116 (Metcalf, W. Gene 138, pp. 1-7 (1994)). This specific origin allows the plasmid to be maintained in donor host cells (e.g., BUN20) that have a genomically integrated pir116 allele, but not in recipient cells lacking this allele. Other important features in the donor vector include oriT, a skeletal marker (kanR), and a constitutive gRNA expression cassette (gRNA). T1 or gRNA T2 ), and a swapping region. Within the swapping region are two pairs of specific gRNA target sites, a double-selectable cassette, and a common long homology sequence (300 bp) for each suture round.

[0263] The entry recipient vector includes an origin of replication (ColE1 or pLacIQ-p15A), a skeletal marker (GmR), and a swapping region, the swapping region consisting of a homology sequence, two gRNA target sites, and a dual-selectable cassette. Some recipient host cells have helper plasmids containing a rhamnose-inducible λ-red recombinant system and an arabinose-inducible Cas9 endonuclease. A temperature-sensitive mutant derivative of the origin of replication (pSC101-ori TS This provides a convenient means of curing the helper plasmid at 42°C when assembly is complete (Hashimoto, TJBacteriol. 127, pp. 1561-1563 (1976)). Other recipient host cells (RE1133) contain a genetically integrated tetracycline-inducible λ-red recombination system and a cuminate-inducible Cas9 endonuclease.

[0264] In each round, after the donor vector is transferred into the recipient cells, both λ-red and Cas9 are induced in a recipient cell-dependent manner in the presence of arabinose and rhamnose or cuminate and tetracycline. T1 Cas9 is guided into both the donor and recipient plasmids to generate double-strand breaks. DNA breaks have been shown to strongly stimulate homologous recombination (see below) (Kuzminov, A. Microbiol. Mol. Biol. Rev. 63, p. 751 (1999)). The donor-derived fragment contains the DNA of interest, two different gRNA target sites for the next round of assembly, a double-selectable marker, and an H3 homology region. The DNA for assembly is designed so that the first 50 bp is homologous to the last 50 bp of the assembled sequence present on the recipient plasmid. These 50 bp act as one homology arm for double crossover homologous recombination, with the other being H3. This swapping event results in a recombined recipient plasmid containing new DNA from the donor. Three different selections are performed to ensure the accuracy of the recombination. 1) Selection for an antiselectable marker (PheS or SacB), 2) Selection for a positive selectable marker (HygR or NsrR), and 3) Selection for a marker on the recipient skeleton (GmR). Selection for helper plasmids and P araBAD and P rhaBAD Promoter or P km-cymR and P Tet2 Promoter suppression is also performed to prevent excessive recombination and undesirable DNA breaks. By alternating between two different gRNAs and two different dual-selectable cassettes, repeatable in vivo suturing becomes possible, efficiently assembling new DNA fragments in a linear fashion, resulting in the desired sequence with the only theoretical limit being a tolerable plasmid size. With liquid handling and multiplexed pinning robots, this platform is highly scalable, enabling thousands of parallel gene assemblies per round.

[0265] CRISPR-Cas9 can efficiently stimulate in vivo suturing: To test whether the CRISPR / Cas9 system provides precise DNA breaks and promotes homologous recombination, in vivo suturing was completed in the presence or absence of target gRNA, Cas9, and λ-red. Recombinant plasmids were recovered when all three were present (Figure 3), demonstrating that CRISPR / Cas9 facilitates double-strand breaks of DNA and enhances the efficiency of λ-red recombination.

[0266] Assembly of functional fluorescent genes in liquid: To demonstrate the ability to assemble multiple fragments into a functional gene using an in vivo suture system, three pieces of the mEGFP gene were constructed by PCR and cloned into a suitable donor skeleton. The three fragments were then sequentially assembled into an entry recipient vector. After each round and before the next, the recovered clones were examined by both restriction digestion and Sanger sequencing to verify assembly accuracy. Both colony-touch PCR and plasmid extraction were performed to ensure that the helper plasmid was retained in each round. In the final round of assembly, green fluorescence was observed in all colonies (approximately 300) on the selection plate, indicating high assembly fidelity (Figures 7A-7I).

[0267] Effect of homology length on suture fidelity: To test how the fidelity of in vivo sutures depends on the homology length between fragments, seven donor vectors were constructed, each containing a third fragment of an mEGFP assembly with a different homology length to the second fragment. After conjugation and recombination, cells induced from different conjugation / recombination events were plate-seeded on selective medium, and the fractions of fluorescent colonies were counted. The results suggest that in this embodiment, a homology of approximately 40 bp is likely to produce error-free fusion products (Figures 7A-7I).

[0268] Multiplexed assembly of functional fluorescent genes on agar: To demonstrate the ability to assemble functional genes on agar plates, three pieces of the mEGFP gene were constructed by PCR and cloned into appropriate donor skeletons. The three fragments were then sequentially assembled into entry recipient vectors in 96-pin or 384-pin format. After the final round of assembly, green fluorescence was observed at the 96 / 96 position in the 96-pin format and at the 383 / 384 position in the 384-pin format, indicating that the suture fidelity on agar was comparable to that in liquid (Figures 7A-7I). Throughout the assembly process, plasmids were recovered from various colonies (all colonies were scraped) and examined by restriction digestion to verify the accuracy of the assembly and / or the retention of helper plasmids. Typical digestion patterns of the final assembly product indicated high-purity recombinant plasmids free from observable undesirable products (e.g., non-recombinant plasmids). To further characterize suture fidelity, 96 positions from the 384-position assembly were sequenced by Sanger sequencing. Sequencing products were derived from colony PCR of pipette tips in contact with each colony. 94 / 96 colonies were found to contain the correct mEGFP sequence. One colony contained an intermediate product (product of the first round of assembly), and one colony contained a suture error (large deletion).

[0269] Multiplexed Assembly of Genes Derived from Oligonucleotide Pools: To demonstrate the ability to assemble various DNA constructs derived from oligonucleotide pools, nine distinguishable serine / tyrosine recombinases and three distinguishable fluorescent molecules (mPapaya, mPlum, sfGFP) were constructed from a 300 bp oligonucleotide pool purchased from IDT(oPools). Each gene was sequentially assembled from five oligonucleotides stitched together. The oligonucleotide pool was integrated into a fitted donor plasmid, thereby transforming donor bacteria. The bacterial pool was syntactically analyzed in a sequence-validated ordered array using the method described in Example 2 (In vivo DNA Analysis Method). The array of donor cells was sequentially joined to recipient cells, and each gene was assembled with at least three replications. Each assembly was determined to contain the correct DNA sequence by Sanger sequencing (Figure 7G).

[0270] Long DNA Assembly: To demonstrate the ability to assemble and produce long assemblies using long DNA blocks, a 9kb segment of the Saccharomyces cerevisiae genome was reconstructed from three 3kb blocks. 3kb blocks were amplified from the genomic DNA and integrated into fitted donor plasmids, thereby transforming donor bacteria and verifying their sequences. Donor cells were sequentially conjugated to recipient cells. To verify the correct assembly product at each assembly step, recombinant recipient plasmids were purified from recipient cells, linearized with restriction enzymes, and analyzed by gel electrophoresis (Figure 7H). The assembly products were also sequenced correctly by Sanger sequencing.

[0271] [Example 2] Method for in vivo DNA analysis Plasmid sequences: Information regarding the plasmids used in in vivo DNA analysis can be found in Table 5 and Figure 37.

[0272] [Table 5]

[0273] Construction of barcoded donor plasmids: The donor vector was constructed using standard cloning methods. It contains 1) KanR (kanamycin resistance), 2) oriT (transfer origin), 3) R6K oriγ (conditional replication origin dependent on phage-derived pir1 expression), and 4) a swapping region, an I-SceI-H1-H4-I-SceI structure, where I-SceI is the recognition site of the endonuclease SceI, and H1(5'-ttgccctctctcttcattcagggtcatgagaggcacgccattcaaggggagaagtgagatc-3' (SEQ ID NO: 43)) and H4(5'-aagaacttttctatttctgggtaggcatcatcaggagcagga-3' (SEQ ID NO: 44)) are homology regions for recombination. In the swapping region of the donor vector, a selection cassette (HygR-SacB or NsrR-PheS) was cloned between H1 and H4 to generate a donor skeleton plasmid for syntactic analysis. To insert random barcodes into the donor skeletons (pSL438 and pSL439), an oligonucleotide (pXL633) containing a NotI restriction site, a barcode region with 15 random nucleotides, and homology regions to both donor skeletons was ordered from IDT. Using pXL633 paired with pXL585, PCR of the barcode was performed using approximately 1 ng of pSL438 or pSL439 as a template. The resulting PCR products were restriction-digested and ligated into the corresponding donor vectors via the NotI and XmaI sites. Following the same cloning protocol described above, competent donor cells BUN20 were transformed with the ligation product, and barcoded donor clones were selected on LB agar plates containing 50 μg / ml kanamycin (Kan) at 37°C. Transformants were then randomly selected and aligned to generate two barcoded 96-well donor collections, pSL438_BC and pSL439_BC. To identify the barcode sequences within the sequenced donor collections, the barcode-containing regions were amplified by colony-touch PCR using pXL583 and pXL584 primers.The amplicons were then purified and Sanger-sequenced using pXL583. Barcodes were then extracted, and two lists of known donor barcode collections were compiled.

[0274] Construction of barcoded recipient plasmids: Plasmid pSL937, used as a backbone for generating barcoded recipient collections by inserting and aligning random barcodes, was constructed using standard methods from the following sources: 1) Plasmid backbone / origin of replication from pBR322, 2) pUC18-mini-Tn7T-Gm 3 1) Derived GmR (gentamicin resistance marker), 3) homologous sequences H1 and H4, and two I-SceI recognition sites in the H1-I-SceI-I-SceI-H4 structure, 4) pSLC-217 4 Rhamnose-derived rhamnose-induced toxin relE(P rhaBAD -relE) was cloned between two SceI sites. Oligonucleotides containing random barcodes were synthesized by IDT and inserted into pSL937 by restriction digestion and ligation.

[0275] To insert a random barcode into the recipient skeleton (pSL937), an oligonucleotide (pXL631) containing an XhoI restriction site, a barcode region with 20 random nucleotides, and a region homologous to pSL937 was ordered from IDT. Barcodes were generated by PCR using pXL631 paired with pXL154 and approximately 1 ng of pSL937 as a template. The resulting PCR product was digested and ligated to pSL937 using MluI and XhoI restriction sites. The ligation reaction was carried out overnight at 16°C with a molar ratio of 3:1 between the barcode insert and the vector. The ligation product was then used to develop the spectinomycin-resistant helper plasmid pML104. 1Competent BUN21 cells containing the specified barcode were transformed. Barcoded recipient clones were selected at 30°C on LB agar plates containing 50 μg / ml spectinomycin (Sp), 20 μg / ml gentamicin (Gm), and 2% glucose. Transformants were then randomly selected and aligned in 96-well plates. The barcode sequences at each position in the aligned recipient collection were identified by sequencing. All 841 barcodes could be clearly identified. These barcodes were rearranged into eight new 96-well plates so that a unique barcode was present at each position.

[0276] Aligned mating: Each barcoded donor plate (two 96-position plates) was mated with each barcoded recipient plate (eight 96-position plates). The donor barcode collection was grown overnight on LB+Kan plates at 37°C. The recipient arrays were grown overnight on LB+Sp+Gm+2% glucose at 30°C. The agar medium for aligned mating contained 0.2% arabinose (Ara) and 0.1 mM IPTG and was preheated at 37°C for 1 hour. Both donor and recipient clones were transferred to the mating plates using a SINGER ROTOR HDA pin pad and grown at 37°C for approximately 3 hours. Each recipient plate was mated with two donor barcoded plates (pSL438_BC and pSL439_BC). The crossed cells were then transferred to a selective LB plate (LB+Ara+Rha+Gm+Hyg) containing 0.2% arabinose, 0.2% rhamnose (Rha), 25 μg / ml gentamicin, and 50 μg / ml hygromycin (Hyg). Recombinant clones were then selected overnight at 37°C.

[0277] Amplicon Sequencing: To extract recombinant plasmids, cells were scraped from selection plates and miniprepped using the Plasmid Plus Mini Kit (QIAGEN). Plasmid DNA was quantified and diluted to approximately 1 ng / μl. This represents approximately 1.5 × 10⁶ unique barcode pairs per 96-array plate. 6 This is a copy. A two-step PCR was performed. First, 4-5 cycles of PCR were performed using OneTaq polymerase (New England Biolabs) with the forward (pBPS_fwr) and reverse (pBPS_rev) primers listed in Table 1. Approximately 1 ng of recombinant plasmid DNA was amplified in a single 50 μl PCR reaction. To increase the multiplexing of sequencing samples, plasmid DNA from specific pairs of mated plates was amplified using specific primer pairs for the first and second PCRs (see Tables 1 and 2). This makes it possible to pool multiple mated plates together in a single sequencing library. The cycle conditions for the first step are shown in Table 6 below.

[0278] [Table 6]

[0279] The primers for the first step of PCR have this general configuration.

[0280] pBPS_fwr:ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNNXXXXXXttcggttagagcggatgtg(Sequence ID 45) pBPS_rev:GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCNNNNNNNNXXXXXXXXXaggtaacccatatgcatggc(Sequence ID 46)

[0281] In these sequences, N corresponds to any random nucleotide and is used in downstream analysis to eliminate asymmetry in counts induced by PCR jackpotting. X corresponds to one of several multiplexing tags, allowing for distinction when different samples are loaded into the same sequencing flow cell. Lowercase sequences correspond to priming sites on recombinant plasmids. Uppercase sequences correspond to Illumina Read 1 or Read 2 sequencing primers. PCR products were purified using a NucleoSpin column (Macherey-Nagel) and eluted in 33 μl of water. The second 23–25 cycles of PCR were performed using PrimeStar HS polymerase (Takara) with 33 μl of purified product from the first PCR as a template and a total volume of 50 μl per tube. The primers for this reaction were the standard Illumina TruSeq dual-index primers (D501-D508 and D701-D712) listed in Tables 1 and 2. The cycle conditions for the second step are shown in Table 7 below.

[0282] [Table 7]

[0283] Next, the PCR products were purified using a NucleoSpin column. Amplicons from each mating plate were uniquely labeled with a custom primer index (first PCR) and a standard Illumina index (second PCR). This quadruple-indexing strategy increased the multiplexing capability for sequencing. The purified amplicons were pooled and paired-end sequencing was performed on Illumina MiSeq, HiSeq, or NextSeq with 25% PhiX genomic DNA spike-in and approximately 800 reads per barcode pair.

[0284] Sequencing Analysis: Donor-recipient dual barcode amplicon sequencing data was analyzed using a custom Python script and Bartender in the following steps. First, Illumina reads were demultiplexed using Illumina indices. Sequences that did not exactly match the two Illumina indices were discarded. The normal representation was then obtained from the demultiplexed sequences. “\D * ?(.GGC|T.GC|TG.C|TGG.)\D{4,7}?AA\D{4,7}?TT\D{4,7}?(.CGG|G.GG|GC.G|GCG.)\D * (Donor barcode) and “\D * ?(.ACA|G.CA|GA.A|GAC.)\D{4,7}?AA\D{4,7}?AA\D{4,7}?TT\D{4,7}?(.TCG|C.CG|CT.G|CTC.)\D * (Recipient barcode) Barcodes were extracted using [a specific method / tool]. Specific molecular identifiers (UMI, N in pBPS_fwr and pBPS_rev) were also extracted based on their expected positions in the Illumina reads. Barcode reads, containing a mixture of true barcode sequences and sequences with errors due to PCR or sequencing, were then aggregated into a consensus sequence using [a specific method / tool]. Each barcode cluster was then examined for replicated UMIs (indicating PCR replicas) using [a specific method / tool], and all replicas were removed to generate the final count for each barcode pair. Fewer than 20 dual barcodes were excluded, many of which were expected to be PCR chimeras (barcodes fused by PCR amplification). The remaining reads were used to elucidate the position of each donor barcode from its corresponding recipient barcode.

[0285] Whole plasmid sequencing on the Oxford Nanopore platform: Recombinant plasmids, including positioning barcodes and oligonucleotides, were extracted as previously described in the amplicon sequencing section. Circular plasmids were linearized with restriction enzyme PmlI (NEB) at 37°C for 2 hours. The linearized products were size-selected by running them onto a 1.2% agarose gel and recovered using the Zymoclean Gel DNA Recovery Kit (Zymoresearch). A sequencing library for the Oxford Nanopore platform was constructed using a ligation sequencing kit (SQK-LSK110, Nanoporetech). 300 ng (approximately 100 fmol) of the linearized recombinant plasmid library was repaired at the ends using NEBNext FFPE Repair Mix and the NEBNext Ultra II End repair / dA-tailing Module (NEB). The nanopore sequencing adapter (AMX-F) was ligated using NEBNext Quick T4 DNA Ligase (NEB). A 30 ng (approximately 10 fmol) library was loaded into a Flongle flow cell (R9.4.1, Oxford Nanopore) to generate reads for recombinant plasmids. The flow cell was run for 16 hours using Miniknow sequencer control software (version 21.11.7, Oxford Nanopore).

[0286] Oxford Nanopore Sequencing Analysis: (2) The sequencing adapter was identified and removed, and the files were separated using sample multiplexing barcodes with "Guppy_Barcoder" from Guppy version 6.0.1+652ffd179. (3) Unwanted sequences that aligned with these contaminants were removed using alignments querying contaminating sequences (sequences that serve as the origin for replication or transport for donor and / or helper plasmids) with "Minimap2" version 2.22-r1110-dirty. (4) Alignments to the expected backbone sequences of recombinant recipient plasmids were generated using "Minimap2" version 2.22-r1110-dirty, and identical backbone sequences were removed from each read using a custom Python script. (5) Using itermae version 0.6.0.1, a ambiguous, standard representation of the sequence surrounding the barcode was used to extract a position-specific barcode from this sequence, which was then aggregated using the message-passing Levenshtein distance approach of Starcode version 1.4. (6) Using these barcodes, a custom shell / oak script was used to separate the plasmid backbone-removed sequences into separate files for each demultiplexed sample and the aggregated barcode sequence. (7) Using the sequences for each barcode in each sample, a multi-sequence alignment was generated using kalign3 version 3.3.1. (8) Using a custom Python script, a single draft consensus sequence was generated from the multi-sequence alignment through a voting process. (9) This draft consensus sequence was refined using racon version 1.5.0, and the consensus was updated based on read matching and sequence quality. (10) This consensus sequence was further refined using "medaka" version 1.5.0 to generate a refined sequence for each positioning barcode in each sample. (11) "itermae" version 0.6.0.1, which includes different normal representations, was again used to extract the payload sequence, the intended target of the realignment project, from the refined region.The refined and positioned payload was analyzed by aligning the raw region from which the skeletal sequence had been removed with the refined region, aligning the payload extracted from the refined region with the intended target sequence, and aligning the raw region from which the skeletal sequence had been removed with all the refined regions generated in the dataset. Alignment was performed using a custom Python script with "Minimap" version 2.22-r1110-dirty or BioPython PairwiseAlignment functionality. Using a custom R script, reads were identified as "on target," i.e., reads with more than 90% identity with the refined region generated for the sample and barcode (i.e., well). A well was classified as "pure" if more than 90% of the raw reads were "on target" with the refined consensus sequence. A sequence was defined as "correct" if the refined payload was exactly identical to one of the intended target sequences, by comparing the length of the refined payload with its alignment to the intended target sequence.

[0287] Construction of a donor plasmid library containing an oligonucleotide pool: Plasmid pSL1071, containing an NsrR-PheS cassette, two I-SceI sites, and two homology regions (H1 and H4) for recombination, was used as a backbone for inserting the oligonucleotide pool into it. The oligonucleotide pool, containing 300 bp oligonucleotides on each side, was designed as follows: GCTTATTCGTGCCGTGTTATGGCGCGCCNN···NNGCGGCCGCGGGCACAGCAATCAAAAGTA (Sequence No. 47) We placed an order with IDT accordingly. GCTTATTCGTGCCGTGTTAT and GGGCACAGCAATCAAAAGTA (SEQ ID NO: 48) are priming sites for forward and reverse primers for amplifying the oligonucleotide pool, GGCGCGCC (SEQ ID NO: 49) and GCGGCCGC (SEQ ID NO: 50) are recognition sites for restriction enzymes AscI and NotI, and NN...NN represents a 244nt sequence randomly selected from the human genome assembly GRCh38. The oligonucleotide pool was amplified using 7ng of template DNA and KAPA HiFi polymerase (Roche) under the cycle conditions described in Table 8.

[0288] [Table 8]

[0289] PCR products were purified using DNA Clean & Concentrator-5 (Zymoresearch). AscI and NotI restriction enzyme recognition sites were used to clone the PCR products into the donor plasmid pSL1071. Digestion of the PCR products with pSL1071 was performed at 37°C for 4 hours. The digested products were then size-selected by running them on a 1.2% agarose gel and recovered using the Zymoclean Gel DNA Recovery Kit (Zymoresearch). Ligation was performed using T4 DNA ligase (NEB) with 25 ng of digested vector and 3.8 ng of insert at 16°C for 15 hours. BUN20 was transformed with the ligated products and conjugated to an array of barcoded recipient plasmids (described above), and the sequences of the constructs at each position in the donor array were determined.

[0290] Results: Barcode array positioning. To validate the accuracy of parsing and positioning, two 96-well plates (pSL438_BC and pSL439_BC) with known donor barcodes were crossed with eight 96-well recipient plates. Data from these 1536 crossing events showed that the correct position was identified and the sequence validated for 93.82%±0.34%, 95.59%±0.27%, and 96.04%±0.21% of donors using 1, 2, and 3 events, respectively (Figures 47A-47D). All failures were due to missing sequencing data. No inaccurate positions were ever identified for donors in the sequencing data. Similar results were obtained when the recipient barcode position was determined from the donor barcode.

[0291] Results: Parsing of oligonucleotide pools To further validate the accuracy of the syntactic analysis, a pool of 100 oligonucleotides was aligned and their sequences were validated. This pool contained a 244-nucleotide sequence randomly selected from the human genome, synthesized as an "oPool" by IDT (Integrated DNA Technologies), and inserted into the inventors' donor plasmid pSL1071 by ligation. BUN20 transformants carrying these plasmids were pooled and then randomly aligned into a total of 20 384-well plates. These aligned bacterial plates were then joined to an aligned collection of recipient barcode strains (the barcode location was known). Recipient cells containing recombinant oligonucleotide-barcode plasmids were pooled. Plasmids were sequenced by nanopore sequencing. Using the sequencing results, for each well in each plate, the consensus sequence of the oligonucleotide, whether the consensus sequence was identical to the expected sequence in the oligonucleotide pool, and whether any other oligonucleotide sequences were present at a low frequency (contaminants) were determined (Figures 47G, 47C, and 47D). Consensus sequences were generated in 5,101 wells (66.4%) out of 7,680 wells available across all plates. Of these consensus sequence wells, 2,329 wells (45.6%) were pure and perfectly matched the target oligonucleotide. These 2,329 perfectly matched oligonucleotides represented 82% of the oligonucleotides expected to be present in the pool.

Claims

1. A method for assembling multiple DNA elements into an assembled DNA element in a recipient cell, (a) A step of bringing a first donor cell containing the first donor plasmid into contact with a recipient cell containing a recipient oligonucleotide, under conditions that transfer the first donor plasmid from the first donor cell to the recipient cell by conjugation, wherein the recipient oligonucleotide is present in the recipient cell plasmid or recipient cell genome. The first donor plasmid comprises, in a sequential order, a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide containing a fragment of a first DNA element (oligo 1), a second homologous recombination region (HR2) containing two homologous recombination regions (HR2.1, HR2.2), and a third endonuclease site (C3), wherein HR2.1 and HR2.2 are adjacent to a non-homologous region (NHR) containing two endonuclease sites (C2.1, C2.2) adjacent to a selectable marker. The recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2, wherein HR3 and HR4 are adjacent to a non-homologous region containing two endonuclease sites (C4.1, C4.2) adjacent to a selectable marker. (b) The first donor plasmid and the recipient oligonucleotide are subjected to endonuclease cleavage using a first endonuclease, the first donor plasmid being cleaved at C1 and C3 and the recipient oligonucleotide being cleaved at C4.1 and C4.2, thereby producing a first donor cassette containing homologous recombination regions HR1 and HR2.2 at each end and a recipient oligonucleotide having corresponding recombination regions HR3 and HR4. (c) The first donor cassette and the recipient oligonucleotide are subjected to conditions that cause homologous recombination to recombine the first donor cassette into the recipient oligonucleotide, thereby providing a first recombined recipient oligonucleotide containing fragments of the first DNA element, following homologous recombination of HR1 and HR3 and HR2.2 and HR4. (d)(i) A step of bringing a second donor cell containing the second donor plasmid into contact with a recipient cell containing the first recombined recipient oligonucleotide, under conditions that the second donor plasmid is transferred from the second donor cell to the first recipient cell by conjugation, The second donor plasmid comprises, in a sequential order, a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a fragment of the second DNA element (oligo 2), a sixth homologous recombination region (HR6) containing two homologous recombination regions (HR6.1, HR6.2), and a sixth endonuclease site (C6), wherein HR6.1 and HR6.2 are adjacent to a non-homologous region (NHR) containing two endonuclease sites (C7.1, C7.2) adjacent to a selectable marker. (e) The second donor plasmid and the recipient oligonucleotide are subjected to endonuclease cleavage using a second endonuclease, the second donor plasmid being cleaved at C5 and C6 and the recipient oligonucleotide being cleaved at C2.1 and C2.2, thereby producing a second donor cassette containing homologous recombination regions HR5 and HR6.2 at each end and a recipient oligonucleotide having corresponding recombination regions HR2.1 and HR4. (f) The recipient oligonucleotide recombined with a second donor cassette is subjected to conditions that cause homologous recombination to recombine with a recipient oligonucleotide recombined with a first donor fragment in recipient cells, thereby providing a second recombined recipient oligonucleotide containing fragments of first and second DNA elements (oligo 1, oligo 2) following homologous recombination of HR5 and HR2.1 and HR6.2 and HR4, and these oligonucleotides form a DNA assembly. Each donor plasmid contains a transport origin (oriT) and a conditional replication origin. The oligonucleotide encoding the guide RNA or the first endonuclease targeting the first, third, and / or fourth endonuclease sites is present on the first donor plasmid and / or in the recipient cell, A method comprising an oligonucleotide encoding a guide RNA or a second endonuclease targeting a second, fifth, and / or sixth endonuclease site, which is present on a second donor plasmid and / or in recipient cells.

2. The method according to claim 1, wherein steps (d) through (f) are repeated one or more times using a third or subsequent donor cell comprising a third or subsequent donor plasmid comprising a third or subsequent oligonucleotide encoding a fragment of a third or subsequent DNA element (oligo 3, oligo 4, ..., oligo N), thereby forming a third or subsequent recombined recipient oligonucleotide comprising fragments of the first, second, and third or subsequent DNA elements, which together form a DNA assembly.

3. The method according to claim 1 or 2, wherein step (a) comprises a plurality of first donor cells each containing a different first donor plasmid, and step (b) comprises a plurality of second, third, or subsequent donor cells each containing a different second, third, or subsequent donor plasmid.

4. The method according to any one of claims 1 to 3, wherein the expression of the first and / or second endonuclease is inducible, and further comprising the step of inducing the expression of the first and / or second endonuclease.

5. The method according to claim 4, wherein the first and / or second endonuclease is selected from RNA-guided endonucleases, homing endonucleases, transcription activator-like effector nucleases, and zinc finger nucleases.

6. The method according to any one of claims 1 to 5, wherein the first, second, or subsequent donor plasmid includes a selectable marker for selecting the first oligonucleotide, the second oligonucleotide, or the subsequent oligonucleotide to integrate into the recipient oligonucleotide.

7. The method according to any one of claims 1 to 6, wherein the recipient oligonucleotide comprises an antiselectable marker for selecting recipient cells that do not contain the first, second, third, or subsequent oligonucleotides.

8. The method according to any one of claims 1 to 7, wherein the donor plasmid or recipient oligonucleotide includes an inducible high-copy origin of replication.

9. The method according to any one of claims 1 to 8, wherein the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating a plasmid of more than 30 kilobases in length.

10. The method according to any one of claims 1 to 9, wherein the donor plasmid or recipient cell contains an oligonucleotide encoding one or more homologous DNA repair genes.

11. The method according to any one of claims 1 to 10, wherein the donor plasmid or recipient cells contain oligonucleotides encoding one or more recombinant-mediated gene manipulation genes.

12. The method according to any one of claims 1 to 11, wherein the first, second, or subsequent homologous recombination (HR) regions and their corresponding HR regions on the recipient oligonucleotide each comprise 20 to 500 base pairs or 20 to 60 base pairs, respectively.

13. The method according to any one of claims 1 to 12, comprising the step of constructing a DNA library using two or more recipient oligonucleotides having compatible homologous recombination regions.

14. The method according to claim 13, which is used to combine genetic regions such as genes, promoters, terminators, and regulatory regions from different species, to construct and / or combine gene regulatory pathways, or to assemble a bacterial array containing plasmids for a screening assay.

15. The method according to any one of claims 1 to 14, wherein, prior to steps (a) and (b) of claim 1, first and second oligonucleotides containing fragments of first and second DNA elements are inserted into first and second donor plasmids.

16. The method according to any one of claims 1 to 15, wherein the DNA assembly comprises at least a portion of a gene, promoter, enhancer, terminator, intron, intergenetic region, barcode, guide RNA (gRNA), or a combination thereof.

17. The method according to any one of claims 1 to 16, wherein the DNA assembly is 100 to 500,000 nucleotides in length.

18. The method according to any one of claims 1 to 17, wherein each first donor cell is located in a first ordered array, and each second, third, or subsequent donor cell is located in a second, third, or subsequent ordered array.

19. The method according to any one of claims 1 to 18, wherein the method generates a combinatorial library containing a plurality of different assembled DNA elements.

20. The method according to any one of claims 1 to 19, wherein the selectable marker is located in a non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions.

21. The method according to any one of claims 1 to 20, wherein the oppositely selectable marker is located in a non-homologous region between HR2.1 and HR2.2, and / or between HR6.1 and HR6.2, and / or between subsequent HR regions.

22. The method according to any one of claims 1 to 21, wherein the conditional origin of replication depends on the presence of an oligonucleotide or conditions for cell proliferation.

23. The method according to any one of claims 1 to 22, wherein the expression of one or more homologous DNA repair genes is inducible.

24. The method according to any one of claims 1 to 23, wherein the method is used for assembling a mutagenerating library or for constructing a combinatorial gRNA library.

Citation Information

Patent Citations

  • Multiplex editing system

    WO2015052231A2

  • Production and tracking of engineered cells with combinatorial genetic modifications

    WO2020163779A1