IN VIVO DNA ASSEMBLY AND ANALYSIS
The in vivo assembly of DNA elements via homologous recombination in recipient cells addresses the inefficiencies of existing methods, enabling efficient and high-throughput assembly of large DNA fragments and improved sequence identification.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
- Filing Date
- 2022-03-04
- Publication Date
- 2026-06-11
AI Technical Summary
Current methods for assembling DNA fragments are expensive, time-consuming, and limited in size and composition, requiring multiple purification steps and various enzymes, while identifying and isolating specific DNA sequences in complex mixtures remains challenging due to complex purification requirements and inefficient sequencing workflows.
A method for in vivo assembly of DNA elements using homologous recombination in recipient cells, involving donor and recipient plasmids with specific endonuclease sites and homologous recombination regions, allowing for the assembly of large DNA elements and barcoding of oligonucleotide sequences, with optional use of endonucleases and selectable/counter-selectable markers.
Enables efficient, high-throughput, and versatile assembly of DNA fragments up to 500,000 nucleotides, reducing the need for multiple enzymes and improving DNA sequence identification and isolation.
Smart Images

Figure 00000000_0000_ABST 
Figure 00000000_0001_ABST
Abstract
Description
RELATED APPLICATIONS
[0001] This application claims priority under 35 USC §119(e) over US Preliminary Application No. 63 / 157,497, filed on March 5, 2021, and US Preliminary Application No. 63 / 157,498, filed on March 5, 2021. SEQUENCE PROTOCOL
[0002] The contents of the sequence log text file named “41243-570001WO_Sequence_Listing_ST25.txt”, which was created on February 15, 2022 and is 24,576 bytes in size. DECLARATION ON THE RIGHTS TO INVENTIONS MADE WITHIN THE FRAMEWORK OF GOVERNMENT-FUNDED RESEARCH AND DEVELOPMENT
[0003] This invention was made with government support under DOE FWP# 100582 of the Department of Energy and NIST IAA P18-630-0001 of the National Institute of Standards and Technology. The government has certain rights to the invention. BACKGROUND
[0004] Recent advances in recombinant oligonucleotide technology have spurred research in traditional biology and bioengineering. A method for assembling DNA fragments via DNA recombination is disclosed in US2004126883A1. However, oligonucleotide assembly methods can be expensive and time-consuming, requiring multiple purification steps and various enzymes. Furthermore, current molecular biology techniques are limited in the size and composition of the DNA elements that can be combined. Therefore, methods for assembling DNA elements (e.g., promoters, gene fragments, etc.) are needed to overcome these limitations and circumvent the need for numerous and expensive enzymes (e.g., ligases, etc.). Novel methods are required to efficiently, with high throughput, and in a versatile, parallel manner assemble DNA fragments.
[0005] Advances in sequencing technologies enable the identification of long DNA fragments. However, the identification and isolation of specific DNA sequences in complex mixtures remains a challenge, due in part to complex purification requirements, low sample recovery, or inefficient sequencing workflows.
[0006] Here, solutions are offered for these and other technical problems. BRIEF SUMMARY OF THE INVENTION
[0007] This includes methods for the in vivo assembly of DNA elements and the DNA barcoding of oligonucleotide sequences.
[0008] The invention provides a method for assembling a plurality of DNA elements into a composite DNA element in a recipient cell, the method comprising: (a) contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions to (i) transfer the first donor plasmid from the first donor cell to the recipient cell by conjugation and (ii) recombine the first donor plasmid and the recipient oligonucleotide in the recipient cell by homologous recombination, wherein the recipient oligonucleotide is located in a recipient cell plasmid or the recipient cell genome, and wherein the first donor plasmid comprises, in successive order, a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a first DNA element fragment (Oligo1),a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2) and a third endonuclease site (C3); the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2; thus, after homologous recombination of HR1 with HR3 and HR2.2 with HR4, a first recombinant recipient oligonucleotide is formed, comprising the first DNA element fragment; (b) Contacting a second donor cell containing a second donor plasmid with the recipient cell containing the first recombinant recipient oligonucleotide under conditions to (i) transfer the second donor plasmid from the second donor cell to the first recipient cell by conjugation and (ii) recombine the second donor plasmid and the first recombinant recipient oligonucleotide,to form a second recombinant recipient oligonucleotide in the recipient cell by homologous recombination; wherein the second donor plasmid comprises, in successive order, a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a second DNA element fragment (Oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2), and a sixth endonuclease site (C6); whereby, following homologous recombination of HR5 with HR2.1 and HR6.2 with HR4, a second recombinant recipient oligonucleotide is provided, comprising the first and second DNA element fragments (Oligo1, Oligo2) that form a DNA assembly, wherein an oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third and / or fourth endonuclease site,on the first donor plasmid, and / or wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth, and / or sixth endonuclease site is present on the second donor plasmid. In some embodiments, HR2.1 and HR2.2 flank a non-homologous region comprising one (C2) or two endonuclease sites (C2.1, C2.2); wherein HR3 and HR4 optionally flank a non-homologous region comprising one (C4) or two endonuclease sites (C4.1, C4.2). In embodiments, HR6.1 and HR6.2 flank a non-homologous region comprising one (C7) or two endonuclease sites (C7.1, C7.2). In some embodiments, the DNA assembly comprises at least a portion of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, a guide RNA (gRNA), or a combination thereof.
[0009] In embodiments, step (b) is repeated one or more times with a third or subsequent donor cell comprising a third or subsequent donor plasmid containing compatible HR regions and a third or subsequent oligonucleotide encoding a third or subsequent DNA element fragment (Oligo3, Oligo4, ... OligoN), thereby forming a third or subsequent recombinant recipient oligonucleotide comprising the first, second and a third or subsequent DNA element fragment, which together form a DNA assembly.
[0010] In embodiments, step (a) comprises a plurality of first donor cells, each comprising a different first donor plasmid; and step (b) comprises a plurality of second, third, or subsequent donor cells, each comprising a different second, third, or subsequent donor plasmid; wherein optionally each first donor cell is located at a position in a first ordered array and each second, third, or subsequent donor cell is located at a position in a second, third, or subsequent ordered array; wherein the method optionally generates a combinatorial library comprising a plurality of different composite DNA elements.
[0011] In some embodiments, an oligonucleotide encoding a first endonuclease targeting the first, third, and / or fourth endonuclease site is present on the first donor plasmid and / or in the recipient cell. In some embodiments, an oligonucleotide encoding a second endonuclease targeting the second, fifth, and / or sixth endonuclease site is present on the second donor plasmid and / or in the recipient cell. In embodiments, the expression of the first and / or the second endonuclease is inducible, and the method further comprises the induction of the expression of the first and / or the second endonuclease. In embodiments, the first and / or the second endonuclease is selected from an RNA-directed endonuclease, a homing endonuclease, a transcription activator-like effector nuclease, and a zinc finger nuclease.
[0012] In embodiments, the first, second, or subsequent donor plasmid comprises a selectable marker that selects for the integration of the first, second, or subsequent oligonucleotide into the recipient oligonucleotide; wherein the selectable marker is optionally located within a non-homologous region between HR2.1 and HR2.2 and / or between HR6.1 and HR6.2 and / or between subsequent HR regions. In embodiments, the recipient oligonucleotide comprises a counter-selectable marker that selects against recipient cells that do not include the first, second, third, or subsequent oligonucleotide; wherein the counter-selectable marker is optionally located within a non-homologous region between HR2.1 and HR2.2 and / or between HR6.1 and HR6.2 and / or between subsequent HR regions.
[0013] In some embodiments, the donor plasmid includes a transfer origin.
[0014] In some embodiments, the donor plasmid includes a conditional origin of replication. In some embodiments, the conditional origin of replication depends on the presence of an oligonucleotide or on a condition of cell growth. In some embodiments, the donor plasmid or the recipient oligonucleotide includes an inducible high-copy origin of replication.
[0015] In some embodiments, the donor plasmid or recipient oligonucleotide comprises a replicon that can replicate plasmids with a length of more than 30 kilobases.
[0016] In some embodiments, the donor plasmid or recipient oligonucleotide is an artificial yeast chromosome (YAC), an artificial mammalian chromosome (MAC), an artificial human chromosome (HAC), or an artificial plant chromosome.
[0017] In some embodiments, the donor plasmid or the recipient oligonucleotide is a viral vector.
[0018] In some embodiments, the donor plasmid includes an oligonucleotide that enables conjugation of the plasmid.
[0019] In some embodiments, the donor plasmid or the recipient cell comprises an oligonucleotide encoding one or more homologous DNA repair genes; optionally, the expression of one or more homologous DNA repair genes can be induced.
[0020] In some embodiments, the donor plasmid or the recipient cell comprises an oligonucleotide that codes for one or more recombination-mediated genetic engineering genes.
[0021] In certain embodiments, the donor cell and the recipient cell are independently of each other a bacterial cell, wherein the bacterial cell is optionally E. coli, Vibrio natriegens or V. cholerae.
[0022] In some embodiments, the composite DNA element has a length of 100 nucleotides to 500,000 nucleotides.
[0023] In embodiments, the first, second or subsequent homologous recombination region (HR) and their corresponding HR regions on the recipient oligonucleotide each comprise about 20 to about 500 base pairs, optionally about 50 to 100 base pairs.
[0024] In embodiments, each of the foregoing methods may further comprise one or more steps, namely lysis of the recipient cells; amplification of a composite DNA element; isolation of a composite DNA element; isolation of a recipient oligonucleotide; sequencing of a composite DNA element; and sequencing of a recipient oligonucleotide.
[0025] In certain embodiments, the steps of contacting the first and second or subsequent donor cells with the first recipient cell are performed simultaneously, optionally with only a final donor plasmid comprising a selectable marker or with each donor plasmid comprising a selectable marker that is not present on the recipient oligonucleotide.
[0026] In one embodiment of one of the foregoing methods, the donor plasmid comprising the final DNA element forming part of a composite DNA element includes a barcode-homologous recombination region (BHR) to generate recipient cells, each containing a recombinant recipient oligonucleotide comprising the composite DNA element, the BHR, and a further HR; and the method further comprises: (i) constructing or acquiring an array of barcode donor cells, each containing a barcode donor plasmid comprising an HR homologous to the BHR, a specific barcode oligonucleotide, and a second HR homologous to the further HR of the recombinant recipient oligonucleotide;(ii) Contacting the array of barcode donor cells with an array of receiver cells under conditions to (a) transfer the barcode donor plasmids from the barcode donor cells to the receiver cells by conjugation and (b) recombine the barcode donor plasmids and receiver oligonucleotides in the receiver cells by homologous recombination, thereby generating an array of receiver cells comprising barcode assemblies.
[0027] In embodiments of one of the foregoing methods, each donor plasmid comprises a further pair of specific endonuclease sites CX, CY flanking a homologous barcode recombination region (BHR), and the method further comprises contacting an array of recipient cells, each comprising a DNA compilation, with an array of barcode donor cells, each containing a barcode donor plasmid comprising a pair of HR regions homologous to the BHR and flanking a specific barcode oligonucleotide, to generate an array of recipient cells comprising barcoded compilations.
[0028] In embodiments of one of the foregoing methods, the method further comprises contacting a reset donor cell comprising a reset donor plasmid with a recipient cell comprising a recombinant recipient oligonucleotide, wherein the reset donor plasmid comprises, in sequential order, a homologous recombination region (HRt) homologous to a terminal sequence of the DNA assembly, a reset endonuclease site, a selectable marker, a reset endonuclease site, a homologous recombination region (HRX), and a transfer origin; wherein the recombinant recipient oligonucleotide comprises, in sequential order, a reset endonuclease site, the DNA assembly, a homologous recombination region homologous to HRX (HRXa), and a reset endonuclease site;whereby, following homologous recombination between HRt and the terminal sequence of the DNA assembly and between HRX and HRXa, a reset plasmid is provided that includes the transfer origin and the DNA assembly. In some embodiments, the reset plasmid is located in a donor cell. In some embodiments, the reset plasmid contains a restricted origin of replication that functions in both donor and recipient cells. In some embodiments, the reset donor plasmid is constructed by a method that includes the introduction of an oligonucleotide insert comprising homologous recombination regions HRt, HRX flanking two endonuclease sites (C1, C2) and a counterselectable marker (CM), HRt-C1-CM-C2-HRX; or a library of such oligonucleotide inserts.Enabling the cleavage of endonuclease sites by an endonuclease and introducing a counter-selectable marker at the cleavage sites using homologous recombination.
[0029] In embodiments of one of the aforementioned methods, the recipient oligonucleotide comprises a mobile genetic element capable of transferring a DNA assembly to other cell types, including yeast cells, plant cells, mammalian cells, or other bacterial cells.
[0030] In embodiments of one of the foregoing methods, the method comprises the use of two or more recipient oligonucleotides with compatible homologous recombination regions to construct a DNA library.
[0031] In embodiments of one of the foregoing methods, the donor plasmid oligonucleotide comprises a first linker oligonucleotide homologous to a terminal sequence of a first DNA assembly and a second linker oligonucleotide homologous to a second oligonucleotide. In embodiments, the linker oligonucleotide further comprises an additional DNA element fragment that is not homologous to the first DNA assembly or the second DNA oligonucleotide. In some embodiments, the method is used to assemble a mutagenesis library to combine genetic regions such as genes, promoters, terminators, and regulatory regions from different species to construct and / or combine genetic regulatory pathways, to construct combinatorial gRNA libraries, or to assemble arrays of bacteria containing plasmids for screening tests.
[0032] In embodiments of one of the foregoing methods, prior to steps (a) and (b) the first and second oligonucleotides comprising the first and second DNA element fragments are inserted into the first and second donor plasmids.
[0033] This also provides a method for assembling a large number of DNA elements into a composite DNA element within a recipient cell, the method comprising: (a) Contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions to transfer the first donor plasmid to the recipient cell by conjugation, wherein the recipient oligonucleotide is contained in a recipient cell plasmid or in the recipient cell genome and wherein the first donor plasmid in sequential order a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a first DNA element fragment (Oligo1), a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2) and a third endonuclease site (C3) comprising HR2.1 and HR2.2 flanking a non-homologous region (NHR) comprising two endonuclease sites (C2.1, C2.2) flanking a selectable marker; The recipient oligonucleotide comprises a third homologous recombination region (HR3) that is homologous to HR1, and a fourth homologous recombination region (HR4) that is homologous to HR2.2 and HR3, and HR4 flanks a non-homologous region that includes two endonuclease sites (C4.1, C4.2) flanking a selectable marker; (b) Subjecting the first donor plasmid and the recipient oligonucleotide to endonuclease digestion with a first endonuclease to cleave the first donor plasmid at C1 and C3 and the recipient oligonucleotide at C4.1 and C4.2, thereby producing a first donor cassette comprising homologous recombination regions HR1 and HR2.2 at each end and a recipient oligonucleotide with compatible recombination regions HR3 and HR4; (c) Subjecting the first donor cassette and the recipient oligonucleotide to conditions to recombine the first donor cassette into the recipient oligonucleotide by homologous recombination, thereby providing a first recombined recipient oligonucleotide comprising the first DNA element fragment after homologous recombination of HR1 with HR3 and HR2.2 with HR4; (d) Contacting a second donor cell comprising a second donor plasmid with the recipient cell comprising the first recombinant recipient oligonucleotide under conditions to (i) transfer the second donor plasmid from the second donor cell to the first recipient cell by conjugation, wherein the second donor plasmid in sequential order a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a second DNA element fragment (Oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2) and a sixth endonuclease site (C6), with HR6.1 and HR6.2 flanking a non-homologous region (NHR) comprising two endonuclease sites (C7.1, C7.2) flanking a selectable marker; (e) Subjecting the second donor plasmid and the recipient oligonucleotide to endonuclease digestion with a second endonuclease to cleave the second donor plasmid at C5 and C6 and the recipient oligonucleotide at C2.1 and C2.2, thereby producing a second donor cassette comprising homologous recombination regions HR5 and HR6.2 at each end and a recipient oligonucleotide with compatible recombination regions HR2.1 and HR4; (f) Subjecting the second donor cassette and the recombinant recipient oligonucleotide to conditions to recombine the first donor fragment and the recombinant recipient oligonucleotide in the recipient cell by homologous recombination; which, following homologous recombination of HR5 with HR2.1 and HR6.2 with HR4, provides a second recombinant recipient oligonucleotide comprising the first and second DNA element fragments (Oligo1, Oligo2) that form a DNA assembly; wherein each donor plasmid includes a transfer origin (oriT) and a conditional replication origin; wherein an oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third and / or fourth endonuclease site is present on the first donor plasmid and / or in the recipient cell, and wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth and / or sixth endonuclease site is present on the second donor plasmid and / or in the recipient cell.
[0034] This document describes methods for conjugating barcodes to oligonucleotides, the methods comprising: (a) inserting each oligonucleotide of a mixture of oligonucleotides into a donor plasmid, wherein each donor plasmid optionally comprises, in sequential order, a first endonuclease site (C1), a first homologous recombination region (HR1), a second homologous recombination region (HR2), and optionally a second endonuclease site (C2); wherein each oligonucleotide is inserted between HR1 and HR2, thereby providing a plurality of donor plasmids comprising donor oligonucleotides, each donor plasmid comprising a single donor oligonucleotide from the mixture of oligonucleotides: C1-HR1-oligo-HR2-C2; (b) Transforming a large number of cells with the large number of donor plasmids so that each cell contains a donor plasmid, thereby forming a large number of donor cells;(c) Plating and cultivating the plurality of donor cells, each at a specific position on a first ordered array, thereby providing a first ordered array of donor cells; (d) providing a plurality of recipient cells in a second ordered array, each recipient cell comprising a recipient oligonucleotide comprising, in sequential order, a specific barcode sequence, wherein the specific barcode sequence identifies a position of the recipient cell in the second ordered array, a third homologous recombination region (HR3) homologous to HR1, optionally a third endonuclease site (C3), and a fourth homologous recombination region (HR4) homologous to HR2;(e) Contacting the first ordered array of donor cells with the second ordered array of recipient cells under conditions to (i) transfer the donor plasmids from the donor cells to the recipient cells at appropriate positions on the array by conjugation, (ii) optionally cleave the first, second and third endonuclease site, and (ii) transfer the oligonucleotides from the donor plasmids to the oligonucleotides of the recipient cells by homologous recombination, forming a third array of fusion oligonucleotides, each comprising a specific barcode sequence and a donor oligonucleotide from the oligonucleotide mixture;and (f) optional sequencing of the fusion oligonucleotides and thereby identifying each oligonucleotide in the array by its barcode sequence. The recipient oligonucleotide may be located in a recipient cell plasmid or in the recipient cell genome. The donor plasmid may include a selectable marker between HR1 and HR2 that selects for the integration of the oligonucleotide into the recipient cell oligonucleotide; optionally, the donor plasmid includes a counter-selectable marker. The recipient cell oligonucleotide may include a fourth endonuclease site (C4).
[0035] This document describes methods for identifying an oligonucleotide from a plurality of oligonucleotides, the methods comprising: (a) providing a plurality of donor cells in a first ordered array, each donor cell comprising a donor plasmid, each donor plasmid optionally comprising, in sequential order, a first endonuclease site (C1), a first homologous recombination region (HR1), a specific barcode sequence, a second homologous recombination region (HR2), and optionally a second endonuclease site (C2), the specific barcode sequence identifying a position of the host cell in the first ordered array; (b) providing a plurality of recipient cells, each recipient cell comprising a recipient plasmid comprising, in sequential order, an oligonucleotide from the plurality of oligonucleotides, a third homologous recombination region (HR3) homologous to HR1,optionally comprising a third endonuclease site (C3) and a fourth homologous recombination region (HR4) homologous to HR2; (c) plating and culturing the plurality of recipient cells, each at a specific position on a second ordered array, thereby providing a second ordered array of recipient cells; (d) contacting the first ordered array with the second ordered array under conditions to (i) transfer the donor plasmids from the donor cells to the recipient cells at appropriate positions on the array by bacterial conjugation, (ii) cleave the first, second, and third endonuclease sites, and (iii) transfer the barcode sequences from the donor plasmids to the oligonucleotides of the recipient cells by homologous recombination, thereby forming a third array of fusion oligonucleotides,each comprising a specific barcode sequence and an oligonucleotide from the oligonucleotide mixture; and (e) sequencing the fusion oligonucleotides, thereby identifying each oligonucleotide in the array by its barcode sequence. The recipient oligonucleotide may be located in a recipient cell plasmid or in the recipient cell genome. The donor plasmid may include a selectable marker between HR1 and HR2 that selects for the integration of the barcode sequence into the recipient cell oligonucleotide; optionally, the donor plasmid includes a counter-selectable marker. The recipient cell oligonucleotide may include a fourth endonuclease site (C4). The first, second, and third endonuclease sites may be the same or different. The donor plasmid may include a transfer origin and / or a conditional origin of replication.where the transfer origin is optionally derived from a mobile element and the conditional replication origin optionally depends on the presence of an oligonucleotide or a cell growth condition. The donor plasmid or the recipient plasmid may include a replicon capable of replicating plasmids of at least 30 kilobases in length, optionally derived from a P1-derived artificial chromosome or a bacterial artificial chromosome. The recipient cell donor plasmid or oligonucleotide may include an inducible high-copy replication origin. The recipient cell donor plasmid or oligonucleotide may be a yeast artificial chromosome (YAC), a mammalian artificial chromosome (MAC),The donor plasmid or recipient cell oligonucleotide may include an artificial human chromosome (HAC) or an artificial plant chromosome. The donor plasmid or recipient cell oligonucleotide may include a viral vector. The endonuclease sites may be cleaved by one or more endonucleases encoded by one or more oligonucleotides in the recipient cell and / or in the donor plasmid, the one or more endonucleases optionally being a homing endonuclease or an RNA-directed DNA endonuclease, and optionally being HO. The donor cell or recipient cell may include an oligonucleotide that (i) enables plasmid conjugation; (ii) encodes one or more homologous DNA repair genes; or (iii) encodes one or more recombination-mediated genetic engineering genes. The donor cells, recipient cells, or recombinant recipient cells can be placed on positions on a third ordered array,The barcode sequence can be transferred to a fourth ordered array or a subsequent ordered array. The donor cells and recipient cells can be independent bacterial cells, the bacterial cell being E. coli, Vibrio natriegens, or V. cholerae. The barcode sequence can be approximately four to 100 nucleotides long, optionally being approximately 30 nucleotides long. The oligonucleotide mixture can be a product of a DNA synthesis or assembly technology, selected from chemical coupling, template-independent enzymatic synthesis using polymerase nucleotide conjugates, polymerase cycling assembly, Gibson chew-back anneal repair, ligase cycling reaction, Phi29 polymerase, rolling circle, loop-mediated isothermal (LAMP), strand shift (SDA), helicase-dependent (HAD), or recombinase polymerase (RPA).Nucleic acid sequence-based amplification (NASBA), Golden Gate cloning, MoClo cloning, BioBricks or assembled BioBricks, thermodynamically balanced inside-out synthesis, DNA cloning, ligation-independent cloning, ligation by selection cloning, recombination, yeast composition, PCR, capture by molecular inversion probes or LASSO probes, DropSynth, and enzymatic DNA synthesis. The oligonucleotide mixture can be a product of a pooled mutagenesis technology selected from a polymerase chain reaction technology, including error-prone PCR, PCR with degenerate oligos and regular PCR, chemical or light mutagenesis, in vitro synthesis with a library of editing oligos, in vivo editing, etc. B. MAGE, MAGESTIC, CRISPR, prime editing, retron editing, and base modification with CRISPR, TALENs, and Zing-finger nucleases. The oligonucleotide mixture can contain at least one fragment of genomic DNA, cDNA,The oligonucleotide mixture may include organoDNA or natural plasmid DNA. It may also include captured or amplified DNA derived from gDNA, cDNA, or organelle DNA, such as from a balanced cDNA library, a PCR product like a multiplex PCR product, molecular inversion probes including LASSO probes, capture by annealing or subtractive hybridization, co-transformation and homologous recombination, rolling circle amplification, or LAMP. The oligonucleotide mixture may also include captured or amplified DNA from a plasmid or plasmid library, such as an open reading frame (ORF) library, a promoter library, a terminator library, an intron library, a BAC library, a PAC library, a lentiviral library, a gRNA library, or PCR products.Restriction digestion products or GATEWAY shuttle products. The oligonucleotides of the oligonucleotide mixture can be integrated into a donor plasmid by a process involving co-transformation and recombination, transformation and recombination, or conjugation and recombination. The co-transformation and recombination process can involve the construction of a linear or circular donor plasmid containing a selectable marker and two homologous recombination regions.which are homologous to sequences at the ends of the oligonucleotides in the mixture; co-transformation of the donor plasmid and the oligonucleotides in cells; induction of homologous recombination; and selection for the selectable marker; wherein the procedure is optionally performed with a library or pool of donor plasmids and / or oligonucleotides. The transformation and recombination procedure may include the construction of linear or circular donor plasmids comprising one selectable marker and two homologous recombination regions, each homologous to sequences at the ends of the oligonucleotides in the mixture.wherein the oligonucleotides are present on plasmids within host cells; the transformation of the host cells with the donor plasmids; the induction of homologous recombination; and the selection for the selectable marker. The conjugation and recombination procedure can involve the construction of linear or circular donor plasmids containing a counterselectable marker (-1) flanked by two optional endonuclease sites and two regions for homologous recombination (HR); wherein the donor plasmids are located in donor cells containing a crippled F plasmid that can induce but not conjugate conjugation, and the oligonucleotides of the mixture are located on plasmids in recipient cells, each flanked by HR regions homologous to the HR regions of the donor plasmids and adjacent to at least one selectable marker (+1),to select for the recombination of each oligonucleotide into a donor plasmid; providing a homologous recombinase and, optionally, one or more endonucleases, either in the recipient cells or encoded by the donor plasmids; contacting the donor and recipient cells under conditions to (i) transfer the donor plasmids from the donor cells to the recipient cells by bacterial conjugation and (ii) recombine the donor plasmids and the recipient plasmids by homologous recombination; and selecting for cells that include the selectable marker but not the counterselectable marker. The procedure can be performed with a library of donor and / or recipient plasmids. The oligonucleotides can form a library such as an ORF library, a promoter library, a terminator library, an intron library, a BAC library, a PAC library, a lentiviral library, a gRNA library,The oligonucleotide mixture may comprise a gDNA library, a cDNA library, a protein domain library, a promoter library, a terminator library, a regulatory element library, a structural element library, or a library of DNA variants derived from DNA mutagenesis. The oligonucleotide mixture may comprise arrays of cells containing plasmid libraries, such as a gRNA library, a gDNA library, a cDNA library, an ORF (Open Reading Frame) library, a protein domain library, a promoter library, a terminator library, a regulatory element library, a structural element library, or a library of DNA variants derived from DNA mutagenesis. The oligonucleotide mixture may also comprise arrays of cells containing DNA element fragments for use in the method according to any one of claims 1-35. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1. Schematic representation of an embodiment of the DNA assembly procedure described herein, in which donor plasmid and recipient oligonucleotide elements are shown as shaded boxes. The figure shows three “rounds” of “DNA stitching,” in each of which a new oligonucleotide comprising a DNA element fragment is added to the recipient oligonucleotide. In the figures, the DNA element fragments can be represented before recombination as “Input DNA 1,” “Input DNA 2,” etc., and after recombination as “DNA1,” “DNA2,” etc.; alternatively, they can also be referred to as “Oligo1,” “Oligo2,” etc., or more generally as “DNA blocks.” The boxes labeled C1, C2, etc., refer to endonuclease sites; the boxes labeled HR1, HR2, etc., refer to homologous recombination regions. those with Oligo1, Oligo2, etc.The boxes labeled with a number refer to an oligonucleotide comprising a DNA element fragment; the boxes labeled with a number and a plus or minus sign refer to selectable (+) and counterselectable (-) markers. Not all elements shown in the scheme are required in every embodiment of the methods described herein; for example, the markers and endonuclease sites labeled C2.1, C2.2, C7.1, and C7.2 are optional. Fig. 2A-B. Box A is a schematic map of an exemplary recipient oligonucleotide (in the form of a plasmid) used in the procedures described here. Box B shows two schematic maps of exemplary donor plasmids and helper plasmids. Fig. 3. CRISPR / Cas9 improves DNA assembly, also referred to as "stitching." The number of colonies per 6 × 10 6Cells on the selection plates. Two donor plasmids, only one of which expresses a functional gRNA, were transformed into each of the three recipient strains: (1) BW28705 (no λ-red and no Cas9), (2) BW28705 / pML300 (λ-red, but no Cas9), (3) BW28705 / pSL359 (Cas9 and λ-red). Fig. 4. Schematic representations of exemplary plasmids that can be used in the in vivo DNA assembly or "stitching" procedure described here. In donor cells, a conjugation-competent helper plasmid can contain the genes for plasmid transfer (tra operon). To immobilize the conjugation plasmid itself, the origin of transfer (oriT) is replaced by a selectable marker (+6). A donor plasmid contains a swapping cassette (+1 / -1 or +2 / -2), two homology regions (H2 and H3), four endonuclease cleavage sites (two circles labeled 1 and two circles labeled 2), a selectable scaffold marker (+4), a conditional origin of replication (R6K) that depends on an allele in the donor genome (pir1-116), the oriT sequence, and a gRNA expression cassette (gRNA1 or gRNA2). Fig. 5. Schematic representation of exemplary plasmids that can be used in the in vivo stitching procedures described here. In the recipient cells, the helper plasmid contains a rhamnose-inducible red operon ( PrhaBAD-rot ), an arabinose-inducible Cas9 ( ParaBAD-Cas9 ), an E. coli RecA gene to enhance homologous recombination, a selectable backbone marker (+5) and a curable origin of replication (pSC101 ori TS The recipient plasmid contains two endonuclease cleavage sites (two circles labeled 1), one swapping cassette (+2 / -2), two homology regions (H2 and H3), and one origin of replication (ColE1). +1: HygR; +2: NsrR; -1: SacB; -2: PheS; +3: GmR; +4: KanR; +5: SpR; +6: TcR. Fig. 6. Schematic overview of an exemplary in vivo stitching method. Donor plasmids carrying a DNA fragment (rectangles striped upwards or downwards) are introduced into donor plasmids and donor cells. Donor plasmids are conjugated with recipient cells, and a DNA fragment is transferred from the donor plasmid to the recipient plasmid. The plasmids are cut using CRISPR / Cas9 induced by arabinose. A guide RNA on the donor plasmid (gRNA1 or gRNA2, alternating between assembly rounds) provides a recognition sequence for the cut (circles "1" and "2", alternating between assembly rounds). Homology regions on both the synthesized oligos and the plasmid backbones (H1 and H3 in round 1) promote the recombination induced by rhamnose and seamlessly join the oligos together for gene assembly.Alternating selectable (+1 and +2) and counter-selectable (-1 and -2) markers on donor plasmids enable recursive DNA transfers with a maximum gene length, theoretically determined by the maximum tolerable plasmid size. R6K and ColE1 are origins of replication. +3 and +4 are selectable markers used for plasmid maintenance. Fig. 7A-I. Examples of DNA assembly. Field A shows three donor plasmids, each carrying a fragment of mEGFP, which were sequentially conjugated and stitched together to form a recipient plasmid (3 stitches). Field B shows the fluorescence of colonies from a negative control, a positive control, and the in vivo stitching products after three rounds of assembly in liquid. The colonies represent independent conjugation and recombination events and are 100% fluorescent. Field C shows arranged assemblies of mEGFP in 96- and 384-position formats. All colonies appear fluorescent. Field D shows the percentage of fluorescent colonies after a final round of liquid assembly using a third mEGFP fragment with a different homology length to the second mEGFP fragment.Field E shows representative restriction digests of colonies with different plasmids scraped from agar during a composition. The expected product from the non-recombinant recipient plasmid cannot be observed after selection for recombinants or curing of the helper plasmid (arrow points). Field F shows a scheme of the analysis of Sanger sequencing results from 96 colonies after mEGFP composition. The sequencing products were obtained from a colony PCR of a pipette tip applied to each colony. One colony contained an intermediate product (the product of the first composition round), and one contained a stitching error (a large deletion). Field G shows the fluorescence of colonies from the in vivo stitching products after five composition rounds for two fluorescent genes, mPapaya and sfGFP, and four recombinase genes.The colonies can represent independent conjugation and recombination events. All mPapaya and sfGFP colonies are fluorescent. Panel H is the trace file of the Sanger sequencing of the mPapaya in vivo stitching products after five rounds of composition. Alignment to the expected sequence shows that the composition is 100% accurate and pure. Panel I shows the results of a composition of three ~3 kb fragments, corresponding to a total length of ~9 kb. The recipient plasmids were digested at various stages of composition using restriction enzymes to separate the stitching products from the vector backbone. The digested products were then subjected to agarose gel electrophoresis to verify the size of the stitching products (lanes 1–3). The gel bands corresponding to the stitching products are marked with an arrow. Linearized vector backbones without the stitching products are shown in lanes 5–6.Selectable and counter-selectable markers in the swapping cassette differ between the composition rounds, with the swapping cassette being ~1.5kb longer in the first and third composition rounds than the swapping cassette in the original recipient plasmid or in the second composition round. Fig. 8A-B. Schematic representation of exemplary donor and recipient plasmids at the beginning of the first round of DNA stitching. Box A shows a scheme for both example plasmids, with the donor plasmid containing the first oligonucleotide (1). Box B shows the legend of the shapes used to illustrate the sequences corresponding to the example genome, the positively selectable marker, the negatively selectable marker, the origin of transfer (oriT), the gRNA expression unit (gRNA), and the position barcode; homology for recombination domain (H), inducible lambda red operon (λrot), inducible I-SceI endonuclease, plasmid, inducible endonuclease (Cas9), gRNA target sites, I-SceI target sites, conjugation tra operon, deleted oriT (oriΔ::TcR), temperature-sensitive origin (pSC101 ori), conditional origin of replication (R6K), and recipient of origin of replication (ColE1). The same legend of shapes in Fig. 8B will be used for the Fig. used. Fig. 9. Schematic representation of exemplary starting plasmids for use in the procedures described here: donor plasmid containing the first oligo (1), recipient plasmid and a plasmid containing oriT, which mediates the conjugation of the donor plasmid with the recipient cell. Fig. 10. Diagram of a subsequent step in an exemplary DNA stitching procedure using the method described in Fig. 9 donor and recipient plasmids shown: gRNA1 directs Cas9 in the recipient cells to generate site-specific double-strand breaks on the donor and recipient plasmids (indicated by downward-pointing arrows). Fig. 11. Schematic representation of a subsequent step in the in Fig. 9-10 illustrates the exemplary method of DNA stitching: The dark shaded sequence elements are used here as the homology region for homologous recombination mediated by lambda red. Fig. 12. Schematic representation of a subsequent step in the Fig. 9-11 illustrated exemplary method of DNA stitching: Homologous recombination between the plasmids, which shows where the sequence from the donor plasmid is inserted into the recipient plasmid and how it is oriented, using a λ-red system. Fig. 13. Schematic representation of a subsequent step in the process described in Fig. 9-12 Exemplary DNA stitching procedures shown: The fragment containing the first oligonucleotide is integrated into the recipient plasmid as shown. Fig. 14. Schematic representation of a subsequent step in the process described in Fig. 9-13 illustrative examples of DNA stitching procedures: The plasmid is selected to obtain the +2 positive selectable marker. Fig. 15. Schematic representation of a subsequent step in the process described in Fig. 9-14 illustrative examples of DNA stitching procedures: The plasmid is selected against the loss of the previous counter-selectable marker. Fig. 16. Schematic representation of a subsequent step in the process described in Fig. 9-15 illustrative examples of DNA stitching procedures: The plasmid is additionally selected to obtain the +3 positive selectable marker on the original recipient scaffold. Fig. 17. Schematic representation of a subsequent step in the process described in Fig. 9-16 illustrative examples of DNA stitching procedures: The second donor plasmid, containing the second oligonucleotide, is ready to be incorporated into the previous ligation product (the new recipient plasmid), which contains the first oligonucleotide. Fig. 18. Schematic representation of a subsequent step in the process described in Fig. 9-17 illustrative examples of DNA stitching procedures: The oriT directs the conjugation of the donor plasmid. Fig. 19. Schematic representation of a subsequent step in the process described in Fig. 9-18 illustrates an exemplary DNA stitching procedure: A second gRNA expression directs Cas9 to create double-strand breaks at the locations indicated by the downward-pointing arrows. Fig. 20. Schematic representation of a subsequent step in the in Fig. 9-19 shows an exemplary method of DNA stitching: The highlighted areas (darker shaded areas) are the homology regions for recombination. Fig. 21. Schematic representation of a subsequent step in the process described in Fig. 9-20 illustrative examples of DNA stitching procedures: The fragment containing the first oligonucleotide is integrated into the recipient plasmid as shown. Fig. 22. Schematic representation of a subsequent step in the process described in Fig. 9-21 shows an exemplary procedure for DNA stitching: The second oligonucleotide is incorporated into the recipient plasmid adjacent to the 3' end of the first oligo, creating a new recipient plasmid. Fig. 23. Schematic representation of a subsequent step in the process described in Fig. 9-22 Exemplary DNA stitching procedures shown: The plasmid containing the first and second oligonucleotides is selected to obtain +1 positive selectable marker. Fig. 24. Schematic representation of a subsequent step in the process described in Fig. 9-23 illustrated exemplary procedure of DNA stitching: The plasmid containing the first and second oligonucleotides is selected for the loss of -2 counter-selectable markers. Fig. 25. Schematic representation of a subsequent step in the process described in Fig. 9-24 illustrated exemplary procedures of DNA stitching: The plasmid containing the first and second oligonucleotides is also selected to retain the selectable marker of the scaffold. Fig. 26. Schematic representation of a subsequent step in the process described in Fig. 9-25 illustrated exemplary procedures of DNA stitching: The scheme for a donor plasmid containing oligonucleotide three and the recipient plasmid with oligonucleotides one and two to initiate the third round of DNA stitching. Fig. 27. Schematic representation of a subsequent step in the process described in Fig. 9-26 shows an exemplary procedure for DNA stitching: The oriT plasmid initiates conjugation. Fig. 28. Schematic representation of a subsequent step in the in Fig. 9-27 shows an exemplary method of DNA stitching: gRNA1 directs Cas9 in the recipient cells to generate site-specific double-strand breaks on the donor and recipient plasmids (indicated by downward-pointing arrows). Fig. 29. Schematic representation of a subsequent step in the process described in Fig. 9-28 Exemplary DNA stitching procedures shown: The sequences highlighted here are used as homology regions for lambda-red mediated homologous recombination. Fig. 30. Schematic representation of a subsequent step in the process described in Fig. 9-29 illustrates exemplary DNA stitching procedures: Posthomologous recombination, which shows where and in what orientation the sequence from the donor plasmid is inserted into the recipient plasmid. Fig. 31. Schematic representation of a subsequent step in the in Fig. 9-30 illustrates an exemplary procedure of DNA stitching: The third oligonucleotide is incorporated into the recipient plasmid adjacent to the 3' end of the second oligonucleotide, creating a new recipient plasmid. Fig. 32. Schematic representation of a subsequent step in the process described in Fig. 9-31 illustrated exemplary procedure of DNA stitching: Analogous to round 1 of DNA stitching, the plasmid is selected to obtain the +2 positive selectable marker. Fig. 33. Schematic representation of a subsequent step in the process described in Fig. 9-32 illustrated exemplary procedures of DNA stitching: The plasmid is selected against the loss of the counter-selectable marker -1. Fig. 34. Schematic representation of a subsequent step in the process described in Fig. 9-33 illustrated exemplary procedures of DNA stitching: The plasmid is selected so that it retains the selectable marker of the scaffold. Fig. 35A-B. Field A shows one in the Fig. Figure 9-34 illustrates a step in which subsequent oligonucleotides (up to reaching the upper limit for the total sequence length) can be incorporated in the same manner, alternating between the two processes of conjugation, double-strand break processing, assembly, and selection / counterselection / backbone selection. Box B is a schematic overview of an exemplary in vivo assembly in which two DNA element fragments (represented in the figure as input DNA 1, input DNA 2 before recombination, and DNA 1, DNA 2 after recombination; and which may also be referred to here as oligo 1, oligo 2, etc., or more generally as "DNA blocks") are added in a single conjugation round. In each well, a recipient cell is first conjugated with a first donor cell containing a first donor plasmid, and subsequently with a second donor cell containing a second donor plasmid.A selectable marker introduced by the second donor plasmid and counter-selectable markers on the recipient plasmid and, if applicable, on the first donor plasmid, enable the selection of a recombinant composite product containing oligonucleotides introduced by both donor plasmids. Fig. 36A-D. Box A is a schematic overview of an exemplary procedure for in vivo DNA analysis, as described herein. In this example, the indexing barcodes are located on the recipient plasmid. Box B is a schematic overview of another exemplary procedure for in vivo DNA analysis, as described herein. In this example, the indexing barcodes are located on the donor plasmid and are added to an in vivo DNA assembly product using a homologous region at the end of the assembly product. In this example, the endonuclease sites used are the same as in the in vivo DNA assembly. Box C is a schematic overview of another exemplary procedure for in vivo DNA analysis, as described herein. In this example, the indexing barcodes are located on the donor plasmid and are added to an in vivo DNA assembly product.In this example, the endonuclease target sites (C) and homologous regions (boxes next to C) differ from those used for in vivo DNA assembly. This example allows for multi-step DNA analysis during an assembly. Box D is a schematic overview of the procedure, which includes a plasmid reset, in which a DNA assembly is transferred from a recipient plasmid to a donor plasmid to enable further assembly rounds with larger DNA blocks. In part A of box D, a reset donor plasmid is conjugated into the recipient cell in a donor cell with homology to the start of the DNA assembly on the recipient plasmid. A site-specific endonuclease cleaves at "D" endonuclease target sites in both the reset donor plasmid and the recipient plasmid.Homologous recombination in regions adjacent to the "D" endonuclease target sites shifts the DNA assembly cassette from the recipient plasmid to the reset donor plasmid. The reset donor plasmid is purified from the recipient cell and transformed into new donor cells, where it can be used for further assembly rounds. In part B of panel D, a diagram illustrates a workflow for assembling long DNA constructs. Small DNA blocks can be stitched together to form larger DNA blocks by four rounds of DNA stitching. The large blocks are transferred to donor plasmids and donor cells using reset donor plasmids. The large blocks can then be stitched together to form even larger blocks. Fig. Figure 37 shows schematic maps of exemplary plasmids for use in in vivo DNA analysis. In donor cells, a conjugation-competent helper plasmid contains the genes for plasmid transfer (tra operon). To immobilize the helper plasmid itself, the origin of transfer (oriT) is replaced by a selectable marker (+6). The donor plasmid contains a swapping cassette (+ and -), two homology regions (H1 and H4), two sites for targeted plasmid cleavage (ovals), a selectable scaffold marker (+4), a conditional origin of replication (R6K) dependent on an allele in the donor genome (pir1-116), and the oriT sequence. In recipient cells, the helper plasmid contains a lactin-inducible red operon (P). lac -red), an E. coli RecA gene to enhance homologous recombination, a selectable backbone marker (+5) and a curable, temperature-sensitive origin of replication (pSC101 ori TSThe recipient plasmid contains two endonuclease cleavage sites (two ovals), a negative selectable marker (-3), and two homology regions (H1 and H4). In addition to the two plasmids, the recipient cells also possess an integrated arabinose-inducible endonuclease I-SceI ( ParaBAD-I-SceI ) for generating DNA slices on the target plasmids. +: HygR or NsrR; -: SacB or PheS; +3: GmR; -3: relE; +4: KanR; +5: SpR; +6: TcR. Fig. Figure 38 shows schematic maps of exemplary plasmids for use in in vivo DNA analysis. Selectable markers: HygR, KanR, GmR, SpR. Counter-selectable markers: SacB, relE. Fig. Figure 39A-B shows schematic representations of an exemplary donor plasmid and a recipient plasmid used for DNA parsing. Box A shows the plasmid diagrams where the donor plasmid contains the first oligonucleotide. Box B shows the legend of the shapes used to illustrate the sequences corresponding to a genome, a positively selectable marker, a negatively selectable marker, the origin of transfer (oriT), the gRNA expression unit (gRNA), and the position barcode; homology for recombination domain (H); inducible lambda red operon (λrot); inducible I-SceI endonuclease; plasmid; inducible endonuclease (Cas9); gRNA target sites; I-SceI target sites; conjugation tra operon; deleted oriT (oriΔ::TcR); temperature-sensitive origin (pSC101 ori); conditional origin of replication (R6K); and recipient of the origin of replication (ColE1). The legend in Box B also applies to the Fig. 40-46. Fig. Figure 40 shows a diagram of the donor and recipient plasmids for use in an example procedure described here. Fig. 41 is an image that represents a second step in the example method of Fig. Figure 40 shows: gRNA1 directs Cas9 in the recipient cells to create site-specific double-strand breaks on the donor and recipient plasmids (indicated by downward-pointing arrows). SceI is I-SceI, a homing endonuclease. Fig. Figure 42 shows an image of a step in the example method of the Fig. 40 and Fig. 41: The H1 and H4 sequences are used here as the homology region for lambda-red mediated homologous recombination. Fig. 43 is an image that represents a step in the example procedure of Fig. 40-42: homologous recombination, which shows where and in what orientation the sequence from the donor plasmid is inserted into the recipient plasmid. Fig. 44 is an image that represents a step in the example procedure of Fig. 40-43: The plasmid is selected to obtain the + positive selectable marker. Fig. Figure 45 shows a picture of a step in the example procedure of the Fig. 40-44: The plasmid is counterselected for the loss of the previous counterselectable marker. Fig. Figure 46 shows a step in the FIGs' example procedure. Figures 40-45: The plasmid is further selected to obtain the +3-positive selectable marker on the original recipient backbone. Fig. 47A-D. Panel A shows the results of an experiment to determine the ability of in vivo DNA analysis to correctly identify the sequence at each position of a plate containing arrayed donor cells, each containing a unique DNA barcode. Each arrayed barcode donor was paired with two or three barcode recipient plates, and recombinant cell colonies, each containing one donor and one recipient barcode, were selected on agar pads. The recombinant cells from the plates were pooled, and the duplicate barcodes were sequenced on an Illumina platform. The sequencing data were used to determine the percentage of the barcode donor array that could be correctly indexed (recovery rate) when conjugated with one, two, or three separate barcode recipient arrays. No barcode donor was incorrectly assigned to the wrong position.Panel B shows the results of an experiment to index and sequence a pool of 100 244-base oligonucleotides ordered from IDT as an oPool. The oligonucleotide pool was integrated into donor plasmids, which were then transformed into donor cells. The donor cells were randomly arranged in 384-well plates with an expected density of less than one cell per well. The donor cells were conjugated with barcoded recipient cell arrays, and the recombinant oligonucleotide-barcoded recipient plasmids were sequenced using an Oxford nanopore sequencer. The results of the analysis of two 384-well plates are shown. The shading indicates whether the input DNA was arranged in the well, whether the sequence matched 100% with one of the 244-base sequences in the oPool, and whether the well was clean (i.e., free of any DNA).(only a 244-base sequence could be detected in the well). The positions marked as "100% matching pure well" are normally used for subsequent DNA reconstruction. Box C is a histogram showing the distribution of errors between the consensus sequence determined by Oxford Nanopore sequencing and the expected DNA sequence in the oPool that is closest to the consensus sequence, using the data from the experiment in Box B. Most wells contain an oligonucleotide identical to one of the sequences in the oPool. Box D is a histogram showing the distribution of the number of independent clones obtained for each oligonucleotide that could be indexed, using the data from the experiment in Box B. Fig. Figure 48 is a schematic overview of a DNA assembly workflow for constructing directed combinatorial libraries from a set of input oligonucleotides. Pools of input DNA from multiple sources are integrated into donor plasmids and disassembled into ordered arrays. The ordered arrays are rearranged at user-defined locations on multiple donor plates. The donor plates are sequentially conjugated to a recipient plate to assemble the desired constructs. An input oligonucleotide can be used in multiple compositions by rearranging donor cells containing that oligonucleotide at multiple positions on the donor plates. Fig. Figure 49 is a schematic overview of branched DNA structure. A partial DNA assembly can be extended with multiple DNA blocks if homology regions are present. If no homology regions are present, a "DNA linker" must first be added to the partial DNA assembly. The DNA linker contains the homology to the end of the partial DNA assembly and to the beginning of the subsequent DNA block to be joined. DETAILED DESCRIPTION I. Definitions and related embodiments
[0036] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by a person skilled in the art in the field of the present invention. The following references provide a general definition to the person skilled in the art of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker Ed., 1988); The Glossary of Genetics, 5th ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Unless otherwise stated, the following terms have their commonly assigned meanings.
[0037] The use of an indefinite or definite singular article (e.g., "a," "a," "the," etc.) in this disclosure and in the following claims follows the traditional approach in patents, which means "at least one," unless in a particular case it is clear from the context that the term specifically means one and only one. Likewise, the term "comprehensive" is open and does not exclude additional subject matter, features, components, etc.
[0038] "Optional" means that the event or circumstance described below may or may not occur, and that the description includes cases in which the event or circumstance occurs and cases in which it does not.
[0039] The term "approximately", used before a numerical designation, e.g., temperature, time, quantity, concentration, etc., including a range, denotes approximate values that may vary by (+) or (-) 10%, 5%, 1%, or any sub-range or partial value in between. Preferably, the term "approximately" means that the value may vary by + / - 10%.
[0040] As used herein, the term “comprehensive” or “includes” means that the compositions and processes contain the elements mentioned but do not exclude others. “Consists substantially of” in the definition of compositions and processes means that other elements essential to the combination for the stated purpose are excluded. Thus, a composition consisting substantially of the elements defined herein would not exclude other materials or steps that do not substantially affect the basic and novel feature(s) of the claimed invention. “Consisting of” means that more than trace elements of other components and essential process steps are excluded. Embodiments defined by any of these transitional terms fall within the scope of this disclosure.
[0041] The terms "nucleic acid," "nucleic acid molecule," "nucleic acid sequence," and "polynucleotide" are used interchangeably here and are intended to encompass, but are not limited to, a polymeric form of covalently linked nucleotides of different lengths, either deoxyribonucleotides or ribonucleotides, or their analogues, derivatives, or modifications. Different polynucleotides can have different three-dimensional structures and perform various known or unknown functions.Non-restrictive examples of polynucleotides include a gene, a gene fragment, an exon, an intron, intergenic DNA (including, without limitation, heterochromatic DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, a ribozyme, cDNA, sgRNA, guide RNA, tracrRNA, a recombinant polynucleotide, a branched polynucleotide, a plasmid, a vector, isolated DNA containing a sequence, isolated RNA containing a sequence, a PCR product, a nucleic acid probe, and a primer. Polynucleotides useful in the procedures of disclosure may include natural nucleic acid sequences and variants thereof, artificial nucleic acid sequences, or a combination of such sequences.
[0042] The term "nucleic acid" refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and their polymers in single-, double-, or multiple-stranded form, or their complements; or to nucleosides (e.g., deoxyribonucleosides or ribonucleosides). In some embodiments, the term "nucleic acid" does not include nucleosides. The terms "polynucleotide," "oligonucleotide," "oligo," or similar terms, in their usual sense, refer to a linear sequence of nucleotides. The term "nucleoside," in its usual and common sense, refers to a glycosylamine containing a nucleobase and a five-carbon sugar (ribose or deoxyribose). Unrestricted examples of nucleosides are cytidine, uridine, adenosine, guanosine, thymidine, and inosine. The term "nucleotide" refers in the usual and customary sense to a single unit of a polynucleotide, i.e., a monomer.Nucleotides can be ribonucleotides, deoxyribonucleotides, or modified versions thereof. Examples of polynucleotides include single- and double-stranded DNA, single- and double-stranded RNA, and hybrid molecules with mixtures of single- and double-stranded DNA and RNA. Examples of nucleic acids, such as polynucleotides, considered here include all types of RNA, such as mRNA, siRNA, miRNA, and guide RNA, and all types of DNA, including genomic DNA, plasmid DNA, and minicircle DNA, as well as all fragments thereof. The term "duplex" in the context of polynucleotides refers, in the usual and common sense, to double stranding. Nucleic acids can be linear or branched. For example, nucleic acids can be a linear chain of nucleotides, or they can be branched, such that they comprise one or more arms or branches of nucleotides.Optionally, the branched nucleic acids are repeatedly branched to form higher-order structures such as dendrimers and the like.
[0043] Nucleic acids, including, for example, nucleic acids with a phosphothioate backbone, can contain one or more reactive units. As used here, the term reactive moiety encompasses any group capable of reacting with another molecule, such as a nucleic acid or a polypeptide, through covalent, non-covalent, or other interactions. For example, the nucleic acid may contain a reactive amino acid group that reacts with an amino acid on a protein or polypeptide through a covalent, non-covalent, or other interaction.
[0044] The terms also include nucleic acids, including known nucleotide analogs or modified backbone residues or linkages, whether synthetic, naturally occurring, or not, that have similar binding properties to the reference nucleic acid and are metabolized similarly to the reference nucleotides. Examples of such analogs include phosphodiester derivatives, such as... B, Phosphoramidate, Phosphorodiamidate, Phosphorothioate (also known as phosphothioate with double-bonded sulfur replacing the oxygen in the phosphate), Phosphorodithioate, Phosphonocarboxylic acids, Phosphonocarboxylates, Phosphonoacetic acid, Phosphonoformic acid, Methylphosphonate, Borphosphonate or O-methylphosphoroamidite bonds (see Eckstein, OLIGONUCLEOTIDES AND ANALOGUES: A PRACTICAL APPROACH, Oxford University Press) as well as modifications to the nucleotide bases such as 5-methylcytidine or pseudouridine.; and peptide nucleic acid backbones and compounds. Other analogous nucleic acids include those with positive backbones, non-ionic backbones, modified sugars, and non-ribose backbones (e.g., phosphorodamidate morpholino oligos or locked nucleic acids (LNAs) as known in the art), including those described in U.S. Patent Nos. 5,235,033 and 5,034,506, and Chapters 6 and 7, ASC Symposium Series 580, CARBOHYDRATE MODIFICATIONS IN ANTISENSE RESEARCH, Sanghui & Cook, eds. Nucleic acids containing one or more carbocyclic sugars also fall under a definition of nucleic acids. Modifications of the ribose-phosphate backbone may be made for various reasons, e.g. For example, to increase the stability and half-life of such molecules in a physiological environment or as probes on a biochip.Mixtures of naturally occurring nucleic acids and analogs can be prepared; alternatively, mixtures of different nucleic acid analogs and mixtures of naturally occurring nucleic acids and analogs can also be prepared. In some embodiments, the internucleotide bonds in the DNA are phosphodiesters, phosphodiester derivatives, or a combination of both.
[0045] A "barcode" refers to one or more nucleotide sequences used to identify a cell or a multitude of cells to which the barcode is associated. Barcodes can be 3-1000 or more nucleotides long, preferably 3-250 nucleotides and more preferably 4-40 nucleotides, including any length within these ranges, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 nucleotides long. A barcode is "specific" if the barcode is (statistically speaking) present in approximately one cell within a cell population. The cell containing the barcode can then be expanded to form a clonal plurality of cells, such that every cell in the plurality of cells contains the same barcode.For example, "a multitude of barcoded cells, each barcoded cell containing a single, unique barcode" can refer to a cell population that (statistically speaking) contains a single cell with a specific barcode or a unique combination of barcodes. Alternatively, it can refer to a cell population that contains a multitude of clonal cell populations, where each cell in each clonal population has the same barcode, but the cells in different clonal populations have different barcodes.
[0046] As used here, the term "complement" refers to a nucleotide (e.g., RNA or DNA) or nucleotide sequence capable of base pairing with a complementary nucleotide or sequence. As described here and generally known in the scientific community, the complementary (matching) nucleotide of adenosine is thymidine, and the complementary (matching) nucleotide of guanosine is cytosine. Thus, a complement can comprise a sequence of nucleotides that form a base pair with corresponding complementary nucleotides of a second nucleic acid sequence. The nucleotides of a complement may partially or completely match the nucleotides of the second nucleic acid sequence. If the nucleotides of the complement completely match every nucleotide of the second nucleic acid sequence, the complement forms base pairs with every nucleotide of the second nucleic acid sequence.If the nucleotides of the complement partially match the nucleotides of the second nucleic acid sequence, only some of the nucleotides of the complement form base pairs with nucleotides of the second nucleic acid sequence.
[0047] As described here, the complementarity of sequences can be partial, i.e., only some of the nucleic acids agree in base pairing, or complete, i.e., all nucleic acids agree in base pairing. Thus, two complementary sequences can have a certain percentage of nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity over a certain range).
[0048] The term "gene" is used here in its usual sense and refers to the DNA segment involved in the production of a protein; it includes regions before and after the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). The leader, trailer, and introns contain regulatory elements necessary during the transcription and translation of a gene. Furthermore, a "protein gene product" is a protein expressed by a specific gene.
[0049] The term "expression vector" refers to a nucleic acid molecule that codes for genes and / or regulatory elements required for gene expression. The expression of a gene from a vector, which can be in the form of a plasmid, can occur in cis or trans configuration. When a gene is expressed in cis, the gene and the regulatory elements are encoded by the same plasmid. Trans expression refers to the case where the gene and the regulatory elements are encoded by different plasmids.
[0050] The term "vector" as used here refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is bound. A vector can take the form of a "plasmid," which in this context refers to a linear or circular double-stranded DNA loop into which additional DNA segments can be ligated. Another type of vector is a viral vector, into which additional DNA segments can be ligated into the viral genome. Certain vectors are capable of autonomous replication within a host cell into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) become integrated into the host cell's genome upon introduction and thus replicate along with the host genome.Furthermore, certain vectors are capable of controlling the expression of genes to which they are surgically linked. Such vectors are referred to here as “expression vectors.” In general, expression vectors useful in recombinant DNA techniques often take the form of plasmids. In this description, the terms “plasmid” and “vector” may be used interchangeably, since the plasmid is the most commonly used form of vector. However, the invention is also intended to encompass other forms of expression vectors, such as viral vectors (e.g., replication-defective retroviruses, adenoviruses, and adeno-associated viruses), which perform equivalent functions. Moreover, some viral vectors are capable of targeting a specific cell type, either specifically or non-specifically.Replication-incompetent viral vectors or replication-defective viral vectors refer to viral vectors that are able to infect their target cells and transmit their viral payload, but then do not proceed along the typical lytic pathway that leads to cell lysis and death.
[0051] In accordance with the methods described herein, an oligonucleotide, plasmid, or vector may contain at least one selectable marker. Selectable markers for use in the methods described herein may be any suitable selectable marker. In some embodiments, and without limitation, the selectable marker is HygR, NsrR, ZeoR, TetA, CmR, SpR, GmR, mFabI, TmR, neoR, or kanR. In some embodiments, the selectable marker is HygR. In some embodiments, the selectable marker is NsrR. In some embodiments, the selectable marker is ZeoR. In some embodiments, the selectable marker is TetA. In some embodiments, the selectable marker is CmR. In some embodiments, the selectable marker is SpR. In some embodiments, the selectable marker is GmR. In some embodiments, the selectable marker is mFabI. In some embodiments, the selectable marker is TmR.In some embodiments, the selectable marker is neoR. In some embodiments, the selectable marker is kanR.
[0052] In accordance with the methods described herein, an oligonucleotide, plasmid, or vector may contain at least one cross-selectable marker, such as one that selects for the integration of the second or subsequent oligonucleotide into a recombinant recipient oligonucleotide in the methods described herein for the assembly of a DNA element. Cross-selectable markers for use in the methods described herein may be any suitable cross-selectable marker. In some embodiments, and without limitation, the cross-selectable marker is PheS, SacB, rpsL, tolC, galK, ccdB, tetA, thyA, lacY, gata-1, URA3, relE, mqsR, chpB, vhaV, or tse2. In some embodiments, the cross-selectable marker is PheS. In some embodiments, the cross-selectable marker is SacB. In some embodiments, the cross-selectable marker is rpsL. In some embodiments, the counter-selectable marker is tolC.In some embodiments, the counter-selectable marker is galK. In other embodiments, the counter-selectable marker is ccdB. In other embodiments, the counter-selectable marker is ccdB. In other embodiments, the counter-selectable marker is tetA. In other embodiments, the counter-selectable marker is thyA. In other embodiments, the counter-selectable marker is lacY. In other embodiments, the counter-selectable marker is gata-1. In some embodiments, the counter-selectable marker is URA3. In some embodiments, the counter-selectable marker is relE. In other embodiments, the counter-selectable marker is mqsR. In other embodiments, the counter-selectable marker is chpB. In other embodiments, the counter-selectable marker is vhaV. In some embodiments, the counter-selectable marker is tse2.
[0053] The terms "transfection," "transduction," or "transfection" can be used interchangeably and are defined as a process for introducing a nucleic acid molecule and / or a protein into a cell. Nucleic acids can be introduced into a cell using non-viral or viral methods. The nucleic acid molecule may be a sequence that codes for complete proteins or functional parts thereof. It is typically a nucleic acid vector containing the elements required for protein expression (e.g., a promoter, a transcription start site, etc.). Non-viral transfection methods include any suitable methods that do not use viral DNA or viral particles as a delivery system to introduce the nucleic acid molecule into the cell.Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, heat shock transfection, magnetic sectioning, and electroporation. For viral methods, any suitable viral vector can be used for the procedures described here. Examples of viral vectors include retroviral, adenoviral, lentiviral, and adeno-associated viral vectors. In some cases, nucleic acid molecules are introduced into a cell using a retroviral vector and standard procedures known in the field. The terms "transfection" and "transduction" also refer to the introduction of proteins from the external environment into a cell.Typically, protein transduction or transfection relies on the binding of a peptide or protein capable of crossing the cell membrane to the protein of interest. See, for example, Ford et al. (2001) Gene Therapy 8:1-4 and Prochiantz (2007) Nat. Verfahren 4:119-20.
[0054] The term "promoter" used here refers to a region of DNA that initiates the transcription of a specific gene. Promoters are typically located near the transcription start site of a gene, upstream of the gene, and on the same strand (i.e., 5' on the sensory strand) of DNA. Promoters can be, for example, about 100 to 1000 base pairs long.
[0055] A nucleotide base “position” is designated by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the 5' end. Due to deletions, insertions, truncations, fusions, and the like, which must be considered when determining optimal alignment, the number of amino acid residues in a test sequence, determined by simply counting from the 5' end, generally does not necessarily correspond to the number of the corresponding position in the reference sequence. For example, if a variant has a deletion compared to an aligned reference sequence, there is no nucleotide base in the variant that corresponds to a position in the reference sequence at the location of the deletion. If an insertion is present in an aligned reference sequence, this insertion does not correspond to a numbered nucleotide position in the reference sequence.In the case of shortening or fusion, nucleotide segments may be present in the reference sequence or the adapted sequence that do not correspond to any nucleotide in the corresponding sequence.
[0056] The terms “numbered with reference to” or “corresponding to”, when used in connection with the numbering of a particular polynucleotide sequence, refer to the numbering of the residues of a particular reference sequence when the particular polynucleotide sequence is compared with the reference sequence.
[0057] The term "virus" or "virus particle" is used here in its usual sense in the context of viral transduction. Transduction with viral vectors can be used to insert or modify genes in mammalian cells.
[0058] The terms "genetic modification," "gene editing," "gene editing," "genome editing," "genome engineering," or similar terms used here refer to a type of genetic engineering in which DNA is inserted, removed, altered, or replaced at one or more specific locations in the genome of a cell. A key step in gene editing is the creation of a double-strand break at a specific location within a gene or genome. Examples of gene-editing tools such as nucleases that perform this step include zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, and the CRISPR / Cas system (clustered regularly interspaced short palindromic repeats).
[0059] The term "DNA element" as used here refers to any DNA sequence that can be transferred between cells, for example, between a donor cell and a recipient cell. A DNA element thus includes, but is not limited to, a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, or a gRNA. A DNA element can be a fragment of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, or a gRNA. A DNA element can be a combination of genes, promoters, enhancers, terminators, introns, intergenic regions, barcodes, gRNAs, and fragments of genes, promoters, enhancers, terminators, introns, intergenic regions, barcodes, and gRNAs. In some embodiments, the DNA element is located in a donor plasmid. In other cases, the DNA element is translocated to, or located in, a recipient oligonucleotide.In other cases, the DNA element is transferred from a recipient oligonucleotide to a reset donor plasmid.
[0060] The term "gene editing reagent" as used here refers to components required for gene editing tools and can include enzymes, riboproteins, solutions, cofactors, and the like. For example, gene editing reagents include one or more components required for zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, and the CRISPR / Cas system (clustered regularly interspaced short palindromic repeats system).
[0061] The term "endonuclease" as used here refers to an enzyme or a component of an endonuclease system (e.g., any component of CRISPR, including a gRNA) that possesses endonucleolytic catalytic activity for the cleavage of polynucleotides. For example, an endonuclease or a component thereof can cleave a phosphodiester bond of an oligonucleotide or polynucleotide. An endonuclease cleaves at a phosphodiester bond within or adjacent to its recognition site sequence, which extends over a length of at least 4 bp. Types of endonucleases include, among others, restriction enzymes, AP endonuclease, T7 endonuclease, T4 endonuclease, Bal 31 endonuclease, endonuclease I, micrococcal nuclease, endonuclease II, neurospora endonuclease, S1 endonuclease, P1 nuclease, mung bean nuclease I, DNase I, RNA-directed DNA endonuclease (e.g., CRISPR, including all CRISPR components, e.g., Cas protein, gRNA, etc.).), Homothallic-switching endonuclease, TALENs, zinc finger nucleases and Endo R.
[0062] Cleavage refers to the breaking of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of processes, including, but not limited to, the enzymatic or chemical hydrolysis of a phosphodiester bond. Both single-stranded and double-stranded cleavages are possible, with double-stranded cleavage potentially resulting from two separate single-stranded cleavage events. DNA cleavage can produce either blunt or offset ends. In some embodiments, a complex of guide RNA and a site-specific modifying enzyme is used for targeted double-stranded DNA cleavage.
[0063] The term "CRISPR," or "clustered regularly interspaced short palindromic repeats," is used here in its usual sense and refers to a genetic element that bacteria use as a form of acquired immunity to protect themselves against viruses. CRISPR comprises short sequences derived from viral genomes that have been inserted into the bacterial genome. Cas (CRISPR-associated proteins) process these sequences and cleave matching viral DNA sequences. The CRISPR sequences thus serve as a guide for Cas to recognize and cut DNA that is at least partially complementary to the CRISPR sequence. By introducing plasmids containing Cas genes and specially engineered CRISPR sequences into eukaryotic cells, the eukaryotic genome can be cut at any desired location.
[0064] In the following, the term "Cas9" or "CRISPR-associated protein 9" is used in its ordinary sense and refers to an enzyme that uses CRISPR sequences as a guide to recognize and cut specific DNA strands that are at least partially complementary to the CRISPR sequence. Cas9 enzymes, together with CRISPR sequences, form the basis of a technology known as CRISPR-Cas9, which can be used to edit genes in organisms. This editing technique has a wide range of applications, including basic biological research, the development of biotechnology products, and the treatment of diseases.
[0065] A “CRISPR-associated protein 9”, “Cas9”, “Csn1”, or “Cas9 protein”, as referred to herein, comprises any of the recombinant or naturally occurring forms of the Cas9 endonuclease or variants or homologs thereof that retain the enzyme activity of the Cas9 endonuclease (e.g., within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% activity compared to Cas9). In some embodiments, the variants or homologs exhibit at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a 50-, 100-, 150-, or 200-amino-acid continuous region) compared to a naturally occurring Cas9 protein. In some embodiments, the Cas9 protein is essentially identical to the protein identified by the UniProt reference number Q99ZW2, or to a variant or homolog that is essentially identical to it.Under certain aspects, the Cas9 protein exhibits at least 75% sequence identity with the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. Under certain aspects, the Cas9 protein exhibits at least 80% sequence identity with the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. Under certain aspects, the Cas9 protein exhibits at least 85% sequence identity with the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. Under certain aspects, the Cas9 protein exhibits at least 90% sequence identity with the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. In some cases, the Cas9 protein exhibits at least 95% sequence identity with the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2.
[0066] A “CRISPR-associated endonuclease Cas12a”, “Cas12a”, “Cas12” or “Cas12 protein” referred to herein includes any of the recombinant or naturally occurring forms of the Cas12 endonuclease or variants or homologs thereof that retain the enzyme activity of the Cas12 endonuclease (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to Cas12). In some embodiments, the variants or homologs exhibit at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity over the entire sequence or a part of the sequence (e.g., a 50-, 100-, 150- or 200-continuous amino acid segment) compared to a naturally occurring Cas12 protein.In some aspects, the Cas12 protein is essentially identical to the protein identified by the UniProt reference number A0Q7Q2, or to a variant or homolog that has substantial identity with it.
[0067] A “CRISPR-associated endoribonuclease Cas13a”, “Cas13a”, “Cas13” or “Cas13 protein” referred to herein includes any of the recombinant or naturally occurring forms of the Cas13 endoribonuclease or variants or homologs thereof that retain the enzyme activity of the Cas13 endoribonuclease (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to Cas13). In some embodiments, the variants or homologs exhibit at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity over the entire sequence or a part of the sequence (e.g., a 50-, 100-, 150- or 200-continuous amino acid segment) compared to a naturally occurring Cas13 protein.In some aspects, the Cas13 protein is essentially identical to the protein identified by the UniProt reference number P0DPB8, or to a variant or homolog that has substantial identity with it.
[0068] The term "TALEN," or "transcription activator-like effector nuclease," used here refers to restriction enzymes generated by attaching a DNA-binding domain (e.g., a TAL effector DNA-binding domain) to a nuclease (e.g., FokI). TALENs typically contain a naturally occurring DNA-binding domain comprising multiple modules called TALs or TALEs. Thus, the TALs containing variable diresids confer DNA-binding specificity.
[0069] A "guide RNA" or "gRNA," as described here, refers to an RNA sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and guide the sequence-specific binding of a CRISPR complex to the target sequence. For example, a gRNA Cas can bind directly to the target polynucleotide. In some embodiments, the gRNA comprises the crRNA and the tracrRNA. For instance, the gRNA may contain the crRNA and the tracrRNA hybridized by base pairing. Thus, the two RNAs can be separately encoded as two RNA molecules, a crRNA and a tracrRNA, which then form an RNA / RNA complex through complementary base pairing between the crRNA and the tracrRNA.Under certain aspects, the degree of complementarity between a guide RNA sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher. Under other aspects, the degree of complementarity between a guide RNA sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is at least approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, or 99%.
[0070] Non-restrictive examples of CRISPR enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. of these. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, or derived mutants thereof in these organisms. In some embodiments, the CRISPR enzyme is codon-optimized for expression in a eukaryotic cell. In some embodiments, the CRISPR enzyme directs the cleavage of one or two strands at the site of the target sequence. In some embodiments, the CRISPR enzyme lacks the activity to cleave the DNA strand.
[0071] As used herein, a “zinc finger” is a polypeptide structural motif folded around a bound zinc cation. In embodiments, the polypeptide of a zinc finger has a sequence of the form X3-Cys-X. 2-4 -Cys-X -His-X -His-X 123-54 , where X is any amino acid (e.g., X denotes 2-4 an oligopeptide with a length of 2-4 amino acids). Thus, the term "zinc finger nuclease", as used here, refers to a nuclease that contains a zinc finger motif and a domain capable of inducing breaks in the target DNA.
[0072] The term “homologous recombination” refers to a type of genetic recombination in which information is exchanged between two similar or identical nucleic acid sequences, which may be referred to here as “homology regions.” In some embodiments of the methods described herein, a homology region may, for example, comprise two homologous regions, which may optionally flank a non-homologous region. In embodiments of the methods described herein, an E. coli RecA gene may be used to enhance homologous recombination. “RecA” refers to the bacterial homolog of the family of ubiquitous 38-kDa homologous DNA repair proteins that mediates ATP-dependent homologous recombination in bacteria. In some embodiments of the methods described herein, the donor or recipient cell contains an oligonucleotide encoding one or more homologous DNA repair genes, such as RecA.In some embodiments, the expression of homologous DNA repair genes is induced. In these embodiments, the homologous DNA repair gene is RecA. In other embodiments, the homologous DNA repair genes are the recombination genes Rotα, Rotβ, and Rotγ. Non-restrictive examples of homologous recombination and gene editing methods using various nuclease systems can be found, for example, in US Patent No. 8945839, PCT International Publication No. WO2013 / 163394, and US Patent Applications Nos. 2016 / 0060657, 2012 / 0192298A1, and US2007 / 0042462. These and other known homologous recombination methods can be used in combination with the methods described here.
[0073] The term “transfection” is used here in its usual sense and refers to a process for the deliberate introduction of naked or purified nucleic acids into eukaryotic cells. In some formulations, the term “transfection” may also refer to other processes and cell types, although other terms are often preferred. For example, the term “transformation” is generally used to describe non-viral DNA transfer into bacteria and non-animal eukaryotic cells, including plant cells. For animal cells, transfection is the preferred term. For instance, the term “transduction” is frequently used to describe virus-mediated gene transfer into eukaryotic cells.
[0074] The terms "bacterial conjugation" and "bacterial mating" are interchangeable and refer to a mode of genetic exchange between bacteria. In bacterial conjugation, typically only a portion of the genome of one cell (the donor) and the entire genome of the partner (the recipient cell) are transferred. Gene transfer in bacterial conjugation is therefore usually partial. In some embodiments, bacterial conjugation involves the transfer of non-genomic bacterial DNA from a donor cell to a recipient cell. In some embodiments, bacterial conjugation is mediated by a plasmid. In some embodiments, bacterial conjugation is mediated by exogenous DNA within the bacteria. In some embodiments, the donor and recipient cells are in contact to allow bacterial conjugation to occur.In some embodiments, the donor cell and the recipient cell contain a connecting bridge (e.g., a pilus) to allow bacterial conjugation to take place.
[0075] In some embodiments, the recipient cell or the donor cell contains an oligonucleotide that enables plasmid conjugation. In some embodiments, the oligonucleotide enabling plasmid conjugation is located in the genome of the donor cell. In some embodiments, the oligonucleotide enabling plasmid conjugation is located in a helper plasmid. In some embodiments, the oligonucleotide enabling plasmid conjugation is the Tra operon.In embodiments, the oligonucleotide enabling plasmid conjugation is selected from: IncF1 Tra (traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traR, traS, traT), IncP Tra operon: (trbA, trbB, trbC, trbD, trbE, trbF, trbG, trbH, trbI, trbJ, trbK, trbL, traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO), IncI1 tra operon: (traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traS, traT, traU, traV, traW, traY), pTiC58 tra genes: (traA, traF, traB, traC, traG, traD, traR, traI), and pIJ101: clt, korB. In some embodiments, the oligonucleotide enabling plasmid conjugation is IncF1 Tra (traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traR, traS, traT).In some embodiments, the oligonucleotide that enables plasmid conjugation is the IncP Tra operon: (trbA, trbB, trbC, trbD, trbE, trbF, trbG, trbH, trbI, trbJ, trbK, trbL, traA, traB, traC, traD, traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO). In some embodiments, the oligonucleotide enabling plasmid conjugation is IncI1 tra operon: (traE, traF, traG, traH, traI, traJ, traK, traL, traM, traN, traO, traP, traQ, traS, traT, traU, traV, traW, traY). In some embodiments, the oligonucleotide enabling plasmid conjugation is pTiC58 tra gene: (traA, traF, traB, traC, traG, traD, traR, traI). In some embodiments, the oligonucleotide enabling plasmid conjugation is pIJ101: clt, korB.
[0076] The term "donor cell" here refers to a cell (e.g., a bacterial cell) that transfers genetic material to another cell (e.g., a bacterial cell, a plant cell, etc.). The cell that receives the transferred genetic material is referred to as the "recipient cell."
[0077] The term "donor plasmid" used here refers to DNA from a donor cell (e.g., a bacterial cell) containing an oligonucleotide sequence (e.g., donor DNA, an oligonucleotide with a DNA element) that is to be transferred from the donor cell to a recipient cell (e.g., a bacterial cell, yeast cell, plant cell, etc.). Typically, the donor plasmid is a circular double-stranded DNA sequence that is separate from the genomic DNA. The term "recipient plasmid" therefore refers to the DNA from a recipient cell that receives the donor DNA. In some embodiments, the DNA of the donor plasmid is incorporated into a different DNA sequence than that of a recipient plasmid. Thus, in some embodiments, the donor DNA can be integrated into genomic DNA.
[0078] In some embodiments, the donor plasmid contains a transfer origin. In some embodiments, the transfer origin is a mobile element. In some embodiments, the mobile element is a plasmid. In some embodiments, the plasmid is an IncFI plasmid, an IncPα plasmid, an IncI1 plasmid, a pTiC58 from Agrobacterium tumefaciens, a pAD1 plasmid, an Inc18 plasmid, or an IncH plasmid. In some embodiments, the plasmid is an IncFI plasmid. In some embodiments, the plasmid is an IncPα plasmid. In some embodiments, the plasmid is an IncI1 plasmid. In some cases, the plasmid is a pTiC58 from Agrobacterium tumefaciens. In some embodiments, the plasmid is a pAD1 plasmid. In some cases, the plasmid is an Inc18 plasmid. In some cases, the plasmid is an IncH plasmid. The plasmids are discussed in more detail in Ippen-Ihler, KA, and Minkley, EG, Jr. (1986).The conjugation system of F, the fertility factor of Escherichia coli. Ann. Rev. Genet. 20:593-624; Guiney, DG, and Lanka, E., (1989), Conjugative transfer of IncP plasmids, in: Promiscuous Plasmids of Gramnegative Bacteria (CM Thomas, ed.), Academic Press, London, pp. 27-56.; Catherine ED Rees, David E. Bradley, Brian M. Wilkins, (1987) Organization and regulation of the conjugation genes of IncI1 plasmid ColIb-P9. Plasmid. 18: 223-236; von Bodman SB, McCutchan JE, Farrand SK. (1989) Characterization of the conjugal transfer functions of the Agrobacterium tumefaciens Ti plasmid pTiC58. J. Bacteriol. 171(10):5281-5289.; Clewell DB, Weaver KE. (1989) Sex pheromones and plasmid transfer in Enterococcus faecalis. Plasmid. 21(3):175-84.; Kohler V, Vaishampayan A, Grohmann E. (2018) Broad-host-range Inc18 plasmids: Occurrence, spread and transfer mechanisms. Plasmid. 99:11-21.; Andreas Schlüter, Patrice Nordmann, Rémy A. Bonnin, Yves Millemann, Felix G.Eikmeyer, Daniel Wibberg, Alfred Pühler, Laurent Poirel. (2014) IncH-Type Plasmid Harboring . blaCTX-M-15 , bla DHA-1 , and qnrB4 Genes Recovered from Animal Isolates. Antimicrobial Agents and Chemotherapy 58(7):3768-3773.
[0079] In some embodiments, the transmission originates from a mobile element. In some embodiments, the mobile element originates from a conjugative transposon. In some embodiments, the conjugative transposon is Tn916 from Enterococcus faecalis or CTnDOT from Bacteroides. In some embodiments, the conjugative transposon is Tn916 from Enterococcus faecalis. In some embodiments, the conjugative transposon is CTnDOT from Bacteroides. In some embodiments, the mobile element is an integrating conjugative element. In some embodiments, the mobile element originates from SXT from Vibrio cholerae or R391 from Providencia rettgeri. In some embodiments, the mobile element originates from SXT from Vibrio cholerae. In some embodiments, the mobile element originates from R391 from Providencia rettgeri. The elements are described in more detail in the references Rice LB (1998).Tn916 family conjugative transposons and dissemination of antimicrobial resistance determinants. Antimicrobial agents and chemotherapy, 42(8), 1871-1877; Cheng Q, Paszkiet BJ, Shoemaker NB, Gardner JF, Salyers AA. (2000) Integration and excision of a conjugative transposon from Bacteroides, CTnDOT. J Bacteriol. 182(14):4035-43.; Bianca Hochhut and Matthew K. Waldor. (1999) Site-specific integration of the conjugal Vibrio cholerae SXT element into prfC. Mol. Microbiology. 32(1):99-110.; Böltner D, MacMahon C, Pembroke JT, Strike P, Osborn AM. R391: a conjugatively integrating mosaic of phage, plasmid, and transposon elements. J Bacteriol. 2002;184(18):5158-5169.
[0080] In some embodiments, the donor plasmid contains a conditional origin of replication. In some embodiments, the conditional replicon is R6K-pir, RSF1010 oriV-RepA / B / C, ColE2 P9-RepA, RP4 oriV-trfA, pPS10 oriV-RepA, pSC101 ori-RepC TS., RK2 oriV, bacteriophage P1 ori, plasmid pSC101 origin of replication, bacteriophage lambda ori, pBR322 plasmid, pSU739 plasmid, or pSU300 plasmid. In some embodiments, the conditional replicon is R6K-pir. In some embodiments, the conditional replicon is RSF1010 oriV-RepA / B / C. In some embodiments, the conditional replicon is ColE2 P9-RepA. In some cases, the conditional replicon is RP4 oriV-trfA. In some embodiments, the conditional replicon is pPS10 oriV-RepA. In some embodiments, the conditional replicon is pSC101 ori-RepC. TSIn some embodiments, the conditional replicon is RK2 oriV. In some cases, the conditional replicon is bacteriophage P1 ori. In some embodiments, the conditional replicon is the origin of replication of plasmid pSC101. In some embodiments, the conditional replicon is bacteriophage lambda ori. In some embodiments, the conditional replicon is plasmid pBR322. In some embodiments, the conditional replicon is plasmid pSU739. In some cases, the conditional replicon is plasmid pSU300. The plasmids are described in the references: Metcalf WW, Jiang W, Daniels LL, Kim SK, Haldimann A, Wanner BL. (1996) Conditional replicative and conjugative plasmids carrying lacZ alpha for cloning, mutagenesis, and allele replacement in bacteria. Plasmid. 35(1):1-13.; Scherzinger E, Bagdasarian MM, Scholz P, Lurz R, Rückert B, Bagdasarian M.(1984) Replication of the broad host range plasmid RSF1010: requirement for three plasmid-encoded proteins. Proc Natl Acad Sci US A. 81(3):654-8.; ColE2-P9: Yagura M, Nishio SY, Kurozumi H, Wang CF, Itoh T. (2006) Anatomy of the replication origin of plasmid ColE2-P9. J Bacteriol. 188(3):999-1010.; Ayres EK, Thomson VJ, Merino G, Balderes D, Figurski DH. Precise deletions in large bacterial genomes by vector-mediated excision (VEX). (1993) The trfA gene of the promiscuous plasmid RK2 is essential for replication in several Gram-negative hosts. J Mol Biol. 5;230(1):174-85.; Maestro B, Sanz JM, Díaz-Orejas R, Fernández-Tresguerres E (2003) Modulation of pPS10 host range by plasmid-encoded RepA initiator protein. J Bacteriol.185(4):1367-75.; Hashimoto-Gotoh, T., & Sekiguchi, M. (1977). Temperature sensitivity mutations in the R plasmid pSC101. Journal of Bacteriology, 131(2), 405-412.; Ayres EK, Thomson VJ, Merino G, Balderes D, Figurski DH.Precise deletions in large bacterial genomes by vector-mediated excision (VEX). (1993) The trfA gene of the promiscuous plasmid RK2 is essential for replication in several Gram-negative hosts. J Mol Biol. 5;230(1):174-85. Stenzel TT, Patel P, Bastia D. (1987) The integration host factor of Escherichia coli binds to bent DNA at the origin of replication of the plasmid pSC101. Cell. 5;49(5):709-17.; Sugiura S, Ohkubo S, Yamaguchi K. (1993) Minimal essential origin of plasmid pSC101 replication: requirement of a region downstream of iterons. J Bacteriol. 175(18):5993-6001; Pal SK, Mason RJ, Chattoraj DK (1986) P1 plasmid replication. Role of initiator titration in copy number control. J Mol Biol 20;192(2):275-85.; LeBowitz JH, McMacken R (1984) The bacteriophage lambda O and P protein initiators promote the replication of single-stranded DNA. Nucleic Acids Res. 12(7):3069-3088.; Grindley ND, Kelley WS. (1976) Effects of different alleles of the E.coli K12 pol A gene on the replication of non-transferring plasmids. Mol Gene Genet. 2;143(3):311-8.; Francia, M.V., & García Lobo, J.M. (1996). Gene integration into the Escherichia coli chromosome mediated by the Tn21 integrase (Int21). Journal of bacteriology, 178(3), 894-898; Mendiola MV, de la Cruz F. (1989) Specificity of insertion of IS91, an insertion sequence present in alpha-haemolysin plasmids of Escherichia coli. moles of microbiol. 3(7):979-84.
[0081] In some embodiments, the conditional origin of replication depends on the presence of an oligonucleotide. In some embodiments, the oligonucleotide encodes pir1, pir1-116, repA / repB / repC (RSF1010 replicant), repA (ColE2-P9 replicant), trfA (RP4 replicant), RepA (pSP10 replicant), RepC TS(pSC101 replicate) or a combination thereof. In some embodiments, the oligonucleotide encodes pir1. In some embodiments, the oligonucleotide encodes pir1-116. In some embodiments, the oligonucleotide encodes repA / repB / repC (RSF1010 replicate). In some embodiments, the oligonucleotide encodes repA (ColE2-P9 replicate). In some cases, the oligonucleotide encodes trfA (RP4 replicate). In some embodiments, the oligonucleotide encodes RepA (pSP10 replicate). In some cases, the oligonucleotide encodes RepC. TS (pSC101 replica).
[0082] In some embodiments, the conditional origin of replication depends on a cell growth condition. In some embodiments, the condition is temperature.
[0083] In the methods presented here, the donor plasmid or the recipient oligonucleotide comprises, in some embodiments, a replicon that can replicate plasmids with a length of 20 or 30 kilobases.
[0084] In the methods presented here, the donor plasmid or recipient oligonucleotide comprises, in some embodiments, a replicon capable of replicating plasmids longer than 30 kilobases. In other embodiments, the replicon can replicate plasmids ranging from approximately 30 kilobases to approximately 500 kilobases in length. In other embodiments, the replicon can replicate plasmids approximately 30 kilobases in length. In other embodiments, the replicon can replicate plasmids approximately 50 kilobases in length. In some embodiments, the replicon can replicate plasmids approximately 70 kilobases in length. In other embodiments, the replicon can replicate plasmids approximately 90 kilobases in length. In some embodiments, the replicon can replicate plasmids approximately 100 kilobases in length. In some embodiments, the replicon can replicate plasmids with a length of approximately 120 kilobases.In some embodiments, the replicon can replicate plasmids with a length of approximately 140 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 160 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 180 kilobases. In some embodiments, the replicon can replicate plasmids with a length of approximately 200 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 220 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 240 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 260 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 280 kilobases. In other embodiments, the replicon can replicate plasmids with a length of approximately 300 kilobases.In some embodiments, the replicon can replicate plasmids with a length of approximately 400 kilobases. In some embodiments, the replicon can replicate plasmids with a length of approximately 500 kilobases. The length can be any value or subrange within the specified ranges, including the endpoints.
[0085] In some embodiments, the replicon is derived from a P1-derived artificial chromosome or a bacterial artificial chromosome. In some embodiments, the replicon is derived from an artificial chromosome derived from P1. In some embodiments, the replicon is derived from a bacterial artificial chromosome. In some embodiments, the donor plasmid or the recipient oligonucleotide contains an inducible high-copy origin of replication. In some embodiments, the donor plasmid contains an inducible high-copy origin of replication. In some embodiments, the recipient oligonucleotide contains an inducible high-copy origin of replication.
[0086] In some embodiments, the donor plasmid or recipient oligonucleotide is an artificial yeast chromosome (YAC), an artificial mammalian chromosome (MAC), an artificial human chromosome (HAC), or an artificial plant chromosome. In some embodiments, the donor plasmid is an artificial yeast chromosome (YAC). In some cases, the donor plasmid is an artificial mammalian chromosome (MAC). In some embodiments, the donor plasmid is an artificial human chromosome (HAC). In some embodiments, the donor plasmid is an artificial plant chromosome. In some embodiments, the recipient oligonucleotide is an artificial yeast chromosome (YAC). In some embodiments, the recipient oligonucleotide is an artificial mammalian chromosome (MAC). In some embodiments, the recipient oligonucleotide is a human artificial chromosome (HAC).In some cases, the recipient oligonucleotide is an artificial plant chromosome.
[0087] In some embodiments, the donor plasmid or the recipient oligonucleotide comprises a conjugable vector, which may be a viral vector. In some embodiments, the donor plasmid is a viral vector. In some embodiments, the recipient oligonucleotide is a viral vector. In some embodiments, the viral vector is a retrovirus. In some embodiments, the viral vector is a lentivirus. In some embodiments, the viral vector is an adenovirus. In some cases, the viral vector is an adeno-associated virus. In some cases, the viral vector is a tobacco mosaic virus. In some cases, the viral vector is a baculovirus. In some cases, the viral vector is a herpes simplex virus. In some cases, the viral vector is a smallpox virus. In embodiments, the viral vector is a gammaretrovirus. In some cases, the viral vector is a Sendai virus.
[0088] The term "control" or "control experiment" is used here in its usual sense and refers to an experiment in which the subjects or reagents of the experiment are treated as in a parallel experiment, except that one procedure, reagent, or variable of the experiment is omitted. In some cases, the control is used as a benchmark for evaluating the effects of the experiment.
[0089] A "control sample" or "control value" refers to a sample that serves as a reference, usually a known reference, for comparison with a test sample. For example, a test sample may be taken under a test condition, such as in the presence of a test compound, and compared with samples taken under known conditions, such as in the absence of the test compound (negative control) or in the presence of a known compound (positive control). A control may also represent an average value calculated from a series of tests or results. Experts know that controls can be developed to evaluate any number of parameters. For example, a control may be developed to compare therapeutic benefit based on pharmacological data (e.g., half-life) or therapeutic measures (e.g., comparison of side effects).Experts know which controls are appropriate in a given situation and can analyze the data by comparing them to control values. Controls are also useful for determining the significance of data. For example, if the values for a particular parameter vary significantly between the controls, the deviations in the test samples will not be considered significant.
[0090] The term "contacting" is used here in its usual sense and refers to the process by which at least two different species (e.g., chemical compounds, including biomolecules or cells) come into sufficient proximity to react, interact, or physically touch. However, the resulting reaction product may arise directly from a reaction between the added reagents or from an intermediate formed from one or more of the added reagents within the reaction mixture.
[0091] The term "expression" encompasses every step involved in the production of the polypeptide, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion. Expression can be detected using conventional protein detection methods (e.g., ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, etc.).
[0092] The term "recombinant," when used in reference to a cell, nucleic acid, protein, or vector, means that the cell, nucleic acid, protein, or vector has been modified by the introduction of a heterologous nucleic acid or protein, or by the modification of a native nucleic acid or protein, or that the cell originates from such a modified cell. For example, recombinant cells express genes that are not present in the native (non-recombinant) form of the cell, or they express native genes that are otherwise abnormally, unduly, or not at all expressed. Transgenic cells and plants are those that express a heterologous gene or coding sequence, typically as a result of recombinant processes.
[0093] The terms “origin of transfer” or “oriT” used here refer to a short sequence (up to 500 bp) required for the transfer of DNA from a bacterial host to a recipient during bacterial conjugation.
[0094] The term "curable origin of replication" used here refers to an origin of replication that does not replicate when cells are grown in the presence of certain chemicals or under certain environmental conditions. Under these conditions, plasmids containing a curable origin of replication are lost from the cell. This is how, for example, pSC101 or i TS not at high temperatures and is lost.
[0095] The term “mobile element” used here refers to a type of genetic material that can move within a genome or be transferred between genomes, even between species.
[0096] The term “conjugative transposon” used here refers to integrated DNA elements that self-cleave and form a covalently closed circular intermediate that can be reintegrated into the same cell or transferred to a recipient cell by conjugation.
[0097] The term “integrating conjugative element” used here refers to a group of chromosomally integrated, self-transmittable genetic elements.
[0098] The term “artificial chromosome from P1” used here refers to a DNA construct derived from the bacteriophage P1.
[0099] The term “bacterial artificial chromosome” used here refers to a manipulated DNA sequence used to clone DNA sequences in bacteria.
[0100] The term "recombination-mediated genetic engineering genes" or "recombination genes" refers to genes that assist in the creation of genetic modifications in a DNA sequence. In some embodiments, recombination-mediated genetic engineering genes enable the in vivo construction of constructs in cells (e.g., bacterial cells) without in vitro genetic engineering techniques. In some embodiments, recombination-mediated genetic engineering genes enable genetic modifications without the use of enzymes such as ligases and restriction enzymes. The genes can, for example, participate in the natural process of homologous recombination in a bacterium without the use of conventional molecular biology techniques. In some cases, the recombining genes are lambda red genes. The recombination genes are Rotα, Rotβ, and Rotγ.The genes can, for example, induce homologous recombination at a high rate in bacteria. The expression of one or more recombination genes can be induced in a donor or recipient cell. In some embodiments, the donor or recipient cell contains an oligonucleotide encoding one or more recombination-mediated genetic engineering genes. In some embodiments, the oligonucleotide encoding one or more recombination-mediated genetic engineering genes is located in the donor cell plasmid. In some embodiments, the recombination-mediated genetic engineering genes are inducible. In some cases, the recombination-mediated genetic engineering genes are Rotα, Rotβ, and Rotγ. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the recipient cell genome.In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in a helper plasmid. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the recipient oligonucleotide, which may be in the form of a plasmid. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the genome of the recipient cell.
[0101] The term "inducible high-copy origins of replication" used here refers to a plasmid or vector that contains a high number of origins of replication (e.g., between 150 and 200 copies in the E. coli plasmid pUC) that can be induced by environmental conditions, such as a change in temperature.
[0102] The term "helper plasmid" used here refers to a plasmid containing genes or other DNA elements that a bacterium needs to perform a specific function. A helper plasmid may contain an endonuclease, elements for transferring foreign DNA into the genome, for transferring a plasmid into another cell, or for carrying out homologous recombination. In some cases, the helper plasmid is an IncF1 plasmid, an IncPα plasmid, an IncI1 plasmid, a pTiC58 from Agrobacterium tumefaciens, a cAD1 plasmid, an Inc18 plasmid, a pIJ101 from Streptomyces, or an IncH plasmid. In some embodiments, the helper plasmid is an IncF1 plasmid. In some embodiments, the helper plasmid is an IncPα plasmid. In some embodiments, the helper plasmid is an IncI1 plasmid. In some embodiments, the helper plasmid is a pTiC58 from Agrobacterium tumefaciens. In some embodiments, the helper plasmid is a cAD1 plasmid.In some cases, the helper plasmid is an Inc18 plasmid. In some cases, the helper plasmid is a pIJ101 from Streptomyces. In some cases, the helper plasmid is an IncH plasmid. In some embodiments, the helper plasmid lacks a functional transfer origin. In some embodiments, the helper plasmid contains a selectable marker that selects for the retention of the helper plasmid in the donor cell.
[0103] The term "homing endonuclease" used here refers to an endonuclease that is encoded either as a free-standing gene in an intron sequence, as a fusion with a host protein, or as a self-splicing protein. In contrast to group II restriction enzymes, homing endonucleases catalyze the hydrolysis of DNA at longer recognition sites. Examples of homing endonucleases include LAGLIDAG, GIY-YIG, His-Cys-Box, HNH, PD-(D / E)xK, and Vsr-like / EDxHD. In some embodiments, the homing endonuclease is I-ScaI, PI-SceI, I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, or I-Vdi141I. In some embodiments, the homing endonuclease is I-ScaI.In some embodiments, the homing endonuclease is PI-SceI. In some embodiments, the homing endonuclease is I-AniI. In some embodiments, the homing endonuclease is I-CeuI. In some embodiments, the homing endonuclease is I-ChuI. In some embodiments, the homing endonuclease is I-CpaI. In some cases, the homing endonuclease is I-CpaII. In some embodiments, the homing endonuclease is I-CreI. In some embodiments, the homing endonuclease is I-DmoI. In some embodiments, the homing endonuclease is H-DreI. In some embodiments, the homing endonuclease is I-HmuI. In some embodiments, the homing endonuclease is I-HmuII. In some embodiments, the homing endonuclease is I-LlaI. In some cases, the homing endonuclease is I-MsoI. In some embodiments, the homing endonuclease is PI-PfuI. In some embodiments, the homing endonuclease is PI-PkoII.In some embodiments, the homing endonuclease is I-PorI. In some embodiments, the homing endonuclease is I-PpoI. In some embodiments, the homing endonuclease is PI-PspI. In some embodiments, the homing endonuclease is I-SceI. In some embodiments, the homing endonuclease is I-SceII. In some embodiments, the homing endonuclease is I-SceIII. In some embodiments, the homing endonuclease is I-SceIV. In some embodiments, the homing endonuclease is I-SceV. In some embodiments, the homing endonuclease is I-SceVI. In some embodiments, the homing endonuclease is I-SceVII. In some cases, the homing endonuclease is I-Ssp6803I. In some embodiments, the homing endonuclease is I-TevI. In some embodiments, the homing endonuclease is I-TevII. In some embodiments, the homing endonuclease is I-TevIII.In some embodiments, the homing endonuclease is PI-T1I. In some embodiments, the homing endonuclease is PI-T1II. In some embodiments, the homing endonuclease is I-Tsp061I or I-Vdi141I. In some cases, the homing endonuclease is I-Vdi141I.
[0104] The term “RNA-directed DNA endonuclease” used here refers to any DNA endonuclease that is guided to a target DNA sequence by an auxiliary or guide RNA molecule. Examples of RNA-directed DNA endonucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3 and Csf4, all variants and Homologs thereof.
[0105] The term “HO” or “homothallic switching endonuclease” used here refers to the zinc finger nuclease in Saccharomyces cerevisiae, which is responsible for initiating mating-type conversion.
[0106] A “cell,” as used here, refers to a cell that performs metabolic or other functions sufficient to maintain or replicate its genomic DNA. A cell can be identified by known methods, such as the presence of an intact membrane, staining with a specific dye, the ability to produce offspring, or, in the case of a gamete cell, the ability to unite with a second gamete cell to produce viable offspring. Cells can include prokaryotic and eukaryotic cells. Prokaryotic cells include, among others, bacteria. Eukaryotic cells include, among others, yeast cells and cells derived from plants and animals, such as mammalian, insect (e.g., Spodoptera), and human cells. Cells can be useful if they are naturally non-adherent or have been treated to prevent them from adhering to surfaces, for example, by…by trypsinization.
[0107] The term "donor cell" here refers to a cell (e.g., a bacterial cell) that transfers genetic material to another cell (e.g., a bacterial cell, a plant cell, etc.). The cell that receives the transferred genetic material is referred to as the "recipient cell."
[0108] The term "donor plasmid" as used here refers to DNA from a donor cell (e.g., a bacterial cell) containing an oligonucleotide sequence (e.g., donor DNA, an oligonucleotide with a DNA element or fragment thereof) that is to be transferred from the donor cell to a recipient cell (e.g., a bacterial cell, yeast cell, plant cell, etc.). Typically, the donor plasmid is a circular double-stranded DNA sequence separated from genomic DNA. Therefore, the term "recipient oligonucleotide" can refer to plasmid DNA in a recipient cell that takes up the donor DNA (e.g., through homologous recombination of the donor DNA into the recipient cell); or the term "recipient oligonucleotide" can refer to any oligonucleotide in the recipient cell that takes up the donor DNA, e.g., genomic DNA of the recipient cell.In some embodiments, the DNA of the donor plasmid is taken up by a different DNA than that of a recipient plasmid. In this way, the donor DNA can be integrated into genomic DNA in certain embodiments.
[0109] The term "isolated" in the context of a nucleic acid or protein means that the nucleic acid or protein is essentially free from other cellular components with which it is associated in its natural state. It may, for example, be in a homogeneous state and exist in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemical methods such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. A protein that is the predominant species in a preparation is essentially purified.
[0110] It is assumed that the examples and embodiments described here serve only for illustration and that various modifications or changes are proposed to persons experienced in the field of technology and are to be included in the scope of the attached claims.
[0111] In embodiments of the methods described herein, the donor or recipient cell contains an oligonucleotide encoding one or more homologous DNA repair genes. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the first, second, or subsequent donor plasmid. In some embodiments, the expression of homologous DNA repair genes is induced. In some embodiments, the homologous DNA repair gene is RecA.
[0112] In the methods described here, the donor cell or the recipient cell contains an oligonucleotide that codes for one or more recombination-mediated genetic engineering genes. In some embodiments, the oligonucleotide that codes for one or more recombination-mediated genetic engineering genes is located in the plasmid of the donor cell.
[0113] In the methods described here, the recombination-mediated genetic engineering genes are inducible. In some cases, the recombination-mediated genetic engineering genes are Rotα, Rotβ, and Rotγ. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the genome of the recipient cell. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in a helper plasmid. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the recipient oligonucleotide, which may be in the form of a plasmid. In some embodiments, the oligonucleotide encoding one or more homologous DNA repair genes is located in the genome of the recipient cell.
[0114] In the methods described here, the donor cells, recipient cells, or recombinant recipient cells can be located in an ordered array or in a first or second ordered array. In some embodiments, the donor cells, recipient cells, or recombinant recipient cells can be transferred to positions in a third ordered array, a fourth ordered array, or a subsequent ordered array.
[0115] In the methods described here, the donor cells and recipient cells are bacterial cells. In some embodiments, the recipient cells are not bacterial cells. In some cases, the recipient cells are plant cells. In some embodiments, the recipient cells are yeast cells. In some cases, the recipient cells are mammalian cells. II. Method for assembling a DNA element
[0116] The invention provides a method for assembling a plurality of DNA elements into a composite DNA element in a recipient cell, the method comprising: (a) contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions to (i) transfer the first donor plasmid from the first donor cell to the recipient cell by conjugation and (ii) recombine the first donor plasmid and the recipient oligonucleotide in the recipient cell by homologous recombination, wherein the recipient oligonucleotide is located in a recipient cell plasmid or the recipient cell genome, and wherein the first donor plasmid comprises, in successive order, a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a first DNA element fragment (Oligo1),a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2) and a third endonuclease site (C3); the recipient oligonucleotide comprises a third homologous recombination region (HR3) homologous to HR1 and a fourth homologous recombination region (HR4) homologous to HR2.2; thus, after homologous recombination of HR1 with HR3 and HR2.2 with HR4, a first recombinant recipient oligonucleotide is formed, comprising the first DNA element fragment; (b) Contacting a second donor cell containing a second donor plasmid with the recipient cell containing the first recombinant recipient oligonucleotide under conditions to (i) transfer the second donor plasmid from the second donor cell to the first recipient cell by conjugation and (ii) recombine the second donor plasmid and the first recombinant recipient oligonucleotide,to form a second recombinant recipient oligonucleotide in the recipient cell by homologous recombination; wherein the second donor plasmid sequentially incorporates a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a second DNA element fragment (Oligo2), a sixth homologous recombination region (HR6), the two homologous recombination regions (HR6.1, HR6.2) and a sixth endonuclease site (C6); whereby, following homologous recombination of HR5 with HR2.1 and HR6.2 with HR4, a second recombinant recipient oligonucleotide is provided, comprising the first and second DNA element fragments (Oligo1, Oligo2) that form a DNA assembly, wherein an oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third and / or fourth endonuclease site,on the first donor plasmid and / or wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth and / or sixth endonuclease site is present on the second donor plasmid. In some embodiments, HR2.1 and HR2.2 flank a non-homologous region comprising one (C2) or two endonuclease sites (C2.1, C2.2). wherein HR3 and HR4 optionally flank a non-homologous region comprising one (C4) or two endonuclease sites (C4.1, C4.2). In some embodiments, the recipient oligonucleotide is located in a plasmid of the recipient cell or in the genome of the recipient cell. In some embodiments, the DNA assembly comprises at least a portion of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, a guide RNA (gRNA), or a combination thereof. In embodiments, step (b) is repeated one or more times with a third or subsequent donor cell comprising a third or subsequent donor plasmid containing compatible HR regions and a third or subsequent oligonucleotide containing a third or subsequent DNA element fragment (oligo3, oligo4, ...OligoN) encodes, thereby forming a third or subsequent recombinant recipient oligonucleotide comprising the first, second, and a third or subsequent DNA element fragment, which together form a DNA assembly. In embodiments, step (a) comprises a plurality of first donor cells, each comprising a different first donor plasmid; and step (b) comprises a plurality of second, third, or subsequent donor cells, each comprising a different second, third, or subsequent donor plasmid; wherein optionally each first donor cell is located at a position in a first ordered array and each second, third, or subsequent donor cell is located at a position in a second, third, or subsequent ordered array; wherein the method optionally generates a combinatorial library comprising a plurality of different assembled DNA elements.
[0117] In embodiments, the donor plasmid, comprising the final DNA element forming part of a composite DNA element, includes a barcode-homologous recombination region (BHR) to generate recipient cells, each containing a recombinant recipient oligonucleotide comprising the composite DNA element, the BHR, and a further HR; and the method further comprises: (i) constructing or acquiring an array of barcode donor cells, each containing a barcode donor plasmid containing an HR homologous to the BHR, a specific barcode oligonucleotide, and a second HR homologous to the further HR of the recombinant recipient oligonucleotide;(ii) Contacting the array of barcode donor cells with an array of receiver cells under conditions to (a) transfer the barcode donor plasmids from the barcode donor cells to the receiver cells by conjugation and (b) recombine the barcode donor plasmids and receiver oligonucleotides in the receiver cells by homologous recombination, thereby generating an array of receiver cells comprising barcode assemblies.
[0118] In embodiments, each donor plasmid comprises a further pair of specific endonuclease sites CX, CY flanking a barcode-homologous recombination region (BHR), and the method further comprises contacting an array of recipient cells, each comprising a DNA compilation, with an array of barcode donor cells, each containing a barcode donor plasmid comprising a pair of HR regions homologous to the BHR and flanking a specific barcode oligonucleotide, to generate an array of recipient cells comprising barcoded compilations.
[0119] In embodiments, the DNA assembly methods may further comprise contacting a reset donor cell comprising a reset donor plasmid with a recipient cell comprising a recombinant recipient oligonucleotide, wherein the reset donor plasmid sequentially comprises a homologous recombination region (HRt) homologous to a terminal sequence of the DNA composition, a reset endonuclease site, a selectable marker, a reset endonuclease site, a homologous recombination region (HRX), and a transfer origin, wherein the recombinant recipient oligonucleotide sequentially comprises a reset endonuclease site, the DNA composition, a homologous recombination region homologous to HRX (HRXa), and a reset endonuclease site, whereby, following homologous recombination between the HRt and the terminal sequence of the DNA assembly and a reset plasmid is provided between HRX and HRXa,which includes the origin of transfer and DNA assembly. In some embodiments, the reset plasmid is located in a donor cell. In some embodiments, the reset plasmid contains a restricted origin of replication that functions in both donor and recipient cells. In some embodiments, the reset donor plasmid is constructed by a method that includes the introduction of an oligonucleotide insert comprising homologous recombination regions HRt, HRX flanking two endonuclease sites (C1, C2) and a counterselective marker (CM), HRt-C1-CM-C2-HRX; or a library of such oligonucleotide inserts; enabling an endonuclease to cleave the endonuclease sites; and introducing a counterselective marker at the cleavage sites using homologous recombination.
[0120] This document describes methods for conjugating barcodes to oligonucleotides, the method comprising: (a) inserting each oligonucleotide of a mixture of oligonucleotides into a donor plasmid, wherein each donor plasmid optionally comprises, in sequential order, a first endonuclease site (C1), a first homologous recombination region (HR1), a second homologous recombination region (HR2), and optionally a second endonuclease site (C2); wherein each oligonucleotide is inserted between HR1 and HR2, providing a plurality of donor plasmids comprising donor oligonucleotides, each donor plasmid comprising a single donor oligonucleotide from the mixture of oligonucleotides: C1-HR1-oligo-HR2-C2; (b) Transforming a large number of cells with the large number of donor plasmids so that each cell contains a donor plasmid, thereby forming a large number of donor cells;(c) Plating and cultivating the plurality of donor cells, each at a specific position on a first ordered array, thereby providing a first ordered array of donor cells; (d) providing a plurality of recipient cells in a second ordered array, each recipient cell comprising a recipient oligonucleotide comprising, in sequential order, a specific barcode sequence, wherein the specific barcode sequence identifies a position of the recipient cell in the second ordered array, a third homologous recombination region (HR3) homologous to HR1, optionally a third endonuclease site (C3), and a fourth homologous recombination region (HR4) homologous to HR2;(e) Contacting the first ordered array of donor cells with the second ordered array of recipient cells under conditions to (i) transfer the donor plasmids from the donor cells to the recipient cells at appropriate positions on the array by conjugation, (ii) optionally cleave the first, second and third endonuclease site, and (ii) transfer the oligonucleotides from the donor plasmids to the oligonucleotides of the recipient cells by homologous recombination, forming a third array of fusion oligonucleotides, each comprising a specific barcode sequence and a donor oligonucleotide from the oligonucleotide mixture;and (f) optional sequencing of the fusion oligonucleotides and thereby identifying each oligonucleotide in the array by its barcode sequence. The recipient oligonucleotide may be in a recipient cell plasmid or in the recipient cell genome. The donor plasmid may include a selectable marker between HR1 and HR2 that selects for the integration of the oligonucleotide into the recipient cell oligonucleotide; optionally, the donor plasmid includes a counter-selectable marker. The recipient cell oligonucleotide may include a fourth endonuclease site (C4).
[0121] In another aspect, a method for assembling a DNA element is provided. The procedure comprises: (a) providing a first host cell containing a first donor plasmid comprising, in sequential order: (i) a first endonuclease target site, (ii) a first homologous recombination region, (iii) optionally a first oligonucleotide containing a first DNA element fragment, (iv) a second homologous recombination region, (v) a second endonuclease target site, and (vi) a third endonuclease target site; (b) providing a recipient cell, wherein the recipient cell contains a recipient oligonucleotide comprising (i) a third homologous recombination region and (ii) a third homologous recombination region comprising: (i) a third homologous recombination region, wherein the third homologous region is homologous to the first homologous recombination region, (ii) a fourth endonuclease target site, and (iii) a fourth homologous region.wherein the fourth homologous recombination region is homologous to the second homologous recombination region; and (c) contacting the first host cell with the recipient cell under conditions to (i) transfer the first donor plasmid from the first host cell to the recipient cell by bacterial conjugation, (ii) direct a first endonuclease to the first endonuclease target site and / or the third endonuclease target site and / or the fourth endonuclease target site, thereby generating double-strand breaks in the first donor plasmid and the recipient oligonucleotide, and (iii) recombine the first donor plasmid and the recipient oligonucleotide in the recipient cell by homologous recombination via the first and second homologous recombination regions with the third and fourth corresponding homologous recombination regions, thereby forming a recombined recipient oligonucleotide. The procedure further includes: (d) providing a second host cell,which contains a second donor plasmid comprising, in successive order: (i) a fifth endonuclease target site, (ii) a fifth homologous recombination region homologous to the second homologous region, (iii) a second oligonucleotide containing a second DNA element fragment, (iv) a sixth homologous recombination region homologous to the fourth homologous region, and (v) a sixth endonuclease target site; (e) Contacting the second host cell with the recipient cell containing the recombinant recipient oligonucleotide under conditions to (i) transfer the second donor plasmid from the second host cell to the recipient cell by bacterial conjugation, (ii) express a second endonuclease, (iii) direct the second endonuclease to the second endonuclease target site, the fifth endonuclease target site and the sixth endonuclease target site,(iv) Recombining the second donor plasmid and the recombined recipient oligonucleotide in the recipient cell by homologous recombination via the fifth and sixth homologous recombination sites with the corresponding second and fourth homologous recombination sites, forming a second recombined recipient oligonucleotide containing an assembled DNA element. In embodiments, another part of the fourth homologous region is homologous to the sixth homologous recombination region compared to the part of the fourth homologous region that is homologous to the second homologous recombination region.
[0122] In embodiments, step (a) comprises a plurality of first host cells, each containing a specific first oligonucleotide. In embodiments, step (d) comprises a plurality of second host cells, each containing a specific second oligonucleotide. In embodiments, each first host cell contains a specific plasmid. In embodiments, each second host cell contains a specific plasmid. In embodiments, each first host cell is located in a position within a first ordered array. In embodiments, a plurality of first host cells are located in a position within a first ordered array, thereby forming a pool of first host cells in the first ordered array. In embodiments, each second host cell is located in a position within a second ordered array.In embodiments, several second host cells are located in one position in a second ordered array, thereby forming a pool of second host cells in each position in the second ordered array.
[0123] In embodiments, the first donor cell is located in a first ordered array, the second donor cell in a second ordered array, and one or more subsequent donor cells in one or more subsequent arrays. Thus, in some embodiments, the method presented here generates a variant library containing a multitude of different composite DNA elements. In embodiments, the variant library is generated by 1) independently producing each variant using a first host cell, a second host cell, or a subsequent host cell at a position in a first, second, or subsequent array, or 2) generating a variant pool using a multitude of first host cells, second host cells, or subsequent host cells at a position in a first, second, or subsequent array.In embodiments, the variant library is generated by independently producing each variant using a first host cell, a second host cell, or a subsequent host cell at a position in a first, second, or subsequent array. In some embodiments, the variant library is generated by creating a variant pool using a plurality of first host cells, second host cells, or subsequent host cells at a position in a first, second, or subsequent array. For example, in the method presented here, including its embodiments, the first DNA element and / or the second DNA element can be a DNA barcode or a plurality of DNA barcodes. In embodiments, the method generates a recursive barcoding platform.In embodiments where the first DNA element and / or the second DNA element is a DNA barcode or a plurality of DNA barcodes, the method can be used, for example, to track cell lines.
[0124] For example, the first DNA element and / or DNA elements can be a gRNA or a multitude of gRNAs. Therefore, in some embodiments, the method includes the generation of combinatorial gRNA libraries.
[0125] In some embodiments, the first endonuclease targets the first endonuclease target site. In some embodiments, the first endonuclease targets the third endonuclease target site. In some embodiments, the first endonuclease targets the fourth endonuclease target site. In some embodiments, the second endonuclease targets the second endonuclease target site. In some embodiments, the second endonuclease targets the fifth endonuclease target site. In some embodiments, the second endonuclease targets the sixth endonuclease target site.
[0126] In some embodiments, the DNA element is a gene. In some embodiments, the DNA element is a promoter. In some embodiments, the DNA element is an enhancer. In embodiments, the DNA element is a terminator. In embodiments, the DNA element is an intron. In embodiments, the DNA element is an intergenic region. In embodiments, the DNA element is a barcode. In embodiments, the DNA element is a translation initiation site. In embodiments, the DNA element is a gRNA. In embodiments, the DNA element is a fragment of one of the aforementioned elements.
[0127] In some embodiments, the recipient oligonucleotide is located in a recipient plasmid. In other embodiments, the recipient oligonucleotide is located in the genome of the recipient cell.
[0128] In some embodiments, the second donor plasmid also contains a seventh homologous recombination region and a seventh endonuclease target site between the components of (d) iii) and (d) iv). In some embodiments, the first endonuclease targets the seventh endonuclease target site.
[0129] In some embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA, and the recipient cell contains an oligonucleotide encoding an RNA-gated DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA, and the genome of the recipient cell contains an oligonucleotide encoding an RNA-gated DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA, and the recipient plasmid contains an oligonucleotide encoding an RNA-gated DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide that codes for a gRNA, and a recipient helper plasmid contains an oligonucleotide that codes for an RNA-guided DNA endonuclease.In some embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA, and the donor plasmid contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA, and the recipient genome contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a gRNA, and the recipient plasmid contains an oligonucleotide encoding an inducible RNA-guided DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide that codes for a gRNA, and a recipient helper plasmid contains an oligonucleotide that codes for an inducible RNA-directed DNA endonuclease.
[0130] In some embodiments, the recipient cell contains an inducible gRNA. In some embodiments, the donor cell contains an oligonucleotide encoding an RNA-directed DNA endonuclease. In some embodiments, the recipient cell contains an oligonucleotide encoding an RNA-directed DNA endonuclease. In some embodiments, the expression of the RNA-directed DNA is constitutive. In some cases, the expression of the RNA-directed DNA is inducible.
[0131] In some embodiments, the RNA-directed DNA endonuclease is Cas9. In other embodiments, the RNA-directed DNA endonuclease is Cas10. In other embodiments, the RNA-directed DNA endonuclease is Cpf1. In other embodiments, the RNA-directed DNA endonuclease is C2c1. In other embodiments, the RNA-directed DNA endonuclease is C2c2. In other embodiments, the RNA-directed DNA endonuclease is C2c3. In some embodiments, the RNA-directed DNA endonuclease is Cas12c1. In some cases, the RNA-directed DNA endonuclease is Cas12a. In some cases, the RNA-directed DNA endonuclease is Cas12b. In some embodiments, the RNA-directed DNA endonuclease is Cas12c2. In some cases, the RNA-directed DNA endonuclease is Cas12g. In some cases, the RNA-directed DNA endonuclease is Cas12e.In some embodiments, the RNA-directed DNA endonuclease is Cas12i1. In some embodiments, the RNA-directed DNA endonuclease is Cas12i2.
[0132] In the methods described here, in some embodiments, an oligonucleotide encoding the first endonuclease is located in the donor plasmid. In some embodiments, an oligonucleotide encoding the first endonuclease is the recipient oligonucleotide. In some embodiments, an oligonucleotide encoding the first endonuclease is located in a recipient cell helper plasmid. In some embodiments, an oligonucleotide encoding the first endonuclease is located in the recipient genome. In some embodiments, the expression of the first endonuclease is induced. In some embodiments, the oligonucleotide encoding the second endonuclease is located in the donor plasmid. In some embodiments, an oligonucleotide encoding the second endonuclease is located in the recipient oligonucleotide.In some embodiments, an oligonucleotide encoding the second endonuclease is located in a recipient cell helper plasmid.
[0133] In some embodiments, the first endonuclease and / or the second endonuclease is a homing endonuclease. In embodiments, the homing endonuclease is I-ScaI, PI-SceI, I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI, I-PpoI, PI-PspI, I-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I, or I-Vdi141I. In some embodiments, the homing endonuclease is I-ScaI. In some embodiments, the homing endonuclease is PI-SceI. In some embodiments, the homing endonuclease is I-AniI. In some embodiments, the homing endonuclease is I-CeuI. In some embodiments, the homing endonuclease is I-ChuI. In some embodiments, the homing endonuclease is I-CpaI. In some cases, the homing endonuclease is I-CpaII. In some embodiments, the homing endonuclease is I-CreI.In some embodiments, the homing endonuclease is I-DmoI. In some embodiments, the homing endonuclease is H-DreI. In some embodiments, the homing endonuclease is I-HmuI. In some embodiments, the homing endonuclease is I-HmuII. In some embodiments, the homing endonuclease is I-LlaI. In some cases, the homing endonuclease is I-MsoI. In some embodiments, the homing endonuclease is PI-PfuI. In some embodiments, the homing endonuclease is PI-PkoII. In some embodiments, the homing endonuclease is I-PorI. In some embodiments, the homing endonuclease is I-PpoI. In some embodiments, the homing endonuclease is PI-PspI. In some embodiments, the homing endonuclease is I-SceI. In some embodiments, the homing endonuclease is I-SceII. In some embodiments, the homing endonuclease is I-SceIII.In some embodiments, the homing endonuclease is I-SceIV. In some embodiments, the homing endonuclease is I-SceV. In some embodiments, the homing endonuclease is I-SceVI. In some embodiments, the homing endonuclease is I-SceVII. In some cases, the homing endonuclease is I-Ssp6803I. In some embodiments, the homing endonuclease is I-TevI. In some embodiments, the homing endonuclease is I-TevII. In some embodiments, the homing endonuclease is I-TevIII. In some embodiments, the homing endonuclease is PI-T1I. In some embodiments, the homing endonuclease is PI-T1II. In some embodiments, the homing endonuclease is I-Tsp061I or I-Vdi141I. In some cases, the homing endonuclease is I-Vdi141I.
[0134] In some embodiments, the endonuclease is a transcription activator-like effector nuclease. In some embodiments, the endonuclease is a zinc finger nuclease.
[0135] In the methods presented here, steps (d) to (e) are repeated in one or more iterations, thereby forming one or more consecutive composite DNA elements. In some embodiments, the first, second, or subsequent donor plasmid contains a selectable marker that selects for the integration of the first, second, or subsequent oligonucleotide into the oligonucleotide of the recipient cell. In some embodiments, the first donor plasmid contains a selectable marker that selects for the integration of the first oligonucleotide into the oligonucleotide of the recipient cell. In some embodiments, the second donor plasmid contains a selectable marker that selects for the integration of the second oligonucleotide into the oligonucleotide of the recipient cell.In some embodiments, the subsequent donor plasmid contains a selectable marker that selects for the integration of the subsequent oligonucleotide into the oligonucleotide of the recipient cell.
[0136] The first donor plasmid may contain a selectable marker that selects for the integration of the first oligonucleotide into the recipient oligonucleotide. The selectable marker may be located between the components of (a)(v) and (a)(iv). The second donor plasmid may contain a selectable marker that selects for the integration of the second oligonucleotide into the recipient oligonucleotide.
[0137] In some embodiments, the recipient cell oligonucleotide contains a counterselective marker that selects for the integration of the first oligonucleotide into the recipient cell oligonucleotide. In some embodiments, the recombinant recipient cell oligonucleotide contains a counterselective marker that selects for the integration of the second or subsequent oligonucleotide into the recombinant recipient cell oligonucleotide. Counterselective markers for use in the processes described herein are described above.
[0138] In some embodiments, the assembled DNA element, which can also be called a DNA assembly, is sequenced. In some embodiments, the recipient oligonucleotide is sequenced. In some embodiments, the recombinant recipient oligonucleotide is sequenced. In some embodiments, the recombinant recipient oligonucleotide is a plasmid, wherein the plasmid is linearized, ligated to sequencing adapters, and sequenced. In some embodiments, the assembled DNA element is amplified and sequenced by PCR. In some embodiments, (a) the recipient cells are lysed, (b) the oligonucleotides are digested with one or more endonucleases, (c) the assembled DNA element is isolated, and (d) the assembled DNA element or multiple assembled genes are ligated to sequencing adapters and sequenced. In some embodiments, the assembled DNA element is isolated.In some embodiments, the recombinant receiver oligonucleotide is isolated.
[0139] In embodiments, the composite DNA element has a length of approximately 100 nucleotides to approximately 500,000 nucleotides. The length can be any value or subrange within the specified ranges, including the endpoints.
[0140] In embodiments, the composite DNA element comprises approximately 100 nucleotides, approximately 1,000 nucleotides, approximately 10,000 nucleotides, approximately 20,000 nucleotides, approximately 40,000 nucleotides, approximately 60,000 nucleotides, approximately 80,000 nucleotides, approximately 100,000 nucleotides, approximately 120,000 nucleotides, approximately 140,000 nucleotides, approximately 160,000 nucleotides, approximately 180,000 nucleotides, approximately 20,000 nucleotides, approximately 240,000 nucleotides, approximately 260,000 nucleotides, approximately 280,000 nucleotides, approximately 300,000 nucleotides, approximately 320,000 nucleotides, approximately 340,000 nucleotides, approximately The length can be 360,000 nucleotides, approximately 380,000 nucleotides, approximately 400,000 nucleotides, approximately 420,000 nucleotides, approximately 440,000 nucleotides, approximately 460,000 nucleotides, approximately 480,000 nucleotides, or approximately 500,000 nucleotides. The length can be any value or subrange within the specified ranges, including the endpoints.
[0141] In embodiments, the first, second, or subsequent homology regions and their corresponding first, second, or subsequent homology regions are approximately 20 to approximately 500 base pairs long. The length can be any value or subrange within the specified ranges, including the endpoints.
[0142] In embodiments, the first, second, or subsequent homology regions and the corresponding first, second, or subsequent homology regions are approximately 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500 base pairs. Base pairs long. In embodiments, the first, second, or subsequent homology region and the corresponding first, second, or subsequent homology region have a length of approximately 50 base pairs. The length can be any value or subrange within the specified ranges, including the endpoints. III. Analytical methods
[0143] This document describes a method for identifying an oligonucleotide from a mixture of oligonucleotides. The method comprises: (a) providing a mixture of oligonucleotides, (b) inserting each oligonucleotide into a donor plasmid, each donor plasmid comprising, in successive order, i) a first endonuclease junction, ii) a first homologous recombination region, iii) a second homologous recombination region, and iv) a second endonuclease junction, the oligonucleotide being inserted between the first homologous recombination region and the second homologous recombination region, thereby generating a plurality of donor plasmids, each donor plasmid containing a single oligonucleotide from the oligonucleotide mixture;(c) Transforming a plurality of host cells with the plurality of donor plasmids such that each host cell contains one donor plasmid, thereby forming a plurality of transformed host cells; (d) Plating and cultivating the plurality of transformed host cells on a first ordered array, wherein each transformed host cell generates a colony of clones in the first ordered array;(e) Providing a plurality of receiver cells in a second ordered array, each receiver cell containing a receiver oligonucleotide which includes, in sequential order: (i) a specific barcode sequence, wherein the specific barcode sequence identifies a position of the receiver cell in the second ordered array, (ii) a corresponding first homologous recombination region, wherein the first homologous recombination region is homologous to the corresponding first homologous recombination region, (iii) a third endonuclease site, and (iv) a corresponding second homologous recombination site, wherein the second homologous recombination region is homologous to the corresponding second homologous recombination region, and wherein the first endonuclease site, the second endonuclease site and the third endonuclease site are cleaveable by an endonuclease;(f) Contacting each colony of clones from the first ordered array with a recipient cell at a corresponding position in the second ordered array under conditions which (i) transfer the donor plasmid from the colony of clones to the recipient cell by bacterial conjugation, (ii) cleave the first, second and third endonuclease cleavage sites, (iii) cleave the first, second and third endonuclease cleavage sites and (iv) transfer the donor plasmid into the recipient cell, (i) the donor plasmid is transferred from the clone colony to the recipient cell by bacterial conjugation, (ii) the first, second and third endonuclease cleavage sites are cleaved by the endonuclease and (ii) the oligonucleotide is transferred from the donor plasmid to the oligonucleotide of the recipient cell by homologous recombination, producing a fusion sequence containing the barcode sequence and the oligonucleotide; (g) Sequencing the fusion sequence;and identifying the sequenced oligonucleotide in the first and / or second ordered array of donor cells and / or recipient cells by identifying the barcode sequence.
[0144] Plating and cultivating cells can involve plating and cultivating cells on a surface. This surface can be, for example, a solid medium. Thus, a colony of clones can be a colony of cells on a solid medium. Plating and cultivating cells can also involve plating and cultivating cells in a liquid medium. For example, a single cell can be plated and cultured in a liquid medium. Therefore, a colony of clones can be a colony of cells in a liquid medium.
[0145] The recipient oligonucleotide can be located in a recipient cell plasmid. The recipient oligonucleotide can be located in the recipient cell genome. The donor plasmid can contain a selectable marker between the first homologous recombination region and the second homologous recombination region, which selects for the integration of the oligonucleotide into the recipient cell oligonucleotide. The recipient cell oligonucleotide can contain two endonuclease cleavage sites in step (e)(iii). The method can further include a counter-selectable marker between the two endonuclease cleavage sites, wherein the counter-selectable marker selects for the integration of the oligonucleotide into the recipient oligonucleotide.
[0146] This document describes a method for identifying an oligonucleotide from a mixture of oligonucleotides. The method comprises: (a) providing a plurality of host cells in a first ordered array, each host cell containing a donor plasmid, each donor plasmid containing, in sequential order: (i) a first endonuclease cleavage site, (ii) a first homologous recombination region, (iii) a specific barcode sequence, (iv) a second homologous recombination region, and (v) a second endonuclease cleavage site; wherein the specific barcode sequence identifies a position of the host cell in the first ordered array;(b) Providing a plurality of recipient cells, each recipient cell including a recipient oligonucleotide which includes an oligonucleotide from the plurality of oligonucleotides, each recipient plasmid sequentially including i) the oligonucleotide sequence, ii) a corresponding first homologous recombination region, wherein the first homologous recombination region is homologous to the corresponding first homologous recombination region, iii) a third endonuclease junction, wherein the first endonuclease junction, the second endonuclease junction and the third endonuclease junction can be cleaved by an endonuclease, and iv) a corresponding second homologous recombination region, wherein the second homologous recombination region is homologous to the corresponding second homologous recombination region;(c) Plating and cultivating the plurality of recipient cells on a second ordered array, each recipient cell generating a colony of clones in the second ordered array; (d) contacting a donor cell from the first ordered array with each clone colony at a corresponding position in the second ordered array under conditions that (i) transfer the donor plasmid from a donor cell to the clone colony by bacterial conjugation, (ii) cleave the first, second, and third endonucleases, (iii) cleave the first, second, and third endonuclease cleavage sites by the endonuclease, and (iii) transfer the barcode sequence from the donor plasmid to the oligonucleotide of the recipient cell by homologous recombination, generating a fusion sequence containing the barcode sequence and the oligonucleotide; (e) sequencing the fusion sequence;and (f) identification of the sequenced oligonucleotide in the first and / or second ordered array of recipient cells by identification of the barcode sequence.;
[0147] The recipient oligonucleotide can be located in a plasmid of the recipient cell.
[0148] The recipient oligonucleotide can be located in the genome of the recipient cell. The donor plasmid can contain a selectable marker between the first homologous recombination region and the second homologous recombination region, which selects for the integration of the barcode into the recipient oligonucleotide. The recipient oligonucleotide can contain two endonuclease cleavage sites in step (b)(iii). The procedure can also include a cross-selectable marker between the two endonuclease cleavage sites, which selects for the integration of the barcode into the oligonucleotide of the recipient cell.
[0149] In the methods presented here, in some embodiments the first, second, and third endonuclease interfaces are the same endonuclease interface. In other embodiments, the first, second, and third endonuclease interfaces are different endonuclease interfaces. In some embodiments, the endonuclease comprises multiple endonucleases.
[0150] The endonuclease can be encoded by an oligonucleotide in the recipient cell. The oligonucleotide can be located in the recipient cell's genome. The oligonucleotide can be located in the recipient plasmid. The oligonucleotide encoding an endonuclease can be located in a helper plasmid. The endonuclease can be encoded by an oligonucleotide in the donor plasmid. The endonuclease can be encoded by an oligonucleotide in the donor plasmid. Endonuclease expression can be inducible.
[0151] The endonuclease can be a transcription activator-like effector nuclease. The endonuclease can be a zinc-finger nuclease. The endonuclease can be HO.
[0152] The RNA-directed DNA endonuclease can be a CRISPR system. The endonuclease can be an RNA-directed DNA endonuclease. The RNA-directed DNA endonuclease can be Cas9. The RNA-directed DNA endonuclease can be Cas10. The RNA-directed DNA endonuclease can be Cpf1. The RNA-directed DNA endonuclease can be C2c1. The RNA-directed DNA endonuclease can be C2c2. The RNA-directed DNA endonuclease can be C2c3. The RNA-directed DNA endonuclease can be Cas12c1. The RNA-directed DNA endonuclease can be Cas12a. The RNA-directed DNA endonuclease can be Cas12b. The RNA-directed DNA endonuclease could be Cas12c2. The RNA-directed DNA endonuclease could be Cas12g. The RNA-directed DNA endonuclease could be Cas12e. The RNA-directed DNA endonuclease could be Cas12i1.The RNA-directed DNA endonuclease could be Cas12i2.
[0153] In the methods presented here, the donor plasmid contains, in some embodiments, an oligonucleotide encoding a guide RNA, and the recipient genome contains an oligonucleotide encoding an RNA-directed DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a guide RNA, and the recipient plasmid contains an oligonucleotide encoding an RNA-directed DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a guide RNA, and the recipient helper plasmid contains an oligonucleotide encoding an RNA-directed DNA endonuclease. In some embodiments, the donor plasmid contains an oligonucleotide encoding a guide RNA, and the recipient plasmid contains an oligonucleotide encoding an inducible RNA-directed DNA endonuclease.
[0154] The recipient cell may contain an inducible gRNA. The donor cell may contain an oligonucleotide encoding an RNA-directed DNA endonuclease. The expression of the RNA-directed DNA may be constitutive. The expression of the RNA-directed DNA may be inducible.
[0155] The method may further include the isolation of the donor plasmid. The method may further include the isolation of the recipient plasmid. The method may further include the isolation of the recombinant recipient plasmid. The method may further include the isolation of the sequenced oligonucleotide. The method may further include the isolation of the donor cell, the recipient cell, the recombinant recipient cell, the recipient oligonucleotide, or the recombinant recipient oligonucleotide, or several of these.
[0156] The method can involve combining one or more subsets of colonies. For example, the method can involve combining one or more subsets of colonies and isolating a variety of donor plasmids from the subset of colonies. The method can involve combining one or more subsets of colonies and isolating a variety of recipient plasmids from the subset of colonies. The method can involve combining one or more subsets of colonies and isolating a variety of recombinant recipient plasmids from the subset of colonies. The method can further involve isolating a variety of sequenced oligonucleotides.
[0157] The recipient cell or the donor cell can contain an oligonucleotide that enables plasmid conjugation. This oligonucleotide can be located in the donor cell's genome. It can also be located in a helper plasmid. The oligonucleotide that enables plasmid conjugation can be the tra operon. Other oligonucleotides that enable plasmid conjugation are described above.
[0158] Donor cells, recipient cells, or recombinant recipient cells can be transferred to positions on a third ordered array, a fourth ordered array, or a subsequent ordered array.
[0159] The barcode sequence can be approximately 4 to 50 nucleotides long. The barcode sequence can be approximately 8 to 50 nucleotides long. The barcode sequence can be approximately 12 to 50 nucleotides long. The barcode sequence can be approximately 16 to 50 nucleotides long. The barcode sequence can be approximately 20 to 50 nucleotides long. The barcode sequence can be approximately 24 to 50 nucleotides long. The barcode sequence can be approximately 28 to 50 nucleotides long. The barcode sequence can be approximately 32 to 50 nucleotides long. The barcode sequence can be approximately 36 to 50 nucleotides long. The length of the barcode can be any value or subrange within the ranges specified here, including the endpoints.
[0160] The barcode sequence can have a length of approximately 4 to 36 nucleotides. The barcode sequence can have a length of approximately 4 to 32 nucleotides. The barcode sequence can have a length of approximately 4 to 28 nucleotides. The barcode sequence can have a length of approximately 4 to 24 nucleotides. The barcode sequence can have a length of approximately 4 to 20 nucleotides. The barcode sequence can have a length of approximately 4 to 16 nucleotides. The barcode sequence can have a length of approximately 4 to 12 nucleotides. The barcode sequence can have a length of approximately 4 to 8 nucleotides. The barcode sequence can be approximately 4, 8, 12, 16, 20, 24, 28, 32, 36, or 40 nucleotides long. The barcode sequence can also be approximately 15 nucleotides long.The barcode sequence can be approximately 40 nucleotides long. The length of the barcode can be any value or subrange within the ranges specified here, including the endpoints. Other embodiments
[0161] The invention is further described by the following additional embodiments.
[0162] Embodiment 1: Method for assembling a DNA element, the method comprising: (1) providing a first host cell comprising a first donor plasmid comprising, in the following order: a first endonuclease target site, a first homologous recombination region, a first oligonucleotide comprising a first DNA element fragment, a second homologous recombination region comprising two homologous recombination regions, a second endonuclease target site and a third endonuclease target site;(2) Providing a recipient cell, wherein the recipient cell comprises a recipient oligonucleotide comprising a third homologous recombination region, wherein the third homologous region is homologous to the first homologous recombination region, a fourth endonuclease target site and a fourth homologous region, wherein the fourth homologous recombination region is homologous to the second homologous recombination region, wherein the recipient oligonucleotide is contained in a recipient cell plasmid or the recipient cell genome;and (3) bringing the first host cell into contact with the recipient cell under conditions to (i) transfer the first donor plasmid from the first host cell to the recipient cell by bacterial conjugation, (ii) direct a first endonuclease to the first endonuclease target site and / or the third endonuclease target site and / or the fourth endonuclease target site, thereby generating double-strand breaks in the first donor plasmid and the recipient oligonucleotide, and (iii) recombine the first donor plasmid and the recipient oligonucleotide in the recipient cell by homologous recombination via the first and second homologous recombination regions with the third and fourth corresponding homologous recombination regions, thereby forming a recombined recipient oligonucleotide;(4) Providing a second host cell comprising a second donor plasmid comprising, in successive order, a fifth endonuclease target site, a fifth homologous recombination region homologous to the second homologous region, a second oligonucleotide encoding a second DNA element fragment, a sixth homologous recombination region comprising two homologous recombination regions homologous to the fourth homologous region, and a sixth endonuclease target site;and (5) bringing the second host cell into contact with the recipient cell containing the recombined recipient oligonucleotide under conditions to (i) transfer the second donor plasmid from the second host cell to the recipient cell by bacterial conjugation, (ii) express a second endonuclease, (iii) direct the second endonuclease to the second endonuclease target site, the fifth endonuclease target site and / or the sixth endonuclease target site, (iv) recombine the second donor plasmid and the recombined recipient oligonucleotide in the recipient cell by homologous recombination via the fifth and sixth homologous recombination sites with the corresponding second and fourth homologous recombination sites, forming a second recombined recipient oligonucleotide comprising an assembled DNA element;wherein an oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third and / or fourth endonuclease site is present on the first donor plasmid and / or wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth and / or sixth endonuclease site is present on the second donor plasmid.;
[0163] In further embodiments, step (a) comprises a plurality of first host cells, each cell comprising a different first oligonucleotide, and / or step (d) comprises a plurality of second host cells, each cell comprising a different second oligonucleotide.
[0164] In further embodiments, each first host cell is located at a position in a first ordered array.
[0165] In further embodiments, several first host cells are located in one position in a first ordered array, thereby forming a pool of first host cells in the first ordered array.
[0166] In other embodiments, every second host cell is located at a position in a second ordered array.
[0167] In further embodiments, several second host cells are located in one position in a second ordered array, thereby forming a pool of second host cells in each position in the second ordered array.
[0168] In further embodiments, the first donor cell is located in a first ordered array, the second donor cell in a second ordered array, and one or more subsequent donor cells in one or more subsequent arrays.
[0169] In further embodiments, the method generates a combinatorial library comprising a variety of different composite DNA elements.
[0170] In other embodiments, the first endonuclease targets the first, third, or fourth endonuclease target site.
[0171] In other embodiments, the second endonuclease targets the second, fifth, or sixth endonuclease target site.
[0172] In other embodiments, the DNA element is a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, or a gRNA.
[0173] In other embodiments, the receiver oligonucleotide is located in a receiver plasmid.
[0174] In other embodiments, the receiver oligonucleotide is located in the genome of the receiver cell.
[0175] In further embodiments, the second donor plasmid also comprises a seventh homologous recombination region and a seventh endonuclease target site between the components of (d) iii) and (d) iv).
[0176] In further embodiments, the first endonuclease targets the seventh endonuclease site.
[0177] In further embodiments, the first and second endonucleases are independently selected from an RNA-directed DNA endonuclease, a homing endonuclease, a transcription activator-like effector nuclease, and a zinc finger nuclease.
[0178] In other embodiments, an oligonucleotide encoding the first endonuclease is located in the donor cell or in the recipient cell.
[0179] In further embodiments, the expression of the first and / or second endonuclease can be induced.
[0180] In other embodiments, an oligonucleotide encoding the second endonuclease is located in the donor cell or in the recipient cell.
[0181] In further embodiments, steps (d) to (e) are repeated for one or more iterations, thereby forming one or more successive composite DNA elements.
[0182] In further embodiments, the first, second or subsequent donor plasmid comprises a selectable marker that selects for the integration of the first oligonucleotide, second oligonucleotide or subsequent oligonucleotide into the oligonucleotide of the recipient cell.
[0183] In further embodiments, the first donor plasmid comprises a selectable marker that selects for the integration of the first oligonucleotide into the recipient oligonucleotide, the selectable marker optionally being located between the second and third endonuclease target sites. In some embodiments, the second donor plasmid comprises a selectable marker that selects for the integration of the second oligonucleotide into the recipient oligonucleotide.
[0184] In further embodiments, the recipient cell oligonucleotide comprises a counter-selectable marker that selects for the integration of the first oligonucleotide into the recipient cell oligonucleotide.
[0185] In further embodiments, the donor plasmid comprises a transfer origin, wherein the transfer origin may optionally be from a mobile element.
[0186] In further embodiments, the donor plasmid comprises a conditional origin of replication, wherein the conditional origin of replication may optionally depend on the presence of an oligonucleotide or a state of cell growth.
[0187] In further embodiments, the donor plasmid or recipient oligonucleotide comprises a replicon that can replicate plasmids with a length of more than 30 kilobases.
[0188] In further embodiments, the donor plasmid or recipient oligonucleotide comprises an inducible high-copy origin of replication.
[0189] In further embodiments, the donor plasmid or recipient oligonucleotide is an artificial yeast chromosome (YAC), an artificial mammalian chromosome (MAC), an artificial human chromosome (HAC), or an artificial plant chromosome. In further embodiments, the donor plasmid or recipient oligonucleotide is a viral vector.
[0190] In other embodiments, the donor cell comprises an oligonucleotide that enables the conjugation of plasmids.
[0191] In further embodiments, the donor cell or the recipient cell comprises an oligonucleotide encoding one or more homologous DNA repair genes, wherein the expression of homologous DNA repair genes is optionally induced.
[0192] In further embodiments, the donor or recipient cell comprises an oligonucleotide that codes for one or more recombination-mediated genetic engineering genes.
[0193] In other embodiments, the donor cell and the recipient cell are independently of each other a bacterial cell.
[0194] In further embodiments, the composite DNA element is sequenced, the recipient oligonucleotide is sequenced and / or the recombinant recipient oligonucleotide is sequenced.
[0195] In further embodiments, the recombinant recipient oligonucleotide is a plasmid, and the plasmid is linearized, ligated to sequencing adapters, and sequenced.
[0196] In further embodiments, the assembled DNA element is amplified and sequenced by PCR, optionally (a) the recipient cells are lysed, (b) the oligonucleotides are digested with one or more endonucleases, (c) the assembled DNA element is isolated, and (d) the assembled DNA element or several assembled genes are ligated to sequencing adapters and sequenced.
[0197] In further embodiments, the composite DNA element or recombinant recipient oligonucleotide is isolated.
[0198] In other embodiments, the assembled DNA fragment has a length of 100 nucleotides to 500,000 nucleotides.
[0199] In further embodiments, the first, second or subsequent homology regions and the corresponding first, second or subsequent homology regions are approximately 20 to approximately 500 base pairs long.
[0200] In further embodiments, the first, second or subsequent homology region and the corresponding first, second or subsequent homology region have a length of approximately 50 base pairs.
[0201] This document describes a method for identifying an oligonucleotide from a mixture of oligonucleotides, the method comprising: (a) providing a mixture of oligonucleotides, (b) inserting each oligonucleotide into a donor plasmid, each donor plasmid comprising in successive order: (i) a first endonuclease junction, (ii) a first homologous recombination region, (iii) a second homologous recombination region, and (iv) a second endonuclease junction, the oligonucleotide being inserted between the first homologous recombination region and the second homologous recombination region, thereby generating a plurality of donor plasmids, each donor plasmid comprising a single oligonucleotide from the mixture of oligonucleotides;(c) Transforming a plurality of host cells with the plurality of donor plasmids such that each host cell contains one donor plasmid, thereby forming a plurality of transformed host cells; (d) Plating and cultivating the plurality of transformed host cells on a first ordered array, wherein each transformed host cell generates a colony of clones in the first ordered array;(e) Providing a plurality of receiver cells in a second ordered array, each receiver cell comprising a receiver oligonucleotide comprising, in sequential order: (i) a specific barcode sequence, wherein the specific barcode sequence identifies a position of the receiver cell in the second ordered array, (ii) a corresponding first homologous recombination region, wherein the first homologous recombination region is homologous to the corresponding first homologous recombination region, (iii) a third endonuclease site, and (iv) a corresponding second homologous recombination site, wherein the second homologous recombination region is homologous to the corresponding second homologous recombination region, wherein the first endonuclease site, the second endonuclease site, and the third endonuclease site are cleaved by an endonuclease;(f) Contacting each colony of clones from the first ordered array with a recipient cell at a corresponding position in the second ordered array under conditions which (i) transfer the donor plasmid from the colony of clones to the recipient cell by bacterial conjugation, (ii) cleave the first, second and third endonuclease cleavage sites, (iii) cleave the first, second and third endonuclease cleavage sites and (iv) transfer the donor plasmid into the recipient cell, (i) the donor plasmid is transferred from the clone colony to the recipient cell by bacterial conjugation, (ii) the first, second and third endonuclease cleavage sites are cleaved by the endonuclease and (ii) the oligonucleotide is transferred from the donor plasmid to the oligonucleotide of the recipient cell by homologous recombination, producing a fusion sequence comprising the barcode sequence and the oligonucleotide; (g) Sequencing of the fusion sequence;and (h) identification of the sequenced oligonucleotide in the first and / or second ordered array of donor cells and / or recipient cells by identification of the barcode sequence.;
[0202] The recipient oligonucleotide can be located in a plasmid of the recipient cell.
[0203] The recipient oligonucleotide can be located in the genome of the recipient cell.
[0204] The donor plasmid may include a selectable marker between the first homologous recombination region and the second homologous recombination region, which selects for the integration of the oligonucleotide into the oligonucleotide of the recipient cell.
[0205] The recipient cell oligonucleotide can include two endonuclease cleavage sites in step (e)(iii).
[0206] The method can further include a counter-selectable marker between the two endonuclease interfaces, wherein the counter-selectable marker selects for the integration of the oligonucleotide into the recipient oligonucleotide.
[0207] This document describes a method for identifying one oligonucleotide from a plurality of oligonucleotides, the method comprising: (a) providing a plurality of host cells in a first ordered array, each host cell comprising a donor plasmid, each donor plasmid comprising in successive order: (i) a first endonuclease cleavage site, (ii) a first homologous recombination region, (iii) a specific barcode sequence, (iv) a second homologous recombination region, and (v) a second endonuclease cleavage site; wherein the specific barcode sequence identifies a position of the host cell in the first ordered array;(b) Providing a plurality of recipient cells, each recipient cell comprising a recipient oligonucleotide comprising an oligonucleotide from the plurality of oligonucleotides, each recipient plasmid comprising in sequential order: i) the oligonucleotide sequence, ii) a corresponding first homologous recombination region, wherein the first homologous recombination region is homologous to the corresponding first homologous recombination region, iii) a third endonuclease junction, wherein the first endonuclease junction, the second endonuclease junction, and the third endonuclease junction can be cleaved by an endonuclease, and iv) a corresponding second homologous recombination region, wherein the second homologous recombination region is homologous to the corresponding second homologous recombination region;(c) Plating and cultivating the plurality of recipient cells on a second ordered array, each recipient cell generating a colony of clones in the second ordered array; (d) contacting a donor cell from the first ordered array with each clone colony at a corresponding position in the second ordered array under conditions that (i) transfer the donor plasmid from a donor cell to the clone colony by bacterial conjugation, (ii) cleave the first, second, and third endonucleases, (iii) cleave the first, second, and third endonuclease cleavage sites by the endonuclease, and (iii) transfer the barcode sequence from the donor plasmid to the oligonucleotide of the recipient cell by homologous recombination, resulting in a fusion sequence comprising the barcode sequence and the oligonucleotide; (e) sequencing the fusion sequence;and (f) identification of the sequenced oligonucleotide in the first and / or second ordered array of donor cells and / or recipient cells by identification of the barcode sequence.;
[0208] The recipient oligonucleotide can be located in a plasmid of the recipient cell.
[0209] The recipient oligonucleotide can be located in the genome of the recipient cell.
[0210] The donor plasmid may include a selectable marker between the first homologous recombination region and the second homologous recombination region, which selects for the integration of the barcode into the recipient oligonucleotide.
[0211] The receiver oligonucleotide can include two endonucleotide interfaces in step (b)(iii).
[0212] The methods may include the provision of a cross-selectable marker between the two endonuclease interfaces, which selects for the integration of the barcode into the receiver cell oligonucleotide.
[0213] The first endonuclease interface, the second endonuclease interface, and the third endonuclease interface can be the same endonuclease interface.
[0214] The first endonuclease interface, the second endonuclease interface, and the third endonuclease interface can be different endonuclease interfaces.
[0215] The endonuclease can comprise multiple endonucleases.
[0216] The donor plasmid can include a transfer origin.
[0217] The transmission can be done from a mobile element.
[0218] The donor plasmid may include a conditional origin of replication.
[0219] The conditional origin of replication can depend on the presence of an oligonucleotide.
[0220] The conditional origin of replication can depend on a state of cell growth.
[0221] The donor plasmid or recipient plasmid may include a replicon capable of replicating plasmids with a length of at least 30 kilobases.
[0222] The replicon can originate from an artificial chromosome derived from P1 or from a bacterial artificial chromosome.
[0223] The donor plasmid or the recipient cell oligonucleotide may include an inducible high-copy replication origin.
[0224] The donor plasmid or oligonucleotide of the recipient cell can be an artificial yeast chromosome (YAC), an artificial mammalian chromosome (MAC), an artificial human chromosome (HAC), or an artificial plant chromosome.
[0225] The donor plasmid or the recipient oligonucleotide can be a viral vector.
[0226] The endonuclease can be encoded by an oligonucleotide in the recipient cell.
[0227] The endonuclease can be encoded by an oligonucleotide in the donor plasmid.
[0228] The endonuclease can be a homing endonuclease.
[0229] The endonuclease can be an RNA-directed DNA endonuclease.
[0230] The endonuclease can be HO.
[0231] The procedures may also include the isolation of the donor plasmid.
[0232] The procedures may also include the isolation of the recipient plasmid.
[0233] The procedures may also include the isolation of the recombinant recipient plasmid.
[0234] The procedures may also include the isolation of the sequenced oligonucleotide.
[0235] The donor or recipient cell can contain an oligonucleotide that enables plasmid conjugation.
[0236] The donor cell or the recipient cell can contain an oligonucleotide that codes for one or more homologous DNA repair genes.
[0237] The donor or recipient cell can contain an oligonucleotide that codes for one or more recombination-mediated genetic engineering genes.
[0238] The donor cells, recipient cells, or recombinant recipient cells can be transferred to positions on a third ordered array, a fourth ordered array, or a subsequent ordered array.
[0239] The donor cells and the recipient cells can be bacterial cells, independent of each other.
[0240] The barcode sequence can have a length of approximately 4 to 40 nucleotides.
[0241] The barcode sequence can have a length of approximately 15 nucleotides.
[0242] It is assumed that the examples and embodiments described here serve only for illustration and that various modifications or changes may be proposed to persons experienced in the field of technology and included in the scope of the attached claims. EXAMPLES Example 1: Method for in vivo DNA assembly (English: “Stitching”)
[0243] Bacterial strains: BUN20 [Δlac-169 rpoS(Am) robA1 creC510 hsdR514 ΔuidA(MluI):pir-116 endA(BT333) recA1 F'(lac+ pro+ ΔoriT:tet)] was used as a donor strain (Li, M. et al. Nat. Genet. 37, 311-319 (2005)). BW23474: [Δlac-169 rpoS(Am) robA1 creC510 hsdR514 ΔuidA(MluI):pir-116 endA(BT333) recA1] was used as a host strain for the cloning and propagation of all donor plasmids (Haldimann, A. et al. Proc. Natl. Acad. Sci. 93, 14361 (1996)). BW28705 [lacIQ rrnB3 ΔlacZ4787 hsdR514 Δ (araBAD)567 Δ (rhaBAD)568 galU95 ΔendA9:FRT ΔrecA635:FRT] or RE1133 (Egbert et al., Nucleic Acids Research, vol. 47 (6), April 8 2019, Pages 3244-3256) [cmR::mutS pTet2-gam-bet-exo-dam / tetR::bioA / B ilvG+ dnaG.Q576A lacIQ1 Pcp8-araE ΔaraBAD pConst-araC ΔrecJ ΔxonA Pkm-cymR-Cas9::bioC] were as Recipient strains used for in vivo stitching. DH5α and DH10β were used for cloning recipient plasmids.
[0244] The DNA oligonucleotides used for the first and second PCR steps are listed in Tables 1 and 2. Table 1: Primers for the first PCR step Table 2: Primers for the second PCR step
[0245] Media and chemicals: Luria-Bertani (LB) broth (1% w / v tryptone, 0.5% w / v yeast extract, 1% w / v NaCl) as a complex medium was routinely used for cloning and for the growth of donor and recipient plasmids. To preserve the plasmids, antibiotics were added at the concentrations listed in Table 3. For LB media containing hygromycin, 0.5% w / v sodium chloride was used because hygromycin is salt-sensitive. L-arabinose (0.2% w / v), L-rhamnose (0.2% w / v), anhydrotetracycline (100 ng / ml), 4-isopropylbenzoic acid (Cumate; 15 µg / ml), and isopropyl β-d-1-thiogalactopyranoside (IPTG; 500 µM) were used to promote the plasmids. ParaBAD , PrhaBAD , PTet2 , P km -cymR or P lacIQ to induce selection against the SacB counterselection marker. Sucrose agar plates (0.5% w / v yeast extract, 1% w / v tryptone, 6% w / v sucrose, 1.5% agar) and corresponding amounts of antibiotics were used for selection against the SacB counterselection marker. Cl-Phe agar plates (0.5% w / v yeast extract, 1% w / v NaCl, 0.4% w / v glycerol, 2% w / v agar, 10 mM D, Lp-Cl-Phe) and corresponding amounts of antibiotics were used for the counterselection of PheS Gly. 294 YEG agar plates (0.5% w / v yeast extract, 1% w / v NaCl, 0.4% w / v glucose, 2% w / v agar) and the corresponding amounts of antibiotics were used for the subcloning of recipient plasmids containing the PheS Gly fragment. 294 contain. Table 3: Selection of drug concentrations Droge Konzentration bei der Arbeit Hygromycin B 200 ug / ml Nourseothricin-Sulfat 100 ug / ml Kanamycin 50 ug / ml Gentamicin 20 ug / ml Spectinomycin 50 ug / ml D,L-p-Cl-Phe 200 ug / ml Zeocin 5 ug / ml Chloramphenicol 25 ug / ml Ampicillin 50 ug / ml
[0246] The plasmid sequences used for in vivo DNA stitching are listed in Table 4. Table 4: Plasmids used in in vivo stitching. Name Plasmid Typ HomologeRegionen Austauschkassette (englisch:„swappingcassettes ") Andere Merkmale pSL270 Helfer NA NA PrhaBAD-rot pSL359 Helfer NA NA PrhaBAD-rot, ParaBAD-Cas9 pML300 Helfer NA NA PrhaBAD-rot pSL402 Empfänger H1, H3 PheS-ampR CmR pSL414 Donor H1, H3 GmR-SacB oriT, kanR, P-gRNA J23119 T1 pSL415 Donor H1, H3 GmR-SacB oriT, kanR, P -mock J23119 pSL398 Empfänger H1, H3 ZeoR-SacB GmR pSL485 Donor H1, H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 1-251 pSL486 Donor H3 HygR-SacB oriT, kanR, P-gRNA J23119 T3 ,198-517 pSL488 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 454-720 pSL684 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 520-720 pSL685 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 510-720 pSL510 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 500-720 pSL511 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 490-720 pSL512 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 489-720 pSL681 Donor H3 PheS-ampR oriT, kanR, P-gRNA J23119 T1 ,mEGFP 470-720 pSL1060 Recipient H1, H3 NsrR-PheS GmR pSL1065 Donor H1, H3 HygR-SacB oriT, kanR, P-gRNA J23119 T1 ,mEGFP 1-251 pSL1062 Donor H3 NsrR-PheS oriT, kanR, P-gRNA J23119 T2 ,mEGFP 198-517 pSL1066 Donor H3 HygR-SacB oriT, kanR, P-gRNA J23119 T1 ,mEGFP 454-720 pSL1064 Donor H1,H3 HygR-SacB oriT, kanR, P-gRNA J23119 T1 pSL1063 Donor H3 HygR-SacB oriT, kanR, P-gRNA J23119 T1 pSL1107 Donor H3 NsrR-PheS oriT, kanR, P-gRNA J23119 T2 pSL1086 Recipient H1,H3 NsrR-PheS pLacIQ-p15A, GmR
[0247] Construction of the in vivo DNA stitching system: Construction of the helper plasmid and the host recipient strains. In the related MAGIC cloning system (Li, M. et al. Nat. Genet. 37, 311-319 (2005)), the recipient cells contain a helper plasmid pML300, which expresses an inducible λ-red and a temperature-sensitive origin of replication (pSC101-ori). TS ) hosts. MAGIC recipient cells also have a genomically integrated inducible I-SceI endonuclease allele. To implement recursive cutting and homologous recombination, another helper plasmid containing λ-red and Cas9 is required. To construct such a helper plasmid, pML300 was first digested and ligated via HindIII and NheI into a multi-cloning site (MCS), resulting in pSL270. A DNA fragment that araC-ParaBaAD-Cas9The plasmid pSL270, which contains a rhamnose-inducible λ-red recombination system, was constructed using Gibson assembly and then cloned with PacI and XhoI restriction sites in pSL270 to generate the helper plasmid pSL359. This plasmid contains a rhamnose-inducible λ-red recombination system ( PrhaBAD-rot ), an arabinose-inducible endonuclease ( ParaBAD-Cas9 ), a pSC101-ori ts and a spectinomycin resistance marker (SpR). pSL359 was then transformed into BW28705 to generate the recipient host cells BW28705 / pSL359. RE1133 (Egbert et al. Nucleic Acids Research, vol. 47 (6) 08 April 2019, Pages 3244-3256) [cmR::mutS pTet2-gam-bet-exo-dam / tetR::bioA / B ilvG+ dnaG.Q576A 1acIQ1 Pcp8-araE ΔaraBAD pConst-araC ΔrecJ ΔxonA P] was used as an alternative recipient strain. km -cymR-Cas9::bioC] was used without a helper plasmid. RE1133 contains a tetracycline-inducible λ-red recombination system (pTet2-gam-bet-exo-dam / tetR::bioA / B) and a cumate-inducible Cas9 endonuclease (Pkm-cymR-Cas9::bioC).
[0248] Construction of Swapping Cassettes: A swapping cassette is defined as the DNA segment on both the donor and recipient plasmids that participates in a DNA exchange: The cassette on the recipient plasmid is replaced by the cassette that was originally on the donor plasmid through homologous recombination. To recursively select for cassette exchange in vivo, each cassette is constructed to contain both a selectable and a counter-selectable marker. In each round, a selectable marker is required in the donor cassette and a counter-selectable marker in the recipient cassette. To implement such a dual selection strategy, two different selection cassettes were constructed from the following sources using standard cloning procedures: 1) PheS Gly 294(D,Lp-Cl-Phe sensitivity) (Kast, P. Gene 138, 109-114 (1994)), SacB (sucrose sensitivity) (Pelicic, V. et al. J. Bacteriol. 178, 1197-1199 (1996)), 3) HygR (hygromycin resistance) (Gritz, L et al. Gene 25, 179-188 (1983)) with an EM7 bacterial promoter, 4) NsrR (nourseothricin resistance) Gene 62, 209-217 (1988). A cassette was designed to deliver PheS Gly 294 and NsrR, and the second cassette was designed to contain HygR and SacB. Several other selection cassettes were also constructed to perform a series of experiments to characterize the in vivo stitching system: 5) ZeoR (zeocine resistance) (Drocourt, D. Nucleic Acids Res. 18, 4009-4009 (1990)), 6) ampR (ampicillin resistance), and 7) CmR (chloramphenicol resistance). In some cases, a cassette was designed to contain PheS Gly 294 and contains NsrR, and the second cassette was designed to contain HygR and SacB.
[0249] Construction of the scaffolds for donor and recipient vectors: The donor vector was designed to contain the following key components: kanR, oriT, R6K oriγ and a constitutive gRNA expression cassette driven by a strong bacterial promoter J23119 (Standage-Beier, K. et al. ACS Synth. Biol. 4, 1217-1225 (2015)).The swapping region was reconfigured to generate a T1(F)-H1-T2(R)-T2(F)-H3-T1(R) fragment, where T1 (5'-GGGGCCACTAGGGACAGGATtgg-3' (SEQ ID NR: 37)) and T2 (5'-CAGGGCGGTCACCTCCGTGtgg-3' (SEQ ID NR: 38)) are two specific target sequences for CRISPR-Cas9 cutting, and H1 (5'-CGAGGGCTAGAATTACCTACCGGGCCTCCACCATGCCTGCG-3' (SEQ ID NR: 39), and H3 (5'-GTACGGGCAACCCGAGAAGGCTGAGCCTGGACTCAACGGGTTGCTGGGTGGACT CCAGACTCGGGGCGACGACTTCACGCAGCAAGGGCGTCGAGCGGTCGTGAAAGT CTTAGTACCGCACGTGCCGACTCACTGGGGATATTGCCTGGAGCTGTACCGTTCT GGGGGGGAGGTTGGAGACCTCCTCTCTTCTCACGACTGGACCGCGAGGGCCGCG TTGCCGGTTCCCCCCCAGGCTGAAGAAGGGACAAGGGTACTGTGGCAGGGGGAC GCCCATTCAGCGGCTGGCGCTTT-3'(SEQ ID NR: 40)) are homology sites for homologous recombination, and (F) and (R) indicate whether the DNA fragment is in the forward or reverse orientation (reverse complement).A selection cassette (HygR-SacB or NsrR-PheS) is inserted between T2(R) and T2(F) to generate stitchable donor plasmids. In some cases, the swapping region of the donor plasmids was reconfigured to create a T2(F)-T1(R)-T1(F)-H3-T2(R) fragment with the selection cassette (HygR-SacB or NsrR-PheS) inserted between T1(R) and T1(F). Both gRNA target sites are positioned in the correct orientation to minimize the distance between the double-strand break loci and the homology regions. H3 is a 300 bp synthetic DNA fragment used as a homology arm in all assembly rounds. H1 is a homology region used in the first round of in vivo stitching and can be incorporated into the donor scaffold or introduced as part of the first stitched oligonucleotide. Other homology regions (H2, H4, H5, etc.)The following oligonucleotides are introduced into the donor plasmids as part of subsequent oligonucleotides and overlap with the homology region of the preceding oligonucleotide in a composition to enable seamless stitching. The input-receive vector contains a selectable marker (GmR) and a replication origin (ColE1). The swapping region was modified to an H1-T1(R)-T1(F)-H3 configuration, and a selection cassette (HygR-SacB or NsrR-PheS) was cloned between T1(R) and T1(F).
[0250] Testing of endonuclease cleavage and homologous recombination efficiency: To test whether the CRISPR / Cas9 system enables precise DNA cleavage and promotes homologous recombination, the stitching operation was performed in the presence and absence of a targeted gRNA, Cas9, and λrot. A recipient plasmid, pSL402, was transformed into three different recipient host cells to construct: 1) BW28705 / pSL402 (-λrot / -Cas9), 2) BW28705 / pML300 / pSL402 (+λrot / -Cas9), 3) BW28705 / pSL359 / pSL402 (+λrot / +Cas9). Two different donor plasmids, pSL414 and pSL415, were then transformed into BUN20 to generate BUN20 / pSL414 and BUN20 / pSL415, each containing one functional and one mock gRNA unit. Each of the two donor strains was paired with each of the three different types of recipient cells as described above.The cells were then diluted and plated onto selection plates (Cl-Phe + Gm + Cm + 0.2% glucose) to obtain recombinant clones overnight at 37 °C. The colonies were counted to quantify recombination events.
[0251] Assembly of the mEGFP gene in liquid: Three fragments of an mEGFP gene were generated by PCR. The first fragment, containing a constitutive promoter pJ23100, a ribosome binding site, and nucleotides 1–251, was cloned into a donor backbone to generate pSL485. The second fragment, containing nucleotides 198–517, was cloned into a donor vector to generate pSL486. The third fragment, containing nucleotides 454–720 and the rrnB T1 terminator, was cloned into a donor vector to generate pSL488. All three donor vectors were transformed into BUN20 and cultured overnight on LB+Kan plates at 37 °C. An input / receiver vector pSL398 was transformed into BW28705 / pSL359 and cultured on an LB agar plate with gentamicin, spectinomycin, and glucose. A clone of the donor BUN20 / pSL485 (D1) and the recipient BW28705 / pSL359 / pSL398 (R0) were cultured overnight in suitable liquid media at 37°C and 30°C, respectively.The cells (1 ml) from donor and recipient were centrifuged, mixed, and resuspended in 1 ml of pre-warmed LB + Ara + Rha liquid medium. After an incubation of approximately 4 hours at 30 °C without shaking, the mating cultures were serially diluted, and the cells were plated on 6% Suc + Carb + Gm + Sp + 0.2% glucose to select the recombinant (R1). To confirm a correct R1 clone, a colony PCR was performed, which was then inoculated overnight at 30 °C in LB + Gm + Sp + 0.2% glucose. Freshly cultured donor cells containing pSL486 (D2) were then centrifuged, mixed, and resuspended with R1 in liquid mating medium at 30 °C for approximately 4 hours. Recombinant clones were selected by plating on Cl-Phe + Hyg + Gm + Sp + 0.2% glucose. A correct clone (R2), confirmed by colony PCR, was cultured overnight in LB + Gm + Sp + 0.2% glucose at 30 °C.As in previous rounds, cells D3 (BUN20 / pSL488) and R2 were mixed and resuspended in liquid mating media. Serial dilution was performed, and the cells were plated to 6% sucrose + carbamide + gluconeogenesis + spicata + 0.2% glucose. Plates from each assembly round were imaged under UV light to count the proportion of GFP-fluorescent colonies. After each assembly round, the selected clones were chosen, and the plasmids were purified for diagnostic restriction digestion and Sanger sequencing.
[0252] Testing the Effect of Homology Length on Stitching: To test the effect of homology length on stitching accuracy, a series of plasmids containing different degrees of homology to the second mEGFP fragment in pSL486 were constructed: 1) pSL684 (0 bp), 2) pSL685 (10 bp), 3) pSL510 (20 bp), 4) pSL511 (30 bp), 5) pSL512 (40 bp), 6) pSL681 (53 bp). These plasmids, along with pSL488 (63 bp), were transformed into donor host cells to form a group of D3 cells, which were then paired with recipient cells containing R2. The cells of each pairing were diluted and plated onto selection plates (6% sucrose + carbohydrates + gluconate + spicata + 0.2% glucose). The colonies were reared overnight and examined for fluorescence under UV light. The number of fluorescent and non-fluorescent colonies was counted to calculate the percentage of the correct composition.
[0253] Ordered assembly of mEGFP: All strains were arranged on 384-gauge agar plates. First, BUN20 / pSL1065 (D1) and BW28705 / pSL359 / pSL1060 (R0) were arranged on a pre-warmed mating plate (LB + Ara + Rha) using a Singer Rotor HDA and cultured at 30 °C for approximately 5 hours. The mated cells were then transferred to a first selection plate (Cl-Phe + Hyg + Gm + Sp + 0.2% glucose). The recombinant clones (R1) were enriched overnight at 30 °C before being transferred to a pre-mating plate (LB + Hyg + Gm + Sp + 0.2% glucose) to optimize the growth of the assembled plasmids for the next round. Fresh overnight arrays of BUN20 / pSL1062 (D2) were then paired with R1 on a pairing plate. The paired cells were then transferred to the first selection (LB + Nat + Gm + Sp + 0.2% glucose) to select recombinant clones overnight at 30 °C.Selected clones (R2) were then transferred to a pre-pairing plate (6% Suc + Nat + Gm + Sp + 0.2% glucose). Fresh overnight arrangements of BUN20 / pSL1066 (D3) were then paired with R2 on the pairing plate. The final composition products (R3) were selected after pre-selection on Cl-Phe + Hyg + Gm + Sp + 0.2% glucose and subsequently on LB + Hyg + Gm + Sp + 0.2% glucose. Plates from each composition round were imaged under UV light using a UV transilluminator to monitor GFP fluorescence. During the composition process, selected clones were chosen, and the plasmids were purified for diagnostic restriction digestion and Sanger sequencing.
[0254] Ordered assembly of 12 genes using pooled oligonucleotides: Nine different serine / tyrosine recombinases and three different fluorophores (mPapaya, mPlum, sfGFP) were selected for assembly. To generate a list of oligonucleotides required for each gene combination, a Python script was written that takes a FASTA file containing the genes to be synthesized as input and outputs a list of oligonucleotides to be ordered from a commercial supplier (IDT oPool), with user-defined variables including the synthesized oligonucleotide length, the minimum homologous overlap length between adjacent oligos, and the maximum homologous overlap length. DNA hairpins and / or repeats can disrupt the homologous recombination machinery and reduce assembly accuracy, although no quantitative studies on this effect are known.Therefore, these regions were identified for each gene using the Primer3 Python extension (Untergasser, A. et al. Nucleic Acids Res. 35, W71-W74 (2007)) and a metric for nucleotide distribution uniformity. Each nucleotide position is evaluated, and user-defined thresholds are used to expand the homology region if smaller homology regions are likely to contain interfering elements. Once the oligonucleotides required for each gene assembly have been determined, restriction sites (NotI and AscI) are added to each end. These are used to clone oligonucleotides into donor vectors, and round specific priming sites are added. The priming sites allow oligonucleotides from a specific round of a parallel gene assembly to be amplified together and parsed by the in vivo parsing platform.The round-specific primers are selected from a primer list previously developed to reduce the possibility of cross-reactivity between primers (and thus lower the number of unwanted PCR products) when used for large oligonucleotide pools (Kosuri, S. et al. Nat. Biotechnol. 28, 1295-1299 (2010)). Using this Python script, each gene was split into five oligonucleotides of ~300 bp with homology of 50-70 bp between the subsequent oligonucleotides. The PCR-amplified oligonucleotides were inserted into the donor plasmids by restriction digestion and ligation. For the first round of assembly of each gene, the oligonucleotides were amplified and cloned into the donor backbone pSL1064, which contains a 40 bp long H1 start region homologous to that in the input recipient plasmid pSL1060. The oligonucleotides that were added in subsequent odd-numbered stitching rounds (e.g.,Oligonucleotides to be added (e.g., 3, 5, 7, 9) were cloned into pSL1063, which contains the same elements as pSL1064 except for the absence of an H1 homology region. Oligonucleotides to be added in equal rounds (e.g., 2, 4, 6, 8) were cloned into pSL1071. The PCR-amplified oligonucleotides and donor plasmids were digested with AscI and NotI for 4 hours at 37°C. The digested products were then size-selected and purified by gel extraction using Zymoclean Gel DNA Recovery Kits. 0.02 pmol of the digested donor plasmid and 0.06 pmol of the digested oligonucleotides were ligated by mixing with 1 µL of T4 ligase and incubating at 22°C for 1 hour. The ligated donor plasmids were transferred into BUN20 donor strains using standard bacterial transformation protocols. 2 µl of the ligation product were added to 50 µl of chemically compatible BUN20 donor strains.The mixture was then incubated on ice for 30 minutes, heat-shocked at 42 °C for 30 seconds, and incubated again on ice for 3 minutes. The cells were resuspended in 950 ml of NEB SOC recovery medium and recovered for 1 hour at 37 °C. The cells were then plated onto selection plates (LB + Hyg for odd-numbered donor plasmids, LB + Nat for even-numbered donor plasmids) and incubated overnight at 37 °C. The resulting colonies containing cloned oligonucleotides were randomly selected and arranged in 96-well plates. The chained oligonucleotide libraries were parsed using an in vivo DNA parsing system and their sequence was verified. Each plate was paired with two different recipient barcode plates. In the oligonucleotide plates in the odd composition rounds, the donor backbone (pSL1063 or pSL1064) contains the HygR-SacB cassette.After mating with BPS recipient arrays on LB + Ara + IPTG agar for approximately 3 hours at 37°C, the recombinant plasmids were selected on LB + Hyg + Gm + Rha + Ara plates overnight at 37°C. For the even oligonucleotide arrays containing the NsrR-PheS cassette, the recombinant plasmids were selected on LB + Nat + Gm + Rha + Ara plates after mating with BPS collections overnight at 37°C. To assemble the 12 genes, BUN20 donor strains carrying the first-round oligonucleotides were first mated with the RE1133 recipient strain carrying the recipient plasmid pSL1086. 50 µl of each overnight culture were mixed, centrifuged for one minute at 8000 rpm, resuspended in 50 µl of LB, and incubated for 30 minutes at 37 °C. The paired cells were then plated onto LB+aTC+cumate plates (cumate = 4-isopropyl benzoate) and incubated for 4 hours at 37 °C to induce Cas9 and λ-red.To isolate cells carrying the recombinant recipient plasmid with the first-round oligonucleotide, the paired cells were spread onto LB+Hyg+Gm+IPTG agar plates and incubated overnight at 37 °C. Non-recombinant donor plasmids are quickly removed because the R6Kγ origin on the donor plasmid is pir in the recipient strain. +RE1133 is non-functional. The colonies from the selection plates were further purified by selection on LB+Hyg+Gm+IPTG+4CP to remove any remaining non-combined pSL1086 recipient plasmids. To assemble the second round of oligonucleotides, purified colonies carrying the recombinant recipient plasmid with the first-round oligonucleotide were paired with BUN20 donor strains carrying the second-round oligonucleotides. The same procedure as for the first assembly was used, except that the cells carrying the recombinant recipient plasmids were selected on LB+Nat+Gm+IPTG and further purified on LB+Nat+Gm+IPTG+6% sucrose. The same composition procedures were used for all subsequent composition steps, with LB+Hyg+Gm+IPTG+6% sucrose being used for selection in odd-numbered compositions and LB+Nat+Gm+IPTG for even-numbered compositions.This process was repeated five times until all 12 genes were fully assembled (. Fig. 7G). The sequences of the assembled products were determined using Sanger sequencing ( Fig. 7H) and an Oxford Nanopore MinION sequencer were checked for accuracy.
[0255] Assembly of a 9 kb fragment: A 9 kb DNA block from chromosome II positions 41489 to 50489 of the Saccachromyces cerevisiae strain BY4741 was assembled. Using the same Python script employed for the assembly of the 12 genes, the 9 kb block was split into three ~3 kb DNA blocks with 50–75 bp homology between the subsequent DNA blocks. These three DNA blocks were PCR-amplified using genomic DNA from the yeast strain BY4741 as a DNA template. The genomic DNA was extracted using the MasterPure Yeast DNA Purification Kit. The first, second, and third DNA blocks were inserted into the donor plasmids pSL1064, pSL1063, and pSL1107, respectively, via AscI / NotI restriction digestion and T4 ligation. The resulting ligated products were then transformed into BUN20 donor strains using standard bacterial transformation techniques.After Sanger sequencing of the donor plasmids, the DNA blocks were assembled into the recipient strain RE1133 / pSL1086. A donor strain carrying the first DNA block was paired with RE1133 / pSL1086 and cultured for 4 hours at 37 °C on LB+aTC+cumate. Cells carrying the recombinant recipient plasmid were selected on LB+Hyg+Gm+IPTG and further purified on LB+Hyg+Gm+IPTG+6% sucrose. The resulting colonies were then paired with donor strains carrying the second DNA block, selected on LB+Nat+Gm+IPTG on recombinant recipient plasmid, and purified on LB+Nat+Gm+IPTG+4CP. Finally, the resulting colonies were paired with donor strains carrying the third DNA block, selected on recombinant recipient plasmid on LB+Hyg+Gm+IPTG, and purified on LB+Hyg+Gm+IPTG+6% sucrose.The sequences of the assembled products were then checked for the expected sequence using the Oxford Nanopore MinION sequencer and gel electrophoresis (. Fig. 7I).
[0256] Amplicon sequencing for parsing an oligonucleotide library: To extract the recombinant plasmids, cells were scraped from the selection plates and prepared using the Plasmid Plus Mini Kit (QIAGEN). The plasmid DNA was then quantified and diluted to ~1 ng / µl, corresponding to approximately 1.5 e⁶ copies per specific barcode-barcode pair on a 96-cell array mating plate. A two-step PCR was performed as described (Levy, SF et al. Nature 519, 181–186 (2015)) with modifications. First, a PCR with 4–5 cycles using OneTaq polymerase (New England Biolabs) was performed using the forward (pBPS_fwr) and reverse (pBPS_rev) primers listed in Table 1. Approximately 1 ng of recombinant plasmid DNA was amplified in a single 50 µl PCR reaction. The primers for the first PCR step have this general configuration:
[0257] The Ns in these sequences correspond to any nucleotide and are used in downstream analysis to eliminate count biases caused by PCR jackpotting. The Xs correspond to one of several multiplexing tags (e.g., the multiplexing tags in Table 1 above), which allows different samples to be distinguished when loaded into the same sequencing flow cell. Examples of multiplexing tags are the underlined sequences in Table 1. The lowercase sequences correspond to the priming sites on the recombinant plasmids. The uppercase sequences correspond to the Illumina Read-1 or Read-2 sequencing primers. The PCR products were purified using NucleoSpin columns (Macherey-Nagel) and eluted in 33 µl of water.A second PCR of 23–25 cycles was performed using PrimeStar HS polymerase (Takara), employing 33 µl of the purified product from the first PCR as a template and a total volume of 50 µl per tube. The primers used for this reaction were the standard Illumina TruSeq dual-indexed primers (D501–D508 and D701–D712) listed in Table 2. The PCR products were subsequently purified using NucleoSpin columns. The amplicons from each pairing plate were uniquely labeled with the user-defined multiplexing tags as well as the standard Illumina indices. This quadruple-indexing strategy not only increases the multiplexing capacity of the sequencing library but also benefits downstream analysis of amplicon chimeras. Purified amplicons were pooled and sequenced with paired ends on an Illumina MiSeq (2 x 300 bp) with 25% PhiX DNA spike-in. The sequence reads were clustered into barcodes using Bartender. (Zhao, L.), Bioinformatics 34, 739-747 (2017)).
[0258] Design of a recursive in vivo stitching technology: The in vivo stitching system utilizes the bacterial conjugation machinery and homologous recombination of lambda red. The donor vector contains a conditional origin of replication from R6K oriγ, which depends on a functional trans-acting factor π encoded by the pir1 gene or its relaxed copy-number control version, pir1-116 (Metcalf, W. Gene 138, 1-7 (1994)). This particular origin allows the plasmid to be maintained in donor host cells carrying a genomically integrated pir116 allele (e.g., BUN20), but not in recipient cells lacking this allele. Other important features of the donor vectors include an oriT, a backbone marker (kanR), and a constitutive gRNA expression cassette (gRNA). T1 or gRNA T2) and the swapping region. Within the swapping region are two pairs of specific gRNA target sites, a double selectable cassette, and a common long homology sequence (300 bp) for each round of stitching.
[0259] The input-receiver vector contains an origin of replication (ColE1 or pLacIQ-p15A), a backbone marker (GmR), and the swapping region, which consists of homology sequences, two gRNA target sites, and the dual selectable cassette. Some recipient host cells possess a helper plasmid containing a rhamnose-inducible λ-red recombination system and an arabinose-inducible Cas9 endonuclease. A temperature-sensitive mutant derivative of the origin of replication (pSC101-ori) is also present. TS) offers a convenient way to harden the helper plasmid at 42 °C once assembly is complete (Hashimoto, TJ Bacteriol. 127, 1561-1563 (1976)). Other recipient host cells (RE1133) contain a genomically integrated tetracycline-inducible λ-red recombination system and a cumate-inducible Cas9 endonuclease.
[0260] In each round, after the donor vector has been transferred into the recipient cells, both λ-red and Cas9 are induced in the presence of arabinose and rhamnose or cumate and tetracycline, depending on the recipient cell. The constitutively expressed gRNA T1Cas9 directs both the donor and recipient plasmids to generate double-strand breaks. DNA breaks have been shown to strongly stimulate homologous recombination (see below) (Kuzminov, A. Microbiol. Mol. Biol. Rev. 63, 751 (1999)). The donor fragment contains DNA of interest, two different gRNA target sites for the next round of assembly, a dual selectable marker, and the H3 homology region. The DNA for assembly is designed such that the first 50 bp are homologous to the last 50 bp of the assembled sequences on the recipient plasmid. These 50 bp serve as one homology arm for homologous double-cross recombination, with the other arm being H3. This swapping event results in a recombinant recipient plasmid containing new donor DNA.To ensure the accuracy of recombination, three different selection methods are used: 1) selection against the counter-selectable marker (PheS or SacB), 2) selection for the positively selectable marker (HygR or NsrR), and 3) selection for a marker on the recipient backbone (GmR). Selection for the helper plasmid, as well as suppression of the ParaBAD and PrhaBAD promoters or the Pkm-cymR and P. Tet2Promoters are also enforced to prevent hyperrecombination and unwanted DNA breaks. The alternating use of two different gRNAs and two different dual-selectable cassettes enables recursive in vivo stitching to efficiently and linearly assemble new DNA fragments, resulting in the desired sequences, with the only theoretical limit being the tolerable plasmid size. With fluid management and multiplex pinning robots, this platform is highly scalable, allowing thousands of parallel gene assemblies per round.
[0261] CRISPR-Cas9 can efficiently stimulate in vivo stitching: To test whether the CRISPR / Cas9 system enables precise DNA cleavage and promotes homologous recombination, in vivo stitching was performed in the presence or absence of a targeted gRNA, Cas9, and λ-red. A recombinant plasmid was restored when all three were present ( Fig. 3), which suggests that CRISPR / Cas9 facilitates DNA double-strand breaks to improve λ-red recombination efficiency.
[0262] Assembly of a functional fluorescent gene in liquid: To demonstrate the ability to assemble multiple fragments into a functional gene using the in vivo stitching system, three parts of the mEGFP gene were constructed by PCR and cloned into suitable donor backbones. Three fragments were then successively inserted into a recipient vector. After each round, the resulting clones were analyzed by both restriction digestion and Sanger sequencing to verify the accuracy of the assembly before the next round. To ensure that the helper plasmid was preserved in each round, both colony contact PCR and plasmid extraction were performed. Green fluorescence was observed in all colonies (~300) on the selection plates in the final round of assembly, indicating high assembly accuracy. Fig. 7A-I).
[0263] The influence of homology length on stitching fidelity: To test how the accuracy of in vivo stitching depends on the length of homology between fragments, seven donor vectors were constructed containing the third fragment of the mEGFP composition with varying lengths of homology to the second fragment. After conjugation and recombination, cells derived from different conjugation / recombination events were plated onto selection media, and the proportion of fluorescent colonies was counted. The results show that, in this example, a homology length of 40 bp is likely to produce error-free fusion products. Fig. 7A-7I).
[0264] Multiplexer assembly of a functional fluorescent gene on agar: To demonstrate the capability of assembling a functional gene on agar plates, three fragments of the mEGFP gene were constructed by PCR and cloned into suitable donor backbones. The three fragments were then successively inserted into an input-receiver vector in either 96- or 384-pin format. After the final assembly round, green fluorescence was observed at 96 / 96 positions in the 96-pin format and at 383 / 384 positions in the 384-pin format, indicating that the stitching accuracy on agar is comparable to that in liquid. Fig.7A-7I). During the assembly process, plasmids were obtained from various colonies (the entire colony was scraped) and examined by restriction digestion to verify the accuracy of the assembly and / or the retention of the helper plasmid. Typical digestion patterns of the final assembly products showed a clean recombinant plasmid without any identifiable unwanted products (e.g., a non-recombinant plasmid). To further characterize the stitching accuracy, 96 positions from a 384-position assembly were sequenced by Sanger sequencing. The sequencing products were obtained from a colony PCR of a pipette tip that came into contact with each colony. 94 / 96 colonies were found to contain the correct mEGFP sequence. One colony contained an intermediate (the product of the first assembly round), and one contained a stitching error (a large deletion).
[0265] Multiplexer assembly of genes from oligonucleotide pools. To demonstrate the ability to assemble a variety of DNA constructs from oligonucleotide pools, we constructed nine different serine / tyrosine recombinases and three different fluorophores (mPapaya, mPlum, sfGFP) from pools of 300 bp oligonucleotides acquired from IDT (oPools). Each gene was assembled from five oligonucleotides linked in series. The oligonucleotide pools were integrated into the appropriate donor plasmid, transformed into donor bacteria, and the bacterial pools were disassembled into sequence-verified arrays using the methods described in "Example 2: Procedures for in vivo DNA analysis." The donor cell arrays were serially conjugated with recipient cells to assemble each gene with at least three replications. Each array was found to contain the correct DNA sequence by Sanger sequencing ( Fig.7G).
[0266] Assembly of long DNA. To demonstrate the ability to assemble long DNA blocks and produce long assemblies, we reconstructed a 9 kb segment of the Saccachromyces cerevisiae genome from three 3 kb blocks. The 3 kb blocks were amplified from genomic DNA, integrated into the corresponding donor plasmid, transformed into donor bacteria, and sequence-verified. The donor cells were serially conjugated with the recipient cells. To verify the correct assembly product at each assembly step, the recombinant recipient plasmid was purified from the recipient cells, linearized with a restriction enzyme, and analyzed by gel electrophoresis. Fig. 7H). The sequence correctness of the composition products was also verified by Sanger sequencing. Example 2: Methods for in vivo DNA analysis
[0267] Plasmid sequences: Information on the plasmids used for in vivo DNA parsing is given in Table 5 and Fig. 37 can be found. Table 5: Plasmids used in in vivo parsing name Plasmid type Homology Regions swapping cassette Other features pML104 helper N / A N / A P lac -rot, recA pSL937 Recipient H1, H4 PrhaBAD-relE GmR pSL438 Donor H1, H4 HygR-SacB oriT, KanR pSL439 Donor H1, H4 HygR-SacB oriT, KanR pSL1071 Donor H1, H4 NsrR-PheS oriT, KanR
[0268] Construction of barcoded donor plasmids: The donor vector was constructed using standard cloning methods. It contains 1) KanR (kanamycin resistance), 2) oriT (transfer origin), 3) R6K oriγ (conditional replication origin dependent on phage-derived pir1 expression) and 4) swapping region, a configuration I-SceI-H1-H4-I-SceI, where I-SceI is the recognition site of the endonuclease SceI and H1 (5'-ttgccctctctcttcattcagggtcatgaggcacgccattcaaggggagaagtgagatc-3'(SEQ ID NR: 43)) and H4 (5'-aagaacttttctatttctgggtaggcatcatcaggagcagga-3' (SEQ ID NR: 44)) are the homology regions for recombination. In the swapping region of the donor vectors, a selection cassette (HygR-SacB or NsrR-PheS) was cloned between H1 and H4 to generate donor backbone plasmids for parsing.To insert random barcodes into the donor backbones (pSL438 and pSL439), an oligonucleotide (pXL633) containing a NotI restriction site, a barcode region with 15 random nucleotides, and a region homologous to both donor backbones was ordered from IDT. pXL633, paired with pXL585, was used for PCR of the barcodes with ~1 ng of either pSL438 or pSL439 as a template. The resulting PCR products were restriction-digested and ligated into the corresponding donor vector via NotI and XmaI sites. Following the same cloning protocol as described above, the ligation products were transformed into competent BUN20 donor cells, and the barcoded donor clones were selected on LB agar plates containing 50 µg / ml kanamycin (Kan) at 37 °C. The transformants were then randomly selected and arranged to generate two 96-well collections of barcoded dispensers: pSL438_BC and pSL439_BC.To identify the barcode sequences in the arranged donor collections, the regions containing the barcodes were amplified by colony touch PCR using pXL583 and pXL584 as primers. The amplicons were then purified and sequenced using pXL583 according to Sanger. Subsequently, the barcodes were extracted to generate two lists of known donor barcode collections.
[0269] Construction of barcoded recipient plasmids: The plasmid pSL937, used as a backbone for inserting random barcodes to generate the arranged and barcoded recipient collection, was constructed from the following sources using standard procedures: 1) Plasmid backbone / origin of replication from pBR322, 2) GmR (gentamicin resistance marker) from pUC18-mini-Tn7T-Gm 3 3) the homology sequences H1 and H4 and two I-SceI recognition sites in an H1-I-SceI-I-SceI-H4 configuration, 4) a rhamnose-inducible toxin relE ( PrhaBAD-relE) from pSLC-217 4 was cloned between two SceI sites. Oligonucleotides with random barcodes were synthesized by IDT and inserted into pSL937 by restriction digestion and ligation.
[0270] To insert random barcodes into the receiver backbone (pSL937), an oligonucleotide (pXL631) containing an XhoI restriction site, a barcode region with 20 random nucleotides, and a region homological to pSL937 was ordered from IDT. pXL631, paired with pXL154, was used to generate barcodes via PCR using ~1 ng of pSL937 as a template. The resulting PCR products were digested and ligated using the MluI and XhoI restriction sites in pSL937. Ligation reactions were performed at a 3:1 molar ratio of barcode insert to vector overnight at 16 °C. The ligation products were then transformed into competent BUN21 cells expressing a spectinomycin-resistant helper plasmid pML104. 1The barcoded recipient clones were selected on LB agar plates containing 50 µg / ml spectinomycin (Sp), 20 µg / ml gentamicin (Gm), and 2% glucose at 30 °C. The transformants were then randomly selected and arranged in 96-well plates. The barcode sequences at each position in the arranged recipient collections were identified by sequencing. A total of 841 barcodes were reliably identified. These barcodes were rearranged in 8 new 96-well plates so that each position contained a unique barcode.
[0271] Ordered pairing: Each barcoded donor plate (two plates with 96 positions) was paired with each barcoded recipient plate (eight plates with 96 positions). The donor barcode collections were grown overnight at 37 °C on LB + Kan plates; the recipient arrays were grown overnight at 30 °C on LB + Sp + Gm + 2% glucose. The array pairing agar contained 0.2% arabinose (Ara) and 0.1 mM IPTG and was preheated for 1 hour at 37 °C. Both donor and recipient clones were transferred to the pairing plates using SINGER ROTOR HDA pads and grown for approximately 3 hours at 37 °C. Each recipient plate was paired with two barcoded donor plates (pSL438_BC and pSL439_BC). The paired cells were then transferred to the LB selection plates containing 0.2% arabinose, 0.2% rhamnose (Rha), 25 µg / ml gentamicin and 50 µg / ml hygromycin (Hyg) (LB + Ara + Rha + Gm + Hyg).The recombinant clones were then selected overnight at 37 °C.
[0272] Amplicon sequencing: To extract the recombinant plasmids, cells were scraped from the selection plates and prepared using the Plasmid Plus Mini Kit (QIAGEN) mini. The plasmid DNA was quantified and diluted to ~1 ng / µl, which corresponds to approximately 1.5 × 10⁻⁶ 6Copies of each specific barcode-barcode pair are obtained per 96-plate array. A two-step PCR was performed. First, a PCR of 4 to 5 cycles was carried out using OneTaq polymerase (New England Biolabs) with the forward (pBPS_fwr) and reverse (pBPS_rev) primers listed in Table 1. ~1 ng of recombinant plasmid DNA was amplified in a single 50 µl PCR reaction. To increase the multiplexing of sequencing samples, a specific pair of primers was used for the first and second PCRs (see Tables 1 and 2) to amplify the plasmid DNA from a specific pair of paired plates, allowing pooling of multiple paired plates into a single sequencing library. The cycle conditions for the first step are listed in Table 6 below. Table 6: Cycle conditions cycle temperature Time 1 x 94°C 10 seconds 3 x 94°C 15 seconds 55°C 20 seconds 68°C 20 seconds 1 x 68°C 5 min
[0273] Primers for the first step of PCR have this general configuration:
[0274] The Ns in these sequences correspond to an arbitrary nucleotide and are used in downstream analysis to remove count biases caused by PCR jackpotting. The Xs correspond to one of several multiplexing tags that allow different samples to be distinguished when loaded into the same sequencing cell. The lowercase sequences correspond to the priming sites on the recombinant plasmids. The uppercase sequences correspond to the Illumina Read-1 or Read-2 sequencing primers. The PCR products were purified using NucleoSpin columns (Macherey-Nagel) and eluted in 33 µl of water. A second PCR of 23–25 cycles was performed using PrimeStar HS polymerase (Takara), employing 33 µl of the purified product from the first PCR as a template and a total volume of 50 µl per tube.The primers used for this reaction were the Illumina TruSeq dual-indexed primers (D501-D508 and D701-D712) listed in Tables 1 and 2. The cycle conditions for the second step are listed in Table 7. Table 7 Cycle conditions for the second step cycle temperature Time 1 x 98°C 3 min 23 x 98°C 10 seconds 69°C 5 seconds 72°C 20 seconds 1 x 72°C 1 minute
[0275] The PCR products were then purified using NucleoSpin columns. The amplicons from each pairing plate were uniquely labeled with the user-defined primer indices (first PCR) as well as the standard Illumina indices (second PCR). This quadruple indexing strategy increases the multiplexing capacity for sequencing. Purified amplicons were pooled and sequenced at the paired end with approximately 800 reads per barcode-barcode pair on an Illumina MiSeq, HiSeq, or NextSeq with 25% PhiX genomic DNA spike-in.
[0276] Sequencing analysis: Donor-receiver dual barcode amplicon sequencing data were analyzed using adapted Python scripts and Bartender in the following steps. First, the Illumina reads were demultiplexed based on the Illumina indices. All sequences without an exact match of two Illumina indices were discarded. The barcodes were extracted from the demultiplexed sequences using the regular expressions “\D*?(.GGC|T.GC|TG.C|TGG.)\D{4,7}?AA\D{4,7}?TT\D{4,7}?(.CGG|G.GG|GC.G|GCG.)\D*” (donor barcode) and “\D*?(.ACA|G.CA|GA.A|GAC.)\D{4,7}?AA\D{4,7}?AA\D{4,7}?TT\D{4,7}?(.TCG|C.CG|CT.G|CTC.)\D*” (receiver barcode). Unique molecular identifiers (UMIs, the Ns in pBPS_fwr and pBPS_rev) were also extracted based on their expected position in the Illumina reads.The barcode reads, which contained a mixture of genuine barcode sequences and sequences with PCR or sequencing errors, were subsequently clustered into consensus sequences using Bartender. Each barcode cluster was then examined with Bartender for replicating UMIs (indicating PCR duplicates), and all duplicates were removed to determine the final count of each barcode pair. Duplicate barcodes with fewer than 20 reads were excluded, as many of them were likely PCR chimeras (barcodes fused by PCR amplification). The remaining reads were used to determine the position of each donor barcode relative to each corresponding recipient barcode.
[0277] Whole plasmid sequencing on the Oxford Nanopore platform: Recombinant plasmids containing positioning barcodes and oligonucleotides were extracted as described in the section on amplicon sequencing. Circular plasmids were linearized with the restriction enzyme PmlI (NEB) at 37°C for 2 hours. The linearized products were selected for size by passing through a 1.2% agarose gel and recovered using the Zymoclean Gel DNA Recovery Kit (Zymoresearch). The Ligation Sequencing Kit (SQK-LSK110, Nanoporetech) was used to generate sequencing libraries for the Oxford Nanopore platform. 300 ng (~100 fmol) of a linearized recombinant plasmid library underwent end repair using the NEBNext FFPE Repair Mix and the NEBNext Ultra II End Repair / dA-tailing Module (NEB). The nanopore sequencing adapters (AMX-F) were ligated with NEBNext Quick T4 DNA Ligase (NEB).30 ng (~10 fmol) of the library were loaded into a Flongle flow cell (R9.4.1, Oxford Nanopore) to generate reads for recombinant plasmids. The flow cell was operated for 16 hours using the Miniknow sequencer control software (version: 21.11.7, Oxford Nanopore).
[0278] Oxford Nanopore sequencing analysis; (2) Sequencing adapters were identified and removed, and files were separated by sample multiplexing barcodes using "guppy_barcoder" from Guppy version 6.0.1+652ffd179; (3) Alignment to query contaminant sequences (sequence of replication or transfer origin for the donor and / or helper plasmids) using "minimap2" version 2.22-r1110-dirty was used to remove unwanted sequences aligned with these contaminants; (4) Alignment to the expected backbone sequences of the recombinant recipient plasmid was performed using "minimap2" version 2.22-r1110-dirty, and a custom Python script was used to trim the identical backbone sequence from each read; (5) Position-specific barcodes were generated from this sequence using fuzzy regular expressions for the sequence surrounding the barcode with "itermae" version 0.6.0.1. Extracted and then clustered using the message-passing Levenshtein distance approach in "starcode" version 1.4; (6) These barcodes were used to separate the plasmid backbone-removed sequences into separate files for each demultiplexed sample and clustered barcode sequence using custom shell / awk scripts; (7) The sequences for each barcode in each sample were used to perform a multiple-sequence alignment using "kalign3" version 3.3.1; (8) A custom Python script was used to generate a single draft consensus sequence from the multiple-sequence alignment through a voting process; (9) This draft consensus sequence was used with "racon" version 1.5.0 to update the consensus based on read agreements and sequence qualities; (10) This consensus sequence was further refined using "medaka" version 1.5.0 polished (English: “polished”) to generate a polished sequence for each position barcode in each sample; (11) The payload sequence, which is the intended target of the rearraying project, was extracted from the polished region using “itermae” version 0.6.0.1 with various regular expressions. The polished and positioned payloads were analyzed by matching the raw regions removed from the backbone sequence with the polished regions, by matching the payload extracted from the polished regions with the intended target sequences, and by matching the raw regions removed from the backbone sequence with all polished regions generated in the dataset. Alignments were performed using “minimap” version 2.22-r1110-dirty or a custom Python script using the BioPython PairwiseAlignment functionality.A custom R script was used to identify reads as "on-target": those that were >90% identical to the polished region generated for that sample and barcode (i.e., well). Wells were classified as "pure" if >90% of the raw reads were "on-target" with the polished consensus sequence. The sequences were compared to the length of the polished payload and its alignment with an intended target sequence and defined as "correct" if the polished payload perfectly matched one of the intended target sequences.
[0279] Construction of donor plasmid libraries with oligonucleotide pools: The plasmid pSL1071, containing the NsrR-PheS cassette, two I-SceI sites, and two homology regions for recombination (H1 and H4), was used as the backbone into which the oligonucleotide pool was inserted. An oligonucleotide pool with one-handed 300-bp oligonucleotides was ordered from IDT according to the design. where GCTTATTCGTGCCGTGTTAT and GGGCACAGCAATCAAAAGTA (SEQ ID NR: 48) are priming sites for the forward and reverse primers for oligonucleotide pool amplification; GGCGCGCC (SEQ ID NR: 49) and GCGGCCGC (SEQ ID NR: 50) are recognition sites for the restriction enzymes AscI and NotI; and NN...NN denotes the 244-nt sequences randomly selected from the human genome unit GRCh38. Oligonucleotide pool amplification was performed using 7 ng template DNA and KAPA HiFi polymerase (Roche) under the cycle conditions described in Table 8. Table 8: Cycle conditions cycle temperature Time 1 x 95°C 3 min 14 x 98°C 20 seconds 53°C 15 seconds 72°C 15 seconds 1 x 72°C 1 minute
[0280] The PCR products were purified using DNA Clean & Concentrator-5 (Zymoresearch). To clone the PCR products into the donor plasmid pSL1071, the recognition sites of the restriction enzymes AscI and NotI were used. The digestion reaction of PCR products and pSL1071 was performed at 37°C for 4 hours. The digested products were then sorted by size in a 1.2% agarose gel and recovered using the Zymoclean Gel DNA Recovery Kit (Zymoresearch). Ligation was performed with 25 ng of the digested vectors and 3.8 ng of the inserts using T4 DNA ligase (NEB) at 16°C for 15 hours. The ligation products were transformed into BUN20 and conjugated with arrays of barcoded recipient plasmids (see above) to determine the sequence of the construct at each position in the donor array.
[0281] Results: Positioning of the barcode arrays. To validate the accuracy of the parsing and positioning, each of the two known 96-well donor barcode plates (pSL438_BC and pSL439_BC) was paired with eight 96-well receiver plates. Based on the data from these 1536 pairing operations, it was found that the correct position of the donor could be identified and sequenced in 93.82% ± 0.34%, 95.59% ± 0.27%, and 96.04% ± 0.21% of cases on 1, 2, and 3 operations, respectively. Fig. 47A-47D). All failed attempts were due to a lack of sequencing data. No incorrect donor position was ever identified in the sequencing data. Similar results were found when determining the position of receiver barcodes from donor barcodes. Results: Parsing of an oligonucleotide pool
[0282] To further validate the parsing accuracy, we compiled and sequence-verified a pool of 100 oligonucleotides. This pool contained 244-nucleotide sequences randomly selected from the human genome, synthesized as an "oPool" by IDT (Integrated DNA Technologies), and inserted into our donor plasmid pSL1071 by ligation. BUN20 transformants carrying these plasmids were pooled and then randomly arranged into a total of twenty 384-well plates. These arranged plates containing bacteria were then conjugated with an arranged collection of recipient barcode strains (the barcode positions are known). Recipient cells carrying recombinant oligonucleotide barcode plasmids were pooled. The plasmids were sequenced using nanopore sequencing.The sequencing results were used to determine the following for each well in each plate: the consensus sequence of the oligonucleotide, whether the consensus sequence is identical to an expected sequence in the oligonucleotide pool, and whether other oligonucleotide sequences are present at low frequencies (contamination). Fig. 47B, Fig. 47C, Fig. 47D). Consensus sequences were generated for 5,101 wells out of 7,680 wells (66.4%) available in all plates. Of these consensus sequence wells, 2,329 wells (45.6%) were pure and perfectly matched a target oligo. These 2,329 perfectly matching oligos represented 82% of the oligonucleotides expected in the pool.
Claims
[1] Method for assembling a plurality of DNA elements into a composite DNA element in a recipient cell, the method comprising: (a) Contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions to (i) transfer the first donor plasmid from the first donor cell to the recipient cell by conjugation and (ii) recombine the first donor plasmid and the recipient oligonucleotide in the recipient cell by homologous recombination, wherein the recipient oligonucleotide is contained in a plasmid of the recipient cell or in the genome of the recipient cell, and wherein the first donor plasmid in sequential order includes a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a first DNA element fragment (Oligo1), a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2), and a third endonuclease site (C3); the recipient oligonucleotide comprises a third homologous recombination region (HR3) that is homologous to HR1, and a fourth homologous recombination region (HR4) that is homologous to HR2.2; which, following homologous recombination of HR1 with HR3 and HR2.2 with HR4, provides a first recombinant recipient oligonucleotide comprising the first DNA element fragment; (b) Contacting a second donor cell, comprising a second donor plasmid, with the recipient cell, comprising the first recombinant recipient oligonucleotide, under conditions to (i) transfer the second donor plasmid from the second donor cell to the first recipient cell by conjugation and (ii) recombine the second donor plasmid and the first recombinant recipient oligonucleotide to form a second recombinant recipient oligonucleotide in the recipient cell by homologous recombination, wherein the second donor plasmid sequentially includes a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a second DNA element fragment (Oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2), and a sixth endonuclease site (C6); which, following homologous recombination of HR5 with HR2.1 and HR6.2 with HR4, provides a second recombinant recipient oligonucleotide comprising the first and second DNA element fragments (Oligo1, Oligo2) that form a DNA assembly; wherein an oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third and / or fourth endonuclease site is present on the first donor plasmid and / or wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth and / or sixth endonuclease site is present on the second donor plasmid. [2] Method according to claim 1, wherein step (b) is repeated one or more times with a third or subsequent donor cell comprising a third or subsequent donor plasmid comprising compatible HR regions and a third or subsequent oligonucleotide encoding a third or subsequent DNA element fragment (Oligo3, Oligo4, ... OligoN), thereby forming a third or subsequent recombinant recipient oligonucleotide comprising the first, second and a third or subsequent DNA element fragment, which together form a DNA assembly. [3] Method according to claim 1 or 2, wherein step (a) comprises a plurality of first donor cells, each comprising a different first donor plasmid; and step (b) comprises a plurality of second, third, or subsequent donor cells, each comprising a different second, third, or subsequent donor plasmid; wherein optionally each first donor cell is located in a position in a first ordered array and each second, third, or subsequent donor cell is located in a position in a second, third, or subsequent ordered array; wherein the method optionally generates a combinatorial library comprising a plurality of different composite DNA elements. [4] Method according to claim 1, wherein the expression of the first and / or the second endonuclease is inducible and the method further comprises inducing the expression of the first and / or the second endonuclease. [5] Method according to claim 1, wherein the first and / or the second endonuclease is selected from an RNA-directed endonuclease, a homing endonuclease, a transcription activator-like effector nuclease and a zinc finger nuclease. [6] Method according to claim 1, wherein the first, second or subsequent donor plasmid comprises a selectable marker which selects for the integration of the first oligonucleotide, second oligonucleotide or subsequent oligonucleotide into the recipient oligonucleotide. [7] Method according to claim 1, wherein the recipient oligonucleotide comprises a counterselectable marker which selects against recipient cells that do not contain the first, second, third or subsequent oligonucleotide. [8] Method according to claim 1, wherein the donor plasmid comprises a conditional origin of replication. [9] Method according to any one of claims 1 to 8, wherein the donor plasmid or recipient oligonucleotide comprises an inducible high-copy origin of replication. [10] Method according to any one of claims 1 to 9, wherein the donor plasmid or recipient oligonucleotide comprises a replicon capable of replicating plasmids with a length of more than 30 kilobases. [11] Method according to any one of claims 1 to 10, wherein the donor plasmid or the recipient cell comprises an oligonucleotide encoding one or more homologous DNA repair genes; wherein the expression of one or more homologous DNA repair genes is optionally inducible. [12] Method according to any one of claims 1 to 11, wherein the donor plasmid or the recipient cell comprises an oligonucleotide encoding one or more recombination-mediated genetic engineering genes. [13] Method according to any one of claims 1 to 12, wherein the composite DNA element has a length of 100 nucleotides to 500,000 nucleotides. [14] Method according to any one of claims 1 to 13, wherein the first, second or subsequent region of homologous recombination (HR) and its corresponding HR regions on the recipient oligonucleotide each comprise about 20 to about 500 base pairs, optionally about 50 to 100 base pairs. [15] Method according to claim 1, wherein the method comprises the use of two or more recipient oligonucleotides with compatible homologous recombination regions to construct a DNA library. [16] Method according to claim 1, wherein the method is used to assemble a mutagenesis library, to combine genetic regions such as genes, promoters, terminators and regulatory regions from different species, to construct and / or combine genetic regulatory pathways, to construct combinatorial gRNA libraries or to assemble arrays of bacteria containing plasmids for screening tests. [17] Method according to claim 1, wherein prior to steps (a) and (b) of claim 1 the first and the second oligonucleotide comprising the first and the second DNA element fragment are inserted into the first and the second donor plasmid. [18] Method for assembling a plurality of DNA elements into a composite DNA element in a recipient cell, the method comprising: (a) Contacting a first donor cell comprising a first donor plasmid with a recipient cell comprising a recipient oligonucleotide under conditions to transfer the first donor plasmid from the first donor cell to the recipient cell by conjugation, wherein the recipient oligonucleotide is contained in a plasmid of the recipient cell or in the genome of the recipient cell, and wherein the first donor plasmid in sequential order a first endonuclease site (C1), a first homologous recombination region (HR1), a first oligonucleotide comprising a first DNA element fragment (Oligo1), a second homologous recombination region (HR2) comprising two homologous recombination regions (HR2.1, HR2.2), and a third endonuclease site (C3) comprising HR2.1 and HR2.2 flanking a non-homologous region comprising two endonuclease sites (C2.1, C2.2) flanking a selectable marker; the recipient oligonucleotide comprises a third homologous recombination region (HR3) that is homologous to HR1 and a fourth homologous recombination region (HR4) that is homologous to HR2.2, wherein HR3 and HR4 flank a non-homologous region comprising two endonuclease sites (C4.1, C4.2) that flank a selectable marker; (b) Subjecting the first donor plasmid and the recipient oligonucleotide to endonuclease digestion with a first endonuclease to cleave the first donor plasmid at C1 and C3 and the recipient oligonucleotide at C4.1 and C4.2, producing a first donor cassette with homologous recombination regions HR1 and HR2.2 at each end and a recipient oligonucleotide with compatible recombination regions HR3 and HR4; (c) Exposing the first donor cassette and the recipient oligonucleotide to conditions to recombine the first donor cassette into the recipient oligonucleotide by homologous recombination, thereby providing a first recombined recipient oligonucleotide comprising the first DNA element fragment following homologous recombination of HR1 with HR3 and HR2.2 with HR4; (d) Contacting a second donor cell, comprising a second donor plasmid, with the recipient cell, comprising the first recombinant recipient oligonucleotide, under conditions to transfer the second donor plasmid from the second donor cell to the first recipient cell by conjugation, wherein the second donor plasmid sequentially contains a fifth endonuclease site (C5), a fifth homologous recombination region (HR5) homologous to HR2.1, a second oligonucleotide encoding a second DNA element fragment (Oligo2), a sixth homologous recombination region (HR6) comprising two homologous recombination regions (HR6.1, HR6.2), and a sixth endonuclease site (C6) comprising HR6.1 and HR6.2 flanking a non-homologous region comprising two endonuclease sites (C7.1, C7.2) flanking a selectable marker. (e) Subjecting the second donor plasmid and the recipient oligonucleotide to endonuclease digestion with a second endonuclease to cleave the second donor plasmid at C5 and C6 and the recipient oligonucleotide at C2.1 and C2.2, producing a second donor cassette with homologous recombination regions HR5 and HR6.2 at each end and a recipient oligonucleotide with compatible recombination regions HR2.1 and HR4; (f) Exposing the second donor cassette and the recombinant recipient oligonucleotide to conditions to recombine the first donor fragment and the recombinant recipient oligonucleotide in the recipient cell by homologous recombination; which, following homologous recombination of HR5 with HR2.1 and HR6.2 with HR4, provides a second recombinant recipient oligonucleotide comprising the first and second DNA element fragments (Oligo1, Oligo2) that form a DNA assembly; where each donor plasmid includes a transfer origin (oriT) and a conditional replication origin; wherein an oligonucleotide encoding a guide RNA or a first endonuclease targeting the first, third and / or fourth endonuclease site is present on the first donor plasmid and / or in the recipient cell, and wherein an oligonucleotide encoding a guide RNA or a second endonuclease targeting the second, fifth and / or sixth endonuclease site is present on the second donor plasmid and / or in the recipient cell. [19] Method according to claim 1 or 18, wherein the DNA assembly comprises at least a part of a gene, a promoter, an enhancer, a terminator, an intron, an intergenic region, a barcode, a guide RNA (gRNA) or a combination thereof. [20] Method according to claim 18, wherein the DNA assembly has a length of 100 nucleotides to 500,000 nucleotides [21] Method according to any one of claims 1 to 20, wherein the first donor cell is located in a position in a first ordered array and each second, third or subsequent donor cell is located in a position in a second, third or subsequent ordered array. [22] Method according to any one of claims 1 to 21, wherein the method generates a combinatorial library comprising a plurality of different composite DNA elements. [23] Method according to any one of claims 1 to 22, wherein the selectable marker is located within a non-homologous region between HR2.1 and HR2.2 and / or between HR6.1 and HR6.2 and / or between subsequent HR regions. [24] Method according to any one of claims 1 to 23, wherein the counter-selectable marker is located within a non-homologous region between HR2.1 and HR2.2 and / or between HR6.1 and HR6.2 and / or between subsequent HR regions. [25] Method according to any one of claims 1 to 24, wherein the conditional origin of replication depends on the presence of an oligonucleotide or on a condition of cell growth. [26] Method according to any one of claims 1 to 25, wherein the expression of one or more homologous DNA repair genes is inducible. [27] Method according to any one of claims 1 to 26, wherein the method is used to assemble a mutagenesis library or to construct combinatorial gRNA libraries.
Citation Information
Patent Citations
Super-size adeno-associated viral vector harboring a recombinant genome larger than 5.7 kb
US20070042462A1
Method for genome editing
US20120192298A1
Methods and compositions for the targeted modification of a genome
US20160060657A1
Uncharged morpholino-based polymers having achiral intersubunit linkages
US5034506A
Alpha-morpholino ribonucleoside derivatives and polymers thereof
US5235033A