Mutant vectors for improved host transformation
A high-throughput method identifies Rep mutants to modulate plasmid copy numbers, addressing the limitations of natural copy numbers in binary vectors, enhancing transformation efficiencies in host cells.
Patent Information
- Application Number
- PCT/US2025/032788
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-07
- Filing Date
- 2025-06-06
- Publication Date
- 2025-12-11
AI Technical Summary
Current binary vectors used in Agrobacterium-mediated transformation rely on natural copy numbers of origins of replication, limiting the precision and efficiency of transformation outcomes, with no reports on identifying and isolating non-native copy number mutants to tailor these outcomes.
Development of a high-throughput method to identify replication initiator protein (Rep) mutations that modulate the copy number of plasmids, specifically through screening methods to increase or decrease the number of plasmid copies, leading to improved transformation efficiencies in plant and fungal cells.
The method enables the identification of Rep mutants that enhance transformation efficiencies by altering plasmid copy numbers, resulting in improved transient and stable transformation outcomes in various host cells, including plants and fungi.
Smart Images

Figure US2025032788_11122025_PF_FP_ABST
Abstract
Description
Mutant Vectors for Improved Host TransformationCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority benefit of U.S. Provisional Application No. 63 / 657,678, filed June 7, 2024, which is incorporated by reference for all purposes.STATEMENT AS TO RIGHTS TO INVENTIONS MADE UNDER FEDERALLY SPONSORED RESEARCH AND DEVELOPMENT
[0002] The invention was made with government support under Contract No. DE-AC02- 05CH11231 awarded by the U.S. Department of Energy. The government has certain rights in the invention.BACKGROUND
[0003] Agrobacterium-mediated transformation (AMT) of plants and other eukaryotes relies on the presence of a plasmid contained within the bacterium that possesses a transfer DNA (T- DNA) element within the coding sequence so that DNA can be mobilized from an Agrobacterium to the host. This plasmid, called a binary vector, has been the subject of much engineering to improve its transformation properties. A neglected avenue of engineering, however, has been modulating the copy number of different origins of replication (ori) to improve desired transformation outcomes. Every current binary vector used in transformation relies on the natural copy number of the origin within that bacterium, even though it may be optimized for transformation efficiency or for generating quality transformation events within plants. There is no report on identifying and isolating non-native copy number mutants of a binary vector that can tailor transformation outcomes precisely.SUMMARY
[0004] The present disclosure provides a high throughput method to identify replication initiator protein (Rep) mutations that either increase or decrease the copy number of different origins of replication within Agrobacterium. As different origins behave differently between bacterial species, a critical aspect in this disclosure is developing copy number selections that worked in Agrobacterium itself. Through the screening method disclosed herein, various Rep mutants were identified that result in changes in the copy number of the plasmid, and improved transformation efficiencies of cells such as plant and fungal cells.
[0005] In one aspect, the present disclosure provides a polynucleotide comprising a nucleic acid sequence encoding a mutant replication initiator protein (Rep) that provides increased or decreased plasmid copy number compared to the wildtype Rep protein. In some embodiments, the mutant increases plasmid copy number relative to wildtype Rep protein. In some embodiments, the mutant Rep has at least 95% identity to a reference amino acid sequence selected from the group consisting of SEQ ID NOS: 1, 2, 3, and 4; and the polynucleotide comprises at least one substitution in the nucleic acid sequence encoding the mutant Rep protein that results in a stop codon that truncates the mutant protein or results in an amino acid substitution relative to the reference amino acid sequence.
[0006] In some embodiments, the reference amino acid sequence is SEQ ID NO: 1 and the polynucleotide comprises at least one substitution in the nucleic acid sequence encoding the mutant Rep protein that results in an amino acid substitution relative to the reference amino acid sequence. In some embodiments, the mutant Rep protein increases copy number of the vector comprising a compatible origin of replication relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 1.
[0007] In some embodiments, the mutant Rep comprises at least one substitution selected from the group consisting of R106H, G11 ID, S126Y, G13 ID, K187I, N198H, T199N, H201L, V218G, E220K, E222V, A223V, R227S, K229N, A246V, G256D, and K310Q. In some embodiments, the mutant Rep comprises a substitution R106H. In some embodiments, the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 2. In some embodiments, the polynucleotide further comprises a pVSl origin of replication (ori). In some embodiments, the present disclosure provides a plasmid vector comprising the polynucleotide described above. In some embodiments, the vector is a binary vector.
[0008] In some embodiments, the reference amino acid sequence is SEQ ID NO: 2 and the polynucleotide comprises at least one substitution in the nucleic acid sequence encoding the mutant Rep protein that results in a stop codon that truncates the mutant protein or results in an amino acid substitution relative to the reference amino acid sequence. In some embodiments, the mutant Rep protein increases copy number of the vector comprising a compatible origin of replication relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 2.
[0009] In some embodiments, the mutant Rep comprises at least one substitution selected from the group consisting of RUG, S20F, E22_stop, Q60_stop, Truncation, P63H, R75C, S136T, E170K, Q173R, N174D, N181H, V191A, A250S, T251M, R271L, I288L, and I305K.In some embodiments, the mutant Rep comprises a substitution S20F. In some embodiments, the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 3. In some embodiments, the polynucleotide further comprises an RK2 ori. In some embodiments, the present disclosure provides a plasmid vector comprising the polynucleotide described above. In some embodiments, the vector is a binary vector.
[0010] In some embodiments, the at least one substitution in the nucleic acid sequence encoding the mutant Rep protein encodes a mutant Rep protein comprising a substitution relative to SEQ ID NO: 3; and the mutant Rep protein increases copy number of the vector relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 3.
[0011] In some embodiments, the mutant Rep comprises at least one substitution selected from the group consisting of E23K, R28C, A37P, E56K, P70T, E90K, R119S, R125L, S137L, D145N, N150I, V152F, A154P, K155N, F160L, R165W, P166S, W172R, and D181Y. In some embodiments, the mutant Rep comprises a substitution E90K. In some embodiments, the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 4. In some embodiments, the polynucleotide further comprises a pSa ori. In some embodiments, the present disclosure provides a plasmid vector comprising the polynucleotide described above. In some embodiments, the vector is a binary vector.
[0012] In some embodiments, the at least one substitution in the nucleic acid sequence encoding the mutant Rep protein encodes a mutant Rep protein comprising a substitution relative to SEQ ID NO: 4; and the mutant Rep protein increases copy number of the vector relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 4.
[0013] In some embodiments, the mutant Rep comprises at least one substitution selected from the group consisting of K66Q, T74M, V91M, K92N, P96L, V99L, S100L, V128I, D129E, H130N, D132V, D134N, L138F, G141D, T148A, and E182V. In some embodiments, the mutant Rep comprises a substitution El 82V. In some embodiments, the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 5. In some embodiments, the polynucleotide further comprises a BBR1 ori. In some embodiments, the present disclosure provides a plasmid vector comprising the polynucleotide described above. In some embodiments, the vector is a binary vector.
[0014] In another aspect, the present disclosure provides a method of screening for mutations in a wildtype Rep protein that increases the copy number of a vector comprising an origin ofreplication for which the wildtype Rep protein initiates replication. In some embodiments, the method comprises:(a) generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprises a vector comprising a nucleic acid sequence encoding a different mutant Rep and a nucleic acid sequence encoding an antibiotic resistance gene operably linked to an inducible promoter;(b) distributing aliquots of the library to individual compartments to be assayed in an array of test compartments, wherein the array tests a range of concentrations of antibiotic and a range of concentration of an agent that induces the inducible promoter such that each aliquot of the array of test compartments is treated with a different concentration of antibiotic and inducing agent relative to the other members of the array of test compartments;(c) detecting cells that survive in individual compartments of the array of test compartments compared to counterpart control compartments comprising wildtype Rep treated with the different concentrations of antibiotic and inducing agent;(d) processing cells of (c) that survive for high throughput sequence analysis; and(e) identifying compartments in which mutant plasmids are enriched compared to a control compartment comprising an aliquot of the mutant library, wherein the control compartment was not subject to treatment with antibiotic and inducing agent, thereby identifying Rep mutants that increase vector copy number.
[0015] In some embodiments, processing comprises a tagmentation reaction to barcode polynucleotides encoding different Rep mutant proteins. In some embodiments, the antibiotic resistance gene is a gentamycin resistance gene, the inducible promoter is a salicylic acid inducible promoter and the inducing agent in salicylic acid. In some embodiments, the salicylic acid inducible promoter is an NahR promoter. In some embodiments, the Rep mutant library is generated using an error-prone PCR (epPCR). In some embodiments, each combination of concentration of antibiotic and concentration of inducing agent is tested in triplicate using three separate compartments.
[0016] In another aspect, the present disclosure provides a method of screening for mutations in a wildtype Rep protein that decrease the copy number of a vector comprising an origin of replication for which the wildtype Rep protein initiates replication. In some embodiments, the method comprises:(a) generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprises a vector comprising a nucleic acid sequence encoding a different mutant Rep and a nucleic acid sequence encoding a sacB gene operably linked to a sucrose-inducible promoter;(b) subjecting aliquots of the library distributed to individual compartments in an array of test compartments to different concentrations of sucrose;(c) detecting cells that survive in individual compartments of the array of test compartments compared to counterpart control compartments treated with the different concentrations of sucrose;(d) processing cells of (c) that survived for high throughput sequence analysis; and(e) identifying compartments in which mutant plasmids are depleted compared to a control compartment comprising an aliquot of the mutant library, wherein the control compartment was not subject to treatment with sucrose, thereby identifying Rep mutants that decrease vector copy number.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figs, la-lb show a directed evolution pipeline to generate plasmid copy number variants via next-generation sequencing, a, Schematic of RepA mutagenesis method for the 4 ORIs. A selection vector was designed with a gentamicin resistance gene driven by a salicylic acid inducible promoter to enable selection in a checkerboard assay. The RepA protein was mutagenized with epPCR, and -100,000 mutants per ORI were pooled to create a mutant library, b, Example checkerboard data from pVSl (other ORI data in Fig. 6) along with a depiction of the population sequencing strategy. Selection conditions within the checkerboard assay that were only permissible for the mutant library are shown in purple; conditions that permitted growth for both the WT and mutant populations are shown in brown.
[0018] Figs. 2a-2d show enriched residue sites after copy number variant selection. Left: AlphaFold structures of the RepA proteins from each of the 4 ORIs screened. Right: primary structure of each protein with each box corresponding to a single residue. The shading of these boxes corresponds to the fold change enrichment of a mutation at this residue within the selected population compared to the unselected control. Any residue that had an enrichment above a cutoff threshold for yielding the top -20 residues is marked with an “X”, representing the sites with the strongest selective pressure. These residues are depicted as a purple orb on the AlphaFold models. These models focus on the structured dimerization interface, and unstructured N and C-terminal tails were trimmed for compaction. These tails along with a quality of model analysis are depicted in Fig. 8.
[0019] Figs. 3a-3b show screening RepA mutants for improved N. benthamiana transient transformation, a, Schematic of the transient expression assay in N. benthamiana used to screen the 71 mutants b, Boxplots depicting the measured GFP output in N. benthamiana leaf disks for each construct (n=48). Each point on the plot is the measured GFP fluorescence intensity of a single leaf disk. Boxes shaded in green were determined to have significantly higher GFP expression than the WT origin construct by a Tukey’s HSD test (p<0.05); light gray boxes have non-significant differences while brown boxes have significantly lower expression levels. The dashed horizontal line in each plot corresponds to the median GFP output value for the WT of each ORI. Letters above each boxplot correspond to the significance group as computed by Tukey's HSD test.
[0020] Figs. 4a-4f show relationship between copy number, Agrobacterium growth rate, and transient transformation, a, A box plot depicting the measured copy numbers of all mutants and the WT for each ORI as measured by dPCR. Each point represents the average of 3 biological replicates (full data in Fig. 10). The WT copy number is depicted as a red diamond, b, A box plot depicting the biological triplicate average of the growth rate for every mutant and WT form of each ORI. The WT growth rates are depicted as a red diamond, c-f, Regressions displaying the relationship between binary vector copy number and GFP output from the N benthamiana assay (left plot), growth rate and GFP output (middle plot), and binary vector copy number and growth rate (right plot) for RK2 (c), pSa (d), pVSl (e), and BBR1 (f). The Adjusted R2value is displayed beneath each regression along with a p-value for the regression as derived from the F-statistic. Regressions that were not found to have significant predictive power were plotted as a dashed red line.
[0021] Figs. 5a-5b show binary vector origin mutants improve stable transformation of both plants and fungi, a, Example plates for the pVSl and RK2 Arabidopsis thaliana stable transformation experiment. The bar graph reports the observed transformation efficiency (recovered plants / total seeds) with the total number of recovered plants placed above each bar. b, Example plates for the Rhodosporidium toruloides stable transformation experiment. A 1 : 10 dilution plate is shown for the pVSl examples due to the high colony density for some R106H mutant plates. The bar graph reports the average number of colonies observed per plate with individual points representing the colony number per independent transformation replicate.
[0022] Fig. 6 shows high-copy origin selection assay results. Depicted are the growth results from the checkerboard assays conducted for the 4 ORIs. For each ORI, 2 96-well blocks weregrown for the WT strain and the mutant library. Each plot represents a superimposition of the wells that permitted bacterial growth. Wells that permitted growth for both the WT and mutant library are shown in brown. Wells that only grew for the mutant library and that were lethal for the WT strain are depicted in green. For each ORI, the selection condition that was used for sequencing the population is outlined in red.
[0023] Figs. 7a-7d show comprehensive results of mutant enrichment selection. A heat map for each residue of each RepA protein from the 4 ORIs of RK2 (a), pVSl (b), BBR1 (c), and pSa (d) is shown. Intensity of the shading corresponds to the degree of mutation enrichment at that codon in the selection conditions used from the checkerboard assay vs the unselected mutant population. Codons that were under heavy positive selection are more heavily shaded. Residues that met the cutoff threshold for each ORI have a cyan triangle and comprise the mutants used in this study. Residues that fall on the putative dimerization interface as calculated by any residue within 4 A of the partner monomer in the AlphaFold model are marked with a green square. A hypergeometric test was conducted for each RepA protein to determine if selected mutation sites were enriched on the dimerization interface, and the p-values for this test are reported. All RepA proteins were determined to be significantly enriched.
[0024] Figs. 8a-8d show full length AlphaFold models with PLDDT scores. Depicted are full length AlphaFold modeled dimers of RK2 (a), pVSl (b), BBR1 (c), and pSa (d). Models are shaded according to their PLDDT scores for each residue corresponding to the predicted modeling quality. All 4 RepA proteins have well-modeled cores around the predicted dimerization interface with poorly-modeled / unstructured C and N-terminal tails.
[0025] Figs. 9a-9d show comparison of standardized pCM2 GFP expression of best origin mutants. A standardized expression vector consisting of the constitutive pCM2 promoter driving eGFP was constructed for each WT ORI and the highest expressing mutant found in the N. benthamiana assay. A comparison of the WT and mutants for each ORI is shown in both N. benthamiana (a) and Lactuca sativa (b) (N = 48 leaf disks from 6 plants). All mutants were significantly higher expressing in N. benthamiana compared to their WT forms. The mutant forms of RK2 and pSa were significantly higher expressing in Lactuca sativa while the increases observed for pVSl and BBR1 were not significantly different in this experiment. A subset of RK2 mutants with variable expression levels along with the WT form were assayed in N. benthamiana using the EHA105 and GV3101 strains of A. tumefaciens to determine if the mutant ORIs could function in another A. tumefaciens strain (c). The significantincrease / decrease observed in EHA105 relative to WT RK2 was mimicked in the less virulent GV3101 strain. A schematic of the uniform vector used is displayed in (d).
[0026] Figs. lOa-lOb show plasmid copy number and growth rate quantification. An ordered bar graph depicting the quantified copy numbers (a) and growth rates (b) of each mutant and WT origin. For each, the individual copy number or growth rate calculated for each biological replicate (N=3) is depicted as a black dot. Bars are shaded according to their origin.
[0027] Fig. 11 shows copy number validation of RK2 origin. As our measured copy number for WT RK2 differed from values previously reported in the literature, a verification experiment was conducted using 2 total DNA extraction kits (Qiagen Blood and Tissue and Qiagen Powerlyzer) with both stationary and log phase growing tumefaciens. A 26,000 well partition plate was used for greater data resolution, and as there are only 24 wells on this plate, a single replicate was omitted from Group 4 to leave room for a negative control for both primer pairs. These results confirm with high confidence that the RK2 ORI used in this study maintains a copy number of around 1 during active growth.
[0028] Figs. 12a-12d show correlation between copy number, growth rate, transient expression output and mutant enrichment in initial selection. The relationship between fold change enrichment of mutant residues from analyzed NGS data and tobacco transient expression output (a), strain growth rate (b), and copy number (c) are plotted. The Pearson R statistic and p-value of each regression are reported in (d).
[0029] Figs. 13a-13d show arabidopsis transformation overview. The WT and E80K mutant of pSa were also assayed in A. thaliana using the same vectors as the tobacco transient expression screen, and the results of this experiment are shown, a, an example plate from the pSa transformation for both WT and E80K b, quantification of recovered plants from -26,500 planted seeds along with the calculated transformation efficiency c, summary of all transformation results for A. thaliana including for 2 additional pVSl mutants along with the top-performing R106H mutant, d, a summary table of the A. thaliana experiments. As a slightly different seed weight was used for each construct, the estimated seed number for each is shown. This number was used to calculate the transformation efficiency.
[0030] Figs. 14a-14d show validation of A. thaliana transformation improvement with the pVSl R106H Mutant using a pCH5::AwZ> Construct, a) Representative plates of A. thaliana recovered on a MS-hygromycin medium for the pVSI WT and pVSl_R106H constructs b) A bar graph depicting the transformation efficiency of both constructs when planting -25,000seeds. The total number of recovered plants is shown above each bar. c) Representative plants displaying a range of expression levels of the Ruby cassette, demonstrating recovered plants are transgenic. Phenotypically green plants grow on hygromycin and have a faint red tint in their roots and trichomes, d) A table recording the exact seed weights and corresponding seed number planted for each construct along with the total number of recovered plants and the transformation efficiency.DETAILED DESCRIPTIONI. Introduction
[0031] The present disclosure provides polynucleotides and vectors comprising a nucleic acid sequence encoding a mutant replication initiator protein (Rep) that modulates copy number of a vector that comprises an origin of replication from which the wildtype Rep protein is capable of initiating DNA replication. In some instances, the vector is a binary vector. In typical embodiments, the Rep mutant increases vector copy number in host cells. In some instances, the Rep mutant decreases vector copy number in host cells. The present disclosure also provides methods of screening Rep mutants that modulate, e.g., increase, copy number of a vector that comprises an origin of replication for which the wildtype Rep protein can initiate DNA replication and / or Rep mutants that improve Agrobacterium-mediated transformation efficiency of a host cell, such as a plant or fungal host cell, compared to the corresponding wildtype Rep protein.II. Terminology
[0032] As used herein, the singular forms "a,” "and” and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a plant" includes a plurality of such plants and reference to "a pathogen" includes reference to one or more pathogens.
[0033] The use of “or” means “and / or” unless stated otherwise. Similarly, “comprise,” “comprises,” “comprising” “include,” “includes,” and “including” are interchangeable and not intended to be limiting.
[0034] It is to be further understood that where descriptions of various embodiments use the term “comprising” those skilled in the art would understand that in some specific instances, an embodiment can be alternatively described using language “consisting essentially of’ or “consisting of’.
[0035] “Nucleic acid” refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form. The term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2- O-methyl ribonucleotides, peptide-nucleic acids (PNAs).
[0036] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. The term nucleic acid is used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.
[0037] The terms “polypeptide,” “peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residue is an analog or mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides can be modified, e.g., by the addition of carbohydrate residues to form glycoproteins. The terms “polypeptide,” “peptide” and “protein” include glycoproteins, as well as non-glycoproteins.
[0038] The term “amino acid” refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, carboxy glutamate, and O-phosphoserine. Amino acid analogs refer to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group. Exemplary amino acid analogs include, e.g., homoserine, norleucine, methionine sulfoxide, and methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refer to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that function in a manner similar to a naturally occurring amino acid.
[0039] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes (A, T, G, C, U, etc.).
[0040] The terms “identical” or “percent identity,” in the context of two or more polypeptide or polynucleotide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acids or nucleotides that are the same (e.g., at least 70%, at least 75%, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher) identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region as measure by manual alignment and visual inspection or using a BLAST or BLAST 2.0, which are described in Altschul et al. (1990) J. Mol. Biol. 215: 403-410 and Altschul et al. (1977) Nucleic Acids Res. 25: 3389-3402, respectively, comparison algorithm (for nucleotide sequences) with default parameters; or BLASTP with default parameters for amino acid sequences. Software for BLAST analyses is publicly available through the National Center for Biotechnology Information (NCBI) web site. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al, supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. Sequence identity can be performed using BLAST or BLAST 2.0 (nucleic acids), or BLASTP9 (amino acids), using default parameters.
[0041] The definition of sequence identity also refers to the complement of a test sequence. IN some embodiments, identity exists over a region that is at least 50 amino acids ornucleotides in length or at least 75 amino acids or nucleotides in length, and in some instance over a region that is at least 100 amino acids or nucleotides in length. In some embodiments, the sequences are at least 75% identical, typically at least 80% or 85% identical, over the entire length of the coding region. In some embodiments, the sequences are at least 90% identical, or at least 95% identical, over the entire length of the coding region.
[0042] The terms “corresponding to”, “determined with reference to”, or “numbered with reference to” when used in the context of the identification of a given amino acid residue in a polypeptide sequence of interest, refers to the position of the residue in a specified reference amino acid sequence when the polypeptide sequence of interest is maximally aligned and compared to the reference sequence. Thus, for example, an amino acid residue in a Rep protein mutant “corresponds to” an amino acid in the sequence SEQ ID NO: 1 when the residue in the mutant aligns with the amino acid in SEQ ID NO: 1 when the mutant polypeptide sequence is optimally aligned to SEQ ID NO: 1. The polypeptide that is aligned to the reference sequence need not be the same length as the reference sequence. Similarly, “corresponding to,” “determined with reference to,” or “numbered with reference to” when used in the context of the identification of a given nucleotide in a nucleic acid sequence of interest, refers to the position of the nucleotide in a specified polynucleotide reference sequence when the nucleic acid sequence of interest is maximally aligned and compared to the reference sequence.
[0043] A "substitution," as used herein, denotes the replacement of one or more amino acids or nucleotides by different amino acids or nucleotides, respectively.
[0044] The term “plasmid” or “vector” will be understood to include any extrachromosomal covalently continuous double-stranded nucleic acid molecule. As disclosed herein, a vector comprises at a minimum an origin of replication, a selectable marker gene, and one or more restriction sites. The term “restriction site” will be understood to mean a sequence of bases in a DNA molecule that is recognized by a restriction enzyme. A region of nucleotides containing more than one restriction sites can be referred as a multiple cloning cite (MCS). A “recombinant vector” refers to a plasmid carrying a gene or polypeptide-encoding nucleic acid of interest inserted into the backbone of the vector, often within the restriction site or the MCS.
[0045] The term “copy number” is the number of molecules of a particular type of plasmid on or in a cell or part of a cell. This will be understood to describe a characteristic of a recombinant expression construct present in a host cell in greater than a single copy per cell.Most plasmids are classified by the terms “multiple copy number,” “low copy number” or “high copy number,” which describes the ratio of plasmid / chromosome molecules.
[0046] The term “host cell” is meant that a cell that contains a plasmid or a vector and supports the replication or expression of the plasmid or the vector. Host cells include, but are not limited to, bacterial cells, e.g., E. coli cells, plant cells, fungal cells, and other Eucaryotic cells.
[0047] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice of the disclosed methods and compositions, the exemplary methods, devices and materials are described herein.III. Description of Illustrative Embodiments
[0048] The present disclosure is based, at least in part, on the inventor’s development of a high throughput screening method to identify mutations in a Rep protein that modulate copy number of a vector, e.g., in Agrobacterium tumefaciens, compared to the corresponding wildtype Rep protein. As used herein, “Rep protein” refers to a replication initiation protein that binds to and initiates replication from an origin of replication present in a plasmid, such as a DNA transfer vector. As used herein, reference to “ori’ refers to an origin of replication. Reference to “ORI” in the context of the identification of Rep protein mutants in the Examples section refers to a unit comprising an origin of replication and open reading frame of the Rep protein, wherein the Rep protein binds to the origin of replication to initiate the replication.
[0049] A Rep protein may have a narrower range of origins of replication for which it is capable of supporting replication or may support replication of a broad range of origins (e.g., a Rep protein from pVSl or pRK2). As used herein, a Rep protein is considered to be “compatible” with an origin of replication from which it can initiate replication.
[0050] Using the high throughput selection method of the present disclosure, the inventors identified multiple Rep mutants that can modulate copy number of plasmids that comprise compatible origins of replication that improve plant transformation by modulating copy number of a vector, such as a binary vector. Thus, for example, these Rep mutants allow for either more or less transfer-DNA (T-DNA) to be introduced into eukaryotic hosts, e.g., plants cells, depending on the desired transformation outcome. Accordingly, such vectors, e.g., binaryvectors, encoding Rep mutants as described herein can be useful in the transformation of plant cells and non-plant eukaryotes, such as filamentous fungi.
[0051] A vector comprising a Rep mutant of the present disclosure is discussed herein primarily in the context of a binary vector comprising the Rep protein provided in cis to a compatible origin of replications. It is understood by one of skill in the art, however, that the Rep mutant can also be provide in trans by another vector, e.g., a ternary vector in an Agrobacterium transformation system.A. Agrobacterium-mediated transformation (AMT)
[0052] In one aspect, the present disclosure provides a polynucleotide encoding a Rep mutant in a binary vector wherein the polynucleotide can improve transformation efficiency of the binary vector to a plant or non-plant cell. In some embodiment, the transformation is a transient transformation. In some embodiment, the transformation is a stable transformation. In some embodiments, the transformation is an Agrobacleriiim-mQ(H\&iQ transformation (AMT)1. T-DNA binary system
[0053] A “transfer DNA (T-DNA) binary system” or “T-binary system” is a system consisting of two vectors typically used to make transgenic plants. Both vectors are artificial vectors that have been derived from the naturally occurring Ti plasmid found in bacterial species of the genus Agrobacterium, such as A. tumefaciens. A first vector is a “T-binary vector” or “binary vector” or “binary plasmid” comprising a T-DNA region and other components needed for selection, replication, and gene expression. The components of a binary vector are detailed further in next section. The second vector is called “vzr helper plasmid” or “vzr helper vector” which carries vir genes that originated from the Ti plasmid of Agrobacterium. These genes encode a series of proteins that cleave the binary vector at the left and right border sequences, and facilitate transfer and integration of T-DNA to the plant's cells and genomes, respectively. When the binary vector and the vir helper plasmid are both present in the same Agrobacterium cell, proteins encoded by the vir helper plasmid act in trans on the T-DNA border repeat elements to mediate processing, secretion, and host genome integration of the sequence between the left and right border repeat elements.2. Plants
[0054] The transformation of plant cells by using Agrobacterium commonly involves incubating the cells or tissues with the bacteria. The type of plant tissues commonly used is embryonic callus cultures due to the known genotype compared to seedlings, regeneration potential and the stability of the regenerated plants. A callus refers to a plant tissue formed from a wound site or cut plant surface.
[0055] In some embodiments, the Agrobacterium-mediated transformation process comprises an additional step after exposing the plant tissue to Agrobacterium. In some embodiments, the step comprises using glass beads or performing sonication to weaken the barrier of the plant tissue and improve the efficiency of Agrobacterium-mediated gene delivery to the plant cells.
[0056] As disclosed herein, the AMT can transfer a binary vector to a plant cell or a nonplant eukaryote. In some embodiments, the host cell is a plant cell. Suitable plants cells include any that can be transformed by d / Y>Aac7c / ' / z / / 77-mediated transformation, including common model plants such as Nicotiana benthamiana or Arabidopsis thaliana and many other plants cells, including dicots and monocots. In some embodiments, the host cell is a fungal cell, such as an ascomycete, a basidiomycete, a zygomycete, a chytrid fungus, or other fungal cell outside of the Dikarya. Examples of fungal cell which can be used for AMT, include but not limited to, as those described in Idnurm, A et al. Fungal Biol Biotechnol 4, 6 (2017), which is herein incorporated by reference. In some embodiments, the fungal cell is from an oleaginous yeast such as Rhodosporidium toruloides.B. Binary Vector
[0057] As indicated above, a “transfer DNA (T-DNA) binary vector” or “binary vector” or “binary plasmid” is a shuttle vector that replicates in multiple hosts (e.g., Escherichia coli and Agrobacterium) and can be used for delivery of T-DNA from a bacterium (e.g., Agrobacterium) into a host cell (e.g., a plant cell). Representative series of binary vectors are listed in Table 1 (Murai, N. (2013) American Journal of Plant Sciences, 4, 932-939).
[0058] As disclosed herein, a binary vector usually comprises a T-DNA region containing a cloning site, flanked by the left and right border regions, at least one origin of replication (ori) associated with a replication initiator protein (Rep), and one or more selectable marker genes. Each component of the binary vector is further detailed below. In some embodiments, thebinary vector further comprises other components such as elements for gene expression (e.g., promoters and polyA signals), stability sequences, and elements for monitoring the recombinant protein in the transformed host cell (e.g., a reporter gene).Table 1. Representative Agrobacterium T-DNA binary vectors1. T-DNA region
[0059] In some embodiments, the T-DNA region comprises one or more restriction sites for the subsequent insertion of the gene of interest. Such cloning sites, e.g., MCS, are flanked by T-DNA border repeat sequences, in direct orientation with one another. In some embodiments, both left and right T-DNA border repeats comprise a sequence of 20-25 bp length. In some embodiments, the T-DNA border repeats differ in length. Further, as understood by one in the art, T-DNA border repeat sequences comprise natural variations (e.g., pseudo border repeats). In some instances, reference to a “T-DNA region” also refers to a sequence comprising the gene / sequence of interest and T-DNA border repeat sequence(s) after the gene / sequence of interest is inserted into the cloning site).2. Origins and Rep proteins
[0060] The present disclosure provides a polynucleotide comprising a nucleic acid sequence encoding a mutant replication initiator Rep that modulates copy number of a vector comprising a compatible origin of replication. In some instances, the Rep mutant increases vector copy number. In other instances, the Rep mutant decreases vector copy number. Examples of Rep proteins include the Rep protein from pVSl, e.g., SEQ ID NO: 1; an RK2 Rep protein, e.g., of SEQ ID NO: 2, a pSa Rep protein, e.g., SEQ ID NO: 3, a BBR1 Rep protein, e.g., SEQ ID NO: 4. Examples of compatible origins of replication for each Rep protein are: the ori sequence from pVSl, e.g., SEQ ID NO: 13; an RK2 ori sequence, e.g., of SEQ ID NO: 14, a pSa ori sequence, e.g., SEQ ID NO: 15, a BBRl ori sequence, e.g., SEQ ID NO: 16.
[0061] In some embodiments, the Rep mutant comprises a sequence at least 95% identity to an amino acid sequence of a wild-type (WT) Rep wherein the mutant Rep protein modulates copy number, e.g., increases, copy number, of a vector that comprises a compatible origin of replication, is encoded by a mutated Rep nucleic acid comprising a nucleotide substitution that results in an amino acid substitution or a stop codon that deletes or truncates the protein. In some embodiments, the WT Rep is from a plasmid pVSl, RK2, pSa, or BBR1. In some embodiments, the mutant Rep comprises a sequence at least 95% identical to any one of SEQ ID NOs: 1-4 in which the Rep protein has at least one mutation that modulates copy number as described in the present disclosure. In some embodiments, the polynucleotide further comprises an origin of replication (ori). In some embodiments, the ori is from the same plasmid of the Rep. For example, the polynucleotide comprises a nucleic acid sequence encoding a mutant Rep derived from plasmid pVSl and a pVSl ori.
[0062] In some embodiments, the Rep mutant comprises a sequence at least 95% identical to an amino acid sequence of a pVSl Rep that modulates copy number of a plasmid comprising a pVSl origin of replication or another compatible origin of replication. In some embodiments, the pVSl Rep mutant comprises at least one substitution relative to SEQ ID NO: 1. In some embodiments, the mutant pVS 1 Rep comprises at least one substitution selected from the group consisting of R106H, G111D, S126Y, G131D, K187I, N198H, T199N, H201L, V218G, E220K, E222V, A223 V, R227S, K229N, A246V, G256D, and K310Q relative to SEQ ID NO: 1. In some embodiments, the mutant pVSl Rep comprises a substitution R106H. In some embodiments, the mutant pVSl Rep comprises at least one substitution selected from those set forth in Table 2.
[0063] In some embodiments, the polynucleotide comprising the nucleic acid sequence encoding a Rep protein of the present disclosure further comprises a nucleic acid sequence of a pVSl origin of replication (ori). An illustrative pVSl ori sequence is provided in SEQ ID NO: 13. In some embodiments, the pVSl ori comprises a sequence at least 80%, 85%, or 90% identical to SEQ ID NO: 13. In other instances, the pVSl ori comprises at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 13; or comprises SEQ ID NO: 13.Table 2. Exemplary pVSl Rep substitutions.
[0064] In some embodiments, the Rep mutant comprises a sequence at least 95% identical to an amino acid sequence of RK2 Rep that modulates copy number of a plasmid comprising an RK2 origin of replication or another compatible origin of replication. In some embodiments, the RK2 Rep mutant comprises at least one substitution relative to SEQ ID NO: 2. In some embodiments, the mutant RK2 Rep comprises at least one substitution selected from the group consisting of RUG, S20F, E22_stop, Q60_stop, Truncation, P63H, R75C, S136T, E170K, Q173R, N174D, N181H, V191A, A250S, T251M, R271L, I288L, and I305K relative to SEQ ID NO: 2. In some embodiments, the mutant RK2 Rep comprises a substitution S20F. In some embodiments, the mutant RK2 Rep comprises at least one substitution selected from Table 3.
[0065] In some embodiments, the polynucleotide comprising the nucleic acid sequence encoding a Rep protein of the present disclosure further comprises a nucleic acid sequence of an RK2 origin of replication (ori). An illustrative RK2 ori sequence is provided in SEQ ID NO: 14. In some embodiments, the RK2 ori comprises a sequence at least 80%, 85%, or 90% identical to SEQ ID NO: 14. In other instances, the RK2 ori comprises at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 14; or comprises SEQ ID NO: 14.Table 3. Exemplary RK2 Rep substitutions.
[0066] In some embodiments, the Rep mutant comprises a sequence at least 95% identical to an amino acid sequence of pSa Rep that modulates copy number of a plasmid comprising a pSa origin of replication or another compatible origin of replication. In some embodiments, the pSa Rep mutant comprises at least one substitution relative to SEQ ID NO: 3. In some embodiments, the mutant pSa Rep comprises at least one substitution selected from the group consisting of E23K, R28C, A37P, E56K, P70T, E90K, R119S, R125L, S137L, D145N, N150I, V152F, A154P, K155N, F160L, R165W, P166S, W172R, and D181Y relative to SEQ ID NO: 3. In some embodiments, the mutant pSa Rep comprises a substitution E90K. In some embodiments, the mutant pSa Rep comprises at least one substitution selected from Table 4.
[0067] In some embodiments, the polynucleotide comprising the nucleic acid sequence encoding a Rep protein of the present disclosure further comprises a nucleic acid sequence of a pSa origin of replication (ori). An illustrative pSa ori sequence is provided in SEQ ID NO: 15. In some embodiments, the pSa ori comprises a sequence at least 80%, 85%, or 90% identical to SEQ ID NO: 15. In other instances, the pSa ori comprises at least 95%, 96%, 907%, 98%, or 99% identity to SEQ ID NO: 15; or comprises SEQ ID NO: 15.Table 4. Exemplary pSa Rep substitutions.
[0068] In some embodiments, the Rep mutant comprises a sequence at least 95% identical to an amino acid sequence of BBR1 Rep that modulates copy number of a plasmid comprising a BBR1 origin of replication or another compatible origin of replication. In some embodiments, the BBR1 Rep mutant comprises at least one substitution relative to SEQ IDNO: 4. In some embodiments, the mutant BBR1 Rep comprises at least one substitutionselected from the group consisting of K66Q, T74M, V91M, K92N, P96L, V99L, S100L, V128I, D129E, H130N, D132V, D134N, L138F, G141D, T148A, and E182V relative to SEQ ID NO: 4. In some embodiments, the mutant BBR1 Rep comprises a substitution El 82V. In some embodiments, the mutant BBR1 Rep comprises at least one substitution selected from Table 5.
[0069] In some embodiments, the polynucleotide comprising the nucleic acid sequence encoding a Rep protein of the present disclosure further comprises a nucleic acid sequence of a BBR1 origin of replication (ori). An illustrative BBR1 ori sequence is provided in SEQ ID NO: 16. In some embodiments, the BBR1 ori comprises a sequence at least 80%, 85%, or 90% identical to SEQ ID NO: 16. In other instances, the BBR1 ori comprises at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 16; or comprises SEQ ID NO: 16.Table 5. Exemplary BBR1 Rep substitutions.
[0070] The ability of a mutant Rep protein to modulate, e.g., increase, copy number of a vector comprising a compatible origin of replication that expresses the mutant Rep can be tested using any number of assays. In some embodiments, a direct measure of copy number can be assessed, e.g., using a quantitative assay, such as an amplification assay, e.g., PCR. In some instance, copy number is increased by at least 20%, 30%, 40%, 50%, 100%, or greater compared to control wildtype plasmid.
[0071] In other embodiments, a surrogate assay can be used as an indirect measure of copy number. For example, in some instances, a binary vector that expresses a mutant Rep protein of the present disclosure can be used in an Agrobacterium transformation system to evaluate transformation efficient of a plant cells, such as Nicotiana benthamiana, in comparison to a corresponding control binary vector comprising a nucleic acid encoding the wildtype (WT) Rep, e.g., using a reporter gene construct.3. Selectable marker genes
[0072] In some embodiments, a binary vector further comprises one or more selectable markers gene. In some embodiments, the binary vector comprises a bacterial selectable marker gene for selection of the binary vector in E. coli and Agrobacterium. In some embodiments, the bacterial selectable marker gene is an antibiotic resistance gene that is functionally against an antibiotic such as a kanamycin, rifampicin, or hygromycin in the bacterial culture. In some embodiments, the binary vector comprises a plant selectable marker gene to confirm successful transfer and subsequent integration of the T-DNA region derived from a particular binary vector into the plant genome. In typical embodiments, a binary vector can comprise both bacterial and plant selectable marker gene. In some embodiments, the selectable marker gene can be chemical selectable, visual, or assayable by some other method, such as enzyme assay, ELISA, and the like. Other sequences that facilitate the T-DNA transfer process may be incorporated within the binary vectors, including, but not limited to, an overdrive sequence(see, for example, Peralta et al. (1986) EMBO J. 5: 1137-1142; Zyprian and Kado (1990) Plant Mol. Biol. 15:245-256) and a T-DNA transfer stimulator sequence (Hansen et al. (1992) Plant Mol. Biol. 20: 113-122).
[0073] In accordance with the present invention, a selectable marker gene can be either a positive or negative selectable marker gene. A positive selectable marker confers selective advantage to the host organism, such as antibiotic resistance as described above, which allows the host organism (e.g., Agrobacterium or a plant) to survive antibiotic selection. In some instances, e.g., in applications to identify mutant Rep proteins, a negative selectable marker that eliminates or inhibits growth of host cells (e.g., an Agrobacterium or a plant cell) upon selection may be employed. An example would be thymidine kinase, which makes the host sensitive to ganciclovir selection. Another example of the negative selectable marker is a bacillus subtilis levansucrase (sacB) gene. In some embodiments, a selectable marker gene can serve as both positive and negative markers by conferring an advantage to the host under one condition, but inhibiting growth under a different condition. An example would be an enzyme that can complement an auxotrophy (positive selection) and be able to convert a chemical to a toxic compound (negative selection). In some instances, a binary vector can carry a selectable marker gene which serve both positive and negative selections. In some instances, a binary vector can carry a positive selectable marker gene and a negative selectable marker gene.B. Methods of screening Rep mutants
[0074] In one aspect, the present disclosure provides a method of screening Rep mutants that modulate vector copy number. In some embodiments, a screening method of the present disclosure identfies Rep mutants that increase vector copy number. Alternatively, a screening method may be performed to identify Rep mutants that decrease vector copy number.1. Rep mutants that increase vector copy number
[0075] In some embodiments, a high copy number binary vector can increase the overall transformation efficiency of a binary vector, which may be a factor in increasing transient expression in hosts, or facilitating transformation in traditionally recalcitrant hosts such as filamentous fungi or monocot plants.
[0076] In some embodiments, the present disclosure provides a method of screening Rep mutants that increase vector copy number. In some embodiments, the method comprises: generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprisesa vector comprising a nucleic acid sequence encoding a different mutant Rep protein, e.g., generated by error-prone PCR, and an antibiotic resistance gene operably linked to an inducible promoter. Non limiting examples of antibiotic resistance genes that can be employed for screening include the genes that encode a protein conferring resistance to kanamycin, spectinomycin, hygromycin, gentamicin, apramycin, chloramphenicol, carbenicillin, Zeocin, erythromycin, or any derivative thereof. In some embodiments, the antibiotic resistance gene is a gentamycin resistance gene. Non limiting examples of inducible promoters include cumate-inducible promoters, salicylic acid-inducible promoters, naringenin-inducible promoters, vanillate-inducible promoters, PcaU-inducible promoters, TetR-inducible promoters, sugar-inducible promoters, and quorum sensing promoters regulated by CinRAM and LuxR. Examples of inducible promoters that can be designed into the binary vector, include but not limited to, as those described in Layla A Schuster, et al. Nucleic Acids Research, Volume 49, Issue 12, Pages 7189-7202, incorporated by reference. In some embodiments, the inducible promoter is a salicylic acid inducible promoter and the inducing agent in salicylic acid. In some embodiments, the salicylic acid inducible promoter is an NahR promoter.
[0077] In some embodiments, the method comprises performing an assay, e.g., a “checkerboard” assay, comprising distributing aliquots of the library to individual compartments, e.g., wells of a multi-well plate, to be assayed in an array of test compartments, wherein each compartment present in the array of test compartments is subjected to a different concentration of antibiotic and inducible agent, relative to the other members of the array, such that the array tests a range of concentrations of antibiotic in combination with a range of concentrations of an agent that induces the inducible promoter. As disclosed herein, concentrations of antibiotic and inducible agent can be vary depending on the types of the antibiotics and inducible agents. In some embodiments, the range of concentrations of an antibiotic or inducible agent for selection is from a working concentration of the antibiotic or inducible agent to 150x above the working concentration. For example, the concentration range for gentamicin selection can be from 20mg / L - 3000mg / L.
[0078] In some embodiments, the Rep mutant library is generated using an error-prone PCR (epPCR) to introduce random mutations. Buffer compositions can be controlled to limit the frequency of mis-incorporation of nucleotide bases introduced into the sequence. In typical embodiments, the substitution frequency is 1 - 3 base pair substitutions per kilobase of DNA.
[0079] In some embodiments, the method is used to screen for pVSl Rep mutants that increase vector copy number. In some instances, the vector further comprises a pVSl ori, thus providing the pVS 1 Rep mutant on the same plasmid as the pVS 1 origin of replication to initiate plasmid replication. In some embodiments, the pVSl ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 13 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 13. In other instances, the cell having a nucleic acid sequence that expresses a pVSl Rep mutant further comprises a separate plasmid comprising a pVSl ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the pVSl Rep mutant and the other vector comprises a pVSl ori sequence. In such an example, the pVSl Rep mutant is provided in trans to the pVSl ori to initiate plasmid replication.
[0080] In some embodiments, the method is used to screen RK2 Rep mutants that increase vector copy number. In some instances, the vector further comprises an RK2 ori, thus providing the RK2 Rep mutant on the same plasmid as the RK2 origin of replication to initiate plasmid replication. In some embodiments, the RK2 ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 14 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 14. In other instances, the cell having a nucleic acid sequence that expresses an RK2 Rep mutant further comprises a separate plasmid comprising an RK2 ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the RK2 Rep mutant and the other vector comprises an RK2 ori sequence. In such an example, the RK2 Rep mutant is provided in trans to the RK2 ori to initiate plasmid replication.
[0081] In some embodiments, the method is used to screen pSa Rep mutants that increase vector copy number. In some instances, the vector further comprises a pSa ori, thus providing the pSa Rep mutant on the same plasmid as the pSa origin of replication to initiate plasmid replication. In some embodiments, the pSa ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 15 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 15. In other instances, the cell having a nucleic acid sequence that expresses a pSa Rep mutant further comprises a separate plasmid comprising a pSa ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the pSa Rep mutant and the other vector comprises a pSa ori sequence. In such an example, the pSa Rep mutant is provided in trans to the pSa ori to initiate plasmid replication.
[0082] In some embodiments, the method is used to screen BBR1 Rep mutants that increase vector copy number. In some instances, the vector further comprises a BBR1 ori, thus providing the BBR1 Rep mutant on the same plasmid as the BBR1 origin of replication to initiate plasmid replication. In some embodiments, the BBR1 ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 16 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 16. In other instances, the cell having a nucleic acid sequence that expresses a BBR1 Rep mutant further comprises a separate plasmid comprising a BBR1 ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the BBR1 Rep mutant and the other vector comprises a BBR1 ori sequence. In such an example, the BBR1 Rep mutant is provided in trans to the BBR1 ori to initiate plasmid replication.
[0083] In some embodiments, the method further comprises sequencing the Rep mutants from the survived cells. In some embodiments, the method further comprises sequencing the Rep mutants before performing the screening assay. In some embodiments, the sequencing is a next-generation Illumina deep-sequencing. In some embodiments, the method further comprises comparing the Rep mutant sequences before and after the screening assay selection, thereby identifying Rep mutants that increase copy number of the vector in the cell.2. Rep mutants that reduce vector copy number
[0084] In some embodiments, a low copy binary vector will be useful in increasing the probability of a "quality event" in plant transformation. Quality events are defined as a transformation event where only a single copy of the T-DNA is integrated into the host genome.
[0085] In some embodiments, the present disclosure thus provides a method of screening Rep mutants that decrease vector copy number. In some embodiments, the method comprises: generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprises a vector comprising a nucleic acid sequence encoding a different mutant Rep and a sacB gene comprising a sucrose-inducible protein, e.g., the sacB gene promoter. The sacB gene is a negative selective marker encoding a protein that becomes toxic when expressed in the presence of sucrose. A binary vector with a Rep mutant that survives at higher concentrations of sucrose while sacB is expressed will be selected as a low copy number binary vector.
[0086] In some instances, a method to identify Rep mutants that reduce vector copy number thus comprises contacting the cells with varying concentrations of sucrose; detecting cells thatsurvive, and identifying mutants that are enriched in the surviving populations, e.g., by high throughput sequencing, thereby identifying Rep mutants that decrease vector copy number.
[0087] In some embodiments, the method is used to screen pVSl Rep mutants that reduce vector copy number. In some instances, the vector further comprises a pVSl ori, thereby providing the pVSl Rep mutant on the same plasmid as the pVSl ori to initiate plasmid replication. In some embodiments, the pVSl ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 13 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 13. In other instances, the cell having a nucleic acid sequence that expresses a pVSl Rep mutant further comprises a separate plasmid comprising a pVSl ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the pVSl Rep mutant and the other vector comprises a pVSl ori sequence. In such an example, the pVSl Rep mutant is provided in trans to the pVSl ori to initiate plasmid replication.
[0088] In some embodiments, the method is used to screened RK2 Rep mutants that reduce vector copy number. In some instances, the vector further comprises an RK2 ori, thereby providing the RK2 Rep mutant on the same plasmid as the RK2 ori to initiate the plasmid replication. In some embodiments, the RK2 ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 14 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 14. In other instances, the cell having a nucleic acid sequence that expresses an RK2 Rep mutant further comprises a separate plasmid comprising an RK2 ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the RK2 Rep mutant and the other vector comprises an RK2 ori sequence. In such an example, the RK2 Rep mutant is provided in trans to the RK2 ori to initiate plasmid replication.
[0089] In some embodiments, the method is used to screen pSa Rep mutants that reduce vector copy number. In some instances, the vector further comprises a pSa ori, thereby providing the pSa Rep mutant has on the same plasmid as the pSa ori to initiate the plasmid replication. In some embodiments, the pSa ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 15 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 15. In other instances, the cell having a nucleic acid sequence that expresses a pSa Rep mutant further comprises a separate plasmid comprising a pSa ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acidsequence encoding the pSa Rep mutant and the other vector comprises a pSa ori sequence. In such an example, the pSa Rep mutant is provided in trans to the pSa ori to initiate plasmid replication.
[0090] In some embodiments, the method is used to screen BBR1 Rep mutants that reduce vector copy number. In some instances, the vector further comprises a BBR1 ori, thereby providing the BBR1 Rep mutant on the same plasmid as on the BBR1 ori to initiate the plasmid replication. In some embodiments, the BBR1 ori comprises a sequence at least 80%, 85%, or 90% identity to SEQ ID NO: 16 or comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 16. In other instances, the cell having a nucleic acid sequence that expresses a BBR1 Rep mutant further comprises a separate plasmid comprising a BBR1 ori. For example, the cell can comprise two vectors wherein one vector comprises a nucleic acid sequence encoding the BBR1 Rep mutant and the other vector comprises a BBR1 ori sequence. In such an example, the BBR1 Rep mutant is provided in trans to the BBR1 ori to initiate plasmid replication. .3. Rep mutants that improve transformation efficiency
[0091] As disclosed herein, not all mutations result in desirable outcomes as there is eventually a tradeoff between copy number and transformation. If copy number becomes too high, the plasmid likely becomes a burden on the cell, and transformation suffers, and if the copy number becomes too low, there are not enough plasmids for efficient transformation. Thus, simply identifying a high-copy mutant is not necessarily sufficient to improve transformation efficiency.
[0092] In some embodiments, the present disclosure provides a method of screening Rep mutants that improve vector transformation efficiency, wherein the method comprises: generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprises a vector comprising a reporter gene and a nucleic acid sequence encoding a different mutant Rep; and detecting cells with enhanced expression level of the reporter protein in comparison to cells comprising a nucleic acid sequence encoding a wild type (WT) Rep, thereby identifying the Rep mutants that improve vector’s transformation efficiency. In some embodiments, the reporter gene encodes a protein that generates a detectable signal such as a colorimetric signal, e.g, a fluorescent signal, e.g, green fluorescent protein, or a chemiluminescent signal, or any other detectable signal that is conveniently measured.IV. Examples
[0093] The present disclosure will be described in greater detail by way of specific examples. The following examples are offered for illustrative purposes only, and are not intended to limit the invention in any manner. Those of skill in the art will readily recognize a variety of noncritical parameters which can be changed or modified to yield essentially the same results.Example 1:
[0094] This example illustrates an improved ^grotoctezvz / zzz-mediated transformation via binary vector copy number engineering.
[0095] The copy number of a plasmid is critically linked to its functionality, yet there have been few attempts to systemically modify this property across broad-host range origins. Using a directed evolution approach, we generated copy number variants of four commonly used broad-host range origins (pVSl, RK2, pSa, and BBR1). When these mutations were introduced into binary vectors used for Agrobacterium-mediated transformation (AMT), we observed improved transient transformation of Nicotiana benthamiana in all four origins, creating a novel tool to screen the impact of copy number variance on downstream transformation. For the best-performing origin, pVSl, higher copy variants were isolated that increase stable transformation efficiencies by 60-100% in Arabidopsis thaliana and 390% in the oleaginous yeast Rhodosporidium toruloides. Our work provides an easily deployable framework to generate plasmid copy variants that will enable greater precision in prokaryotic genetic engineering, in addition to improving AMT efficiency.Introduction
[0096] Agrobacterium-mediated transformation (AMT) is an indispensable tool in both plant and fungal biotechnology for transgene insertion into target cells1 3. By replacing the native oncogenic genes used by the bacterial pathogen Agrobacterium tumefaciens with user-defined DNA sequences, the innate DNA transfer capacity of this bacteria can be exploited for untargeted DNA insertions into diverse plant, fungal, and mammalian cell lines4 7. Initial improvements to AMT involved relocating two DNA sequences known as the left and right border (LB and RB) from the native ~200kb tumor-inducing (Ti) plasmid to a smaller helper plasmid known as a binary vector8. Any DNA contained between the LB and RB will be mobilized as a transfer DNA (T-DNA) and inserted into target cells with help from virulence (vzr) genes contained within the Ti plasmid, enabling tractable engineering by cloning differentsequences into the T-DNA1. The simplicity of changing the transgene target by customizing the sequence between the LB and RB has made AMT a vital tool for agricultural biotechnology, bioenergy crop engineering, and synthetic biology.
[0097] Despite the genetic revolution that AMT initiated, efficient transformation is still a considerable bottleneck for genetic engineering in most plant species. Decades of optimization of AMT have identified induction conditions9,10, strains of tumefaciens4, and enhancements to vzr-gene expression that have increased AMT efficiency in numerous plant species11. In addition to the transgene-harboring binary vector, some protocols have utilized a second introduced plasmid, termed a ternary vector, which over-expresses various genes involved in Agrobacterium virulence to improve AMT efficiencies in recalcitrant plants such as maize11. Such work has inspired more recent research to completely refactor the Ti-plasmid which harbors the vzr-genes12, laying the groundwork for fine-tuned engineering of individual virulence components to further enhance AMT. Even with these improvements, transformation efficiencies remain low for most genetic backgrounds and limit the scale and throughput of genetic engineering projects, presenting a need for improved tools to better control and enhance the AMT process.
[0098] While many improvements to AMT have focused on either altering vzr-gene expression or sequences within the T-DNA1,2,13,14, relatively little work has been done to engineer the binary vector backbone itself. Binary vectors consist of a T-DNA region, a bacterial selectable marker, and an origin of replication (ORI) to enable proper replication and maintenance of the plasmid in Agrobacterium. The ORI, which dictates both host range and copy number, is known to impact AMT efficiency in plants15 l 7. Zhi et al. postulated a direct relationship between binary vector copy number and maize transformation efficiency through a comparison of 3 binary vector ORIs, suggesting higher copy numbers could be used to improve AMT15. A similar comparison between 4 ORIs conducted by Oltmanns et al. in Arabidopsis thaliana and a different cultivar of maize uncovered a more complicated relationship that implicated the strain of A. tumefaciens, the ORI, and the plant genetic background as factors that independently influence transformation efficiencies16. Notably, however, neither study evaluated the impact of copy number variants within the same ORI, making it difficult to isolate the impact of copy number alone while controlling for intrinsic differences inherent to each ORI. While there is evidence that different ORI copy numbers influence AMT efficiency, there has never been an attempt to systematically modify this variable for the binary vector.
[0099] In prokaryotic synthetic biology, it has long been known that the copy number of a plasmid can dramatically impact engineered metabolic pathways and the functionality of synthetic circuits18 22. As plasmid copy number influences transgene expression magnitude and metabolic load, previous work has been done to engineer plasmid copy number, particularly in narrow host-range ORIs commonly used in E. coli such as pMBl and pSClOl derivatives19,23. For broad host-range ORIs that are used within tumefaciens, engineering efforts have been extremely limited with only a few mutants isolated18 24 27. To evaluate the impact of plasmid copy number across multiple broad host-range ORIs used for AMT, a high throughput screen is needed to identify and characterize numerous copy number variants. To date, however, no generally applicable method exists that can be applied to diverse ORIs to systematically screen for copy number diversity.
[0100] In order to specifically evaluate the impact of binary vector copy on AMT, we sought to generate copy number variants in four broad host-range origins (RK2, pVSl, pSa, and BBR1) used in Agrobacterium. To accomplish this, we leveraged a high-throughput growth- coupled selection assay to rapidly identify ORI mutations that influence copy number. Across all four origins, we were able to utilize a transient expression assay in Nicotiana benthamiana to rapidly screen for mutants that could improve AMT efficiency. Using top candidate plasmid variants from this screen, we were able to significantly improve stable transformation in both Arabidopsis thaliana and the oleaginous yeast Rhodosporidium toruloides. Thus, our work demonstrates the significant impact of binary vector backbone engineering on AMT across kingdoms, creating an easily deploying strategy to improve transformation.Results and DiscussionDeveloping a high-throughput, directed evolution pipeline to diversify plasmid copy number
[0101] As previous work suggested a relationship between binary vector copy number and AMT transformation efficiency15, we sought to develop a method to systematically select for higher copy mutants across diverse ORIs. Despite numerous intrinsic differences, many ORIs utilize analogous but non-homologous replication initiation proteins to control plasmid replication and copy number. These non-homologous RepA proteins share a general function of binding in cis to a motif within the ORI to recruit various endogenous bacterial factors to enable plasmid replication28. Previous studies have demonstrated that mutations within RepA proteins can alter plasmid copy number for diverse ORIs including pSClOl19, BBR118,RK226’27, and pVSl29. To leverage this general property, the entire RepA ORF was randomly mutagenized using error-prone PCR for each ORI. These mutagenized ORFs were then built into a selection vector, and -100,000 colonies were pooled per ORI to create mutant libraries.
[0102] To select for higher copy number variants within these libraries, we utilized a directed evolution assay that coupled plasmid copy number to bacterial antibiotic tolerance (Fig. 1). This survival-coupled selection identified WT-lethal conditions that were permissible to the mutant population, enriching higher-copy mutants within this population (Fig. 6). For each selected library, the entire population’s RepA ORFs were sequenced using an Illumina MiSeq alongside unselected library controls to evaluate the selective pressure of each RepA residue on survival.
[0103] This whole-population sequencing approach enabled quantification of mutant enrichment at every residue of diverse RepA proteins (Fig. 2). By comparing mutant distribution frequencies between the unselected and selected population, residues that contributed to the survival of the mutants in WT-lethal conditions can be identified. For the RK2 and BBR1 ORIs, large fold-change enrichments for specific nucleotide positions within the selected population were found that reached 63.8 and 26.7 fold over the unselected population. The pVSl and pSa ORIs had notably lower fold change enrichments reaching 7.2 and 4.4, respectively. Across the 4 ORIs, we observed >2-fold enrichment for 38, 118, 126, and 150 nucleotide positions for pSa, pVSl, BBR1, and RK2. Mutations that were significantly enriched in the selected population compared to the unselected control were putatively associated with a higher-copy phenotype (Table 6).Table 6. Ref mutations and GFP expression phenotypes
[0104] Mapping selected residues onto a RepA AlphaF old-generated model showed a significant enrichment of mutation sites on the predicted dimerization interfaces for all 4 ORIs (Fig. 7, Fig. 8). Some RepA proteins are postulated to regulate plasmid copy number through a “handcuffing” process in which monomeric RepA promotes plasmid replication while the dimerized form inhibits further polymerase activity, finalizing plasmid copy number30. It has been hypothesized that mutations that weaken the affinity of RepA to itself may reduce dimerization and enable additional replication, resulting in a higher final plasmid copy number19, and the enrichment of selected residues on the dimerization interface for the 4 non- homologous RepA proteins screened supports this notion. From this dataset, we chose ~20 residues per ORI that were found to be highly enriched for a mutation corresponding to an amino acid substitution to further characterize copy number and AMT performance in a Nicotiana benthamiana transient expression assay.In planta screening of ORI copy number mutants reveal enhanced transient transformation
[0105] Residues that were found to be significantly enriched in our selection were cloned into uniform plant expression binary vectors. For a given ORI, identical binary vectors consisting of a constitutive plant promoter driving GFP were created, varying by a single SNP within the RepA ORF corresponding to one of the selected mutations. These vectors were transformed into the EHA105 strain of A. tumefaciens and used to screen the impact of single RepA SNPs in a transient expression assay in N benthamiana (Fig. 3a). Because the binary vectors for a given ORI were identical except for a single SNP in RepA, differences in plant GFP output were attributed to the impact of this SNP on AMT efficiency. A total of 71 candidates across all origins were screened, and mutants for all 4 ORIs were identified that significantly increased GFP output relative to their WT forms (Fig. 3b).
[0106] We found the pSa ORI to have the largest fold-change increase in expression with the E90K mutant having a 6.9x increase in GFP output relative to WT pSa. Furthermore, pSa hadthe greatest number of mutants that showed improvement over the WT form with 13 out of the 19 mutants tested having significantly higher expression. This ORI also had the largest dynamic range of GFP expression which spanned 28-fold over the 19 total mutants, significantly extending above and below WT output levels. These results are consistent with the known low WT copy number of this ORI in Agrobacterium which potentially enables large improvements to be made by increasing the copy number16.
[0107] The RK2 origin also showed a substantial increase in plant GFP output with the S20F mutant showing a 5.4x increase in expression. Notably for this ORI, 6 of the 19 mutants screened had extremely low GFP expression coupled with an observed slow growth phenotype, forming 2 distinct groups of high or low-expressing mutants rather than the gradient of outputs observed in other ORIs.
[0108] Of the 4 ORIs used in this study, pVSl had the highest GFP output in its WT form (Fig. 9). As RK2 was used in a pilot experiment and found to increase GFP expression by 5.4x above WT levels for the best mutant, a lower strength constitutive promoter, pCL231, was used to screen pVSl mutants to prevent saturating the machine’s GFP channel. A 2. lx increase in GFP output was observed for the best-performing pVSl mutant, R106H. The BBR1 ORI was also found to have a high level of expression in its WT form, so the same pCL2 promoter was used to drive GFP expression in planta. For the 16 BBR1 mutants screened, 10 had significantly lower GFP outputs than the WT form, making this the only ORI with a majority of mutants having decreased expression. The top-performing mutant, El 82V, showed a 1.7x increase in GFP output relative to WT BBR1. As no previous work had established the copy number of BBR1 in A. tumefaciens, this result suggests that increasing the copy number may not be as beneficial for this origin.
[0109] Collectively, our results demonstrate that manipulation of the binary vector ORI greatly expands the dynamic range of T-DNA delivery to target cells, creating an additional orthogonal variable to regulate transient expression that is independent of promoter and Agrobacterium strain selection. All ORIs screened had mutants with variable expression levels that both significantly increased or decreased GFP output relative to WT levels, indicating selection of ORI variants can be used to dynamically tailor transient expression levels by up to 28-fold for a given ORI. The ability for mutants to increase transient expression in an additional strain of A. tumefaciens was also validated using the strain GV3101, and top-performing mutants were assayed in another transient plant system, Lactuca sativa (Fig. 9). Taken together,these results demonstrate that manipulation of the binary vector ORI can enable user-controlled manipulation of T-DNA transfer, refining our control over AMT in a transient system.Plasmid copy number and Agrobacterium growth rate both contribute to AMT in ORI- specific manners
[0110] To determine the relationship between binary vector copy number and N. benthamiana transient expression efficiency, we used digital PCR (dPCR) to quantify the plasmid copy number of each construct (Fig. 10). While growing the cultures for copy number quantification, we observed that some samples had notably different growth rates compared to the WT ORIs. Thus, we also quantified the growth rates of A. tumefaciens harboring WT or mutant ORIs in a time course plate growth assay (Fig. 10). While previous work postulated that increased binary vector copy numbers would correspond to higher AMT efficiencies due to increased T-DNA delivery to target cells15, our results demonstrate a more complex relationship between binary vector copy number, strain growth rate, and AMT-mediated GFP output (Fig. 4).[OHl] Of the ORIs tested in this study, pSa showed the most direct relationship between plasmid copy number and GFP output. The WT copy number of 4.5 increased up to 49 copies per cell, and no plateau was observed for the regression line between copy number and GFP output, suggesting the copy number can be increased further to enhance AMT performance. No significant relationship was observed between copy number and growth rate, or between growth rate and GFP output for pSa, making copy number the only measured variable of importance.
[0112] Unlike the trend observed for pSa, RK2 exhibited an increase in GFP output with higher copy numbers, but beyond an optimal level, excessively high copy numbers led to a decline in performance. The WT RK2 copy number of 1.2 increased to a maximum of 18 copies per cell in the I305K mutant, but this construct showed a notably lower AMT performance. Mid-copy numbers — i.e., between 5-15 copies — showed the best results for RK2, making the optimal copy number for this ORI lower than that of pSa, thus demonstrating ORI-specific impacts on AMT that fall outside of copy number alone. The growth rates of RK2 mutants with higher copy numbers exceeded the WT growth rate until around 7 copies per cell, after which the growth rates declined. There was a direct correlation between growth rate and GFP output, suggesting a tight relationship between these 3 variables for this ORI. The measured RK2 WT copy number of 1.2 was substantially lower than the value of 7-10 previously reported in theliterature26,27. We thus conducted a verification experiment to determine the WT copy number using 3 ODs (0.25, 0.5, and stationary phase) and 2 extraction kits, and we produced a consistent result of a WT copy number of around 1 per cell. (Fig. 11).
[0113] In a similar manner to RK2, increased copy numbers of pVSl increased GFP output until an optimal number which, if exceeded, saw declines in transient GFP expression. The WT copy number of 9.5 increased to a maximum of 66 copies, and copy numbers between 30-40 produced the highest GFP outputs. There was a direct negative relationship between copy number and growth rate with high copy mutants growing more slowly on average than the WT ORI. Unlike RK2, however, there was no observed relationship between growth rate and AMT performance, and slower growing mutants outperformed their WT counterpart.
[0114] BBR1 was the only ORI that did not follow our hypothesis, and no relationship was found between plasmid copy number and GFP output. The WT form of BBR1 replicates at the rather high 52 copies per cell in Agrobacterium, and increasing this value further saw no benefit to AMT performance. Rather than copy number, growth rate was the only measured variable of importance with faster-growing mutants tending to outperform their slower-growing counterparts. Interestingly, plasmid copy number had no relationship to growth rate, implying the mutants grew at different rates independent of the copy number of the plasmid. The top performing mutant, El 82V, had a copy number of 8, demonstrating that lower copy numbers can potentially outperform the WT form for BBR1.
[0115] We also sought to understand whether mutation fold-change enrichment in our initial screen could predict downstream AMT performance. When we assessed the Pearson's correlation of mutation enrichment with tobacco GFP, growth rate, and copy number, only weak correlations between GFP and enrichment, as well as copy number and enrichment, were found to be significant for the pSa origin alone, suggesting the fold-change enrichment within the checkerboard selection assay is not predictive of any downstream character (Fig. 12).
[0116] Taken as a whole, our results demonstrate that increased copy numbers improve AMT efficiencies for 3 of the 4 ORIs screened. Excessively high copy numbers likely impose some form of a metabolic burden on a cell as previously described20, resulting in deleterious effects for pVSl and RK2 that would likely emerge for higher copies of pSa. Growth rate of the host bacteria is also an important variable for some but not all ORIs and is the only measured variable of importance for BBR1. The intrinsic differences between ORIs and their interaction with host machinery in vivo likely contribute to the unique relationships observed for each ORI.Our findings demonstrate that there are no unifying principles that can explain how different ORI copy numbers correlate with AMT efficiency, highlighting the need for empirical, in vivo testing of each mutant variant to evaluate what is otherwise an unpredictable impact on AMT performance.Superior mutants from the N. benthamiana screen increase stable transformation of Arabidopsis thaliana and Rhodosporidium toruloides
[0117] To evaluate if the results from the transient tobacco screen could be translated into stable transformation systems for both plants and fungi, stable transformation experiments were conducted for Arabidopsis thaliana and Rhodosporidium toruloides. As pVSl typically achieves the highest transformation efficiency of commonly used ORIs, we selected 3 mutants and the WT form of pVSl to determine if this highly efficient ORI could be improved. Additionally, the top mutant and WT of RK2 and pSa were selected to demonstrate potential improvement in other ORIs. For the A. thaliana transformation, 20 plants were floral dipped per construct, and the seeds were bulked into a single stock. A seed weight corresponding to -26,500 seeds was selected on kanamycin media and assayed for total plant recovery. The pVSl R106H mutant increased the A. thaliana stable transformation efficiency by 60% over WT pVS 1. (Fig. 5). RK2 and pSa had more dramatic increases at 2800% and 280% respectively but with lower total transformation efficiencies than pVS 1. The 2nd and 3rd highest expressing pVSl mutants were also tested and found to increase stable transformation by 35% and 24% (Fig- 13)
[0118] To further validate this result, a construct delivering a T-DNA conferring hygromycin resistance along with a constitutively expressed Ruby reporter was used to floral dip 20 additional plants for the WT and R106H forms of pVSl. From an additional -25,000 seeds, the R106H mutant improved the transformation efficiency in this experiment by 103% over WT, and the red pigmentation of the plantlets confirmed their status as transgenic (Fig. 14).
[0119] For the oleaginous yeast R toruloides, the top constructs from pVSl and RK2 were used alongside their WT counterparts. Both mutants improved the stable transformation efficiency, by 390% (pVSl) and 510% (RK2) compared to their WT ORIs. Taken together, these results demonstrate that mutants with increased T-DNA delivery to target cells — as identified in the N. benthamiana screen — also have increased stable transformation efficiencies, making N. benthamiana a viable platform for rapidly screening ORI mutant impact on AMT.Conclusions
[0120] Using a directed evolution selection assay, this study created and characterized the largest library of binary vector copy number variants to date, surpassing variant levels found for these ORIs in any other bacterial species. As the 4 ORIs engineered are from unique incompatibility groups, these variants will likely be useful for various prokaryotic engineering efforts. RepA proteins from these ORIs are distinct in their form, regulation, and function within the host, and therefore, the growth-coupled mutational pipeline is likely amenable to any ORI that replicates with a RepA protein, creating a high-throughput tool to engineer numerous broad and narrow host range plasmids. As plasmid copy number can dramatically impact genetic engineering in bacteria, the libraries of characterized copy number mutants will likely serve to enable more precise genetic control in non-model bacteria where previously no such variants existed.
[0121] The copy number libraries created in this study enable us to study the impact of binary vector copy number on AMT at a resolution not previously possible. We have demonstrated that single SNPs in a binary vector backbone can vary transient expression levels by 28-fold for a given ORI within tobacco, demonstrating a large gradient of T-DNA delivery to target cells. We have demonstrated that the results from this tobacco transient expression screen can be translated to improvements in stable transformation efficiency in both plants and fungi. Thus, engineering single SNPs into the binary vector backbone can notably improve AMT efficiency, presenting an easily deployable method for enhancing downstream transformation.
[0122] For such transformation, single T-DNA insertions are often desired, and previous studies have determined that different ORIs that replicate at varying copy numbers alter the average number of T-DNA insertions per cell16. This library should enable a user to titrate variable amounts of T-DNA into a cell to find the ideal copy number for both optimized transformation efficiency and desired insertion number. Future work should focus on characterizing the nearly 2 orders of magnitude range of copy numbers to determine the relationship between binary vector copy number, stable transformation efficiency, and quality event generation.
[0123] From a gene editing perspective, the number of repair templates delivered to a cell is thought to limit the efficiency of homology directed repair (HDR)32,33. Transient delivery of increased levels of repair template using a higher-output mutant along with CRISPR-Cas9 or another targeted nuclease may enhance HDR efficiencies, and higher-copy mutants from eachORI can be used as tools to evaluate this. Further extensions of this work through stacking individual mutations and exploring different amino acid substitutions at enriched residues should enable an even greater range of copy numbers and AMT performances, providing a route for further optimization. This is particularly true for pSa which did not have an observed plateau for GFP output with higher copy numbers, suggesting this number can be raised even further to improve AMT. Taken as a whole, we have created a toolkit of copy number variants derived from single RepA SNPs that can successfully tailor plant transient expression levels as well as improve stable transformation in plant and fungal systems, expanding and refining our control over the critical process of AMT.MethodsMedia, chemicals, and culture conditions
[0124] Routine bacterial cultures were grown in Luria-Bertani (LB) Miller medium (BD Biosciences, USA). E. coli was grown at 37°C and A. tumefaciens at 28°C shaking at 200RPM unless otherwise noted. Cultures were supplemented with kanamycin (50 mg / L, Sigma Aldrich, USA), gentamicin (30 mg / L, Fisher Scientific, USA), spectinomycin (100 mg / L, Sigma Aldrich, USA), or rifampicin (100 mg / L, Teknova, USA) when indicated. All other compounds unless otherwise specified were purchased through Sigma Aldrich.Strains and plasmids
[0125] All bacterial strains and plasmids used in this work are listed in Table 7. All strains and plasmids created in this work are viewable through the public instance of the JBEI registry. (https: / / public-registry.jbei.org / folders / 2528). All plasmids generated in this paper were designed using Device Editor and Vector Editor software while all primers used for the construction of plasmids were designed using j 5 software34 36. Plasmids were assembled via Gibson Assembly® using standard protocols37, Golden Gate Assembly using standard protocols38, or restriction digest followed by ligation with T4 ligase as previously described39. Plasmids were routinely isolated using the Qiaprep Spin Miniprep kit (Qiagen, USA), and all primers were purchased from Integrated DNA Technologies (IDT, Coralville, IA). Plasmid sequences were verified using whole plasmid sequencing (Primordium Labs, Monrovia, CA). Agrobacterium was routinely transformed via electroporation as described previously40.Table 7. Bacterial strains and plasmids used in this work.Mutating RepA proteins via Error Prone PCR (epPCR)
[0126] For each origin used in this study, the RepA or RepA-like protein was randomly mutagenized using epPCR according to the methodology previously described41. Briefly, the RepA ORFs were amplified with a high-fidelity Phusion polymerase (NEB - M0530S) using primers to add Bsal restriction sites to the 5’ and 3’ ends of the product. Bands of the proper size were gel purified (Zymo Research - D4007), and the resulting product was used as a template. An epPCR master mix was made, containing MnCh and unequal nucleotide ratios (10 mM TrisCi [pH = 8.3], 50 mM KC1, 7 mM MgCl2, 1 mM dCTP, 1 mM dTTP, 0.2 mM dATP, 0.2 mM dGTP, 0.5 mM MnCh, 1 uM forward and reverse primers, 20 pg / uL purified RepA template, 1 uL Taq polymerase [Thermo - EP0401] per 50 uL reaction). For each reaction, a 12-cycle amplification was run to constrain the mutation number in each product strand to around 1-3 mutations. The appropriate bands were excised and gel extracted, and the resulting product digested with Bsal and cleaned via column purification (NEB - T1030S)Constructing Mutant Libraries
[0127] To make selection vectors, the gentamicin resistance gene (gentamicin-3- acetyltransferase) driven by a salicylic acid inducible promoter was originally cloned into pGingerBS-NahR42, and then variants were created for each of the pVSl, pSa, and RK2 origins using a Gibson-like assembly with NEBuilder HiFi DNA Assembly (NEB - E2621L). The entire vector except the RepA protein was then amplified in a Phusion PCR reaction with primers containing PaqCI restriction overhangs designed to have complementary sticky ends with the epPCR products from above. These PCR products were gel purified and digested withPaqCI before being column purified and ligated with the digestion product from the epPCRs using T7 DNA ligase (NEB - M0318S).
[0128] After 30 minutes of room temperature incubation, the ligation reaction was column purified, and 1 uL was electroporated into A. coli per reaction (NEB - C3020K). Following a 1 hour recovery in the supplied medium, the electroporated samples were placed into 50 mL of LB + spectinomycin. A 1 : 100 dilution was made for each flask and was plated onto solidified LB + spectinomycin to allow for an estimate of library size. Around 5 flasks were prepared for each origin in order to create a library size of at least 100,000 independent transformants. After an overnight growth at 37 °C, 5 mL aliquots from each flask were mini -prepped (Qiagen - 27104), and all samples from each origin were combined into single tubes, comprising the mutant vector library of -100,000 mutants.
[0129] 1 uL of these libraries per reaction were then electroporated into A. tumefaciens C58C1 electrocompetent cells of 4. tumefaciens. Using the same method as above with 50 mL recovery flasks, cells were recovered at 30°C overnight, and enough flasks were prepared to obtain a library size of around 150,000 mutants. 5 mL from each flask were combined and made into glycerol stocks for use in the higher copy number selection screen.Selecting higher copy mutants from constructed libraries
[0130] To attempt selection for higher plasmid copy number mutants with a checkerboard assay, an entire 1 mL glycerol stock of each of the mutant libraries was grown in 50 mL of LB + spectinomycin (50 mg / L) overnight at 30°C along with a culture of the WT strain for each ORI. For each 96-well block, 100 mL of LB + spectinomycin was inoculated with saturated culture in a 1 :200 ratio with a volume of 500 uL. Gentamicin was added to all wells in the plate following a serial dilution row-wise from rows A to H with a dilution factor of 2, starting from 3000 mg / L and ending with 23.4 mg / L. Salicylic acid (SA) was added to all wells in the plate following a serial dilution column-wise from columns 1 to 12 with a dilution factor of 2, starting from 5 uM and ending with 2.44 nM.
[0131] The resulting 96-well blocks were then covered in vent film and incubated shaking overnight at 30 °C. Growth was determined by calculating the ODeoo of each well the following day using a Genesys 50 (Thermo Scientific), and the mutant population was compared to its WT counterpart to determine gentamicin-SA combinations that were lethal to the WT ORI but that permitted growth of the mutant population. 2 selection conditions that were WT lethal per ORI were chosen and grown in triplicate in 10 mL cultures of LB + spectinomycin + gentamicin+ salicylic acid inoculated with a fresh glycerol stock of the mutant population. A culture containing just WT plasmid was also prepared for each selection condition as a negative control. Additionally, a triplicate of cultures containing the mutant library without gentamicin + salicylic acid selection was grown as an unselected control.
[0132] The following selection conditions were used: pVSl) 750 mg / L gentamicin, 5 uM salicylic acid and 375 mg / L gentamicin, 156 nM salicylic acid, pSa) 1500 mg / L gentamicin, 5 uM salicylic acid, and 750 mg / L gentamicin, 625 nM salicylic acid, BBR1) 2250 mg / L gentamicin, 5 uM salicylic acid, and 1000 mg / L, 78 nM salicylic acid. RK2) 1500 mg / L gentamicin, 5 uM salicylic acid. The overnight cultures were plasmid-prepped according to a modified Qiagen protocol:(https: / / www.qiagen.com / us / resources / resourcedetail?id=95083ccb-9583-489e-b215- 99bd91 c0604e&lang=en).
[0133] These samples were then used for the following Tagmentation procedure to barcode the Rep A ORFs to allow for identification of enriched mutations within the selected population relative to the unselected control.Sequencing mutant library survivors to identify enriched RepA mutations
[0134] The extracted plasmids from the selected populations were sequenced using an Illumina MiSeq to identify enriched mutations that may have contributed to their survival in WT-lethal conditions. The entire RepA ORF + 150 bp up and downstream were PCR amplified and gel purified. Purified samples were then fragmented into random 600 bp pieces and barcoded using the Illumina bead-linked transposon Tagmentation kit (Illumina - 20060059) according to the protocol provided by the manufacturer. The concentration of the barcoded fragments was then measured using a Qubit Fluorometer with the Qubit dsDNA HS Assay Kit (Invitrogen, USA). Library size analysis was conducted using a Bioanalyzer (Agilent Technologies, USA). To enhance the quality of Illumina reads for regions with low variations in the conserved PCR amplicons, 20% PhiX Control v3 DNA (Illumina FC-110-3001) was added to the library. The libraries were evenly pooled, and the library, along with the PhiX mixture at a concentration of 18 pM, was loaded onto the MiSeq platform using the MiSeq Reagent Kit v3 (paired-end, 2x300 bp, Illumina Inc., USA).
[0135] Mapping of Illumina reads was done as described previously43. Briefly, quality control for the reads from this run was performed using FastQC vO.11.8 and MultiQC vl .10.1. Read trimming was performed using trimmomatic v0.3944. Trimmed reads were aligned to theircorresponding reference file using Bowtie v2.4.545. Mutations in sequencing reads were identified using pysamstats vl.1.2 and pysam v0.18.0. All code from this work is publicly available on GitHub (https: / / github.com / shih-lab / origin_story), and all sequencing reads are available on the NCBI SRA under BioProject accession PRJNA1031697.Building plant GFP expression vectors with enriched RepA mutations from the MiSeq data
[0136] To construct GFP expression vectors, promoters of variable strengths derived from the PCONS suite of constitutive plant promoters were utilized31. Expression vectors containing a GFP construct driven by the constitutive plant promoters pCM2 (for RK2 and pSa) or pCL2 (pVSl and BBR1) along with a 35S::KanR component for stable-line selection were made using Gibson assembly. RK2 was run as a pilot for this project using the stronger pCM2 promoter to drive GFP. Due to the high GFP output of the WT forms of pVSl and BBR1, a very large fold-change increase in GFP expression could potentially saturate the GFP channel of the plate reader, and as such, the weaker pCL2 was used for these origins. As pSa had a weaker starting WT expression, pCM2 was used for this origin. To demonstrate the performance of the highest performing mutants expressed from the same promoter for all origins, an additional set of vectors were cloned with pCM2::GFP for pVSl and BBR1 (Fig. 9). For the Ruby vector described in Fig. 14, the strongly expressing pCH5 promoter was used to drive the Ruby reporter46.
[0137] Single SNPs were introduced into the RepA coding sequence via PCR with the SNP of interest designed into a primer overhang. Proper PCR band lengths were verified in a high throughput manner using an Agilent ZAG (zero-gel electrophoresis). PCR products were column purified and assembled using NEBuilder HiFi DNA Assembly (NEB - E2621L). The resulting assemblies were transformed via heat shock into XL 1 -Blue E. coli and grown overnight at 37°C on solidified LB + kanamycin plates. Single colonies were selected, grown in a 4 mL LB + kanamycin liquid culture, and mini-prepped, and the integrity of the plasmid sequence including the proper SNP was verified using whole-plasmid sequencing from Primordium Labs (https: / / www.primordiumlabs.com / ). Verified plasmids were then transformed into the EHA105 strain of 4. tumefaciens via electroporation and selected on plates of solidified LB + kanamycin (50 mg / L) and rifampicin (100 mg / L) at 28°C.Screening origin mutants in N. benthamiana
[0138] A previously established method for transiently expressing genes in Nicotiana benthamiana was used to screen the impact of the mutant ORIs in planta47. Single colonies of transformed EHA105 were selected and used to inoculate 5 mL of LB + kanamycin and rifampicin for overnight growth shaking at 200 RPM at 28°C. The following morning, 10 mL of fresh LB + kanamycin + rifampicin was added to each vial, and the cultures were grown for 2 hours. An aliquot from each culture was taken and used to measure the ODeoo. lOmL of each culture were then centrifuged at 4000 RPM for 15 minutes. The supernatant was removed, and cells were diluted to an ODeoo of 1.0 using an appropriate volume of tobacco infiltration buffer (10 mM MgCh, 10 mM MES, 200 uM acetosyringone [added fresh], pH = 5.7). Resuspended cultures were allowed to induce for 2 hours at room temperature while gently rocking.
[0139] N. benthamiana were grown at 25°C under long day (16 hours light, 8 hours darkness) conditions of 150 umoles / m2S PAR (400 nm-700 nm wavelength) lighting. Sunshine #4 growing mixture supplemented with Osmocote® was used as the planting medium. For each origin mutant, 6 4-week old N. benthamiana plants were syringe-infiltrated on the 4thand 5thleaf from the top of the plant. Infiltrated plants were watered and returned to the growth room for 3 days. After this 3-day period, 4 leaf disks of infiltrated tissue per leaf were taken using a standard 6mm hole puncher, and these disks were placed abaxial-side up into a 96-well culture plate filled with 320 uL of water such that the disks were floating on the surface. These 96- well plates were then assayed for GFP fluorescence using a BioTek Synergy Hl microplate reader set for 483 nm excitation and 512 nm emission conditions. GFP fluorescence readings for each disk were analyzed in R to compare mutant performances compared to the WT origins.Lactuca sativa transient expression assays
[0140] Buttercrunch lettuce plants were grown in 18-cell flats at 25°C under long day conditions of 150 umoles / m2S PAR using the same Sunshine #4 + Osmocote® planting medium. 5-week old plants were infiltrated with a blend of 2 EHA105 strains consisting of the same GFP binary vector strains as the tobacco experiments and a strain harboring a binary vector for the P19 gene. P19 is a viral gene that suppresses plant gene silencing, and coinfiltration of this construct aids in transient expression, particularly in more recalcitrant backgrounds48. These 2 strains were brought to an ODeoo of 1.0 and mixed in a 1 : 1 ratio before inducing in the tobacco infiltration medium mentioned above for 2 hours. The 5thleaf of each plant was infiltrated, and a 4 day incubation period at 25°C was used prior to harvesting leaf disks for GFP quantification.Copy number quantification
[0141] Plasmid copy number was determined using nanoplate digital PCR (dPCR). QIAcuity EvaGreen PCR kit (Cat. No. / ID: 250111) was used for all dPCR reactions. Samples were transferred into an 8.5K 96-well QIAcuity nanoplate and loaded into a QIAcuity One system. dPCR reactions were conducted with a standard protocol (2 minutes 95 °C followed by 40 cycles of 15 seconds at 95°C, 15 seconds at 56°C, 15 seconds at 72 °C). The nanoplate was imaged with an exposure duration of 150 ms and gain of 2. Images were analyzed with the QIAcuity Software Suite. The concentrations of plasmids were calculated using Poisson statistical methods by the QIAcuity Software Suite. For each biological replicate, 2 dPCR reactions were run-one targeting the single copy rpoB gene found within A. tumefaciens C58 and one targeting the kanamycin resistance (KanR) gene located on the binary vector. The final copy number was determined as the ratio of KanR:r / ?o copies which served as a measurement of the number of plasmid counts to genomic counts respectively. The following primers were used for the dPCR reactions: KanR F gatcatcctgatcgacaagaccgg. KanR R ctgccgagaaagtatccatcatggc. rpoB F GAGTACCGGAATCTCGTCAAAGCC. rpoB R CGAAGATCTCTACGGCAACTACCTGG.Arabidopsis thaliana stable transformation
[0142] Arabidopsis thaliana was transformed using the floral dip method of transformation as previously described49. Arabidopsis seeds were sprinkled over the surface of wet Sunshine #4 growing mixture and allowed to grow for 12 days. After this period, 18-pot flats were prepared with wet Sunshine #4 growing mixture with Osmocote®, and 5 Arabidopsis seedlings were transplanted per pot in an X pattern. Flats were grown at 22°C under short day conditions (8 hours light, 16 hours darkness) of 150 umoles / m2S PAR light. After 3 weeks, the flats were moved to a 22 °C chamber with long day lighting conditions of the same intensity. The first floral meristem from each plant was excised to promote axillary shoot growth. After 3 more weeks, all plants began flowering with multiple inflorescences, and these were used for the floral dip procedure.
[0143] The top performing origin, pVSl, along with lower expressing origins RK2 and pSa were selected for an analysis of impact of enhanced binary vector copy number on stable transformation. The same plasmids used in the tobacco screen were used for this experiment as they contained a 35S::KanR component that allows for selection of T1 plants with kanamycin. These plasmids were electroporated into the A. tumefaciens strain GV3101 andselected on solidified plates of LB + rifampicin, kanamycin, and gentamicin. 300 mL cultures of each strain were grown overnight at 30°C in LB + rifampicin, kanamycin, and gentamicin. These cultures were centrifuged for 20 minutes at 4000 RPM, and the supernatant was discarded and replaced with 300 mL of a floral dip medium consisting of 5% sucrose with 0.02% Silwet (50 g sucrose and 200 uL Silwet brought to 1 L with water).
[0144] Pots containing 5 A. thaliana plants were gently inverted, and the inflorescences were submerged into the Agrobacteria solution for 15 seconds with light agitation. All pots dipped in the same construct were then laid sideways in an empty planting tray and were covered with a plastic lid that had been misted with water. Plants were left at room temperature overnight in this humidity chamber before being returned to an upright position in the 22°C long-day chamber. One week after the initial floral dip, this process was repeated with the same plants to further expose newly grown buds to the bacteria. For each construct, 4 pots containing a total of 20 plants were dipped, and following the second dip, a single stake was placed into one pot, and the numerous inflorescences were gently tied together in a single mass. 2 weeks after the 2nddip, a paper bag was used to cover the inflorescence heaps, and the plants were moved to a drying rack to die and desiccate.
[0145] Following desiccation, seeds were sifted multiple times to separate them from any silique debris. 200 mg of seed per sample were measured and placed into tubes for the plating experiment. Seeds were then sterilized by first being submerged in 70% ethanol for 1 minute followed by shaking in a solution of 50% bleach from a 12.5% sodium hypochlorite stock with a drop of Triton-20 surfactant for 7.5 minutes before rinsing with sterile water 3 times. A moderately dense layer of seeds that were distributed in close by not overlapping proximity to each other was spread onto the surface of selection plates consisting of MS salts + 1 g / L MES brought to a pH of 5.7 along with 50 mg / L of kanamycin and 200 mg / L timentin to kill residual A. tumefaciens and 8 g / L agar for solidification. Each 200 mg of seeds covered around 6 5-inch plates. The plates were then wrapped in vent tape and placed into a long day 22°C chamber and allowed to grow for 3 weeks before quantification of successfully growing non-chlorotic plants.Rhodosporidium toruloides transformation
[0146] To construct a Rhodosporidium toruloides expression vector, the plasmid JPUB 01352350was modified to contain the WT or mutant forms of the pVSl or RK2 ORIs (R106H and S20F mutations respectively). Origin variants from this base vector were createdusing Gibson assembly and were electroporated into the EHA105 strain of A. tumefaciens. Single colonies were used to inoculate LB + kanamycin and rifampicin for overnight growth at 28 °C. Agrobacterium-mediated transformation of R toruloides was conducted as previously described51.
[0147] Selection was conducted on plates containing nourseothricin, and for each transformation, plates with a 1 : 10 dilution along with full-concentration plating of remaining cells were made. Plates were incubated at 30°C for 2 days and then imaged using a AnalytikJena UVP GelSolo followed by colony count quantification using Fiji (https : / / imagej .net / software / fij i / downloads)V. References1. Nester, E. W. Agrobacterium: nature’s genetic engineer. Front. Plant Sci. 5, 730 (2014).2. Gelvin, S. B. Agrobacterium-Mediated Plant Transformation: the Biology behind the “Gene-Jockeying” Tool. Microbiol. Mol. Biol. Rev. 67, 16-37 (2003).3. Thompson, M. G. el al. Agrobacterium tumefaciens: A Bacterium Primed for Synthetic Biology. BioDesign Research 2020, 8189219 (2020).4. Hwang, H.-H., Yu, M. & Lai, E.-M. Agrobacterium-mediated plant transformation: biology and applications. Arabidopsis Book 15, e0186 (2017).5. de Groot, M. J., Bundock, P., Hooykaas, P. J. & Beijersbergen, A. G. Agrobacterium tumefaciens-mediated transformation of filamentous fungi. Nat. Biotechnol. 16, 839-842 (1998).6. Kunik, T. el al. Genetic transformation of HeLa cells by Agrobacterium. Proc Natl Acad Sci USA 98, 1871-1876 (2001).7. Zambryski, P. et al. Ti plasmid vector for the introduction of DNA into plant cells without alteration of their normal regeneration capacity. EMBO J. 2, 2143-2150 (1983).8. Lee, L.-Y. & Gelvin, S. B. T-DNA binary vectors and systems. Plant Physiol. 146, 325- 332 (2008).9. Dillen, W. et al. The effect of temperature on Agrobacterium tumefaciens-mediated gene transfer to plants. Plant J. 12, 1459-1463 (1997).10. Winans, S. C., Kerstetter, R. A. & Nester, E. W. Transcriptional regulation of the virA and virG genes of Agrobacterium tumefaciens. J. BacterioL 170, 4047-4054 (1988).11. Anand, A. et al. An improved ternary vector system for Agrobacterium-mediated rapid maize transformation. Plant Mol. Biol. 97, 187-200 (2018).12. Thompson, M. et al. Genetically refactored Agrobacterium-mediated transformation. BioRxiv (2023) doi: 10.1101 / 2023.10.13.561914.13. Komari, T. etal. Binary vectors and super-binary vectors. Methods Mol. Biol. 343, 15-41 (2006).14. Altpeter, F. et al. Advancing crop transformation in the era of genome editing. Plant Cell 28, 1510-1520 (2016).15. Zhi, L. et al. Effect of Agrobacterium strain and plasmid copy number on transformation frequency, event quality and usable event quality in an elite maize cultivar. Plant Cell Rep. 34, 745-754 (2015).16. Oltmanns, H. et al. Generation of backbone-free, low transgene copy plants by launching T-DNA from the Agrobacterium chromosome. Plant Physiol. 152, 1158-1166 (2010).17. Vaghchhipawala, Z. et al. RepB C-terminus mutation of a pRi-repABC binary vector affects plasmid copy number in Agrobacterium and transgene copy number in plants. PLoS ONE 13, e0200972 (2018).18. Tao, L., Jackson, R. E. & Cheng, Q. Directed evolution of copy number of a broad host range plasmid for metabolic engineering. Metab. Eng. 7, 10-17 (2005).19. Thompson, M. G. et al. Isolation and characterization of novel mutations in the pSClOl origin that increase copy number. Sci. Rep. 8, 1590 (2018).20. Rouches, M. V., Xu, Y., Cortes, L. B. G. & Lambert, G. A plasmid system with tunable copy number. Nat. Commun. 13, 3908 (2022).21. Cook, T. B. et al. Genetic tools for reliable gene expression and recombineering in Pseudomonas putida. J. Ind. Microbiol. BiotechnoL 45, 517-527 (2018).22. Lee, T. S. et al. BglBrick vectors and datasheets: A synthetic biology platform for gene expression. J. Biol. Eng. 5, 12 (2011).23. Wadood, A., Dohmoto, M., Sugiura, S. & Yamaguchi, K. Characterization of copy number mutants of plasmid pSClOl. J. Gen. Appl. Microbiol. 43, 309-316 (1997).24. Itoh, Y., Soldati, L., Leisinger, T. & Haas, D. Low- and intermediate-copy-number cloning vectors based on the Pseudomonas plasmid pVSl. Antonie Van Leeuwenhoek 54, 567-573 (1988).25. Murai, N. Review: Plant Binary Vectors of Ti Plasmid in Agrobacterium tumefaciens with a Broad Host-Range Replicon of pRK2, pRi, pSa or pVS 1. Am. J. Plant Sci. 04, 932-939 (2013).26. Fang, F. C., Durland, R. H. & Helinski, D. R. Mutations in the gene encoding the replication-initiation protein of plasmid RK2 produce elevated copy numbers of RK2 derivatives in Escherichia coli and distantly related bacteria. Gene 133, 1-8 (1993).27. Durland, R. H., Toukdarian, A., Fang, F. & Helinski, D. R. Mutations in the trfA replication gene of the broad-host-range plasmid RK2 result in elevated plasmid copy numbers. J. Bacteriol. 172, 3859-3867 (1990).28. del Solar, G., Giraldo, R., Ruiz-Echevarria, M. J., Espinosa, M. & Diaz-Orejas, R. Replication and control of circular bacterial plasmids. Microbiol. Mol. Biol. Rev. 62, 434- 464 (1998).29. Heeb, S. et al. Small, stable shuttle vectors based on the minimal pVSl replicon for use in gram -negative, plant-associated bacteria. Mol. Plant Microbe Interact. 13, 232-237 (2000).30. Gasset-Rosa, F. et al. Negative regulation of pPSlO plasmid replication: origin pairing by zipping-up DNA-bound RepA monomers. Mol. Microbiol. 68, 560-572 (2008).31. Zhou, A. et al. A suite of constitutive promoters for tuning gene expression in plants. ACS Synth. Biol. 12, 1533-1545 (2023).32. Hahn, F., Eisenhut, M., Mantegazza, O. & Weber, A. P. M. Homology-Directed Repair of a Defective Glabrous Gene in Arabidopsis With Cas9-Based Gene Targeting. Front. Plant Sci. 9, 424 (2018).33. Aird, E. J., Lovendahl, K. N., St Martin, A., Harris, R. S. & Gordon, W. R. Increasing Cas9-mediated homology-directed repair efficiency through covalent tethering of DNA repair template. Commun. Biol. 1, 54 (2018).34. Ham, T. S. et al. Design, implementation and practice of JBELICE: an open source biological part registry platform and tools. Nucleic Acids Res. 40, el41 (2012).35. Chen, J., Densmore, D., Ham, T. S., Keasling, J. D. & Hillson, N. J. DeviceEditor visual biological CAD canvas. J. Biol. Eng. 6, 1 (2012).36. Hillson, N. J., Rosengarten, R. D. & Keasling, J. D. j5 DNA assembly design automation software. ACS Synth. Biol. 1, 14-21 (2012).37. Gibson, D. G. et al. Enzymatic assembly of DNA molecules up to several hundredkilobases. Nat. Methods 6, 343-345 (2009).38. Engler, C., Kandzia, R. & Marillonnet, S. A one pot, one step, precision cloning method with high throughput capability. PLoS ONE 3, e3647 (2008).39. Green, M. R. & Sambrook, J. Molecular Cloning: A Laboratory Manual (Fourth Edition), Volume 1, 2 & 3. 2000 (Cold Spring Harbor Laboratory Press, 2012).40. Kaman-Toth, E., Pogany, M., Danko, T., Szatmari, A. & Bozso, Z. A simplified and efficient Agrobacterium tumefaciens electroporation method. 3 Biotech 8, 148 (2018).41. Wilson, D. S. & Keefe, A. D. Random mutagenesis by PCR. Curr. Protoc. Mol. Biol. Chapter 8, Unit8.3 (2001).42. Pearson, A. N. et al. The pGinger Family of Expression Plasmids. Microbiol. Spectr. 11, e0037323 (2023).43. Waldburger, L. et al. Transcriptome architecture of the three main lineages of agrobacteria. mSystems 8, e0033323 (2023).44. Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114-2120 (2014).45. Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357-359 (2012).46. He, Y., Zhang, T., Sun, H., Zhan, H. & Zhao, Y. A reporter for noninvasively monitoring gene expression and plant transformation. Hortic. Res. 7, 152 (2020).47. Belcher, M. S. et al. Design of orthogonal regulatory systems for modulating gene expression in plants. Nat. Chem. Biol. 16, 857-865 (2020).48. Jay, F., Brioudes, F. & Voinnet, O. A contemporary reassessment of the enhanced transient expression system based on the tombusviral silencing suppressor protein Pl 9. Plant J. 113, 186-204 (2023).49. Clough, S. J. & Bent, A. F. Floral dip: a simplified method for Agrobaclerium-m dx&i transformation of Arabidopsis thaliana. Plant J. 16, 735-743 (1998).50. Geiselman, G. M. et al. Conversion of poplar biomass into high-energy density tricyclic sesquiterpene jet fuel blendstocks. Microb. Cell Fact. 19, 208 (2020).51. Zhang, S. et al. Engineering Rhodosporidium toruloides for increased lipid production. Biotechnol. Bioeng. 113, 1056-1066 (2016).VI. Illustrative Rep Protein Sequences and origins of replication:SEQ ID NO: 1 pVSl Rep ProteinVSGRKPSGPVQIGAALGDDLVEKLKAAQAAQRQRIEAEARPGESWQAAADRIRKESRQPPAAGAPSIRKPPKGDEQPDFFVPMLYDVGTRDSRSIMDVAVFRLSKRDRRAGEV IRYELPDGHVEVSAGPAGMASVWDYDLVLMAVSHLTESMNRYREGKGDKPGRVFRPHVADVLKFCRRADGGKQKDDLVETCIRLNTTHVAMQRTKKAKNGRLVTVSEGEALISRYKIVKSETGRPEYIEIELADWMYREITEGKNPDVLTVHPDYFLIDPGIGRFLYRLARRAAGKAEARWLFKTIYERSGSTGEFKKFCFTVRKLIGSNDLPEYDLKEEAGQAGPILVMRYRNLIEGEASAGSSEQ ID NO: 2 RK2 Rep ProteinMNRTFDRKAYRQELIDAGFSAEDAETIASRTVMRAPRETFQSVGSIVQQATAKIERDS VQLAPPALPAPSAAVERSRRLEQEAAGLAKSMTIDTRGTMTTKKRKTAGEDLAKQV SEAKQAALLKHTKQQIKEMQLSLFDIAPWPDTMRAMPNDTARSALFTTRNKKIPREA LQNKVIFHVNKDVKITYTGAELRADDDELVWQQVLEYAKRTPIGEPITFTFYELCQD LGWSINGRYYTKAEECLSRLQATAMGFTSDRVGHLESVSLLHRFRVLDRGKKTSRCQVLIDEEIVVLFAGDHYTKFIWEKYRKLSPTARRMFDYFSSHREPYPLKLETFRLMCG SDSTRVKKWREQVGEACEELRGSGLVEHAWVNDDLVHCKRSEQ ID NO: 3 pSa Rep ProteinMPKNNKAPGHRINEIIKTSLALEMEDAREAGLVGYMARCLVQATMPHTDPKTSYFE RTNGIVTLSIMGKPSIGLPYGSMPRTLLAWICTEAVRTKDPVLNLGRSQSEFLQRLGM HTDGRYTATLRNQAQRLFSSMISLAGEQGNDFGIENVVIAKRAFLFWNPKRPEDRAL WDSTLTLTGDFFEEVTRSPVPIRIDYLHALRQSPLAMDIYTWLTYRVFLLRAKGRPFV QIPWVALQAQFGSSYGSRARNSPELDDKARERAERAALASFKYNFKKRLREVLIVYP EASDCIEDDGECLRIKSTRLHVTRAPGKGARIGPPPT-SEQ ID NO: 4 BBR1 Rep ProteinMEISMATQSREIGIQAKNKPGHWVQTERKAHEAWAGLIARKPTAAMLLHHLVAQM GHQNAVVVSQKTLSKLIGRSLRTVQYAVKDLVAERWISVVKLNGPGTVSAYVVND RVAWGQPRDQLRLSVFSAAVVVDHDDQDESLLGHGDLRRIPTLYPGEQQLPTGPGEEPPSQPGIPGMEPDLPALTETEEWERRGQQRLPMPDEPCFLDDGEPLEPPTRVTLPRR-SEQ ID NO: 5 pVSl Rep ORF gtgagcggtcgcaaaccatccggcccggtacaaatcggcgcggcgctgggtgatgacctggtggagaagttgaaggccgcgcag gccgcccagcggcaacgcatcgaggcagaagcacgccccggtgaatcgtggcaagcggccgctgatcgaatccgcaaagaatcc cggcaaccgccggcagccggtgcgccgtcgattaggaagccgcccaagggcgacgagcaaccagattttttcgttccgatgctctat gacgtgggcacccgcgatagtcgcagcatcatggacgtggccgttttccgtctgtcgaagcgtgaccgacgagctggcgaggtgatc cgctacgagcttccagacgggcacgtagaggtttccgcagggccggccggcatggccagtgtgtgggattacgacctggtactgatg gcggtttcccatctaaccgaatccatgaaccgataccgggaagggaagggagacaagcccggccgcgtgttccgtccacacgttgc ggacgtactcaagttctgccggcgagccgatggcggaaagcagaaagacgacctggtagaaacctgcattcggttaaacaccacgc acgttgccatgcagcgtacgaagaaggccaagaacggccgcctggtgacggtatccgagggtgaagccttgattagccgctacaag atcgtaaagagcgaaaccgggcggccggagtacatcgagatcgagctagctgattggatgtaccgcgagatcacagaaggcaaga acccggacgtgctgacggttcaccccgattactttttgatcgatcccggcatcggccgttttctctaccgcctggcacgccgcgccgcaggcaaggcagaagccagatggttgttcaagacgatctacgaacgcagtggcagcaccggagagttcaagaagttctgtttcaccgtg cgcaagctgatcgggtcaaatgacctgccggagtacgatttgaaggaggaggcggggcaggctggcccgatcctagtcatgcgcta ccgcaacctgatcgagggcgaagcatccgccggttcctaaSEQ ID NO: 6 RK2 Rep ORF atgaatcggacgtttgaccggaaggcatacaggcaagaactgatcgacgcggggttttccgccgaggatgccgaaaccatcgcaag ccgcaccgtcatgcgtgcgccccgcgaaaccttccagtccgtcggctcgatagtccagcaagctacggccaagatcgagcgcgaca gcgtgcaactggctccccctgccctgcccgcgccatcggccgccgtggagcgttcgcgtcgtctcgaacaggaggcggcaggtttg gcgaagtcgatgaccatcgacacgcgaggaactatgacgaccaagaagcgaaaaaccgccggcgaggacctggcaaaacaggtc agcgaagccaagcaggccgcgttgctgaaacacacgaagcagcagatcaaggaaatgcagctttccttgttcgatattgcgccgtgg ccggacacgatgcgagcgatgccaaacgacacggcccgctctgccctgttcaccacgcgcaacaagaaaatcccgcgcgaggcg ctgcaaaacaaggtcattttccacgtcaacaaggacgtgaagatcacctacaccggcgccgagctgcgggccgacgatgacgaact ggtgtggcagcaggtgttggagtacgcgaagcgcacccctatcggcgagccgatcaccttcacgttctacgagctttgccaggacct gggctggtcgatcaatggccggtattacacgaaggccgaggaatgcctgtcgcgcctacaggcgacggcgatgggcttcacgtccg accgcgttgggcacctggaatcggtgtcgctgctgcaccgcttccgcgtcctggaccgtggcaagaaaacgtcccgttgccaggtcct gatcgacgaggaaatcgtcgtgctgtttgctggcgaccactacacgaaattcatatgggagaagtaccgcaagctgtcgccgacggc ccgacggatgttcgactatttcagctcgcaccgggagccgtacccgctcaagctggaaaccttccgcctcatgtgcggatcggattcc acccgcgtgaagaagtggcgcgagcaggtcggcgaagcctgcgaagagttgcgaggcagcggcctggtggaacacgcctgggtc aatgatgacctggtgcattgcaaacgctagSEQ ID NO: 7 pSa Rep ORF atgcctaagaacaacaaagcccccggccatcgtatcaacgagatcatcaagacgagcctcgcgctcgaaatggaggatgcccgcga agctggcttagtcggctacatggcccgttgccttgtgcaagcgaccatgccccacaccgaccccaagaccagctactttgagcgcacc aatggcatcgtcaccttgtcgatcatgggcaagccgagcatcggcctgccctacggttctatgccgcgcaccttgcttgcttggatatgc accgaggccgtgcgaacgaaagaccccgtgttgaaccttggccggtcgcaatcggaatttctacaaaggctcggaatgcacaccgat ggccgttacacggccacccttcgcaatcaggcgcaacgcctgttttcatccatgatttcgcttgccggcgagcaaggcaatgacttcgg cattgagaacgtcgtcattgccaagcgcgcttttctattctggaatcccaagcggccagaagatcgggcgctatgggatagcaccctca ccctcacaggcgatttcttcgaggaagtcacccgctcaccggttcctatccgaatcgactacctgcatgccttgcggcagtctccgcttg cgatggacatttacacgtggctgacctatcgcgtgttcctgttgcgggccaagggccgccccttcgtgcaaatcccttgggtcgccctg caagcgcaattcggctcatcctatggcagccgcgcacgcaactcgcccgaactggacgataaggcccgagagcgggcagagcgg gcagcactcgccagcttcaaatacaacttcaaaaagcgcctacgcgaagtgttgattgtctatcccgaggcaagcgactgcatcgaag atgacggcgaatgcctgcgcatcaaatccacacgcctgcatgtcacccgcgcacccggcaagggcgctcgcatcggcccccctccg acttgaSEQ ID NO: 8 BBR1 Rep ORF atggagataagcatggccacgcagtccagagaaatcggcattcaagccaagaacaagcccggtcactgggtgcaaacggaacgca aagcgcatgaggcgtgggccgggcttattgcgaggaaacccacggcggcaatgctgctgcatcacctcgtggcgcagatgggcca ccagaacgccgtggtggtcagccagaagacactttccaagctcatcggacgttctttgcggacggtccaatacgcagtcaaggacttg gtggccgagcgctggatctccgtcgtgaagctcaAcggccccggcaccgtgtcggcctacgtggtcaatgaccgcgtggcgtggg gccagccccgcgaccagttgcgcctgtcggtgttcagtgccgccgtggtggttgatcacgacgaccagGacgaatcgctgttgggg catggcgacctgcgccgcatcccgaccctgtatccgggcgagcagcaactaccgaccggccccggcgaggagccgcccagccag cccggcattccgggcatggaaccagacctgccagccttgaccgAaacggaggaatgggaacggcgcgggcagcagcgcctgcc gatgcccgatgagccgtgttttctggacgatggcgagccgttggagccgccgacacgggtcacgctgccgcgccggtagSEQ ID NO: 9 pVSl ORI ctaagagaaaagagcgtttattagaataatcggatatttaaaagggcgtgaaaaggtttatccgttcgtccatttgtatgtgcatgccaacc acagggttcccctcgggatcaaagtactttgatccaacccctccgctgctatagtgcagtcggcttctgacgttcagtgcagccgtcttct gaaaacgacatgtcgcacaagtcctaagttacgcgacaggctgccgccctgcccttttcctggcgttttcttgtcgcgtgttttagtcgcat aaagtagaatacttgcgactagaaccggagacattacgccatgaacaagagcgccgccgctggcctgctgggctatgcccgcgtca gcaccgacgaccaggacttgaccaaccaacgggccgaactgcacgcggccggctgcaccaagctgttttccgagaagatcaccgg caccaggcgcgaccgcccggagctggccaggatgcttgaccacctacgccctggcgacgttgtgacagtgaccaggctagaccgc ctggcccgcagcacccgcgacctactggacattgccgagcgcatccaggaggccggcgcgggcctgcgtagcctggcagagccg tgggccgacaccaccacgccggccggccgcatggtgttgaccgtgttcgccggcattgccgagttcgagcgttccctaatcatcgac cgcacccggagcgggcgcgaggccgccaaggcccgaggcgtgaagtttggcccccgccctaccctcaccccggcacagatcgc gcacgcccgcgagctgatcgaccaggaaggccgcaccgtgaaagaggcggctgcactgcttggcgtgcatcgctcgaccctgtac cgcgcacttgagcgcagcgaggaagtgacgcccaccgaggccaggcggcgcggtgccttccgtgaggacgcattgaccgaggcc gacgccctggcggccgccgagaatgaacgccaagaggaacaagcatgaaaccgcaccaggacggccaggacgaaccgtttttca ttaccgaagagatcgaggcggagatgatcgcggccgggtacgtgttcgagccgcccgcgcacgtctcaaccgtgcggctgcatgaa atcctggccggtttgtctgatgccaagctggcggcctggccggccagcttggccgctgaagaaaccgagcgccgccgtctaaaaag gtgatgtgtatttgagtaaaacagcttgcgtcatgcggtcgctgcgtatatgatgcgatgagtaaataaacaaatacgcaaggggaacg catgaaggttatcgctgtacttaaccagaaaggcgggtcaggcaagacgaccatcgcaacccatctagcccgcgccctgcaactcgc cggggccgatgttctgttagtcgattccgatccccagggcagtgcccgcgattgggcggccgtgcgggaagatcaaccgctaaccgt tgtcggcatcgaccgcccgacgattgaccgcgacgtgaaggccatcggccggcgcgacttcgtagtgatcgacggagcgccccag gcggcggacttggctgtgtccgcgatcaaggcagccgacttcgtgctgattccggtgcagccaagccattacgacatatgggccacc gccgacctggtggagctggttaagcagcgcattgaggtcacggatggaaggctacaagcggcctttgtcgtgtcgcgggcgatcaaa ggcacgcgcatcggcggtgaggttgccgaggcgctggccgggtacgagctgcccattcttgagtcccgtatcacgcagcgcgtgag ctacccaggcactgccgccgccggcacaaccgttcttgaatcagaacccgagggcgacgctgcccgcgaggtccaggcgctggccgctgaaattaaatcaaaactcatttgagttaatgaggtaaagagaaaatgagcaaaagcacaaacacgctaagtgccggccgtccgag cgcacgcagcagcaaggctgcaacgttggccagcctggcagacacgccagccatgaagcgggtcaactttcagttgccggcggag gatcacaccaagctgaagatgtacgcggtacgccaaggcaagaccattaccgagctgctatctgaatacatcgcgcagctaccagag taaatgagcaaatgaataaatgagtagatgaattttagcggctaaaggaggcggcatggaaaatcaagaacaaccaggcaccgacgc cgtggagtgccccatgtgtggaggaacgggcggttggccaggcgtaagcggctgggttgcctgccggccctgcaatggcactgga acccccaagcccgaggaatcggcgtgagcggtcgcaaaccatccggcccggtacaaatcggcgcggcgctgggtgatgacctggt ggagaagttgaaggccgcgcaggccgcccagcggcaacgcatcgaggcagaagcacgccccggtgaatcgtggcaagcggcc gctgatcgaatccgcaaagaatcccggcaaccgccggcagccggtgcgccgtcgattaggaagccgcccaagggcgacgagcaa ccagattttttcgttccgatgctctatgacgtgggcacccgcgatagtcgcagcatcatggacgtggccgttttccgtctgtcgaagcgtg accgacgagctggcgaggtgatccgctacgagcttccagacgggcacgtagaggtttccgcagggccggccggcatggccagtgt gtgggattacgacctggtactgatggcggtttcccatctaaccgaatccatgaaccgataccgggaagggaagggagacaagcccg gccgcgtgttccgtccacacgttgcggacgtactcaagttctgccggcgagccgatggcggaaagcagaaagacgacctggtagaa acctgcattcggttaaacaccacgcacgttgccatgcagcgtacgaagaaggccaagaacggccgcctggtgacggtatccgaggg tgaagccttgattagccgctacaagatcgtaaagagcgaaaccgggcggccggagtacatcgagatcgagctagctgattggatgta ccgcgagatcacagaaggcaagaacccggacgtgctgacggttcaccccgattactttttgatcgatcccggcatcggccgttttctct accgcctggcacgccgcgccgcaggcaaggcagaagccagatggttgttcaagacgatctacgaacgcagtggcagcaccggag agttcaagaagttctgtttcaccgtgcgcaagctgatcgggtcaaatgacctgccggagtacgatttgaaggaggaggcggggcagg ctggcccgatcctagtcatgcgctaccgcaacctgatcgagggcgaagcatccgccggttcctaatgtacggagcagatgctagggc aaattgccctagcaggggaaaaaggtcgaaaaggtccgtttcctgtggatagcacgtacattgggaacccaaagccgtacattggga accggaacccgtacattgggaacccaaagccgtacattgggaaccggtcacacatgtaagtgactgatataaaagagaaaaaaggc gatttttccgcctaaaactctttaaaacttattaaaactcttaaaacccgcctggcctgtgcataactgtctggccagcgcacagccgaag agctgcaaaaagcgcctacccttcggtcgctgcgctccctacgccccgccgcttcgcgtcggcctatcgcggccgctggccgctcaa aaatggctggcctacggccaggcaatctaccagggcgcggacaagccgcgccgtcgccactcgaccgccggcgcccacatcaag gcaccctgcctcgcgcgtttcggtgatgacggtgaaaacctctgacacatgcagctcccggagacggtcacagcttgtctgtaagcgg atgccgggagcagacaagcccgtcagggcgcgtcagcgggtgttggcgggtgtcggggcgcagccatgacccagtcacgtagcg atagcggagtgtatactggcttaactatgcggcatcagagcagattgtactgagagtgcaccatatgcggtgtgaaataccgcacagat gcgtaaggagaaaataccgcatcaggcgctcttccgcttcctcgctcactgactcgctgcgctcggtcgttcggctgcggcgagcggt atcagctcactcaaaggcggtaatacggttatccacagaatcaggggataacgcaggaaagaacatgtgagcaaaaggccagcaaa agSEQ ID NO: 10 RK2 ORI agcactagtgtagaaaagatcaaaggatcttcttgagatcctttttttctgcgggggatcaggaccgctgccggagcgcaacccactca ctacagcagagccatgtagggccgccggcgttgtggatacctcgcggaaaacttggccctcactgacagatgaggggcggacgttgacacttgaggggccgactcacccggcgcggcgttgacagatgaggggcaggctcgatttcggccggcgacgtggagctggccag cctcgcaaatcggcgaaaacgcctgattttacgcgagtttcccacagatgatgtggacaagcctggggataagtgccctgcggtattga cacttgaggggcgcgactactgacagatgaggggcgcgatccttgacacttgaggggcagagtgctgacagatgaggggcgcacc tattgacatttgaggggctgtccacaggcagaaaatccagcatttgcaagggtttccgcccgtttttcggccaccgctaacctgtcttttaa cctgcttttaaaccaatatttataaaccttgtttttaaccagggctgcgccctgtgcgcgtgaccgcgcacgccgaaggggggtgccccc ccttctcgaaccctcccggcccgctaacgcgggcctcccatccccccaggggctgcgcccctcggccgcgaacggcctcacccca aaaatggcagccacgtagaaagccagtccgcagaaacggtgctgaccccggatgaatgtcagctactgggctatctggacaaggga aaacgcaagcgcaaagagaaagcaggtagcttgcagtgggcttacatggcgatagctagactgggcggttttatggacagcaagcg aaccggaattgccagctggggcgccctctggtaaggttgggaagccctgcaaagtaaactggatggctttcttgccgccaaggatctg atggcgcaggggatcaagatcgacggatcgatccggggaattaattccggggcaatcccgcaaggagggtgaatggcagccacgt agaaagccagtccgcagaaacggtgctgaccccggatgaatgtcagctactgggctatctggacaagggaaaacgcaagcgcaaa gagaaagcaggtagcttgcagtgggcttacatggcgatagctagactgggcggttttatggacagcaagcgaaccggaattgccagc tggggcgccctctggtaaggttgggaagccctgcaaagtaaactggatggctttcttgccgccaaggatctgatggcgcaggggatc aagatcgacggatcgatccggggaattaattccggggcaatcccgcaaggagggtgaatgaatcggacgtttgaccggaaggcata caggcaagaactgatcgacgcggggttttccgccgaggatgccgaaaccatcgcaagccgcaccgtcatgcgtgcgccccgcgaa accttccagtccgtcggctcgatagtccagcaagctacggccaagatcgagcgcgacagcgtgcaactggctccccctgccctgccc gcgccatcggccgccgtggagcgttcgcgtcgtctcgaacaggaggcggcaggtttggcgaagtcgatgaccatcgacacgcgag gaactatgacgaccaagaagcgaaaaaccgccggcgaggacctggcaaaacaggtcagcgaagccaagcaggccgcgttgctga aacacacgaagcagcagatcaaggaaatgcagctttccttgttcgatattgcgccgtggccggacacgatgcgagcgatgccaaacg acacggcccgctctgccctgttcaccacgcgcaacaagaaaatcccgcgcgaggcgctgcaaaacaaggtcattttccacgtcaaca aggacgtgaagatcacctacaccggcgccgagctgcgggccgacgatgacgaactggtgtggcagcaggtgttggagtacgcgaa gcgcacccctatcggcgagccgatcaccttcacgttctacgagctttgccaggacctgggctggtcgatcaatggccggtattacacg aaggccgaggaatgcctgtcgcgcctacaggcgacggcgatgggcttcacgtccgaccgcgttgggcacctggaatcggtgtcgct gctgcaccgcttccgcgtcctggaccgtggcaagaaaacgtcccgttgccaggtcctgatcgacgaggaaatcgtcgtgctgtttgct ggcgaccactacacgaaattcatatgggagaagtaccgcaagctgtcgccgacggcccgacggatgttcgactatttcagctcgcac cgggagccgtacccgctcaagctggaaaccttccgcctcatgtgcggatcggattccacccgcgtgaagaagtggcgcgagcaggt cggcgaagcctgcgaagagttgcgaggcagcggcctggtggaacacgcctgggtcaatgatgacctggtgcattgcaaacgctag gcagggccttgtggggtcagttccggctgggggttcagcaSEQ ID NO: 11 pSa ORI cttaaactaacgaacgtaaataaggaggatagacatggacatgcctcgtattaaaccgggtcagcgtgttatgatggcactgcgtaaaa tgattgcaagcggtgaaatcaaaagtggtgaacgtattgcagaaattccgaccgcagcagcactgggtgttagccgtatgccggttcgt atcgcactgcgttcactggaacaagaaggtctggttgttcgtctgggtgcacgtggttatgcagcccgtggtgttagcagcgatcagattcgtgatgcaattgaagttcgtggtgttctggaaggttttgcagcacgtcgtctggcagaacgtggtatgacctcagaaacccatgcacgt tttgttgtactgattgcagaaggtgaagcactgtttgcagccggtcgcctgaatggtgaagatctggatcgttatgccgcatataatcagg catttcatgataccctggttagcgcagcaggtaatggtgcagttgaaagcgcactggcacgtaatggttttgaaccgtttgcagcagccg gtgcactggccctggatctgatggacctgtctgccgaatatgaacatctgctggcagcacatcgtcagcatcaggcagttctggatgca gttagctgtggtgatgccgaaggtgcagaacgtattatgcgtgatcatgcactggcagcaattcgtaatgcaaaagtttttgaagcagca gcaagcgcaggcgcaccgctgggtgcagcatggtcaattcgtgcagattgaccatggattcttcgtctgtttcactggccgtcgttttac aacgtcgtgacactttagattgatttaaaacttcatttttaatttaaaaggatctaggtgaagatcctttttgataatcggcgctgtgcgctatg gcgatctggaagtgtcgggacgaggggcgtcccggcggtaaattcgacgtgcctgcggcgcgtcgaacaaggggcgtgcccggt gtcagttcgtggcatttttcgaggcgcgacgccatttccaaggctcctgagcattcgggtctgaccaaaggccgagccgttggcggcg ggcctctttttcatactcgtacatctgcgcgtcggttggtttgccgtaataacggtaagcccaggccatgccttcttgaaccatgatcgcat tgatgttggtgagttgtgtttggccgccggggtattgcaacggcgcgtaaacgaccccaagagtgcggccataccgatcaacctcttttt cggtcacttgaacctcttggcgaaaggtcaagtcggcgagccgttggcgagcacgggagccgaaggcttggccgctttccggtgcgt caatatcggccaatctcacgcggatggtctgacggttcaccaaaacgtcgatagtgtcaccgtcaaggattcggacgacttcaccccg gaagtcggcccaagcgggcacactgacgattaggacgacagcggccgcgaccgcgcgaagggcggcaagggcgcttttcattgtt tgcctcctgttttcaagacggctgtgagattggcgacctgctctttgagggcttccacctgaccttgcagactggcggcgcgctcgatgg cctccttggcctgtttgcgggcctcgatagcctcgttatcgcgttgggtgagcttttccatgcagcggtttagttcttcgccgctgcggcgt ttcacttcggcaagctggtcggccagcttgtcgcgctcgcgttccataggttcgagctgattcactcgttcgcggagctggtcgttttcgc gggtgaaggtgtcggctagttcgattgcttcggcaagctgctggctgatggccgctttgtcggcctcgatctgtttccgatcttcgtcaaa ccgggcgttggcgtgcgccagggcgatagcccatagcgcattgccaagctcggcaagatgctcgttgactgcaaccggcaatgggt ctgatgagggcagggtggcggtcttgcggtttttccattcagccattgcatcggaaatggttgtgaagctaccgcttccgagtttcttgcg cacggcggccaaagtgggccggatgccttcggcgtccagttcgtcggctgctcgccaaatgtcttgtttagtgattgccattcttgcggg cctctgtactgtagtatgttgtatgatactacatactacaacaatttaacagagccatcttggaatctggtgtctctgcgcctataattctgga acagctactttccgaacgactcctgcgttgatcggaaatccagaagcccgagaggttgccgcctttcgggctttttctttttcaaaaaaaa aatttataaaacgatctgttgcggccgccgggttgtgggcaaaggcgctggcgctcgacggtgggcaaccgcttgcggttgtccacg ggcggagccggtgcgcgtagcgcattgtccacaagccaagggcgaccaataattgatatatatattcataattgaaaagctaattgaac atactacttgctgtaactacttgccggagcgaggggtgtttgcaagctgttgatctgaaagggctattagcgttctcacgtgcctttttgatt agcgatttcacgtgaccttattagcgatttcacgtactccgattagcgatttcacgtaccctgattagcgatttcacgtggatagtttttggag cgggccggaaagccccgtgaatcaaggctttgcggggcattagcggtttcacgtggataactaccctctatccacaggcttccgggga taaaaaagcccgctcgacggcgggctgttggatgggaaggcttgaccaagccaagcgtagcgttggcctggtcaagtcggagggg ggccgatgcgagcgcccttgccgggtgcgcgggtgacatgcaggcgtgtggatttgatgcgcaggcattcgccgtcatcttcgatgc agtcgcttgcctcgggatagacaatcaacacttcgcgtaggcgctttttgaagttgtatttgaagctggcgagtgctgcccgctctgccc gctctcgggccttatcgtccagttcgggcgagttgcgtgcgcggctgccataggatgagccgaattgcgcttgcagggcgacccaag ggatttgcacgaaggggcggcccttggcccgcaacaggaacacgcgataggtcagccacgtgtaaatgtccatcgcaagcggaga ctgccgcaaggcatgcaggtagtcgattcggataggaaccggtgagcgggtgacttcctcgaagaaatcgcctgtgagggtgagggtgctatcccatagcgcccgatcttctggccgcttgggattccagaatagaaaagcgcgcttggcaatgacgacgttctcaatgccgaag tcattgccttgctcgccggcaagcgaaatcatggatgaaaacaggcgttgcgcctgattgcgaagggtggccgtgtaacggccatcg gtgtgcattccgagcctttgtagaaattccgattgcgaccggccaaggttcaacacggggtctttcgttcgcacggcctcggtgcatatc caagcaagcaaggtgcgcggcatagaaccgtagggcaggccgatgctcggcttgcccatgatcgacaaggtgacgatgccattggt gcgctcaaagtagctggtcttggggtcggtgtggggcatggtcgcttgcacaaggcaacgggccatgtagccgactaagccagcttc gcgggcatcctccatttcgagcgcgaggctcgtcttgatgatctcgttgatacgatggccgggggctttgttgttcttaggcatgttgttcc ctccccggcatggtgatggttggtctagtgtttgtgggtttgatgttccggcgtttgatgaacaggcgcaaggtgtgagggctgacgcct aacaactcggctgcgcgactttgcggcaagccaaggttcacgtatgcctgtacttcatcaatacggctgtccagcttcaaggcgctcga tttgctgcccttgggtcgcccgagcgtcttgccgcgctctctggcgacttgtagcgcctcggtggtacgtgcctgaatgaaatgccgctc gatctgtgcagccaagccaagcacggttgccatgatgtcgctttgtaggctgccgtccatgatgatcttctgtttggtcacatggacgatt aggccgcgctcgctcgccgctttgagaatttccaaggcggcgagggcggaaccggcaatgcgcgtaatctccggcgtcagtagcac gtcgccacgctcggccttttcgatgattgctccgagcttgcgcttgcgccagtcctttgctctgctggcaatttcttcctcgatctgtagcg gcgcgaagcctttggcgttcgcgtattcgagcaaaccgtatttttggttttccgggtcttggccgtcacgcgaaacccggagataggcat agtattttggcatttgcagggaaaacgtcagattcggttaaacatgcctcattctagcgcagattaaataggaattaaataccctgtagcg gtatagataaaacgttggtttgattaccgcctttgagtgagctgataccgctcgccgcagccgaacgaccgagcgcagcgagtcagtg agcgaggaagcctgcataacgcgaagtaaSEQ ID NO: 12 BBR1 ORI ctggcgctgggcctgtttctggcgctggacttcccgctgttccgtcagcagcttttcgcccacggccttgatgatcgcggcggccttggc ctgcatatcccgattcaacggccccagggcgtccagaacgggcttcaggcgctcccgaaggtctcgggccgtctcttgggcttgatcg gccttcttgcgcatctcacgcgctcctgcggcggcctgtagggcaggctcatacccctgccgaaccgcttttgtcagccggtcggcca cggcttccggcgtctcaacgcgctttgagattcccagcttttcggccaatccctgcggtgcataggcgcgtggctcgaccgcttgcggg ctgatggtgacgtggcccactggtggccgctccagggcctcgtagaacgcctgaatgcgcgtgtgacgtgccttgctgccctcgatgc cccgttgcagccctagatcggccacagcggccgcaaacgtggtctggtcgcgggtcatctgcgctttgttgccgatgaactccttggc cgacagcctgccgtcctgcgtcagcggcaccacgaacgcggtcatgtgcgggctggtttcgtcacggtggatgctggccgtcacgat gcgatccgccccgtacttgtccgccagccacttgtgcgccttctcgaagaacgccgcctgctgttcttggctggccgacttccaccattc cgggctggccgtcatgacgtactcgaccgccaacacagcgtccttgcgccgcttctctggcagcaactcgcgcagtcggcccatcgc ttcatcggtgctgctggccgcccagtgctcgttctctggcgtcctgctggcgtcagcgttgggcgtctcgcgctcgcggtaggcgtgctt gagactggccgccacgttgcccattttcgccagcttcttgcatcgcatgatcgcgtatgccgccatgcctgcccctcccttttggtgtcca accggctcgacgggggcagcgcaaggcggtgcctccggcgggccactcaatgcttgagtatactcactagactttgcttcgcaaagt cgtgaccgcctacggcggctgcggcgccctacgggcttgctctccgggcttcgccctgcgcggtcgctgcgctcccttgccagcccg tggatatgtggacgatggccgcgagcggccaccggctggctcgcttcgctcggcccgtggacaaccctgctggacaagctgatgga caggctgcgcctgcccacgagcttgaccacagggattgcccaccggctacccagccttcgaccacatacccaccggctccaactgcgcggcctgcggccttgccccatcaatttttttaattttctctggggaaaagcctccggcctgcggcctgcgcgcttcgcttgccggttgga caccaagtggaaggcgggtcaaggctcgcgcagcgaccgcgcagcggcttggccttgacgcgcctggaacgacccaagcctatg cgagtgggggcagtcgaaggcgaagcccgcccgcctgccccccgagcctcacggcggcgagtgcgggggttccaagggggca gcgccaccttgggcaaggccgaaggccgcgcagtcgatcaacaagccccggaggggccactttttgccggagggggagccgcgc cgaaggcgtgggggaaccccgcaggggtgcccttctttgggcaccaaagaactagatatagggcgaaatgcgaaagacttaaaaat caacaacttaaaaaaggggggtacgcaacagctcattgcggcaccccccgcaatagctcattgcgtaggttaaagaaaatctgtaatt gactgccacttttacgcaacgcataattgttgtcgcgctgccgaaaagttgcagctgattgcgcatggtgccgcaaccgtgcggcaccc taccgcatggagataagcatggccacgcagtccagagaaatcggcattcaagccaagaacaagcccggtcactgggtgcaaacgg aacgcaaagcgcatgaggcgtgggccgggcttattgcgaggaaacccacggcggcaatgctgctgcatcacctcgtggcgcagat gggccaccagaacgccgtggtggtcagccagaagacactttccaagctcatcggacgttctttgcggacggtccaatacgcagtcaa ggacttggtggccgagcgctggatctccgtcgtgaagctcaAcggccccggcaccgtgtcggcctacgtggtcaatgaccgcgtgg cgtggggccagccccgcgaccagttgcgcctgtcggtgttcagtgccgccgtggtggttgatcacgacgaccagGacgaatcgctg ttggggcatggcgacctgcgccgcatcccgaccctgtatccgggcgagcagcaactaccgaccggccccggcgaggagccgccc agccagcccggcattccgggcatggaaccagacctgccagccttgaccgAaacggaggaatgggaacggcgcgggcagcagc gcctgccgatgcccgatgagccgtgttttctggacgatggcgagccgttggagccgccgacacgggtcacgctgccgcgccggtag cacttgggttgcgcagcaacccgtaagtgcgctgttccagactatcggctgtagccgcctcgccgccctataccttgtctgcctccccg cgttgcgtcgcggtgcatggagccgggccacctcgacctgaatggaaSEQ ID NO: 13 pVSl ori (Rep binding motif) tgtggatagcacgtacattgggaacccaaagccgtacattgggaaccggaacccgtacattgggaacccaaagccgtacattgggaa ccggtcacacatgtaagtgactgatataaaagagaaaaaaggcgatttttccgcctaaaactctttaaaacttattaaaactcttaaaaccc gcctggcctgtgcataSEQ ID NO: 14 RK2 ori (Rep binding motif) ggccgccggcgttgtggatacctcgcggaaaacttggccctcactgacagatgaggggcggacgttgacacttgaggggccgactc acccggcgcggcgttgacagatgaggggcaggctcgatttcggccggcgacgtggagctggccagcctcgcaaatcggcgaaaa cgcctgattttacgcgagtttcccacagatgatgtggacaagcctggggataagtgccctgcggtattgacacttgaggggcgcgacta ctgacagatgaggggcgcgatccttgacacttgaggggcagagtgctgacagatgaggggcgcacctattgacatttgaggggctgt ccacaggcagaaaatccagcatttgcaagggtttccgcccgtttttcggccaccgctaacctgtcttttaacctgcttttaaaccaatattta taaaccttgtttttaaccagggctgcgccctgtgcgcgtgaccgcgcacgccgaaggggggtgcccccccttctcgaaccctcccgg cccgctSEQ ID NO: 15 pSa ori (Rep binding motif)gcgctcgacggtgggcaaccgcttgcggttgtccacgggcggagccggtgcgcgtagcgcattgtccacaagccaagggcgacca ataattgatatatatattcataattgaaaagctaattgaacatactacttgctgtaactacttgccggagcgaggggtgtttgcaagctgttg atctgaaagggctattagcgttctcacgtgcctttttgattagcgatttcacgtgaccttattagcgatttcacgtaSEQ ID NO: 16 BBR1 ori (Rep binding motif) gcggccaccggctggctcgcttcgctcggcccgtggacaaccctgctggacaagctgatggacaggctgcgcctgcccacgagctt gaccacagggattgcccaccggctacccagccttcgaccacatacccaccggctccaactgcgcggcctgcggccttgccccatca atttttttaattttctctggggaaaagcctccggcctgcggcctgcgcgcttcgcttgccggttggacaccaagtggaaggcgggtcaag gctcgcgcagcgaccgcgcagcggcttggccttgacgcgcctggaacgacccaagcctatgcgagtgggggcagtcgaaggcga agcccgcccgcctgccccccgagcctcacggcggcgagtgcgggggttccaagggggcagcgccaccttgggcaaggccgaag gccgcgcagtcgatcaacaagccccggaggggccactttttgccggagggggagccgcgccgaaggcgtgggggaaccccgca ggggtgcccttctttgggcaccaaagaactagatatagggcgaaatgcgaaagacttaaaaatcaacaacttaaaaaaggggggtac gcaacagctcattgcggcaccccccgcaatagctcattgcgtaggttaaagaaaatctgtaattgactgccacttttacgcaacgcataa ttgttgt
[0148] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. All publications, patents, patent applications, and sequence accession numbers cited herein are hereby incorporated by reference for the subject matter for which they are specifically cited.
Claims
WHAT IS CLAIMED IS:
1. A polynucleotide comprising a nucleic acid sequence encoding a mutant replication initiator protein (Rep) that modulates copy number of a vector compared to wildtype Rep protein, wherein the mutant Rep has at least 95% identity to a reference amino acid sequence selected from the group consisting of SEQ ID NOS: 1, 2, 3, and 4; and the polynucleotide comprises at least one substitution in the nucleic acid sequence encoding the mutant Rep protein that results in a stop codon that truncates the mutant protein or results in an amino acid substitution relative to the reference amino acid sequence.
2. The polynucleotide of claim 1, wherein the reference amino acid sequence is SEQ ID NO: 1 and the polynucleotide comprises at least one substitution in the nucleic acid sequence encoding the mutant Rep protein that results in an amino acid substitution relative to the reference amino acid sequence; wherein the mutant Rep protein increases copy number of the vector comprising a compatible origin of replication relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 1.
3. The polynucleotide of claim 2, wherein the mutant Rep comprises at least one substitution selected from the group consisting of R106H, G111D, S126Y, G131D, K187I, N198H, T199N, H201L, V218G, E220K, E222V, A223V, R227S, K229N, A246V, G256D, and K310Q.
4. The polynucleotide of claim 3, wherein the mutant Rep comprises a substitution R106H.
5. The polynucleotide of claim 2, wherein the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 2.
6. The polynucleotide of any one of claims 2-5, wherein the polynucleotide further comprises a pVSl origin of replication (ori).
7. A plasmid vector comprising the polynucleotide of any one of claims 2- 6.
8. The plasmid vector of claim 7, wherein the vector is a binary vector.
9. The polynucleotide of claim 1, wherein the reference amino acid sequence is SEQ ID NO: 2 and the polynucleotide comprises at least one substitution in the nucleic acid sequence encoding the mutant Rep protein that results in a stop codon that truncates the mutant protein or results in an amino acid substitution relative to the reference amino acid sequence; wherein the mutant Rep protein increases copy number of the vector comprising a compatible origin of replication relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 2.10 The polynucleotide of claim 9, wherein the mutant Rep comprises at least one substitution selected from the group consisting of R11G, S20F, E22_stop, Q60_stop, Truncation, P63H, R75C, S136T, E170K, Q173R, N174D, N181H, V191A, A250S, T251M, R271L, I288L, and I305K.
11. The polynucleotide of claim 10, wherein the mutant Rep comprises a substitution S20F.
12. The polynucleotide of claim 9, wherein the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 3.
13. The polynucleotide of any one of claims 9-12, wherein the polynucleotide further comprises an RK2 ori.
14. A plasmid vector comprising the polynucleotide of claim any of claims 9-13.
15. The plasmid vector of claim 14, wherein the vector is a binary vector.
16. The polynucleotide of claim 1, wherein the at least one substitution in the nucleic acid sequence encoding the mutant Rep protein encodes a mutant Rep protein comprising a substitution relative to SEQ ID NO: 3; and the mutant Rep protein increases copy number of the vector relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 3.
17. The polynucleotide of claim 16, wherein the mutant Rep comprises at least one substitution selected from the group consisting of E23K, R28C, A37P, E56K, P70T,E90K, R119S, R125L, S137L, D145N, N150I, V152F, A154P, K155N, F160L, R165W, P166S, W172R, and D181Y.
18. The polynucleotide of claim 17, wherein the mutant Rep comprises a substitution E90K.
19. The polynucleotide of claim 16, wherein the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 4.
20. The polynucleotide of any one of claims 16-19, wherein the polynucleotide further comprises a pSa ori.
21. A plasmid vector comprising the polynucleotide of any one of claims 16-20.
22. The plasmid vector of claim 21, wherein the vector is a binary vector.
23. The polynucleotide of claim 1, wherein the at least one substitution in the nucleic acid sequence encoding the mutant Rep protein encodes a mutant Rep protein comprising a substitution relative to SEQ ID NO: 4; and the mutant Rep protein increases copy number of the vector relative to a wildtype Rep protein comprising amino acid sequence SEQ ID NO: 4.
24. The polynucleotide of claim 23, wherein the mutant Rep comprises at least one substitution selected from the group consisting of K66Q, T74M, V91M, K92N, P96L, V99L, S100L, V128I, D129E, H130N, D132V, D134N, L138F, G141D, T148A, and E182V.
25. The polynucleotide of claim 24, wherein the mutant Rep comprises a substitution El 82V.
26. The polynucleotide of claim 23, wherein the mutant Rep comprises at least one substitution selected from the substitutions set forth in Table 5.
27. The polynucleotide of any one of claims 23-26, wherein the polynucleotide further comprises a BBR1 ori.
28. A plasmid vector comprising the polynucleotide of any one of claims29. The plasmid vector of claim 28, wherein the vector is a binary vector.
30. A method of screening for mutations in a wildtype Rep protein that increases the copy number of a vector comprising an origin of replication for which the wildtype Rep protein initiates replication, wherein the method comprises:(a) generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprises a vector comprising a nucleic acid sequence encoding a different mutant Rep and a nucleic acid sequence encoding an antibiotic resistance gene operably linked to an inducible promoter;(b) distributing aliquots of the library to individual compartments to be assayed in an array of test compartments, wherein the array tests a range of concentrations of antibiotic and a range of concentration of an agent that induces the inducible promoter such that each aliquot of the array of test compartments is treated with a different concentration of antibiotic and inducing agent relative to the other members of the array of test compartments;(c) detecting cells that survive in individual compartments of the array of test compartments compared to counterpart control compartments comprising wildtype Rep treated with the different concentrations of antibiotic and inducing agent;(d) processing cells of (c) that survive for high throughput sequence analysis; and(e) identifying compartments in which mutant plasmids are enriched compared to a control compartment comprising an aliquot of the mutant library, wherein the control compartment was not subject to treatment with antibiotic and inducing agent, thereby identifying Rep mutants that increase vector copy number.
31. The method of claim 30, wherein processing comprises a tagmentation reaction to barcode polynucleotides encoding different Rep mutant proteins.
32. The method of claim 30 or 31, wherein the antibiotic resistance gene is a gentamycin resistance gene, the inducible promoter is a salicylic acid inducible promoter and the inducing agent in salicylic acid.
33. The method of claim 32, wherein the salicylic acid inducible promoter is an NahR promoter.
34. The method of any one of claims 30-33, wherein the Rep mutant library is generated using an error-prone PCR (epPCR).
35. The method of any one of claims 30-34, wherein each combination of concentration of antibiotic and concentration of inducing agent is tested in triplicate using three separate compartments.
36. A method of screening for mutations in a wildtype Rep protein that decrease the copy number of a vector comprising an origin of replication for which the wildtype Rep protein initiates replication, wherein the method comprises:(a) generating a Rep mutant library that comprises a plurality of cells, wherein each cell comprises a vector comprising a nucleic acid sequence encoding a different mutant Rep and a nucleic acid sequence encoding a sacB gene operably linked to a sucrose-inducible promoter;(b) subjecting aliquots of the library distributed to individual compartments in an array of test compartments to different concentrations of sucrose;(c) detecting cells that survive in individual compartments of the array of test compartments compared to counterpart control compartments treated with the different concentrations of sucrose;(d) processing cells of (c) that survived for high throughput sequence analysis; and(e) identifying compartments in which mutant plasmids are depleted compared to a control compartment comprising an aliquot of the mutant library, wherein the control compartment was not subject to treatment with sucrose, thereby identifying Rep mutants that decrease vector copy number.
Citation Information
Patent Citations
Geminivirus resistant transgenic plants
US20090229013A1