Methods and compositions useful for genetically modifying plant cells

The BiBi system with dual binary vectors in Agrobacterium strains addresses the limitations of Agrobacterium T-DNA transfer by increasing transgene expression in plant cells, facilitating efficient introduction of complex genetic pathways and enhancing production of valuable molecules.

WO2025227114A1PCT designated stage Publication Date: 2025-10-30RGT UNIV OF CALIFORNIA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/026509
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current methods for genetic transformation of plant cells using Agrobacterium T-DNA transfer lack a predictive theoretical framework, leading to limitations in the number of strains that can infect each plant cell and potential cross-talk between transgenes, which is critical for metabolic engineering and gene libraries.

Method used

A composition comprising a first and second binary vector with different origins of replication in Agrobacterium, along with a vir helper plasmid, to introduce multiple genes into plant cells using BiBi strains that carry two binary vectors, overcoming competition between bacterial strains and increasing transgene expression.

Benefits of technology

The BiBi system significantly enhances the number of transgenes expressed in each plant cell, allowing for more efficient introduction of complex genetic pathways and improved production of molecules like high-value secondary metabolites and protein complexes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025026509_30102025_PF_FP_ABST
    Figure US2025026509_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides for a library of binary vectors comprising a first library of first binary vectors, and a second library of second binary vectors, which in combination comprise genes of interests of a biosynthetic pathway or a catabolic pathway. The present invention provides for a method for introducing the genes of interest encoding for the enzymes of a biosynthetic pathway or a catabolic pathway into a target host cell comprising: (a) providing a library of host cells, (b) contacting the library of host cells to a target host cell resulting in introducing the genes of interest into the target host cell, and (c) identifying one or more the target host cell that are capable of synthesizing or catabolizing the compound of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory Methods and compositions useful for genetically modifying plant cells Inventors: Patrick M. Shih, Simon Alamos, Matthew Szarzanowicz, Mitchell G. Thompson, Liam Kirkpatrick CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application Ser. No. 63 / 639,031, filed April 26, 2024, which is incorporated by reference in its entirety. STATEMENT OF GOVERNMENTAL SUPPORT

[0002] The invention was made with government support under Contract Nos. DE-AC02- 05CH11231 awarded by the U.S. Department of Energy. The government has certain rights in the invention. REFERENCE TO SEQUENCE LISTING

[0003] Reserved. FIELD OF THE INVENTION

[0004] The present invention is in the field of genetic transformation of eukaryotic cells. BACKGROUND OF THE INVENTION

[0005] Bacteria in the group Agrobacterium are plant pathogens that can transfer a piece of DNA known as the transfer-DNA (T-DNA) into the plant cell nucleus. The T-DNA is part of a tumor-inducing plasmid (pTi) which carries virulence associated vir genes necessary to sense, respond to and infect plant cells. Although the genes encoded within the T-DNA are necessary for tumor formation they are otherwise dispensable for all other steps of pathogenesis, up to and including DNA transfer. This makes it possible to engineer the T- DNA to carry any gene of interest which is then expressed in plant cells. It is hard to overstate how revolutionary this unique aspect of the pathogenesis of Agrobacterium has been for plant genetic engineering and synthetic biology. For example, by combining multiple Agrobacterium strains, each coding for a different enzyme in their T-DNA, it is possible to quickly reconstitute complex metabolic pathways composed of multiple steps in planta. This ability to mix and match genes in a combinatorial fashion simply by combining different engineered Agrobacterium strains has been a powerful way to rapidly iterateAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory through the design-build-test cycle required for plant synthetic biology. More recently, cultures of hundreds of thousands of Agrobacterium strains carrying different T-DNAs have been used to measure plant gene libraries in parallel using next generation sequencing. These are but a few examples of a growing repertoire of experiments of increasing sophistication that rely on Agrobacterium T-DNA transfer. Interpreting and optimizing these experiments depend on a quantitative understanding of Agrobacterium DNA transfer, in particular the number of strains that transform each plant cell. In the case of metabolic engineering, it is desirable to maximize the number of strains that infect each plant cell to ensure that all enzymes are coexpressed. The opposite is true for gene libraries of trans-acting elements where, to avoid cross-talk between library transgenes, it is paramount to ensure that each cell gets at most one construct. A fundamental aspect of any engineering discipline is having a predictive and quantitative model based on first principles that describes the behavior of the system in question, yet, despite recent progress, we still lack a predictive theoretical framework to capture agrobacterium T-DNA transfer in single plant cells. SUMMARY OF THE INVENTION

[0006] The present invention provides for a composition comprising a first binary vector and second binary vector. Each binary vector has a different origin of replication in Agrobacterium. Suitable origins of replication in Agrobacterium include, but are not limited to, pVS1, BBR1, oriV, repA, OriA, and StaA origins of replication. In some embodiments, the first binary vector and the second binary vector have different origin of replication. In some embodiments, the first binary vector has a pVS1 origin of replication, and the second binary vector has a BBR1 origin of replication. In some embodiments, the composition further comprises a vir helper plasmid.

[0007] Each binary vector comprises one or more of the following: (1) An origin of replication or ORI for E. coli or OriE, or for another bacterium: a particular element on the plasmid for starting its replication in E. coli, or in another bacterium. This component is useful for allowing maintenance of the vector in E. coli, or in another bacterium. (2) An origin of replication or ORI for Agrobacterium or OriA: a particular site on the plasmid for starting its replication in Agrobacterium. (3) Multiple cloning sites or MCS: this region containing restriction enzyme sites to allow the insertion of the gene of interest. (4) Plant selectable marker: to allow for selection of the transgenic plants. (5) Bacterial selectable marker: to allow for selection of the transformed bacteria. (6) Promoter: a site to driveAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory transcription of a gene of interest. The first binary vector comprises a first gene of interest, and the second binary vector comprises a second gene of interest. (7) Poly(A) signals: an element containing poly-A, important to produce a protein. (8) Reporter: a sequence encoding a particular protein with a specific function for monitoring the recombinant protein, such as β-glucuronidase or GUS, luciferase or LUC, or Green Fluorescent Protein or GFP), or any other fluorescent protein described herein. In some embodiments, the first binary vector and the second binary vector comprise a fluorescent protein having a different color from each other. In some embodiments, each first binary vector and / or each second binary vector comprises: (1) an origin of replication or ORI for another bacterium, (2) an origin of replication or ORI for Agrobacterium, (3) a multiple cloning sites (MCS), (4) a plant selectable marker, (5) a bacterial selectable marker, (6) a promoter capable of expressing a gene of interest, and (7) a poly(A) signal. In some embodiments, each first binary vector comprises a first fluorescent protein which fluorescences with a first color, and each second binary vector comprises a second fluorescent protein which fluorescences with a second color, wherein the first color and the second color are different.

[0008] The present invention provides for a host cell comprising a first binary vector and a second binary vector of the present invention. In some embodiments, the host cell further comprises a vir helper plasmid. In some embodiments, the host cell is an Agrobacterium species. In some embodiments, the Agrobacterium species is Agrobacterium tumefaciens. An Agrobacterium strain comprising the first binary vector and the second binary vector is also known as BiBi strain. In some embodiments, the host cell is a bacterium that is not A. tumefaciens / fabrum, for example, Escherichia coli or a Rhizobium cell, such as Rhizobium rhizogenes.

[0009] In some embodiments, each vector of the binary vector pair has a different selectable marker, such as different antibiotic resistance markers or genes. Exemplary antibiotic resistance markers or genes include, but are not limited, by the following: kanamycin resistance, spectinomycin resistance, ampicillin resistance, chloramphenicol resistance, tetracycline resistance, rifampicin resistance, carbenicillin resistance, gentamicin resistance, hygromycin resistance, and the like.

[0010] In some embodiments, each binary vector comprises a gene of interest. The present invention provides for a library of binary vectors. In some embodiments, the library of binary vectors comprises a first library of first binary vectors, and a second library of second binaryAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory vectors, wherein the first library of first binary vectors and the second library of second binary vectors in combination comprise genes of interests of enzymes (and optionally transporters) of a biosynthetic pathway or a catabolic pathway. In some embodiments, the first library of first binary vectors comprises about half of the genes of the biosynthetic pathway or catabolic pathway; and the second library of second binary vectors comprise about the other half of the genes of the biosynthetic pathway or catabolic pathway. The present invention provides for a library of host cells wherein each host cell comprises a first binary vector and a second binary vector from a library of binary vectors of the present invention, such that the library of host cells comprises all of the genes of interest encoding all of the enzymes (and optionally transporters) of a biosynthetic pathway or a catabolic pathway. In some embodiments, the library of host cells is a library of BiBi strains.

[0011] In some embodiments, the library comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 binary vectors, wherein each binary vector comprises a gene of interest which encodes an enzyme or transporter, wherein the enzymes are of a biosynthetic pathway for producing a compound of interest, or catabolic pathway to breaking down a compound of interest. In some embodiments, the gene encodes a tranporter, such as a cell transmembrane tranporter or an organelle transporter, such as a plastid or mitochondrial transporter. In some embodiments, the transporter is for transporting a compound or molecule through a membrane, such as importing a compound or molecule into a cell or organelle, or exporting a compound or molecule out of a cell or organelle. In some embodiments, the transporter is for importing a compound or molecule into a cell or organelle, wherein the compound or molecule is a precursor for the biosynthesis of the compound of interest. In some embodiments, the transporter is for exporting a compound or molecule out of a cell or organelle, wherein the compound or molecule is the compound of interest.

[0012] In some embodiments, the compound of interest is a compound or molecule that naturally synthesized or catabolized by the target host cell. In some embodiments, the precursor(s) for synthesizing the compound of interest are naturally synthesized by the target host cell, or the target host cell has been genetically modified to synthesize the precursor(s). In some embodiments, the compound of interest is a biofuel.

[0013] In some embodiments, the vir helper plasmid comprises a nucleic acid encoding refactored minimized set of Agrobacterium virulence genes. In some embodiments, the Agrobacterium virulence genes are operatively linked to one or more promoters. In someAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory embodiments, the vir helper plasmid comprises features or components described by U.S. Provisional Patent Application Ser. No.63 / 588,661, hereby incorporated by reference.

[0014] The present invention provides for a method for introducing the genes of interest encoding for the enzymes of a biosynthetic pathway or a catabolic pathway into a target host cell comprising: (a) providing a library of host cells of the present invention, (b) contacting the library of host cells to a target host cell resulting in introducing the genes of interest into the target host cell, and (c) identifying one or more the target host cell that are capable of synthesizing or catabolizing the compound of interest. In some embodiments, the contacting step comprises a plurality of target host cells. In some embodiments, all of the target host cells are of the same species, or on the same plant, or on a plurality of plants of the same species. The target host cell is a eukaryotic cell, such as a plant cell or fungal cell. In some embodiments, the genes of interest are stably integrated into the genome of the target host cell.

[0015] Agroinfiltration is a powerful and widespread method in both basic and applied plant science. We have discovered that competition between individual bacterial limits the ability of the population to introduce genetic complexity into plant cells. Herein we should that we can overcome this limitation by introducing multiple binary vectors into individual Agrobacterium cells. Our system will allow for better metabolic production from transiently expressed, genetically complex engineered pathways in plants.

[0016] Agrobacterium is used to transfer transgenes encoded in plasmids called ‘binary vectors’ into plant cells. This makes it possible to use plants infiltrated with agrobacterium as a chassis for the in planta biosynthesis of molecules of biotechnological interest such as high- value secondary metabolites and protein complexes like antigens and antibodies.

[0017] These molecules often require the expression of multiple (10 or more) transgenes, making it necessary to use multiple agrobacterium strains, each one responsible for delivering a different transgene. We discovered that different agrobacterium strains compete with each other in terms of transgene delivery, limiting the number of transgenes that can be expressed with this approach. To bypass this competition, we engineered strains of bacteria that carry two binary vectors, which we term ‘BiBi’ strains. We show that using BiBi strains one can substantially increase the number of different transgenes that are expressed in each plant cell. In addition, we developed a mathematical framework to predict the performance of ourAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory system. This advance becomes critical for biosynthetic pathways that require the expression of tens of different proteins, as demonstrated by the biosynthesis of Glucoraphanin.

[0018] We have successfully demonstrated that the “BiBi” system can increase the efficiency in which complex genetic pathways can be introduced into tobacco. A major milestone that remains in order to demonstrate utility to industry writ large would be to apply this method to complex small molecules of interest and show that 1) production of the molecule can only be observed using the BiBi method and 2) utilization of the BiBi system dramatically improves the titer of target molecules.

[0019] The invention can be used to improve the yields on the transient production of complex molecules (i.e. pharmaceutical natural product derivatives) in a target host cell, such as tobacco. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The foregoing aspects and others will be readily appreciated by the skilled artisan from the following description of illustrative embodiments when read in conjunction with the accompanying drawings.

[0021] Figure 1. Live cell fluorescence microscopy setup to quantify Agrobacterium T-DNA delivery as a function of bacterial density in plant tissues. (A) The expression of transgenes delivered into plant cells by Agrobacterium may depend on a number of processes, some of which can depend on the bacterial density. (B) Schematic of the experimental setup to dissect the steps in (A). A transgenic N. benthamiana line carrying a ubiquitously expressed nuclear-localized BFP transgene is used for agroinfiltration experiments.3 strains are mixed and infiltrated into the BFP-NLS line: two reporter strains carrying T-DNAs coding for a constitutively expressed nuclear-localized fluorescent proteins (GFP or RFP) and the Hygromycin resistance hpt gene in the same T-DNA and an empty vector strain whose T- DNA codes only for hpt. RB = T-DNA right border, LB = left border. (C) Schematic of the reporter strain titration experiments. Along the vertical axis, the infiltration OD of the reporter strains is kept constant while the total culture OD is increased using the empty vector strain. Along the horizontal axis, the reporter strains are titrated while the total OD is kept constant using the empty vector (EV) strain. Both reporter strains, GFP and RFP, are mixed in equal ratios in each titration. (D) Example maximum intensity projections of live fluorescence microscopy images obtained with the setup from (B) and (C). Scale bar = 500Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory μm.

[0022] Figure 2. The transformation of plant cells by Agrobacterium follows some but not all aspects of a Poisson distribution. (A) Mean fraction of transformable nuclei expressing the GFP fluorescence reporter as a function of the infiltration OD of the GFP-NLS reporter strain for 6 infiltration mixes with increasing total ODs. Each given total OD was achieved by adding empty vector cells to each mix as shown in Figure 1. Error bars are SEM. The solid curves correspond to the best fit using Equation 4 with alpha as the free fitting parameter. The value of alpha \pm its standard deviation obtained from these fits is shown on the right next to the goodness of each fit (R^2 values) (see also Fig.7). (B) The expected fraction of nuclei coexpressing GFP and RFP if the two reporters are independent (y-axis) is plotted against the observed fraction of nuclei expressing both GFP and RFP (x-axis). (C) The alpha parameters obtained from the fits in (A) are plotted against the total infiltration mix OD. The solid lines correspond to a linear fit to the log of these alpha values. Note that because the y- axis is shown in a log-scale, alpha is an exponentially decaying function of the total OD. As a reference, the alpha estimate from Carlson 2023 is also shown. Error bars show the standard error of the fits. (D) Left: The average fraction of cells expressing GFP from (A) is plotted against the scaled OD of the reporter strain. Also shown is the RFP data from the same set of experiments. Shown on top of these data are the fraction of transformed cells under variable OD where EV cells were not added to the mix. Right: same as left except that the OD of the reporter strains was multiplied by a scaling factor e^{mOD_{tot}}, where m is the slope of the linear fits in (C) and OD_tot is the total OD of the infiltration mix. The solid line shows the Poisson prediction for alpha = c, where c is the y-intercept of the linear fits in (C). Error bars show the standard deviation. In (A),(B), and (D) N = 4, 7, 5, 8, 6, and 5 plants (one image per plant) for total ODs 0.05-3, respectively. N = 5 plants for the variable OD data in (D).

[0023] Figure 3. Competition is independent of pathogenesis and attachment and can be explained by a limiting host metabolite. (A) Cartoon of the metabolite competition model. Top: the reporter and the competitor strain start off in a nonpathogenic state and become capable of transformation upon consumption of a metabolite. Bottom: when the EV strain is added it deprives the reporter strain of the limiting metabolite. (B) Cartoon of the spatial competition model. Top: each plant cell has a limited number of slots on its surface where bacteria can make contact. Bottom: Increasing the OD of the EV strain leaves fewerAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory transformation slots available for the reporter strain. (C) Fit of the models in (A) and (B) to the data in Figure 2, Panel A. (D) Tissue-level GFP fluorescence of agroinfiltrated leaves measured with a plate reader.3 different strains carrying a GFP transgene under promoters of different strengths were titrated in buffer and infiltrated alone. Shown on top is where the tissue fluorescence should saturate according to each of the models in (A) and (B). (E) Fitted values of the transformation probability constant ɑ (y-axis) as a function of the total OD (x- axis). To obtain ɑ the reporter strains were titrated as in Figure 2, Panel A, except that using C58C1 instead of EV to keep the total OD constant at 0.1, 0.5, and 2.0 (see also Fig.11). (F) Leaf GFP fluorescence measured with a plate reader for a pCM2-GFP reporter strain infiltrated at an OD of 0.025. This reporter strain was combined with a second strain at an OD of 0.075 or 1.975 to achieve a final total OD of 0.1 and 2.0, respectively. Buffer only was used as a control. In (D) and (F) N = 24 (6 plants, 16 leaf punches / plant). In (E) N = 5 plants with one image per plant for each of the total ODs, 0.1, 0.5 and 2.0. In (C)-(E) error bars are SEM.

[0024] Figure 4. Secreted proteins encoded in the Ti plasmid allow different strains to cooperatively achieve higher levels of T-DNA expression in infected plant cells. (A) Top: the reporter strain is infiltrated alone or combined with a second strain and the tissue-level fluorescence is measured. Bottom: box plots showing the GFP fluorescence of Tobacco leaves infiltrated with a reporter strain carrying a pCM2-GFP transgene in its T-DNA at a constant OD of 0.025. This reporter strain was co-infiltrated with the EV strain (yellow boxes) or C58C1 cells (purple boxes) to a final total OD of 0.1 or 2. An infiltration using the reporter strain alone diluted in the infiltration buffer was used as a control (light blue). (B) Top: Cartoon showing the possible cellular basis for the results in (A) where adding a second strain increases the fraction of cells transformed by the reporter strain. Bottom: box plots showing the fraction of nuclei expressing GFP for the same infiltration conditions as in (A). (C) Top: Cartoon showing an alternative cellular basis for the results in (A) where adding a second strain increases the fluorescence of those cells transformed by the reporter strain. Bottom: box plots showing the average GFP fluorescence of nuclei that were detected in the GFP channel in (B) (for (B) and (C) N=9 plants, one image per plant). (D) An RFP reporter strain was infiltrated alone with buffer or in combination with a second helper strain. All the helper strains share the GV3101 genetic background and carry a GFP T-DNA. ΔVirE12 is a knockout for the VirE1 and VirE2 genes. The ΔVirE12 complement strain carries a previously described complementation construct coding for the VirE1 and VirE2 genes in aAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory plasmid. Box plots show the tissue-level RFP fluorescence. (E) Leaves were infiltrated with the GFP and RFP reporter strains mixed at an OD of 0.002 each, with (bottom) or without (top) the addition of the EV strain at an OD of 0.046. Density scatter plots show the GFP (x- axis) and RFP (y-axis) fluorescence of nuclei detected as expressing both reporters. Shown in the top right corner are Pearson's correlation coefficients (⍴). Top scatter N = 775 nuclei, bottom scatter N = 1175 nuclei. Nuclei were pooled from N=10 plants with two images per plant. (F) Mean ± standard deviation of tissue-level fluorescence. Across all measurements, the OD of a GFP reporter strain was kept constant at 0.025 and the OD of a second coinfiltrated strain (EV or C58C1) was titrated to achieve varying total ODs, from 0.025 to 2. In (A)-(D) the numbers on top of box plots show the p-value of a two-sided Student's t-test rounded to the second decimal place. In (A), (D), and (F), N=48 (6 plants with 8 measurements per plant).

[0025] Figure 5. Characterization of plasmid co-delivery in BiBi strains carrying two binary vectors per cell shows increased transgene expression diversity per unit of OD. (A) Maximum intensity projection images of BFP-NLS tobacco leaves infiltrated with reporter bacteria at an OD of 0.002 and EV at an OD of 0.498 (total of OD 0.5, BFP not shown). Top: reporter bacteria correspond to two strains, GFP-NLS in the pVS1 plasmid and RFP-NLS in the BBR1 plasmid, each at an OD of 0.001. Bottom: a single BiBi reporter strain carrying two plasmids, GFP-NLS in pVS1 and RFP-NLS in BBR1. Blue arrowheads indicate example nuclei expressing both fluorescent proteins. White and black arrowheads indicate nuclei expressing only RFP or GFP, respectively. Scale bar = 500 µm. (B) Cartoon of the BiBi codelivery model. T-DNA expression is modeled as a two step process. First, the bacterium has to contact the plant cell, an event that follows Poisson-like statistics. Next, this bacterium can transfer the pVS1 plasmid with a probability p and the BBR1 plasmid with a probability r. Delivery of both plasmids is independent of the other, making the co-delivery probability p x r. (C) Fraction of leaf epidermis cells expressing GFP or RFP as a function of reporter bacteria infiltration OD. The total OD was kept constant at 0.5 using EV. Solid lines show the best fit of the model in (B) which assumes that in all cases strains share the same values of ɑ, p, and r. (D, left) Expected fraction of cells coexpressing GFP and RFP if delivery of these transgenes is independent plotted as a function of the observed fraction of cells coexpressing both reporters (as in Fig.2, Panel B). The dashed line shows y=x. The solid line shows the best fit of the model in (B) to the BiBi data using the same parameter values as in (C). The dotted line shows the model prediction if p=r=1.0. (D, right) Fraction of cells coexpressingAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory both reporters as a function of the infiltration OD of reporter bacteria. The solid line shows the model fit using the same parameters as (C).

[0026] Figure 6. Testing the implications of agrobacterium competition for long engineered biosynthetic pathways. (A) Theoretical prediction of the fraction of cells coexpressing all N plasmids if these plasmids are harbored by N regular strains (dashed lines) or N / 2 BiBi strains (solid lines) for different total OD of the mixture. (B) Overview of the glucoraphanin biosynthesis reconstitution experiment in tobacco.14 transgenes need to be transiently expressed via agroinfiltration for the efficient biosynthesis of glucoraphanin from endogenous methionine as a precursor in tobacco leaves. These transgenes encode for 13 enzymes and one plastid transporter. All reactions occur in the cytosol except for the ones catalyzed by enzymes 3, 4, and 5 which are plastidial. (C) Plasmids used to clone each of the 14 genes in (B). (D) The plasmids in (C) are used to create 4 different agrobacterium mixes. The pVS1 mix consists of 14 strains carrying one pVS1 plasmid per strain. Similarly, the BBR1 mix is composed of 14 strains each of which carries one of the 14 transgenes in a BBR1 plasmid. BiBi mixes consist of 7 strains carrying two plasmids per cell. In BiBi mix 1 the odd-numbered transgenes are launched from pVS1 and even ones from BBR1. The opposite arrangement of transgenes and backbones was used for BiBi mix 2. (E) The BBR1 mix and BiBi mix 1 from (D) were infiltrated into tobacco leaves at 4 different ODs, 0.05, 0.1, 0.5, and 2.0 and glucoraphanin production was quantified. Shown is the normalized glucoraphanin concentration (right y-axis) and the predicted fraction of cells coexpressing all 14 transgenes (left y-axis) as a function of the infiltration OD. The dashed line shows the prediction from (A) assuming that the fraction of cells expressing the 14 transgenes is linearly proportional to glucoraphanin production. Red diamonds and green triangles show the mean ± standard deviation of N=5 plants. Inset: data and predictions in a log scale for the left and right y-axes.

[0027] Figure 7. Related to Figure 2: OD titration results for the RFP reporter strain. Shown is the RFP-NLS mean ± SEM data from the same set of experiments shown in Figure 2, Panel A. As in Figure 2, Panel A, the solid lines correspond to Poisson fits using Equation 4. Also shown on the right are the parameter fits for ɑ and the goodness of fit (R2).

[0028] Figure 8. Related to Figure 2: Estimation of the fraction of transformable cells. (A) Representative Maximum intensity projection images of a BFP-NLS tobacco leaf infiltrated with saturating ODs of GFP-NLS and RFP-NLS reporter cells (0.5 each). (B) details from theAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory merged image in (A). Empty arrowheads indicate nuclei that do not express GFP or RFP and are considered to be non-transformable by Agrobacterium in our setup. These nuclei are much smaller than other epidermis nuclei and tend to occur in pairs or clusters. (C) Example of the analysis used to obtain the fraction of transformable cells from images like the ones in (A). For each image the fraction of BFP nuclei that express GFP or RFP as a function of reporter strain OD was fitted to a sigmoidal function. The fraction at which the fits plateau is considered the ‘fraction of transformable cells’ of this image. (D) The fraction transformable was estimated as in (C) for all images. The plot on the left shows the mean ± SEM of this estimated fraction for GFP and RFP as a function of the total infiltration OD. The histogram on the left shows the distribution of the estimated values across the dataset. The mean of this distribution is used as the fraction of transformable cells throughout the article.

[0029] Figure 9. Related to Figure 2: Imaging the same set of samples in a confocal setting with overall better resolution supports the accuracy of widefield microscopy. A subset of the samples used in Figure 2 were also imaged in a laser scanning confocal microscope right after acquiring the widefield fluorescence images. In the confocal experiments a higher magnification and better signal quality setting were used, at the expense of much slower data acquisition. The fraction of cells transformed by GFP-NLS or RFP-NLS was calculated using the same pipeline as for the widefield images. Finally, the results from both experiments were plotted against one another as scatter plots.

[0030] Figure 10. Related to Figure 2: Differences in GFP and RFP fluorescence at the single nucleus level cannot explain why fewer nuclei are detected when EV is added at high ODs. The GFP-NLS and RFP-NLS reporter strains were infiltrated at OD 0.02 each with the addition of the EV strain at OD 0.46 or 1.96 for a final OD of 0.5 or 2.0. Images were acquired using confocal microscopy. (A) The pixel intensity in the GFP channel was calculated for each nucleus using BFP fluorescence to determine the nucleus volumes. In parallel, nuclei were classified as expressing GFP or not based on our image analysis pipeline. The violin plots show the fluorescence intensity in the green channel of nuclei classified as transformed by the GFP-NLS (green) or untransformed (gray). (B) Same as (A) but using the RFP channel. (C) scatter plots of single nucleus fluorescence in the green and red channels for all nuclei detected in the BFP channel. Each dot corresponds to a single nucleus and they are color-coded depending on whether they were assigned to the transformed or untransformed categories by our computational pipeline.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0031] Figure 11. Related to Figure 3: Comparison of transformation rates by reporter strains using EV or C58C1 as a competitor strain. (A) The reporter strains were titrated in equal ratios using a third strain to keep the total OD constant, as described in Figure 1, Panel C. The third competitor strain used was either EV or C58C1. Shown is the mean ± SEM of the fraction of transformable cells expressing each reporter when EV (empty markers) or C58C1 (filled markers) was used. Dashed and solid lines correspond to Poisson fits to the data for EV and C58C1, respectively. (B) The values of ɑ estimated from the fits in (A) are shown as a function of the total OD, as in Figure 2, Panel C, and Figure 3, Panel E. The EV data in (A) and (B) are the same as in Figure 2.

[0032] Figure 12. Related to Figure 5: The transformation rate of T-DNAs launched from BBR1 is similar to that of pVS1 when BBR1 strains are mixed in the absence of reporter pVS1 strains.

[0033] Figure 13. Related to Figure 5: Fraction of cells that express RFP given that they also express GFP and fit of the BiBi model to the data. As in Figure 5, reporter strains were titrated keeping the total OD constant at 0.5 using the EV strain. Shown is the fraction of nuclei that express GFP that also express RFP (y-axis) as a function of the OD of a BiBi strain or two regular strains combined. The model in Figure 5, Panel B was simultaneously fitted to the data shown here as well as data in Figure 5, Panels C, D, and E. Shown is the best model fit (solid red and blue lines). DETAILED DESCRIPTION OF THE INVENTION

[0034] Before the invention is described in detail, it is to be understood that, unless otherwise indicated, this invention is not limited to particular sequences, expression vectors, enzymes, host microorganisms, or processes, as such may vary. It is also to be understood that the terminology used herein is for purposes of describing particular embodiments only, and is not intended to be limiting.

[0035] In this specification and in the claims that follow, reference will be made to a number of terms that shall be defined to have the following meanings:

[0036] The terms "optional" or "optionally" as used herein mean that the subsequently described feature or structure may or may not be present, or that the subsequently described event or circumstance may or may not occur, and that the description includes instancesAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory where a particular feature or structure is present and instances where the feature or structure is absent, or instances where the event or circumstance occurs and instances where it does not.

[0037] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither or both limits are included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0038] The term “about” refers to a value including 10% more than the stated value and 10% less than the stated value.

[0039] As used herein, the term "promoter" refers to a polynucleotide sequence capable of driving transcription of a DNA sequence in a cell. Thus, promoters used in the polynucleotide constructs of the invention include cis- and trans-acting transcriptional control elements and regulatory sequences that are involved in regulating or modulating the timing and / or rate of transcription of a gene. For example, a promoter can be a cis-acting transcriptional control element, including an enhancer, a promoter, a transcription terminator, an origin of replication, a chromosomal integration sequence, 5' and 3' untranslated regions, or an intronic sequence, which are involved in transcriptional regulation. These cis-acting sequences typically interact with proteins or other biomolecules to carry out (turn on / off, regulate, modulate, etc.) gene transcription. Promoters are located 5' to the transcribed gene, and as used herein, include the sequence 5' from the translation start codon.

[0040] A "constitutive promoter" is one that is capable of initiating transcription in nearly all cell types, whereas a "cell type-specific promoter" initiates transcription only in one or a few particular cell types or groups of cells forming a tissue. In some embodiments, the promoter is secondary cell wall-specific and / or fiber cell-specific. A "fiber cell-specific promoter" refers to a promoter that initiates substantially higher levels of transcription in fiber cells asAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory compared to other non-fiber cells of the plant. A "secondary cell wall-specific promoter" refers to a promoter that initiates substantially higher levels of transcription in cell types that have secondary cell walls, e.g., lignified tissues such as vessels and fibers, which may be found in wood and bark cells of a tree, as well as other parts of plants such as the leaf stalk. In some embodiments, a promoter is fiber cell-specific or secondary cell wall-specific if the transcription levels initiated by the promoter in fiber cells or secondary cell walls, respectively, are at least 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 50-fold, 100-fold, 500-fold, 000-fold higher or more as compared to the transcription levels initiated by the promoter in other tissues, resulting in the encoded protein substantially localized in plant cells that possess fiber cells or secondary cell wall, e.g., the stem of a plant. Non- limiting examples of fiber cell and / or secondary cell wall specific promoters include the promoters directing expression of the genes IRX1, IRX3, IRX5, IRX7, IRX8, IRX9, IRX10, IRX14, NST1, NST2, NST3, MYB46, MYB58, MYB63, MYB83, MYB85, MYB103, PAL1, PAL2, C3H, CcOAMT, CCR1, F5H, LAC4, LAC17, CADc, and CADd. See, e.g., Turner et al 1997; Meyer et al 1998; Jones et al 2001; Franke et al 2002; Ha et al 2002; Rohde et al 2004; Chen et al 2005; Stobout et al 2005; Brown et al 2005; Mitsuda et al 2005; Zhong et al 2006; Mitsuda et al 2007; Zhong et al 2007a, 2007b; Zhou et al 2009; Brown et al 2009; McCarthy et al 2009; Ko et al 2009; Wu et al 2010; Berthet et al 2011. In some embodiments, a promoter is substantially identical to a promoter from the lignin biosynthesis pathway. A promoter originated from one plant species may be used to direct gene expression in another plant species.

[0041] A polynucleotide or amino acid sequence is "heterologous" to an organism or a second polynucleotide or amino acid sequence if it originates from a foreign species, or, if from the same species, is modified from its original form. For example, when a polynucleotide encoding a polypeptide sequence is said to be operably linked to a heterologous promoter, it means that the polynucleotide coding sequence encoding the polypeptide is derived from one species whereas the promoter sequence is derived from another, different species; or, if both are derived from the same species, the coding sequence is not naturally associated with the promoter (e.g., is a genetically engineered coding sequence, e.g., from a different gene in the same species, or an allele from a different ecotype or variety, or a gene that is not naturally expressed in the target tissue).

[0042] The term "operably linked" refers to a functional relationship between two or moreAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory polynucleotide (e.g., DNA) segments. Typically, it refers to the functional relationship of a transcriptional regulatory sequence to a transcribed sequence. For example, a promoter or enhancer sequence is operably linked to a DNA or RNA sequence if it stimulates or modulates the transcription of the DNA or RNA sequence in an appropriate host cell or other expression system. Generally, promoter transcriptional regulatory sequences that are operably linked to a transcribed sequence are physically contiguous to the transcribed sequence, i.e., they are cis-acting. However, some transcriptional regulatory sequences, such as enhancers, need not be physically contiguous or located in close proximity to the coding sequences whose transcription they enhance.

[0043] The terms “host cell” of “host organism” is used herein to refer to a living biological cell that can be transformed via insertion of an expression vector.

[0044] The terms "expression vector" or "vector" refer to a compound and / or composition that transduces, transforms, or infects a host cell, thereby causing the cell to express nucleic acids and / or proteins other than those native to the cell, or in a manner not native to the cell. An "expression vector" contains a sequence of nucleic acids (ordinarily RNA or DNA) to be expressed by the host cell. Optionally, the expression vector also comprises materials to aid in achieving entry of the nucleic acid into the host cell, such as a virus, liposome, protein coating, or the like. The expression vectors contemplated for use in the present invention include those into which a nucleic acid sequence can be inserted, along with any preferred or required operational elements. Further, the expression vector must be one that can be transferred into a host cell and replicated therein. Particular expression vectors are plasmids, particularly those with restriction sites that have been well documented and that contain the operational elements preferred or required for transcription of the nucleic acid sequence. Such plasmids, as well as other expression vectors, are well known to those of ordinary skill in the art.

[0045] The terms "polynucleotide" and "nucleic acid" are used interchangeably and refer to a single or double-stranded polymer of deoxyribonucleotide or ribonucleotide bases read from the 5' to the 3' end. A nucleic acid of the present invention will generally contain phosphodiester bonds, although in some cases, nucleic acid analogs may be used that may have alternate backbones, comprising, e.g., phosphoramidate, phosphorothioate, phosphorodithioate, or O-methylphophoroamidite linkages (see Eckstein, Oligonucleotides and Analogues: A Practical Approach, Oxford University Press); positive backbones; non-Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory ionic backbones, and non-ribose backbones. Thus, nucleic acids or polynucleotides may also include modified nucleotides that permit correct read-through by a polymerase. "Polynucleotide sequence" or "nucleic acid sequence" includes both the sense and antisense strands of a nucleic acid as either individual single strands or in a duplex. As will be appreciated by those in the art, the depiction of a single strand also defines the sequence of the complementary strand; thus the sequences described herein also provide the complement of the sequence. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. The nucleic acid may be DNA, both genomic and cDNA, RNA or a hybrid, where the nucleic acid may contain combinations of deoxyribo- and ribo-nucleotides, and combinations of bases, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine, isoguanine, etc.

[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.

[0047] In some embodiments, the compound of interest is a compound naturally produced by the host cell. In some embodiments, the compound of interest is a compound not naturally produced by the host cell. In some embodiments, the compound of interest is a biofuel or bioproduct, or any other organic compound, and the corresponding biosynthetic enzyme(s) for producing the compound of interest thereof, are described and taught in U.S. Patent Nos. 7,985,567; 8,420,833; 8,852,902; 9,109,175; 9,200,298; 9,334,514; 9,376,691; 9,382,553; 9,631,210; 9,951,345; 10,167,488; 10,273,605; 10,814,724; and 11,660,961; and PCT International Patent Application Nos. PCT / US2014 / 48293, PCT / US2018 / 049609, PCT / US2017 / 036168, PCT / US2018 / 029668, PCT / US2008 / 068833, PCT / US2008 / 068756, PCT / US2008 / 068831, PCT / US2009 / 042132, PCT / US2010 / 033299, PCT / US2011 / 053787, PCT / US2011 / 058660, PCT / US2011 / 059784, PCT / US2011 / 061900, PCT / US2012 / 031025,Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory and PCT / US2013 / 074214 (all of which are incorporated in their entireties by reference). In some embodiments, the compound of interest is a terpene, isoprenoid, carboxylic acid, lactone, trimethylpentanoic acid, 1-deoxyxylulose 5-phosphate, 1-deoxy-D-xylulose 5- phosphate (DXP), fatty acid, or derivatives thereof, alkyl lactone, lactam, isoprenylalkanoate, 3- me t h y l - 2 - b u t e n - 1 - o l , 3 - me t h y l - 3 - b u te n - 1 - o l , an d 3 - me t h y l -b u t a n - 1 -o l , f a t t y ac i d e s t e r , a l p h a - o l e f i n , d i a c i d , d i a mi n e , sesquiterpene,bisabolene, or oxidized aromatic amino acid. In some embodiments, the compound of interest is any product or intermediate in the mevalonate (MVA) pathway, including any compound from acetyl-CoA to mevalonate. In some embodiments, the biosynthetic enzyme(s) are phosphomevalonate decarboxylase (PMD), phosphatase, AtoB, hydroxymethylglutaryl-CoA synthase (HMGS), hydroxymethylglutaryl-CoA reductase (HMGR), and / or mevalonate kinase (MK). In some embodiments, the biosynthetic enzyme(s) is a polyketide synthase.

[0048] In some embodiments, the promoters are each independently constitutive or inducible. In some embodiments, the promoters are promoters described herein.

[0049] The refactored minimized set of Agrobacterium virulence genes at least excludes: (1) the virA and virG genes, as these genes are regulatory genes; (2) the virB1 gene (Berger et al., J. Bacteriol.176(12): 3646-3660, 1994); and, (3) the virE12 gene (which can be replaced with the Agrobacterium rhizogenes GALLS gene (Hodges et al., J. Bacteriol. 191(1): 355-364, 2009).

[0050] In some embodiments, the target host cell is a eukaryotic cell, such as a plant or fungal cell. In some embodiments, the plant is a tobacco plant. In some embodiments, the fungal cell is a Rhodosporidium cell, such as Rhodosporidium toruloides. In some embodiments, the fungal cell is torulosis’s a yeast. In some embodiments, the yeast is Saccharomyces species, such as a Saccharomyces cerevisiae.

[0051] In some embodiments, when the target host cell is a plant cell, the promoter linked to the gene of interest is a constitutive promoter or a tissue-specific promoter. Examples of tissue-specific promoters under developmental control include promoters that initiate transcription only (or primarily only) in certain tissues, such as vegetative tissues, cell walls, including e.g., roots or leaves. A variety of promoters specifically active in vegetative tissues, such as leaves, stems, roots and tubers are known. For example, promoters controlling patatin, the major storage protein of the potato tuber, can be used (see, e.g., Kim, Plant Mol.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory Biol.26:603-615, 1994; Martin, Plant J.11:53-62, 1997). The ORF13 promoter from Agrobacterium rhizogenes that exhibits high activity in roots can also be used (Hansen, Mol. Gen. Genet.254:337-343, 1997). Other useful vegetative tissue-specific promoters include: the tarn promoter of the gene encoding a globulin from a major taro (Colocasia esculenta L. Schott) corm protein family, tarin (Bezerra, Plant Mol. Biol.28:137-144, 1995); the curculin promoter active during taro corm development (de Castro, Plant Cell 4:1549-1559, 1992) and the promoter for the tobacco root-specific gene TobRB7, whose expression is localized to root meristem and immature central cylinder regions (Yamamoto, Plant Cell 3:371-382, 1991).

[0052] Leaf-specific promoters, such as the ribulose biphosphate carboxylase (RBCS) promoters can be used. For example, the tomato RBCS1, RBCS2 and RBCS3A genes are expressed in leaves and light-grown seedlings, only RBCS1 and RBCS2 are expressed in developing tomato fruits (Meier, FEBS Lett.415:91-95, 1997). A ribulose bisphosphate carboxylase promoters expressed almost exclusively in mesophyll cells in leaf blades and leaf sheaths at high levels (e.g., Matsuoka, Plant J.6:311-319, 1994), can be used. Another leaf- specific promoter is the light harvesting chlorophyll a / b binding protein gene promoter (see, e.g., Shiina, Plant Physiol.115:477-483, 1997; Casal, Plant Physiol.116:1533-1538, 1998). The Arabidopsis thaliana myb-related gene promoter (Atmyb5) (Li, et al., FEBS Lett. 379:117-1211996), is leaf-specific. The Atmyb5 promoter is expressed in developing leaf trichomes, stipules, and epidermal cells on the margins of young rosette and cauline leaves, and in immature seeds. Atmyb5 mRNA appears between fertilization and the 16 cell stage of embryo development and persists beyond the heart stage. A leaf promoter identified in maize (e.g., Busk et al., Plant J.11:1285-1295, 1997) can also be used.

[0053] Another class of useful vegetative tissue-specific promoters are meristematic (root tip and shoot apex) promoters. For example, the "SHOOTMERISTEMLESS" and "SCARECROW" promoters, which are active in the developing shoot or root apical meristems, (e.g., Di Laurenzio, et al., Cell 86:423-433, 1996; and, Long, et al., Nature 379:66-69, 1996); can be used. Another useful promoter is that which controls the expression of 3-hydroxy-3-methylglutaryl coenzyme A reductase HMG2 gene, whose expression is restricted to meristematic and floral (secretory zone of the stigma, mature pollen grains, gynoecium vascular tissue, and fertilized ovules) tissues (see, e.g., Enjuto, Plant Cell.7:517- 527, 1995). Also useful are kn1-related genes from maize and other species which showAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory meristem-specific expression, (see, e.g., Granger, Plant Mol. Biol.31:373-378, 1996; Kerstetter, Plant Cell 6:1877-1887, 1994; Hake, Philos. Trans. R. Soc. Lond. B. Biol. Sci. 350:45-51, 1995). For example, the Arabidopsis thaliana KNAT1 promoter (see, e.g., Lincoln, Plant Cell 6:1859-1876, 1994) can be used.

[0054] In some embodiments, the promoter is substantially identical to the native promoter of a promoter that drives expression of a gene involved in secondary wall deposition. Examples of such promoters are promoters from IRX1, IRX3, IRX5, IRX8, IRX9, IRX14, IRX7, IRX10, GAUT13, or GAUT14 genes. Specific expression in fiber cells can be accomplished by using a promoter such as the NST1 promoter and specific expression in vessels can be accomplished by using a promoter such as VND6 or VND7. (See, e.g., PCT / US2012 / 023182 for illustrative promoter sequences). In some embodiments, the promoter is a secondary cell wall-specific promoter or a fiber cell-specific promoter. In some embodiments, the promoter is from a gene that is co-expressed in the lignin biosynthesis pathway (phenylpropanoid pathway). In some embodiments, the promoter is a C4H, C3H, HCT, CCR1, CAD4, CAD5, F5H, PAL1, PAL2, 4CL1, or CCoAMT promoter. In some embodiments, the tissue-specific secondary wall promoter is an IRXl, IRX3, IRX5, IRX8, IRX9, IRX14, IRX7, IRX10, GAUT13, GAUT14, or CESA4 promoter. Suitable tissue-specific secondary wall promoters, and other transcription factors, promoters, regulatory systems, and the like, suitable for this present invention are taught in U.S. Patent Application Pub. Nos.2014 / 0298539, 2015 / 0051376, and 2016 / 0017355.

[0055] One of skill will recognize that a tissue-specific promoter may drive expression of operably linked sequences in tissues other than the target tissue. Thus, as used herein a tissue- specific promoter is one that drives expression preferentially in the target tissue, but may also lead to some expression in other tissues as well.

[0056] In some embodiments, each GOI is operatively linked to a promoter that is activated by the transcription activator. In some embodiments, each GOI is a biosynthetic gene that expresses an enzyme that catalyzes the biosynthesis of a compound of interest, or an intermediate thereof.

[0057] References cited:Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0058] It is to be understood that, while the invention has been described in conjunction with the preferred specific embodiments thereof, the foregoing description is intended to illustrate and not limit the scope of the invention. Other aspects, advantages, and modifications within the scope of the invention will be apparent to those skilled in the art to which the invention pertains.

[0059] All patents, patent applications, and publications mentioned herein are hereby incorporated by reference in their entireties.

[0060] The invention having been described, the following examples are offered to illustrate the subject invention by way of illustration, not by way of limitation.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory Example 1 Quantitative dissection of Agrobacterium pathogenesis uncovers competitive and cooperative interactions modulating DNA transfer

[0061] Agrobacterium pathogenesis involves the transfer of a DNA segment into host plant cells. This feature has made Agrobacterium the cornerstone of both basic and applied plant genetic engineering from its very inception. As the sophistication of the scientific problems that require Agrobacterium-mediated transformation increases, it becomes critical to achieve a quantitative and predictive understanding of DNA transfer at the level of single plant cells. Such a framework would make it possible to better design and interpret experiments and could help identify engineering bottlenecks. Furthermore, contrasting quantitative measurements with the predictions from mathematical models could be a powerful way to reveal unexpected principles of bacterial pathogenesis. Using quantitative live cell imaging we systematically tested whether a classic mathematical model of interactions between individual pathogens and host cells holds true for Agrobacterium infecting leaves of Nicotiana benthamiana. This dialog between theory and experiments uncovered previously unappreciated competitive and cooperative interactions between bacteria. Using genetic perturbations we were able to identify the molecular basis of these interactions. To overcome the constraints imposed by competition between bacteria, we engineered a dual binary vector system termed ‘BiBi’ which allows us to increase the diversity of delivered transgenes without having to increase the number of bacterial cells. Building on these findings, we demonstrate that accounting for bacterial competition can improve plant metabolic engineering of a 14-step biosynthetic pathway.

[0062] At a more basic level, a quantitative dissection of Agrobacterium DNA transfer may allow testing mechanistic hypotheses about bacterial pathogenesis with a resolution that is not possible with other model bacterial pathogens. Indeed, in the case of viruses, contrasting quantitative measurements with predictions made by mathematical models has shed light on unappreciated basic aspects of pathogenesis. This progress has been possible because, like Agrobacterium, viruses can be engineered to transfer genetically-encoded reporters into host cells. This feature should make it possible to use live microscopy to noninvasively quantify specific aspects of the pathogen-host interaction with single cell resolution in living tissues in a way that is not feasible with other kinds of bacteria.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0063] Here, we sought to quantitatively dissect a popular and relatively simple model system of Agrobacterium infection, the transformation of Nicotiana benthamiana (Tobacco) leaf cells by a culture of Agrobacterium infiltrated into the leaf, also known as agroinfiltration. As our starting hypothesis, we asked whether Agrobacterium-mediated transformation can be described as a Poisson process. This probabilistic model, originally proposed for bacteriophage infection by Ellis and Delbrück almost 90 years ago, remains to this day the go-to mathematical framework for the infection of a population of host cells by a population of infectious agents and was recently applied to Tobacco agroinfiltration. This model rests on the key assumption that pathogens infect host cells with a constant probability, independent of the abundance of other bacteria or whether a host cell was previously infected or not. Because bacteria can be notably social organisms, we suspected that this model may not fully capture Agrobacterium pathogenesis. Furthermore, the nature of the host-pathogen interface -an ecological niche rife with conflict and subject to ecological constraints- offers another reason to scrutinize the assumptions that underpin the Poisson model. RESULTS AND DISCUSSION Challenging the Poisson model of agrobacterium-mediated transformation of plant cells.

[0064] A classic textbook approach to describe the infection of a population of host cells by a group of infectious agents is the Poisson distribution. As such, it serves as a natural starting framework to attempt to describe plant cell transformation by Agrobacterium.

[0065] The Poisson distribution: Equation 0)describes the frequency of the number of k occurrences of a random, recurring event within an interval of, for example, space or time. It assumes that the probability of occurrence per interval is a constant, implying that events are completely independent of each other. The only free parameter in Equation 0 is λ, the average number of events per interval, that is, the total number of events divided by the total number of intervals. The original application of the Poisson distribution to pathogenesis is often ascribed to Ellis and Delbruck who used it to model the distribution of the number of bacteriophages infecting single E. coli. Here, the Poisson ‘interval’ corresponds to a single bacterium cell and an ‘event’ is defined as the infection by a single infectious agent. This allows calculating lambda asAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory Equation 1)

[0066] Obtaining λ in this way can prove to be technically challenging because it requires counting the number of infectious agents and host cells in an experiment. When a culture of Agrobacterium is infiltrated into a leaf the experimenter has no control over nor knowledge of the number of plant cells in this complex tissue. For this reason it is convenient to instead treat λ as being proportional to a relative density Equation 2)where means 'proportional to'. The number of bacteria per unit of leaf volume is in turn proportional to the number of bacteria per unit of infiltration buffer volume, which can be measured using an optical density (OD600) measurement. This allows expressing λ in terms of the OD of bacteria in the infiltration buffer Equation 3)with α as a proportionality constant that allows us to convert the infiltration OD, ODi, into the number of Agrobacterium cells per plant cell, λ. Note that α may also encompass other, more subtle aspects such as the fraction of the infiltrated OD corresponding to bacterial cells that are viable and / or potentially capable of transformation.

[0067] Not until very recently, this model had not been explicitly tested in the context of agroinfiltration, although researchers did use it implicitly at an intuitive level. For example, it was assumed that by increasing the density of bacteria a larger fraction of leaf cells would get transformed and that more cells would get transformed by multiple bacteria. To test if the predicted probability distribution from Equation 0 conforms to experiments it would be necessary to experimentally obtain k, namely, the number of different bacteria that infect each plant cell. This is technically very challenging because it would require labeling tens of different T-DNAs to then determine which ones are present in each plant cell. An alternative approach is to ask whether a host cell was infected or not. This corresponds to binarizing the Poisson distribution into two categories: cells that are not infected and cells that are infected at least once. The frequency of cells that are infected at least once can be obtained from the Poisson probability of k>0 given byAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory Equation 4)

[0068] With very minor changes, this formula was the prediction tested in a recent study by Carlson et al. who reported that the Poisson prediction as stated by Equation 4 can describe agroinfiltration. The present study was motivated in part by replicating the findings of this pioneering work and testing whether they can be generalized.

[0069] To challenge this model we decided to aim at its key assumptions. Perhaps the most important assumption of the Poisson model as applied to microorganisms is that they infect host cells independently of one another. This assumption posits that the probability of success of a single infectious agent is a constant, regardless of their total population density: for the purposes of transformation, a single strain of bacteria must be oblivious to the presence and abundance of other strains. The likelihood of infection by a given strain is solely determined by its own density, and not by that of other strains. Second, the independence assumption requires the probability of a host cell to get infected to be independent of whether or how many times it got infected before. It also assumes that all host cells are equally susceptible to infection. Recent work on E. coli bacteriophages demonstrated that this assumption does not always hold since phage infection can make bacteria less prone to be infected by subsequent phages. Given the centrality of these assumptions to the Poisson model we designed an experimental setup to test them.

[0070] To count the number of plant cells expressing a T-DNA we followed a live fluorescence microscopy approach. We generated Agrobacterium strains carrying binary vectors containing within their T-DNA either a green fluorescent protein (sfGFP) or red fluorescent protein (mCherry) fused to a nuclear localization signal (NLS) and driven by the strong Arabidopsis Ubiquitin 10 promoter (hereafter GFP-NLS and RFP-NLS, respectively). This promoter is expressed strongly in tobacco leaves but is silent in bacteria. We refer to strains carrying fluorescent markers as reporter strains (Fig.1, Panel A). To control the total number of bacteria independent of the density of reporter strains we created an empty vector strain (hereafter EV) with the same genetic background and carrying the same binary vector as reporter strains but lacking a fluorescent marker. To label host cells in a manner that is independent of agroinfiltration we generated a stable transgenic Tobacco line with a constitutively expressed nuclear localized tagBFP2 blue fluorescent protein marker (hereafter BFP-NLS). Compared to fluorescent reporters lacking subcellular localization, nuclearAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory markers make it far easier to accurately identify individual cells computationally (Fig.1, Panel A).

[0071] Next, to obtain data that could be directly contrasted with the prediction from Equation 4 and test if agroinfiltration follows a Poisson distribution, we titrated the reporter strains in equal ratios. To test if the rate of transformation by these strains is independent of the abundance of other bacteria we performed these titrations under different total bacterial densities, using the EV strain to reach different total bacterial ODs (Fig.1, Panel B). We infiltrated leaves of the BFP-NLS tobacco line with these different titration mixes and performed live cell fluorescence microscopy in a widefield microscope 3 days after infiltration. The BFP channel was used to identify leaf epidermis nuclei to then determine what fraction of these nuclei were infected by the reporter strains based on whether GFP and / or RFP expression could be detected (Fig.1, Panel C).

[0072] Consistent with previous reports, we found that even at saturating ODs of reporter strains, approximately 56% of epidermis nuclei do not express detectable levels of GFP or RFP. These untransformed nuclei are much smaller and occur in pairs or clusters which is consistent with these cells corresponding to guard cells and other cells in the stomatal lineage that are not susceptible to transformation via agroinfiltration, as previously described (Fig.7). Hence, for the purpose of comparing the fraction of cells expressing RFP or GFP to the Poisson prediction we use 44% of the total number of BFP nuclei as the total number of cells, unless specified. Agrobacterium strains compete with each other for plant cell transformation

[0073] To test if agrobacterium-mediated transformation in tobacco leaves follows a Poisson distribution as predicted by Equation 3 we plotted the number of BFP nuclei for which a corresponding GFP-expressing nucleus was detected as a function of the infiltration OD of the GFP reporter strain. We used the EV strain to keep the total OD constant across titrations. As shown in Figure 1D, the Poisson model from Equation 4 fits these data relatively well. The same applies to nuclei detected in RFP in the same set of experiments (Fig.8). While this is consistent with previous results, so far efforts have focused on a single total OD of bacteria without testing the role of total bacterial density as a variable. We thus sought to use the EV strain to test the effect of the total bacterial density in the transformation rate of reporter strains. The Poisson model predicts that, for a given density of the reporter strains, theAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory fraction of transformed nuclei should be the same regardless of the OD of the EV strain. This is clear from Equation 4, which takes as an input variable the infiltration OD of a single strain, not that of other strains. The abundance of other strains of bacteria such as EV do not factor into this naive Poisson prediction because it assumes that strains are completely independent of each other. To test if this prediction holds up we titrated the EV strain to achieve increasing total culture ODs of 0.05, 0.1, 0.5, 1, 2, and 3. If Equation 4 was sufficient to describe plant cell transformation, infiltrations with different densities of EV cells should produce overlapping curves in terms of the fraction of transformed cells. We found that, while at a given total OD the curves fit Equation 4 relatively well (R^2 between 0.66-0.87), these different curves do not lie on top of each other. At a given OD of the reporter strain, increasing the EV OD reduces the fraction of transformed plant cells. For example, at a GFP reporter strain OD of 0.025, about 80% of nuclei express GFP when the total OD is 0.5 (i.e. EV strain added at an OD is 0.475) but this fraction decreases to ~40% at a total OD of 2 (where EV is added at an OD of 1.975). This decrease in the transformation rate with increasing total OD is reflected in the fact that the estimated values of the probability constant ^^ obtained from these fits differ across total ODs. The fitted ^^ value decreases from ~100 for a total OD of 0.05 to ~6 for a total OD of 3 (Fig.1, Panel D).

[0074] To confirm that this result is not a detection artifact, we imaged a subset of the samples using laser scanning confocal microscopy, which has a better resolution and signal / noise ratio than widefield microscopy. In addition, we used higher magnification and overall, more accurate imaging conditions for this set of images. The results from both microscopes correlate very well (R^2 = 0.91 for GFP and 0.87 for RFP) with about 5% more nuclei detected in both the GFP and RFP channels in the confocal without any obvious bias across total ODs (Fig.9). These differences may stem in part from the field of view within each sample being different between microscopes. It is however possible that the expression level of GFP and RFP rather than the fraction of nuclei transformed with these reporters decreased with increasing OD. Conceivably, this could have led to a detection artifact where fewer nuclei are detected at high total ODs simply because they are not bright enough to pass our detection threshold. To rule out this possibility we used confocal microscopy to measure the pixel fluorescence intensity in the GFP and RFP channels across all BFP nuclei, regardless of whether they were assigned to the transformed or untransformed categories by our computational detection pipeline. The nuclei fluorescence distribution in the GFP channel of nuclei assigned to the GFP-positive category was clearly distinct and barely overlappingAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory from that of nuclei assigned to the GFP-negative group. The same was true for RFP. Moreover, these fluorescence values were largely identical between a total OD of 0.5 and 2 (Fig.10). This analysis rules out that a decrease in GFP or RFP fluorescence makes transformed nuclei harder to detect at higher ODs. We conclude that, although agroinfiltration at any total OD can be described as a Poisson process, the probability of transformation is not a constant but depends on the total density of bacteria, with higher densities leading to a lower probability of transformation. We interpret this as competition between infiltrated bacteria for plant cell transformation, which contradicts one of the key independence assumptions of the Poisson model.

[0075] Having shown that one of the independence assumptions of our starting model does not hold up, we next focused on the second independence requirement, namely that whether a plant cell was transformed by one reporter strain does not affect its likelihood of being transformed by another one. Such correlation between reporter strains could arise from, for example, transformation by one strain affecting the susceptibility of this cell to being transformed by a second strain. A lack of independence between reporter strains could also stem from there being plant cells that are more susceptible to transformation than others. To test if the reporter strains act independently of each other, we measured the frequency of cells co-transformed by GFP and RFP, p(GFP∩RFP), and compared it with its expected value if the two reporter strains are independent of each other, given by the multiplication of their individual frequencies, p(GFP) x p(RFP). We use lowercase p for frequency in contrast to uppercase P for the Poisson probability. Plotting these two frequencies against one another shows that they are very similar across total ODs (Fig.1, Panel E). Although there is small bias towards fewer co-transformed cells than expected if independent, we found that the independence hypothesis holds relatively well (X^2 = 0.0253). Hence, although the probability of transformation per unit of OD (captured by the ^^ parameter) decreases by increasing the OD of other strains of bacteria, at a given total OD, bacteria still can be described as transforming plant cells independently of each other. This result is inconsistent with there being subpopulations of plant cells that have a higher susceptibility to transformation. It is also inconsistent with infection by one bacterium affecting the transformation susceptibility of plant cells.

[0076] Our findings so far may appear contradictory: on the one hand the probability of transformation of a given strain decreases by increasing the total bacterial OD. Yet, bacteriaAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory still transform host cells independently of one another at any OD. These observations suggest a scenario where an increase in the total bacterial OD reduces the chances of transformation per bacterium, but transformation remains a random process at any OD. In this view, the total bacterial OD can be thought of as interacting with or dictating a 'hidden variable' which determines the transformation probability. If such a hidden variable exists, it should be possible to find a scaling factor that accounts for this variable and makes all the transformation curves collapse into a single master curve. To obtain this scaling factor we took advantage of the exponentially decaying relationship between the fitted ^^'s and the total OD, with a slope of m when plotted in a log y-scale. This scaling parameter, which can completely determine the transformation probability, is simplyWhen all the reporter strain infiltration ODs are multiplied by this scaling parameter the curves of the fraction of transformed host cells collapse into a single curve. As expected, this master curve fits well to a Poisson distribution with an ^^ value corresponding to the y-intercept of the fitted ^^'s as a function of the total OD, c ~100. One way to think of c is as the value of the transformation probability constant ^^ in the absence of any competition, when the number of bacteria in the leaf is arbitrarily small. We next wished to test if this scaling factor is generalizable to new experimental settings different from the ones used to find this scaling property. To this end we titrated the reporter strains in the absence of the EV strain, that is, without keeping the total OD constant across titrations. As shown in Figure 2, the data from this set of experiments is also captured by the master curve. This scaling property strongly suggests that there is a relatively simple underlying mechanism by which the chance of successful transformation per unit of reporter strain OD is adjusted as a function of the total bacterial OD.

[0077] We note that the ^^ values previously reported using a single total OD of 0.6 do not lie on the exponential curve of ^^ as a function of OD that we obtained (Fig.2). Our ^^ at OD 0.5 is about 5 times larger than this previously reported value and our ^^ at very low ODs is close to an order of magnitude larger. This has stark implications for experiments requiring that transformed plant cells are contacted by at most one agrobacterium strain. For example, using the previously reported ^^=11 and a total OD of 0.02, the Poisson distribution predicts that only ~10% of the cells that are transformed are contacted by more than one strain. In contrast, using ^^=100 and the same OD, it is predicted that ~70% of the transformed cells receive T- DNAs from 2 or more bacteria strains. Using ^^=100 it is necessary to lower the OD of the mix to 0.002 or less to ensure that most of the plant cells that are transformed are transformedAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory by at most a single strain. Because of these differences, we urge readers caution when picking the right OD for experiments that require minimizing multiple transformation. Competition can be explained by a limiting host resource

[0078] The finding that different Agrobacterium strains can outcompete one another in terms of T-DNA delivery poses the question of what is the nature of this competition. The existence of a scaling factor that, when multiplied by the infiltration OD of the reporter strains, makes their transformation probability independent of the total OD suggests the presence of a hidden variable. We hypothesized that this hidden variable may be related to a limiting host resource that needs to be consumed or occupied by bacteria prior to transformation. By using this limiting resource for itself, the competing EV strain may deprive the reporter strain of its chances of transformation. We reasoned that this limiting resource may correspond to a plant metabolite that bacteria must consume to become infectious. We refer to this picture as the 'metabolite competition model'. As an alternative mechanism we envisioned a 'spatial competition model' where the limiting host resource corresponds to the area on the surface of the plant cell that is accessible for bacteria to make contact and transfer the T-DNA. We note that these two models are not mutually exclusive.

[0079] To explore if these mechanisms are consistent with our data, we first followed a mathematical modeling approach. To model the metabolite competition scenario, we assume that reporter and EV bacteria can exist in two states, infectious and noninfectious. Infiltrated cells start in a noninfectious state but can become infectious inside the leaf upon consumption of a metabolite M. Those bacteria that can make the transition from noninfectious to infectious then go on to transform plant cells following a Poisson process with a single probability constant ^^. This ^^ should be a true constant, regardless of the data, OD of the competing EV strain and should also be equal to c, the rate in the absence of any competition shown in Figure 2. We simulated this system numerically to obtain the OD value of infectious reporter cells at steady state for different initial OD values of reporter and EV strains at the time of infiltration. We then used this steady state value to calculate the fraction of transformed plant cells according to the Poisson model from Equation 4 replacing ^^ by c=100 (for a detailed derivation of this model see Supplemental Calculations: metabolite competition model). As shown in Figure 3, Panel C, this simple model is sufficient to recapitulate the microscopy data from Figure 2.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0080] Next, we asked if the spatial competition model could fit our results. We modeled the number of different bacteria contacting each plant cell as a Multinomial distribution. Here, each plant cell has N 'transformation slots', each corresponding to a multinomial trial with three mutually exclusive possible outcomes: empty, taken by the reporter strain or taken by the EV strain. Whether each of these slots is taken and by which strain depends on the probability of that event. Here, as in the Poisson model, we assume this probability to be proportional to the infiltration OD of each strain. Like the resource competition model, the spatial competition model is also able to fit our microscopy data reasonably well (Fig.3, Panel C) (for a detailed derivation of this model see Supplemental Calculations: spatial competition model).

[0081] To further validate that these models are consistent with our experiments, we asked if they could successfully make quantitative predictions about new data from a different kind of experimental setup. We hypothesized that regardless of the underlying mechanism, competition between bacteria should manifest as saturation in the level of tissue-wide gene expression driven by a reporter strain even when titrated alone, in the absence of the competing EV strain. According to the metabolite limitation model, there should be an OD of the reporter strain at which all of the host resources get taken up such that any additional infiltrated reporter cells become unable to become infectious and thus cannot contribute transcription templates to increase gene expression in the leaf. In the case of the spatial limitation model, saturation arises because there is an OD at which all the transformation slots on the plant cell surface get taken up by bacteria. We refer to this saturation OD as OD_sat. To test this prediction, we titrated a strain carrying a T-DNA with GFP under the medium-strength pCM2 promoter and measured the tissue GFP fluorescence using a plate reader. We found that GFP fluorescence increases roughly linearly with infiltration OD and saturates at OD 0.25-0.5 (Fig.3, Panel D). This plateau could represent a limit imposed by the plant cell that is specific to the pCM2-GFP construct. For example, at the saturation OD the plant cell might run out of trans factors to sustain transcription from the pCM2 promoter or translation of more GFP transcripts. To rule out this possibility we repeated this experiment using two additional promoters, pCL1 and pCH1, which are roughly an order of magnitude weaker or stronger than pCM2, respectively. We found the saturation OD of pCL1 and pCH1 to be very similar to that of pCM2 (Fig.3, Panel D). These results are qualitatively consistent with a limitation imposed by a host resource at the level of bacteria rather than at the level of the plant cell. To test if the tissue-level results are quantitatively consistent withAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory either of our resource competition models, we used each one of them to calculate the value of OD_sat that best fits the microscopy data. The metabolite competition model predicts a value of ODsat of 0.47 ∓0.17 while the spatial competition model predicts OD_sat = 0.3 ∓ 0.1. Both values match relatively well the OD at which the fluorescence from tissue-level experiments saturates (for an explanation of how the saturation OD is derived from each model see Supplemental Calculations).

[0082] To distinguish between metabolic and spatial competition we next sought to genetically alter the capacity of Agrobacterium to attach itself to the plant cell surface. Attachment in Agrobacterium depends on a number of mechanisms, some of which are mediated by or dependent on the Ti plasmid. We therefore first tested if the C58C1 strain, which shares the same chromosomal background as EV but lacks pTi, could out-compete the reporter strains. Using the live imaging and analysis setup described previously we estimated the transformation probability constant ɑ using C58C1 to keep the total OD constant at 0.1, 0.5 and 2. As observed for EV (Fig.2) we found that ɑ decreases exponentially with increasing total OD when C58C1 is the competitor strain (Fig.3, Panel E). We conclude that attachment mechanisms that depend on pTi are dispensable for competition. In addition, since C58C1 lacks the machinery necessary to transform plant cells this result confirms the conclusions from our tissue-level fluorescence experiment, namely that competition occurs at the level of bacteria, upstream of T-DNA delivery (Fig.3, Panel D). Next, to assay for competition in a way that was faster than live cell imaging we used the tissue-level fluorescence driven by a pCM2-GFP strain. Here, the reporter strain was infiltrated at an OD of 0.025 in combination with EV or C58C1 as the competitor strain at an OD of 0.075 or 0.1975 to yield a final bacterial OD of 0.1 or 2.0. Consistent with our live imaging experiments, adding EV or C58C1 at a high OD drastically decreased the leaf GFP fluorescence (Fig.3, Panel F). Given that C58C1 is still able to compete, we used this genetic background to create knockouts in genes that are necessary for the synthesis of two extracellular attachment polymers, unipolar polysaccharide (UPP) and cellulose. The competitive capacity of these mutants was indistinguishable from that of C58C1 (Fig.3, Panel F).

[0083] Taken together, these experiments allow us to discard a number of potential sources of competition. First, they demonstrate that competition between strains does not arise from a scarce resource inside plant cells. In addition, these results show that competition does notAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory stem from a scarcity of available area on the host cell surface at high bacterial densities. We conclude that it is highly likely that Agrobacterium consumes one or more extracellular leaf metabolites that become limiting at high bacterial densities. Identifying these metabolites should be the focus of future research as it should make it possible to increase transformation rates and transgene expression in plant leaves. Agrobacterium cells can act cooperatively via the sharing of extracellular proteins.

[0084] When investigating the basis of competitive interactions, we serendipitously discovered that different Agrobacterium strains can increase the pathogenicity of each other. Specifically, we found that the tissue-level GFP fluorescence of a reporter strain can be increased by adding EV cells at a low OD (Fig.3, Panel F and Fig.4, Panel A). This was intriguing because all our data so far showed that the EV strain out-competes the reporter strains. To better understand the mechanistic basis of this phenomenon we asked whether cooperativity depends on the presence of the Ti plasmid. Adding C58C1 at any OD reduces the tissue-level GFP fluorescence driven by the reporter strain compared to a buffer only control, unlike EV which is able to increase it at low OD (Fig.3, Panel F and Fig.4, Panel A). As expected due to competition, adding C58C1 or EV at a high OD drastically reduces GFP fluorescence (Fig.4, Panel A). The observation that a strain lacking the Ti plasmid can act in a competitive but not a cooperative fashion demonstrates that cooperativity is encoded in this plasmid.

[0085] The tissue-level leaf fluorescence represents an aggregate measurement of a number of distinct molecular processes such as the fraction of transformed host cells, the number of T-DNAs per host cell and the expression level of each of these T-DNAs. To gain further insight into the molecular basis of cooperation we asked whether we could dissect cooperativity into some of these underlying factors. To this end we repeated these infiltrations and used live confocal fluorescence microscopy to determine the fraction of cells expressing the GFP reporter and the level of GFP fluorescence in these cells. We found that these two metrics do not always track with the tissue-level GFP fluorescence. Regarding the fraction of cells transformed with GFP and in line with our previous results, both the EV strain and C58C1 act in a competitive fashion at a high OD (Fig.4, Panel B). At a low OD, neither EV nor C58C1 significantly increased the fraction of detected GFP-expressing cells compared to the buffer alone (Fig.4, Panel B). This result shows that the cooperative effect by the EV strain at the level of tissue-level GFP fluorescence cannot be explained by anAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory increase in likelihood of the GFP reporter strain transforming host cells. We next asked if rather than increasing the probability of transformation in a yes / no binary sense, cooperativity stems from a higher level of GFP expression in cells already transformed by the reporter strain. To test this, we measured the average nuclear fluorescence intensity exclusively among nuclei detected in the GFP channel. For this metric, the EV strain acts in a cooperative but not in a competitive manner since nuclear fluorescence is similar at either low or high OD and both are significantly higher than the buffer control (Fig.4, Panel C). On the other hand, C58C1 was able to compete but not cooperate. Here, competition by C58C1 is somewhat less effective than for the fraction of transformed host cells since at low OD of C58C1 the fluorescence of GFP nuclei is close to that of the buffer control (Fig.4, Panel C). Taken together, these single cell measurements confirm that cooperation requires pTi and show that it arises mainly from an increase in GFP output in cells that have already been transformed by the GFP reporter strain.

[0086] Because a universal aspect of bacterial pathogenesis is the delivery of effector proteins into the host cell, we hypothesized that a given bacterium can benefit from effector proteins delivered from a different bacterium contacting the same plant cell. The fact that all the Agrobacterium secreted pathogenic effectors are encoded in pTi combined with our finding that pTi is required for cooperation was consistent with this idea. Indeed, there is ample evidence for so-called 'extracellular complementation' of Agrobacterium mutants lacking secreted vir genes. This kind of complementation has been reported for other bacterial pathogens. A recent study showed that in the plant pathogen Pseudomonas syringae a battery of pathogenic effectors can be partitioned into different strains, each containing a single effector gene, and still produce pathogenesis as long as the strains are mixed prior to inoculation. However, unlike all these previous reports, in our case the GFP reporter strain whose activity is enhanced and the helper strain responsible for this enhancement are genetically identical, have the full set of effector vir genes and are thus fully capable of pathogenesis on their own. To test if sharing of vir genes explains cooperation, we asked if a GV3101 strain lacking VirE1 and VirE2 could enhance the tissue-level expression of an RFP reporter strain compared to a buffer control. As shown in Figure 4, Panel D the VirE12 mutant strain can still increase the fluorescence driven by the reporter RFP strain but to a lesser extent than GV3101. This decrease in cooperation can be partially restored by a VirE12 complementation transgene. Thus, VirE1 and / or VirE2, both of which are secreted and act inside of the plant cell, partially contribute to cooperation.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0087] These data are consistent with a scenario where transformation of a plant cell by a single agrobacterium can in some cases be limited by the vir gene dosage of this bacterium, perhaps due to non-genetic heterogeneity in vir gene expression. This implies that a single bacterium may also happen to have excess dosage of vir genes. If both kinds of bacteria contact the same plant cell, then the vir-poor bacterium can be aided by the vir-rich bacterium. An alternative explanation is that VirE12 and other genes encoded in pTi modify the extracellular environment at a larger scale and make the mesophyll in general more conducive to agrobacterium pathogenesis. If cooperation occurs at the level of single plant cells, we expect that the level of expression of GFP and RFP T-DNAs delivered by two different strains should correlate in individual nuclei. Further, this correlation should decrease when a third EV strain is added to the culture since vir-poor cells from the GFP and RFP strains can then take advantage of vir genes derived from the abundant EV strain. To test this prediction, we coinfiltrated the GFP and RFP reporter strains at a low OD of 0.002 with or without the EV strain at an OD of 0.046. As predicted, in the absence of EV the GFP and RFP fluorescence correlate at the level of single nuclei (r=0.4) and this correlation decreases by a factor of two in the presence of EV (r=0.23) (Fig.4, Panel E).

[0088] Taken together, these results point to a kind of density-dependent duality in Agrobacterium interactions within the leaf. At low cell densities, strains carrying the Ti plasmid can cooperate with each other to achieve higher expression of genes encoded in their T-DNA compared to each strain alone, at least in part by sharing extracellular vir proteins when they infect the same plant cell. This increase in expression could be the result of vir genes increasing the number of T-DNAs that make it into the nucleus and / or the transcriptional level of these T-DNAs. With increasing ODs, competition between strains outweighs cooperation and the total number of bacteria starts to negatively affect the transformation capacity of any one strain. Unlike cooperation, this competitive effect is independent of the Ti plasmid (Fig.3). Thus, regardless of the OD, the competitive and not the cooperative modality is available to C58C1. To explicitly test this duality, we performed a titration experiment where we kept the OD of the reporter constant at a low OD of 0.025 and titrated increasing ODs of either EV or C58C1. As predicted, when adding EV the tissue- level GFP fluorescence shows a biphasic behavior, progressively increasing with increasing EV OD but decreasing past an OD between 0.1-0.5 (Fig.4, Panel F). Co-infiltrating any amount of C58C1 results in a decrease in GFP fluorescence for all ODs tested (Fig.4, Panel F).Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0089] How to reconcile this cooperative effect with the competition we described in the previous sections? First, it can be argued that since cooperation can only be detected at low ODs, it is present in the background of the experiments in the previous sections. Second, cooperation by pTi-carrying strains manifests at the level of the fluorescence intensity of nuclei detected in the GFP channel but not at the level of the fraction of nuclei expressing GFP. Because our competition experiments were based solely on the modulation of the fraction of nuclei expressing the reporter, they may in this sense be considered as independent of cooperation. BiBi strains carrying two binary vectors per cell can bypass the constraints imposed by bacterial competition.

[0090] Our finding that different strains compete with each for a chance to transform plant cells may pose hard constraints on the use of Agrobacterium in plant engineering. In recent years, mixed cultures of 10-20 Agrobacterium strains have been used to deliver multiple enzyme transgenes in order to reconstitute complex biosynthetic pathways in tobacco leaves. Since each strain carries a different enzyme, the extent of full pathway reconstitution depends on the fraction of plant cells transformed by all strains. To increase the fraction of plant cells that express all enzymes it is necessary to use high total ODs but, as we demonstrated previously, by increasing the total OD the effective OD of each of these strains becomes progressively lower. Hence, competition may impose a limit to how many different transgenes can be coexpressed per cell. In principle, this constraint can be overcome by putting multiple transgenes in one plasmid carried by a single strain. However, this is laborious and, for very long pathways, a technically challenging cloning problem. Having one enzyme per strain is comparatively straightforward and allows mixing and matching enzymes simply by combining different Agrobacterium cultures. In turn, this enables combinatorial, rapid large-scale screening of pathway designs.

[0091] The evidence presented so far was consistent with competition occurring among bacteria, upstream of plant cell contact and transformation. Hence, to bypass competition we sought a way to increase the diversity of expressed T-DNA transgenes without increasing the number of bacteria infiltrated into the leaf, all the while keeping the convenient 'mix and match' feature of agroinfiltration characteristic of the ‘one transgene-one plasmid’ approach. To this end we engineered strains carrying two binary vectors per cell, which we call 'BiBi' strains. One of the plasmids confers Kanamycin resistance and carries the pVS1 origin ofAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory replication and the other confers Spectinomycin resistance and has the BBR1 origin. Hereafter we refer to these binary vectors as pVS1 and BBR1. Although Agrobacterium strains carrying two binary vectors have been reported before, their performance in terms of co-delivery has not been quantitatively characterized.

[0092] To characterize the co-delivery of different T-DNAs in BiBi strains we followed the same live cell fluorescence microscopy approach we used for strains carrying a single binary vector. We infiltrated at a very low OD of 0.002 a BiBi strain carrying GFP-NLS in the BBR1 plasmid and RFP-NLS in the pVS1 plasmid. We compared this experiment to a coinfiltration of two regular strains, each one carrying either one of these vectors. The same OD of reporter bacteria was used in both cases (OD of GFP-NLS+RFP-NLS = OD BiBi) and the total OD was kept constant at 0.5 using the EV strain. A visual inspection of these images revealed that the BiBi infiltration resulted in a much higher proportion of nuclei expressing both GFP and RFP compared to the coinfiltration control, revealing co-delivery of the pVS1 and BBR1 plasmids (Fig.5, Panel A). Although nuclei expressing only one of the fluorescent proteins were much more common in the coinfiltration, they were still present in the BiBi infiltration (Fig.5, Panel A), showing that the two T-DNAs are not always coexpressed. Thus, at a qualitative level BiBi strains do perform as intended, although the T-DNAs in pVS1 and BBR1 are not co-delivered with 100% efficiency. The finding that different T- DNAs launched from BiBi strains tend to be coexpressed albeit not completely motivated us to further understand this process.

[0093] We envisioned a scenario where T-DNA delivery is a two-step process (Fig.5, Panel B). First, the bacterium has to make contact with the plant cell, which based on our previous results, we assume to be Poisson-like (i.e., a random process whose probability is proportional to the OD of the strain). Next, this bacterium can deliver a T-DNA that gets expressed but this step is not always successful, as revealed by the fact that BiBi strains do not always result in coexpression of GFP and RFP in individual nucleus (Fig.5, Panel A). Rather, we propose that T-DNA delivery is a probabilistic step with a probability of p for the pVS1 T-DNA and r for the BBR1 T-DNA. For simplicity we assume that different T-DNAs harbored in the same BiBi strain are delivered completely independently of each other. This is, the probability that a pVS1 T-DNA is delivered is the same in a regular one-plasmid strain and a BiBi strain and the same is true for a BBR1 T-DNA. This implies that the probability of coexpressing the pVS1 and BBR1 T-DNAs given contact is p x r (Fig.5, Panel B).Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory

[0094] To test if this simple model captures the behavior of BiBi strains we first asked if the two T-DNAs are indeed independently delivered. To address this assumption, we asked if the presence of a binary vector affects the expression of a T-DNA delivered from another vector in the same cell in terms of the fraction of host cells expressing this T-DNA. To this end we compared the fraction of plant cells expressing GFP as a function of reporter bacteria OD for a GFP-NLS transgene launched from the pVS1 or the BBR1 vector, with a second vector carrying the RFP-NLS transgene in a different backbone in the same strain or in a different co-infiltrated strain (Fig.5, Panel C). These experiments revealed that delivery of the pVS1- encoded T-DNA is largely independent of whether the BBR1 T-DNA is launched from the same strain or a different infiltrated strain (Fig.5, Panel C). Analogous results were obtained for BBR1 albeit the rate of transformation was lower for this binary vector compared to pVS1 (Fig.5, Panel C). Intriguingly, when two BBR1 strains are mixed alone their rate of transformation increases and becomes similar to that of pVS1. Thus, our model assumptions in terms of independence hold up with r<p<1 except when BBR1 strains are infiltrated alone in which case r=p.

[0095] Next, to further validate our model and constrain the parameter fits we plotted the observed fraction of nuclei expressing both GFP and RFP versus the expectation based on independence (as in Fig.2, Panel B), we find that BiBi strains result in a much larger proportion of nuclei expressing both reporters than expected given independence (Fig.5, Panel D). Yet, this co-delivery frequency is clearly lower from the expectation if pVS1 and BBR1 were always delivered (i.e. if p=r=1, Fig.5, Panel D, left). To further challenge our model, we used it to fit the fraction of nuclei that express both reporters as a function of the infiltration OD of reporter strains (Fig.5, Panel D, right). In addition, we fitted the fraction of GFP nuclei that also express RFP, p(RFP|GFP). We found that all these disparate data can bewell explained by our model with parameters p = 0.62 0.034, r = 0.37 0.03, and ^^ =60.15 5.4 (overall R^2 = 0.53 for all the combined data shown in Fig. 5).

[0096] Qualitatively, these results demonstrate that BiBi strains can be used to increase the fraction of cells that express multiple T-DNAs per unit of OD. Our finding that our two-step transformation model (Fig.5, Panel B) can successfully describe disparate aspects of the data shows that it does capture two hitherto unappreciated aspects of Agrobacterium pathogenesis: that T-DNA expression is a two-step process and that both steps are probabilistic. It may be possible to increase the predictive power of the BiBi model by adding interaction termsAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory between plasmids and / or strains but this risks overfitting to the relatively noisy data obtained from live cell imaging. This mathematical framework has other, more practical implications as well. It allows us to predict to what extent BiBi strains can be deployed to increase the coexpression of multiple transgenes in metabolic engineering setups where 10-20 strains are coinfiltrated. Testing the impact of competition on an engineered metabolic pathway.

[0097] Having shown that the BiBi system can be used to bypass bacterial competition and increase the fraction of plant cells that express multiple plasmids, we sought to explore its implications for metabolic engineering. Recently, tobacco has emerged as the primary platform for the discovery and commercial synthesis of complex, high-value plant metabolites. These molecules tend to require multiple enzymatic steps for their biosynthesis, all of which presumably need to co-occur in the same cell for the complete pathway to be reconstituted. Furthermore, to make yields sufficiently high for discovery and commercial applications, the fraction of cells expressing all enzymes needs to be as large as possible. Our results demonstrate that increasing the infiltration OD of a mix of strains increases the chances that plant cells are transformed by all strains. On the other hand, we discovered that a consequence of increasing the infiltration OD is a decrease in the rate of transformation of each individual strain. These two counteracting forces may influence to what extent it is possible to increase the density of bacteria to maximize the fraction of cells that express all enzymes in a complex engineered metabolic pathway.

[0098] To understand how these dynamics may play out as a function of the total OD and the number of coinfiltrated strains we took advantage of the theoretical scaffold we developed so far. We combined the adjusted Poisson model that incorporates the total OD as a scaling factor with the BiBi co-delivery model from the previous section. Using this framework, we calculated the fraction of leaf cells that express N plasmids when these plasmids are delivered by N regular strains or N / 2 BiBi strains (Fig.6, Panel A). This exercise revealed a number of predictions. First, the model shows that across conditions the fraction of coexpressing cells peaks at a total OD of about 1.1 above which it starts to decrease. This behavior stems from the fact that the transformation probability ɑ decreases exponentially with total OD (Fig.2, Panel C) (for further details see Supplementary calculations: Explaining the predicted decrease in transformation efficiency at high ODs). According to this prediction, packing N plasmids in N / 2 BiBi strains always leads to a higher fraction of cells expressing NAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory transgenes, although this increase is small to negligible when N is low. On the other hand, for N≥14, coexpression using regular strains becomes extremely unlikely but is considerably increased using BiBi strains. To better visualize the benefit of using BiBi strains over regular strains we calculated the ratio of the fraction of cells expressing all transgenes between a mix of BiBi and regular strains.

[0099] To test these predictions, it would be necessary to label ten or more transgenes with single cell resolution. Because of the wide excitation and emission spectra of fluorescent proteins this is not possible using the fluorescence microscopy approach we used previously. We reasoned that an alternative approach would be to quantify a signal that is produced only if a cell expresses all the different T-DNAs in a mix. This signal could be, for example, a metabolite that is the end product of a metabolic pathway composed of multiple enzymes, all of which need to be coexpressed in the same cell for this molecule to be produced. To follow this approach, we focused on the biosynthesis of glucoraphanin, a relatively complex biosynthetic pathway previously reconstituted by our group. Glucoraphanin is a glucosinolate found in cruciferous vegetables such as Broccoli that is thought to have a number of beneficial effects for human health including anti-cancer properties. Efficient glucoraphanin synthesis in tobacco requires the expression of 14 genes: 13 enzymes and one transporter (Fig.6, Panel B). Importantly, focusing on a complex molecule that has received considerable interest in plant metabolic engineering would also allow us to determine to what extent BiBi strains may improve existing bioengineering efforts.

[0100] We cloned each of the 14 genes necessary for efficient glucoraphanin biosynthesis in two different vectors: a binary vector carrying the pVS1 origin of replication and conferring Kanamycin resistance in bacteria and a binary vector carrying the BBR1 origin and conferring Spectinomycin resistance (Fig.6, Panel C). We then used these 28 plasmids to create 2 different BiBi strain mixes. BiBi mix 1 consists of 7 BiBi strains carrying all the odd-numbered genes in the glucoraphanin pathway in the pVS1 vector and the even-numbered ones in the BBR1 vector. BiBi mix 2 carries odd-numbered genes in the pVS1 vector and even-numbered genes in BBR1 (Fig.6, Panel D). We infiltrated each of these mixes at total ODs 0.05, 0.1, 0.5, and 1.0 and extracted glucoraphanin following a previously established protocol. We then quantified glucoraphanin biosynthesis using LC- MS / MS. We were able to reconstitute the biosynthesis pathway using the BBR1 mix and BiBi mix 1, but no glucoraphanin production was detected with the pVS1 mix or with BiBiAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory mix 2. We obtained the same result after all plasmids were resequenced and strains were retransformed. Thus, for unknown reasons one or more of the pVS1 transgenes does not produce a functional product. Regardless, glucoraphanin production in the BBR1 mix and BiBi mix 1 infiltrations allowed us to compare these data with the predictions from Figure 6, Panel A.

[0101] Our naive expectation was that the amount of glucoraphanin should correlate linearly with the predicted fraction of cells coexpressing all 14 transgenes. As a result, we expected BiBi mix 1 to produce significantly more product than the BBR1 mix, particularly at low ODs. The data was consistent with some of these predictions. At a low OD of 0.05 we were able to detect small amounts of glucoraphanin in the BiBi mix but not in the BBR1 mix, which is consistent with a much larger fraction of cells coexpressing the transgenes using BiBi (Fig.6, Panel E). At OD 0.1 we predicted close to 4 orders of magnitude more cells cotransformed by BiBi compared to BBR1 strains. At this OD we detected production in both mixes although the increase of BiBi over BBR1 strains was close to 1 order of magnitude rather than 4 (Fig.6, Panel E). As predicted, at ODs 0.5 and 1.0 glucoraphanin production seemed to saturate for both mixes. However, in contrast to our prediction, product formation was similar between BiBi and BBR1 mixes. Taken together these results demonstrate that BiBi strains can be used to increase production from a transiently expressed metabolic pathway. Yet, the improvement over regular strains was lower than expected due to the regular strains outperforming the prediction.

[0102] The discrepancy between prediction and experiment is most likely due to the presence of nonlinear effects associated with metabolic pathways that are not at play when using fluorescent protein markers. Most saliently, small metabolites can readily move between adjacent plant cells via plasmodesmata. In this way, a group of nearby cells, each one lacking one or more of the glucoraphanin biosynthesis transgenes, may complement each other by sharing pathway intermediates. At high ODs most transformable cells will be transformed by most of the 14 strains, which should facilitate this kind of complementation, bringing the BBR1 mix closer to the BiBi mix. Characterizing this phenomenon in the future may be possible using methods to detect specific intermediates in-situ in single cells. Even though for 14 transgenes it appears that bypassing competition with BiBi strains does not significantly improve production at high ODs, the fact that BiBi performs better at low ODs indicates that this system may provide an advantage for longer metabolic pathways.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory DISCUSSION

[0103] Agroinfiltration has long been one of the cornerstones of plant metabolic engineering and synthetic biology and is quickly becoming a popular platform for testing massive transgene libraries. The nature of these experiments requires a high level of quantitative precision at the level of design and data interpretation. To make this possible the field needs a ‘theory of the experiment’, a mathematical framework to explain how the experimental system unfolds depending on the design settings. Not until very recently little to no attention had been paid to the need for such a framework. The present work fills this gap by providing a predictive model that can be applied to a variety of situations, offering concrete ways to optimize experimental designs.

[0104] Because agroinfiltration as a tool depends on the pathogenic mechanisms of Agrobacterium, we are convinced that there is a very obvious connection between understanding agroinfiltration from an engineering perspective and the basic study of bacterial pathogenesis. Our results demonstrate how both efforts reinforce one another and highlight that when it comes to Agrobacterium the borders between applied and basic efforts are largely artificial. This synergism is only possible because, uniquely among bacteria, Agrobacterium can transfer genetic material into plant cells. This feature makes Agrobacterium an incomparable model system to quantitatively dissect host-bacterial interactions at the single cell level. Indeed, there are a number of tools available to label DNA and RNA at the single molecule level in live plant cells, opening the door to studying pathogenesis with molecular resolution in the future.

[0105] To understand agroinfiltration we followed a ‘theory first’ approach where, rather than using ad-hoc models to explain the data retroactively, the goal is to turn one’s qualitative assumptions and hypotheses into mathematical formulas that make polarizing predictions about future experiments. The benefits of this strategy are multiple. By contrasting quantitative predictions with quantitative measurements, it is possible to uncover often subtle discrepancies, which then invite us to revise our starting assumptions. On the other hand, when predictions do match expectations, they provide a strong proof for one’s grasp of the finer details of the system in question. Finally, these kinds of models make it possible to precisely design future experiments.

[0106] We showed that, while the Poisson model can be used to describeAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory agrobacterium transformation in tobacco leaves, the probability parameter ɑ governing this distribution is not a constant but is affected by competition between bacteria. In addition, we found our estimates for this parameter to be different from previous estimates by a factor of 5-10. While one may be tempted to ascribe this discrepancy to experimental and / or analytical mistakes, we believe that the differences stem from different laboratories using different experimental variables. These variables include strains, reporters, plant and bacteria growth conditions, just to name the most obvious ones. For this reason, if the intention is to minimize the fraction of cells transformed by more than one Agrobacterium, researchers might find it useful to empirically obtain ɑ using their own setting.

[0107] Agrobacterium is often used in plant science to deliver reporter constructs, usually in combination with other strains such as upstream effectors. Our finding that the expression of a transgene delivered by a strain can be increased or decreased depending on the OD of a second strain means that interactions between strains need to be taken into account when designing and interpreting this kind of experiment. As a general rule, it is a good idea to keep the total OD of bacteria constant using an empty vector strain when necessary.

[0108] This ‘theory first’ approach revealed the existence of a number of hitherto unknown phenomena whereby Agrobacterium pathogenesis is modulated in a positive or negative way by the bacterial density. The existence of competitive and cooperative interactions among bacterial pathogens has long been appreciated although its mechanistic basis is only beginning to become clear. Using genetic perturbations, we were able to narrow down the nature of competition to the consumption of a host metabolite. The metabolic complexity of the leaf makes identifying this metabolite very challenging, although it should be possible to leverage genome-wide mutagenesis to identify this elusive molecule. It is not hard to envision a scenario where transgene expression and diversity can be increased simply by adding this limiting metabolite in the infiltration buffer.

[0109] We also identified the cellular and genetic basis for cooperation, which likely arises when a bacterium with a low dosage of vir genes infects the same cell as a bacterium with an excess dosage. This kind of phenotypic heterogeneity among genetically identical bacteria has been found to be responsible for a number of emerging phenomena in bacterial populations, from antibiotic resistance to sporulation. It will be interesting to identify the source of this phenotypic heterogeneity to explore how it can be exploited for engineeringAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory purposes. A future challenge will be to expand models of interactions between strains to integrate competition and cooperation.

[0110] Our finding that competition limits agrobacterium-mediated transient transformation inspired us to find an engineering solution to this problem. We reasoned that, since competition occurs among bacteria, it was necessary to increase the number of binary plasmids without increasing the number of cells, which motivated the development of BiBi strains. Here, following a dialog between theory and experiments we were able to identify yet another previously unappreciated aspect of Agrobacterium pathogenesis, the fact that contact with a plant cell does not always lead to T-DNA expression. Thus, T-DNA delivery and expression must be understood as a stochastic process whose governing mechanisms are yet to be discovered. Our data show that the origin of replication of the binary vector correlates with the rate of delivery. The mechanism behind this connection is far from clear because changing the origin can have complex effects on cellular physiology and the control mechanisms of many origins are poorly understood. Future experiments will be needed to further elucidate the mechanism behind T-DNA delivery in BiBi strains and these efforts should greatly benefit from engineered origins of replication. Learning why sometimes a bacterium fails to deliver one T-DNA but not the other could offer opportunities to modulate the T-DNA delivery process.

[0111] To demonstrate the use of BiBi strains in a real engineering application and further test our predictive understanding of agroinfiltration we tested a 14-strain pathway. We found BiBi strains to improve pathway performance compared to regular strains but the regular strains greatly outperformed our expectation. Much more glucoraphanin was produced by the regular strains than predicted solely from the fraction of cells transformed by all 14 strains. In this sense, despite our efforts, the performance of metabolic pathways during agroinfiltration still remains a ‘black box’. How is it that 14 or more strains can be coinfiltrated to reconstitute a metabolic pathway if there’s only an infinitesimally small predicted fraction of plant cells transformed with all transgenes? Of course, one explanation is that our prediction is at fault. However, given how well our models explain the fluorescent protein reporter data it is hard to believe that they are off by several orders of magnitude in the case of the 14-strain glucoraphanin experiment. We believe that the key to this ‘black box’ lies in plasmodesmata. These intercellular connections allow plant cells to form multicellular symplastic domains where metabolites are shared. Hence, to reconstitute a longAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory metabolic pathway what may matter is what fraction of these multicellular domains expresses all enzymes, not how many individual cells. In turn, this would depend on the size of these domains and the rate of movement of different metabolites within them. Unfortunately, we lack quantitative data for both variables. The intercellular complementation of metabolic pathways in agroinfiltrated tobacco may provide a useful platform to study this basic phenomenon. In combination with BiBi strains and the addition of the elusive limiting metabolite, increasing the rates of metabolite intercellular movement could further improve tobacco as a platform for metabolic engineering.

[0112] A popular alternative to agroinfiltration used to rapidly express transgenes in plant cells is protoplast transformation. Plasmid DNA transfection into protoplasts is routinely used to quantify the level of expression of reporter constructs in transactivation assays. Protoplasts are also being used to express massive gene libraries. Interpreting and optimizing these experiments may benefit from the kind of quantitative dissection we engaged in here, particularly if the goal is to extract single cell data. On the other hand, metabolic engineering has almost exclusively relied on agroinfiltration rather than protoplast transfection. Perhaps the fact that protoplasts lack plasmodesmata connections makes it much harder to reconstitute long pathways via intercellular sharing of intermediates. MATERIALS AND METHODS Plasmids and Agrobacterium strains.

[0113] All plasmids used in this study can be accessed from the JBEI public repository. The BFP-NLS plasmid was created in a previous study. The binary vectors used for the GFP-NLS, RFP-NLS, and the EV strain were based on the pCambia1300 backbone. The BBR1 plasmids were custom made based on the XX backbone. All plasmids were created using standard Gibson assembly or Golden Gate assembly protocols. All plasmids were transformed into the GV3101::pMP90 Agrobacterium strain via electroporation. A list of all the plasmids used in this study can be found in the JBEI registry. Plant growth conditions.

[0114] Nicotiana benthamiana (tobacco) plants were grown in an indoor growth room kept at 60% humidity and 25 ºC temperature under a 16 / 8 light / dark daily cycle and 120 μmol / m2s of light intensity. Plants were grown in Sunshine Mix #4 soil (Sungro)Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory supplemented with 499 Osmocote 14-14-14 fertilizer (ICL) at 5 mL / L and infiltrated with agrobacterium cultures 29 days after sowing. Agrobacterium growth conditions.

[0115] Agrobacterium glycerol stocks were streaked on LB plates containing antibiotics. The day before infiltration, single colonies were grown overnight in liquid LB containing antibiotics, shaking at 30°C. The day of infiltration, cultures were diluted 1:10 in LB with the same antibiotic concentrations and grown for a few more hours under the same conditions until reaching an OD600 of 0.5-1.0. The antibiotic concentrations used for all strains except C58C1 and strains carrying BBR1 plasmids were: 50 μg / ml Rifampicin, 50 μg / ml Kanamycin, and 30 μg / ml Gentamycin. Strains carrying BBR1 plasmids were grown using 50 μg / ml Rifampicin, 100 μg / ml Spectinomycin, and 30 μg / ml Gentamycin. C58C1 was grown without antibiotics. Agroinfiltration.

[0116] Agrobacterium liquid cultures at an OD600 of 0.5-1.0 were spun down at 4K rcf for 10-20 minutes and then resuspended in a similar volume of infiltration buffer (10 mM MES pH5.6, 10~mM MgCl2, 150 μM Acetosyringone). Cultures were incubated in infiltration buffer shaking for 1 hour at room temperature prior to the next OD600measurements. Next, a 1:5 dilution of each culture in infiltration buffer was prepared in 1 mL final volume and the OD600of this dilution was measured using a spectrophotometer. The final OD dilutions for infiltration were then prepared using infiltration buffer. Mixes were infiltrated in the 6th leaf (counting upwards with cotyledons being leaves 1 and 2) within 1 hour from preparation. In the imaging experiments of Fig.1, Fig.2, and Fig.4 and Fig.5 all dilutions of a given total OD were infiltrated in the same leaf. In the imaging experiments all 5 mixes were infiltrated in the same leaf. The position within the leaf of each infiltration was randomized between plants. Imaging: widefield fluorescence microscopy.

[0117] All widefield fluorescence images were taken in a Leica DM6B microscope. Three sequential images were taken for each z-stack of each sample, one for each fluorescent protein (BFP, GFP, and RFP). The filter sets used were: the DAPI filter for BFP (excitation 350 pm 50, emission 460 pm 50), the L5 filter for GFP (excitation 480 pm 40, emission 527Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory pm 30), and the TXR filter for RFP (excitation 560 pm 20, emission 630 pm38). A 5x dry objective was used to acquire 2.641 mm2images of 2048 x 2048 pixels, resulting in a pixel size of 0.78 μm. In each sample, 5 z-sections were acquired every 20 μm. The laser power was set to 17 % for all channels. The camera exposure time was 800 ms. Imaging: Laser-scanning confocal microscopy.

[0118] All confocal images were acquired using a Zeiss LSM710 microscope. In each Z-slice, two sequential scans were used, one for BFP and RFP and another one for GFP. In the first scan, excitation wavelengths were 405 nm using the Diode and 568 nm using the InTune laser. The two emission windows in this first scan were 410-530 nm (for BFP) and 585-630 nm (for RFP). The second scan used the Argon 488 nm laser for excitation and an emission window of 494-581nm for GFP. The frame size was 2048 x 2048 pixels with 1.5 Zoom using a 5x dry objective, resulting in a 1.133 mm2image with a pixel size of 0.55 μm. Scans were performed bidirectionally at a speed set to 9, corresponding to a pixel dwell time of 0.39 μs and a scan time of 7.75 s / slice. Averaging during acquisition was done using 2 lines. The laser intensities and gain settings were chosen to avoid detector saturation and were as follows. Laser power of 1.0 for 405 nm, 4.0 for 568 nm and 1.0 for 488 nm. Gain of 711 for BFP, 589 for RFP and 480 for GFP. The pinhole was set to 1 AU in both scans. Z- slices were taken every 2.5 μm with enough slices to capture all the epidermis nuclei in the field of view. Data shown in the same graphs comes from images taken during the same imaging session to avoid confounding effects stemming from day-to-day variation in laser power. Image analysis: nuclei identification.

[0119] Max-intensity z-projections of each channel were generated and a difference of gaussians (DOG) filter with a sigma of 5 pixels was applied to these max-projections to filter out fuzzy, small and / or irregularly-shaped features and enhance circular features corresponding to nuclei. DOG images were then segmented into objects (i.e. nuclei) and background using a fluorescence threshold value. To find this fluorescence threshold we used the Otsu method included in the skimage Python package. The result of this thresholding is a binary image of nuclei and background. Each of these binary images was visually inspected and compared side by side to the original max projection to check the segmentation by eye. The fluorescence threshold applied to the corresponding DOG image was adjusted wheneverAttorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory necessary. To separate objects that merged with each other we used the watershed algorithm. Finally, an area filter was applied to the BFP binary image mask to remove objects with an area smaller than 40 pixels, which for circular objects corresponds to sim 3.5 pixels in radius or sim 2.5 μm. As a reference, guard cell nuclei (the smallest nuclei in the epidermis) have a radius of sim 4 μm. We also filtered out objects with an area larger than 900 pixels. Image analysis: calculating the fraction of transformed nuclei

[0120] The total number of nuclei used as the denominator to calculate fractions corresponds to the number of objects found in the BFP channel as described in 'Image analysis: nuclei identification'. To determine how many nuclei were found in the GFP channel, we applied the GFP mask to the BFP mask by multiplying these binary images. Next, we applied an area filter to the resulting image using the same area thresholds as for the BFP mask. We then counted the number of objects, which correspond to nuclei detected as expressing GFP. The same procedure was used for RFP. To determine if a nucleus expresses both GFP and RFP we asked if the centroid of a given GFP-detected object overlaps that of an RFP-detected object. Objects whose centroids are closer than 5 pixels are considered as overlapping and counted as expressing both reporters. Image analysis: nucleus fluorescence intensity.

[0121] Confocal Z-stacks were segmented into binary images of nuclei and background as described in 'Image analysis: nuclei identification'. We applied an erosion algorithm to shave off boundary pixels from the edges of binary objects to ensure that only pixels that belong to the nucleus were included. Next, for each binary object, we calculated the mean fluorescence intensity of the pixels in the corresponding max-projected image which overlap with this binary object. For the purpose of this average, we did not include the 5\% brightest and dimmest pixels so as to remove outlier values. Leaf fluorescence measurements using plate reader assays.

[0122] For tissue-level fluorescence measurements, 46 mm leaf disks were cut from each agroinfiltrated leaf using a single-hole puncher and placed on top of 300 µL of tap water in a black, clear-bottom, 96-well plate. A BioTek Synergy H1 plate reader was then used to measure the GFP fluorescence of each leaf disk using an excitation wavelength of 488 nm and an emission wavelength of 520 nm.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory Curve fitting and parameter inference.

[0123] Curve fitting and parameter inference were performed via non-linear least squares using the Python SciPy package. No bounds for alpha were used for fitting Equation 3). The fit was performed using all the data simultaneously.

[0124] While the present invention has been described with reference to the specific embodiments thereof, it should be understood by those skilled in the art that various changes may be made and equivalents may be substituted without departing from the true spirit and scope of the invention. In addition, many modifications may be made to adapt a particular situation, material, composition of matter, process, process step or steps, to the objective, spirit and scope of the present invention. All such modifications are intended to be within the scope of the claims appended hereto.

Claims

Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory What is claimed is:

1. A library of binary vectors comprises a first library of first binary vectors, and a second library of second binary vectors, wherein the first library of first binary vectors and the second library of second binary vectors in combination comprise genes of interests of enzymes and / or transporters of a biosynthetic pathway or a catabolic pathway.

2. The library of claim 1, wherein each first binary vector has a first origin of replication in Agrobacterium, each second binary vector has a second origin of replication in Agrobacterium, wherein the first origin of replication in Agrobacterium is different from the second origin of replication in Agrobacterium.

3. The library of claim 2, wherein the first origin of replication in Agrobacterium, and the second origin of replication in Agrobacterium are independently a pVS1, BBR1, oriV, repA, OriA, or StaA origin of replication.

4. The library of claim 3, wherein the first origin of replication in Agrobacterium is a pVS1 origin of replication, and the second origin of replication in Agrobacterium is BBR1 origin of replication.

5. The library of claim 3, wherein each first binary vector and / or each second binary vector comprises: (1) an origin of replication or ORI for another bacterium, (2) an origin of replication or ORI for Agrobacterium, (3) a multiple cloning sites (MCS), (4) a plant selectable marker, (5) a bacterial selectable marker, (6) a promoter capable of expressing a gene of interest, and (7) a poly(A) signal.

6. The library of claim 3, wherein each first binary vector comprises a first fluorescent protein which fluorescences with a first color, and each second binary vector comprises a second fluorescent protein which fluorescences with a second color, wherein the first color and the second color are different.

7. A library of host cells wherein each host cell comprises a first binary vector and a second binary vector from the library of binary vectors of claim 1, wherein the library of host cells comprises all of the genes of interest encoding all of the enzymes and / or transporters of a biosynthetic pathway or a catabolic pathway.Attorney Docket: 2024-046-02 Lawrence Berkeley National Laboratory 8. A method for introducing the genes of interest encoding for enzymes of a biosynthetic pathway or a catabolic pathway into a target host cell comprising: (a) providing a library of host cells of claim 1, (b) contacting the library of host cells to a target host cell resulting in introducing the genes of interest into the target host cell, and (c) identifying one or more the target host cell that are capable of synthesizing or catabolizing the compound of interest.

Citation Information

Patent Citations

  • Novel methods and constructs for plant transformation

    US20030159184A1

  • Plant transformation with in vivo assembly of a sequence of interest

    US20090265814A1

  • In vivo Assembly of Transcription Units

    US20170191074A1

  • Binary vectors and uses of same

    US20200255845A1