Manipulation of bacterial cellulose synthesis, modification and secretion

By employing chimeric enzymes with tailored domains, the synthesis and modification of cellulose polymers are enhanced, addressing the limitations in existing methods and enabling the production of modified cellulose with desired properties.

WO2025212913A1PCT designated stage Publication Date: 2025-10-09UNIV OF VIRGINIA PATENT FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/022991
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2025-04-03
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current understanding of bacterial cellulose synthesis and modification, particularly in terms of phosphoethanolamine (pEtN) translocation and modification, is limited, and existing methods do not effectively facilitate the production of modified cellulose polymers with desired properties.

Method used

The development of chimeric enzymes comprising specific N-terminal and C-terminal domains, such as those corresponding to SEQ ID NO: 13 and SEQ ID NO: 10, or at least 95% sequence identity thereto, which are introduced into host cells to promote the synthesis and modification of cellulose by recruiting multiple BcsG enzymes, optionally replacing the phosphoethanolamine catalytic domain with an acetyltransferase domain.

Benefits of technology

This approach enables the production of modified cellulose polymers, including phosphoethanolamine cellulose, by enhancing the recruitment and catalytic activity of BcsG enzymes, thereby facilitating controlled synthesis and translocation of cellulose across the cell envelope, and introducing specific modifications like acetylation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000068_0000
    Figure 00000068_0000
  • Figure 00000069_0000
    Figure 00000069_0000
  • Figure 00000070_0000
    Figure 00000070_0000
Patent Text Reader

Abstract

Provided herein are modified celluloses, methods of making and their use.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]ZIMMER-PETN (03018-02) / / 1036.375WO1 MANIPULATION OF BACTERIAL CELLULOSE SYNTHESIS, MODIFICATION AND SECRETION PRIORITY This application claims the benefit of priority to U.S. Provisional Serial No.63 / 574,236, filed April 3, 2024, which is incorporated by reference as if fully set forth herein. GOVERNMENT GRANT SUPPORT This invention was made with government support under R35GM144130 awarded by the National Institutes of Health. The government has certain rights in the invention. INCORPORATION BY REFERENCE OF SEQUENCE LISTING This application contains a Sequence Listing which has been submitted electronically in ST26 format and hereby incorporated by reference in its entirety. Said ST26 file, created on April 2, 2025, is named 1036375WO1.xml and is 36,553 bytes in size. BACKGROUND OF THE INVENTION Cellulose is the most abundant biopolymer on earth. Chemically, cellulose is a linear polysaccharide composed of β-1,4 linked glucose. Individual strands participate in strong hydrogen bonding networks with neighboring strands and contribute to the physical and chemical integrity of plant cell walls and cellulosic materials. Microorganisms are also major producers of cellulose. Bacterial biofilms have major impacts on healthcare, industrial, and infrastructure systems. In a biofilm community, bacteria are encased in and protected by a complex biopolymer meshwork including protein polymers, polysaccharides, and nucleic acids. Cellulose is a common bacterial biofilm component produced by a variety of prokaryotes, such as Escherichia coli (Ec), Salmonella enterica, and Gluconacetobacter / Komagataeibacter xylinus (Gx). SUMMARY In one embodiment, the disclosure provides chimeric enzymes. In some aspects, a chimeric enzyme comprises an N-terminal domain selected from a sequence corresponding to SEQ ID NO: 13 or at least 95% sequence identity thereto operably linked to a heterologous cellulose synthase catalytic domain. In certain embodiments, the chimeric enzyme further comprises a C-terminal domain corresponding to SEQ ID NO: 10 or at least 95% sequence identity thereto. In other aspects, a chimeric enzyme comprises a heterologous catalytic domain while lacking a phosphoethanolamine catalytic domain, with the heterologous catalytic domain ZIMMER-PETN (03018-02) / / 1036.375WO1 optionally being an acetyltransferase domain. Additionally, nucleic acid sequences encoding these chimeric enzymes are provided. In another embodiment, the disclosure provides a method for modifying cellulose. The method comprises providing a host cell capable of synthesizing cellulose; introducing into the host cell an expression construct encoding a chimeric BcsA enzyme, a BcsB protein, and a BcsG enzyme; and culturing the host cell under conditions sufficient to express the encoded proteins such that the chimeric enzyme - comprising an N-terminal domain corresponding to SEQ ID NO: 13 or at least 95% sequence identity thereto and optionally a C-terminal domain corresponding to SEQ ID NO: 10 or at least 95% sequence identity thereto - facilitates recruitment of multiple BcsG enzymes to promote synthesis and translocation of a cellulose polymer. Optionally, the modified cellulose polymer is recovered. In certain aspects, the cellulose is modified by the addition of phosphoethanolamine groups at C6 hydroxyl positions, and in other aspects, a phosphoethanolamine transferase catalytic domain of the BcsG enzyme is replaced with a heterologous catalytic domain such as an acetyltransferase domain. The expression construct may be provided as a vector, and the nucleic acids encoding the chimeric BcsA enzyme, BcsB protein, and BcsG enzyme may be supplied in one or more constructs, with the host cell being a bacterial cell, a plant cell, or an algal cell. In a further embodiment, the disclosure provides an alternative method for modifying cellulose. This method involves providing a host cell capable of synthesizing cellulose; introducing into the host cell an expression construct encoding a chimeric enzyme comprising a BcsG enzyme with a heterologous catalytic domain that lacks a phosphoethanolamine catalytic domain, and optionally an associated chimeric BcsA enzyme with an N-terminal domain corresponding to SEQ ID NO: 13 or at least 95% sequence identity thereto operably linked to a heterologous cellulose synthase catalytic domain, as well as optionally a BcsB protein; culturing the host cell under conditions sufficient to express the encoded proteins; and optionally recovering the modified cellulose polymer. In further aspects, when present, the chimeric BcsA enzyme may further comprise a C-terminal domain corresponding to SEQ ID NO: 10 or at least 95% sequence identity thereto, and the heterologous catalytic domain of the BcsG enzyme may be an acetyltransferase domain. The expression construct is provided as a vector, with the nucleic acids encoding these proteins supplied in one or more constructs, and the host cell is selected from bacterial, plant, or algal cells. In one embodiment, SEQ ID NO: 10 and / or SEQ ID NO: 13 or 95% identity thereto can be introduced into a natively expressed cellulose synthase (e.g., at the 3’ and 5’ end, respectively) so that the modified cellulose synthase associates with a co-expressed BcsG ZIMMER-PETN (03018-02) / / 1036.375WO1 chimera. In other words, one does not have to heterologously express a chimera BcsA if the host organism already contains a cellulose synthase, such is the case for most plants and algae (some tunicates (animals) also express cellulose synthases). BRIEF DESCRIPTION OF THE DRAWINGS The drawings illustrate generally, by way of example, but not by way of limitation, various embodiments discussed herein. FIGS. 1A-1F. BcsA coordinates a BcsG trimer. a, Cartoon illustration of the Ec cellulose synthase complex. The putative cellulose secretion path is shown as a dashed line. b, Low resolution cryo-EM map of the Ec inner membrane-associated cellulose synthase complex (IMC, EMD-23267). c, AlphaFold2-predicted structure of Ec BcsG (AF-P37659-F1) with close-up showing the acidic cavity extending from the putative membrane surface. d, Cryo- EM composite map of the IMC after focused refinements of the periplasmic BcsB hexamer and the BcsG trimer associated with BcsA, respectively. BcsB subunits are colored from light grey to pink, BcsA and trimeric BcsG are colored in shades of blue and yellow, and the associated BcsB TM helix is colored grey. Contour level: 5. e, AlphaFold2-predicted complex of BcsA, the associated BcsB TM helix, and the trimeric BcsG colored as in panel d. f, Co- purification of Ec BcsG (His-tagged) with the N-terminal domain of Ec BcsA (NTD, Strep- tagged) by immobilized metal (IMAC) and Strep-Tactin affinity chromatography followed by size exclusion chromatography (SEC). FIGS.2A-2I. Fig.2: Engineering pEtN cellulose biosynthesis. a, Illustration of the Rs- Ec BcsA chimera. The Rs BcsAB complex is shown as a light and dark grey surface for BcsA and BcsB, respectively (PDB: 4P00). The introduced Ec N- and C-terminal domains are shown as cartoons colored blue and red, respectively. b, Size exclusion chromatography of the chimeric BcsAB-BcsG complex. Inset: Coomassie-stained SDS-PAGE of the peak fraction. c, In vitro catalytic activity of the purified chimeric BcsAB-BcsG complex. Cellulose production was quantified radiometrically by scintillation counting. DPM: Disintegrations per minute. d, Representative 2D class averages of the chimeric BcsAB-BcsG complex. e, Low resolution cryo-EM map (semitransparent surface, contoured at 4.8ϭ) of the chimeric BcsAB-BcsG complex overlaid with the refined map. Insets show a carved map of the chimeric BcsA NTD- BcsG complex as well as a focused refinement of BcsAB, respectively. f, Solid state NMR spectrum of pEtN cellulose produced by the chimeric BcsAB-BcsG complex in vivo (R / E-FG; black line), overlaid with a reference spectrum of pure pEtN cellulose from Ec (dashed red line). g, Western blots of IMVs containing the wild type (WT) RsBcsAB complex alone or the WT or chimeric BcsAB complex (R / E) co-expressed with BcsF and BcsG (FG). BcsA and ZIMMER-PETN (03018-02) / / 1036.375WO1 BcsG are His- and Strep-tagged, respectively. * Indicates an N-terminal degradation product of the chimeric BcsA. h, Polysaccharide analysis by carbohydrate gel electrophoresis (PACE) of in vitro synthesized pEtN cellulose. IMVs shown in panel g were used for in vitro synthesis reactions. Cellulase (BcsZ) released, and Alexa Fluor 647-labeled cello-oligosaccharides are resolved by PACE. Cello-oligosaccharide standards are ANTS (8-aminonaphthalene-1,3,6- trisulfonic acid) labeled and range from mono (Glc) to hexasaccharides (CE2-6). i, Quantification of cellulose biosynthesis by the IMVs shown in panel g containing the indicated BcsA variants and based on incorporation of3H-glucose into the water-insoluble polymer. Error bars in panels c and i represent standard deviations from the means of three replicas. FIGS. 3A-3B. Interactions of BcsC with BcsB. a, Low resolution map of the Bcs complex in the presence of BcsC’s periplasmic domain. The terminal BcsB subunit is colored lightpink and BcsC is colored green and lightblue for TPR#1 and #2, respectively. The map is contoured at 1.2ϭ. Inset: High resolution map of the BcsB-BcsC complex obtained after fusing BcsC’s TPRs #1-4 to the N-terminus of BcsB and focused refinement of the BcsB subunit bound to BcsC (contoured at 4.8ϭ. Density likely representing BcsG’s periplasmic domain (PPD) is encircled. b, Detailed interactions of BcsB and BcsC. FIGS. 4A-4C. Interactions of BcsC with cellulose. a, Low resolution cryo-EM maps of BcsC. Top panel: BcsC in the absence of cellotetraose with the full-length AlphaFold2-BcsC model (AF-P37650-F1) docked into the density. TPR#1 is colored lightblue, TPR#15-19 are colored as in panels b and c. Bottom panel: BcsC in the presence of cellotetratose. The model of the refined C-terminal BcsC fragment is docked into the density with the putative ligand colored blue. b, Close-up views of the putative cellulose binding sites of the refined maps in the absence (apo) and the presence of cellotetraose (shown as sticks colored cyan and red). Both maps are contoured at 10.5ϭ. c, Surface representation of BcsC bound to the putative cellotetraose ligand. FIGS.5A-5F. Cellulase activity is necessary for cellulose secretion. a, Congo red (CR) fluorescence images of Ec macrocolonies expressing the indicating components as part of the Bcs complex. ‘All Bcs’ express the inner membrane complex (IMC) together with BcsZ and BcsC. ΔBcsZ: no BcsZ. b, Carboxymethylcellulose digestion on agar plates using periplasmic Ec extracts. Cel9M and CMCax: Periplasmic extracts of cells expressing the Bcs components with BcsZ replaced by the indicated enzyme. BSA and A. niger: Controls with purified BSA or Aspergillus niger cellulase spotted on the agar plates. Cellulose digestion was imaged after CR staining, resulting in the observed plaques. c, Evaluation of CR fluorescence exhibited by Ec macrocolonies expressing the indicated components as part of the Bcs complex. d, ZIMMER-PETN (03018-02) / / 1036.375WO1 Representative 2D class averages of BcsZ tetramers. e, Model of the BcsZ tetramer overlaid with the cellopentaose-bound crystal structure (PDB: 3QXQ). Cellopentaose is shown as ball- and-sticks in cyan and red. Two-fold symmetry axes are indicated by black ellipses. f, Detailed views of the boxed regions in panel e. CR and cellulase plate assays were repeated at least three times with similar results. FIG.6. Model of cellulose pEtN modification and secretion. BcsA recruits three copies of BcsG to the cellulose biosynthesis site via its N-terminal cytosolic domain. The catalytic domain of BcsG either faces the lipid bilayer to receive a pEtN group or contacts the nascent cellulose chain for modification. BcsC interacts with the terminal BcsB subunit of the semicircle to establish an envelope-spanning complex. Cellulose is guided towards the OM through interactions with the TPR solenoid. Mislocalized cellulose is degraded by BcsZ to facilitate secretion. The cytosolic BcsE and BcsQR as well as BcsF components are omitted for clarity. FIGS. 7A-7B. BcsG is a phosphoethanolamine transferase. a, Comparison of the crystal structures of EptA (PDB: 5FGN) and the catalytic domain of BcsG (PDB: 6PD0) with the AlphaFold-2 predicted full-length structure of BcsG. b, Proposed BcsG reaction mechanism as presented in reference (25). FIGS.8A-8C. Sequence alignment of BcsA and its interaction with BcsG. Sequences were aligned in MUSCLE (43) and displayed in Jalview (44). N- and C-terminal helical elements are shown as cylinders in blue and pink, respectively. b, AlphaFold2-predicted complex of BcsA with three copies of BcsG’s TM segment, colored based on pLDDT score from blue to red for high to low confidence scores. c, Overlay of the generated AlphaFold2 predicted complex of the Ec BcsA-BcsG3 complex (colored blue for BcsA and yellow, violet and wheat for the BcsG subunits) and the model obtained after rigid body docking of the subunits into the experimental cryo EM map (colored gray). (SEQ ID NOS: 1-7) FIG.9. Mass spectrometry of in vitro-synthesized acetylated cellulose. Cellulose was synthesized by BcsA that was coupled to a BcsG-Wssf chimera. The obtained cellulosic material was digested with cellulase prior to MS analysis. The spectrum identifies a cellotetraose unit carrying a single acetyl-group of mass 708.12. In this instance, acetyl-CoA served as acetyl donor. Similar spectra are obtained for acetylated cellohexaose, cellopentaose, and cellotriose. DESCRIPTION Reference will now be made in detail to certain embodiments of the disclosed subject matter. While the disclosed subject matter will be described in conjunction with the enumerated ZIMMER-PETN (03018-02) / / 1036.375WO1 claims, it will be understood that the exemplified subject matter is not intended to limit the claims to the disclosed subject matter. Provided herein are modified celluloses, methods of making and their use. In particular, compositions and methods of producing phosphoethanolamine cellulose and derivatives thereof. Phosphoethanolamine (pEtN) cellulose is a naturally occurring modified cellulose produced by several Enterobacteriaceae. The minimal components of the E. coli cellulose synthase complex include the catalytically active BcsA enzyme, an associated periplasmic semicircle of hexameric BcsB, as well as the outer membrane (OM)-integrated BcsC subunit containing periplasmic tetratricopeptide repeats (TPR). Additional subunits include BcsG, a membrane-anchored periplasmic pEtN transferase associated with BcsA, and BcsZ, a conserved periplasmic cellulase of unknown biological function. While events underlying the synthesis and translocation of cellulose by BcsA are well described, little is known about its pEtN modification and translocation across the cell envelope. Herein it is shown that the N- terminal cytosolic domain of BcsA positions three copies of BcsG near the nascent cellulose polymer. Further, the terminal subunit of the BcsB semicircle tethers the N-terminus of a single BcsC protein to establish a trans-envelope secretion system. BcsC’s TPR motifs bind a putative cello-oligosaccharide near the entrance to its OM pore. Additionally, it shown that only the hydrolytic activity of BcsZ but not the subunit itself is necessary for cellulose secretion, suggesting a secretion mechanism based on enzymatic removal of mis-localized cellulose. Lastly, a pEtN modification of cellulose is introduced in orthogonal cellulose biosynthetic system by protein engineering. Definitions The following definitions are included to provide a clear and consistent understanding of the specification and claims. As used herein, the recited terms have the following meanings. All other terms and phrases used in this specification have their ordinary meanings as one of skill in the art would understand. Such ordinary meanings may be obtained by reference to technical dictionaries, such as Hawley's Condensed Chemical Dictionary 16th Edition, by M. Larranga, R Lewis Sr., and R Lewis, New York, N.Y., 2016. The practice of the present invention will employ, unless otherwise indicated, conventional methods of chemistry, biology, biochemistry, and molecular biology and recombinant DNA techniques within the skill of the art. Such techniques are explained fully in the literature. See, e.g., J. Wertz, J. P. Mercier, and O. Bedue Cellulose Science and Technology (Fundamental Sciences: Chemistry, EPFL Press, 2010); T. Wuestenberg Cellulose ZIMMER-PETN (03018-02) / / 1036.375WO1 and Cellulose Derivatives in the Food Industry: Fundamentals and Applications (Wiley-VCH, 2014); Lehninger, Biochemistry (Worth Publishers, Inc., current addition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (3rdEdition, 2001); Methods In Enzymology (S. Colowick and N. Kaplan eds., Academic Press, Inc.). References in the specification to "one embodiment," "an embodiment," etc., indicate that the embodiment described may include a particular aspect, feature, structure, moiety, or characteristic, but not every embodiment necessarily includes that aspect, feature, structure, moiety, or characteristic. Moreover, such phrases may, but do not necessarily, refer to the same embodiment referred to in other portions of the specification. Further, when a particular aspect, feature, structure, moiety, or characteristic is described in connection with an embodiment, it is within the knowledge of one skilled in the art to affect or connect such aspect, feature, structure, moiety, or characteristic with other embodiments, whether or not explicitly described. As used herein, the term "in some embodiments" refers to embodiments of all aspects of the disclosure, unless the context clearly indicates otherwise. The singular forms "a," "an," and "the" include plural reference unless the context clearly dictates otherwise. Thus, for example, a reference to "a compound" includes a plurality of such compounds, so that a compound X includes a plurality of compounds X. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for the use of exclusive terminology, such as "solely," "only," and the like, in connection with any element described herein, and / or the recitation of claim elements or use of "negative" limitations. The term "and / or" means any one of the items, any combination of the items, or all of the items with which this term is associated. The phrase "one or more" is readily understood by one of skill in the art, particularly when read in context of its usage. For example, one or more substituents on a phenyl ring refers to one to five, or one to four, for example if the phenyl ring is di-substituted. As used herein, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating a listing of items, “and / or” or “or” shall be interpreted as being inclusive, e.g., the inclusion of at least one, but also including more than one of a number of items, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of” or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating ZIMMER-PETN (03018-02) / / 1036.375WO1 exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” As used herein, the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof, are intended to be inclusive similar to the term “comprising.” The term "about" can refer to a variation of ± 5%, ± 10%, ± 20%, or ± 25% of the value specified. For example, "about 50" percent can in some embodiments carry a variation from 45 to 55 percent. For integer ranges, the term "about" can include one or two integers greater than and / or less than a recited integer at each end of the range. Unless indicated otherwise herein, the term "about" is intended to include values, e.g., weight percentages, proximate to the recited range that are equivalent in terms of the functionality of the individual ingredient, the composition, or the embodiment. The term about can also modify the endpoints of a recited range as discuss above in this paragraph. As will be understood by the skilled artisan, all numbers, including those expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and so forth, are approximations and are understood as being optionally modified in all instances by the term "about." These values can vary depending upon the desired properties sought to be obtained by those skilled in the art utilizing the teachings of the descriptions herein. It is also understood that such values inherently contain variability necessarily resulting from the standard deviations found in their respective testing measurements. As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges recited herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof, as well as the individual values making up the range, particularly integer values. A recited range (e.g., weight percentages or carbon groups) includes each specific value, integer, decimal, or identity within the range. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, or tenths. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art, all language such as "up to," "at least," "greater than," "less than," "more than," "or more," and the like, include the number recited and such terms refer to ranges that can be subsequently broken down into sub-ranges as discussed above. In the same manner, all ratios recited herein also include all sub-ratios falling within the broader ratio. Accordingly, specific values recited for radicals, substituents, and ranges, are for illustration only; they do not exclude other defined values or other values within defined ranges for radicals and substituents. ZIMMER-PETN (03018-02) / / 1036.375WO1 One skilled in the art will also readily recognize that where members are grouped together in a common manner, such as in a Markush group, the invention encompasses not only the entire group listed as a whole, but each member of the group individually and all possible subgroups of the main group. Additionally, for all purposes, the invention encompasses not only the main group, but also the main group absent one or more of the group members. The invention therefore envisages the explicit exclusion of any one or more of members of a recited group. Accordingly, provisos may apply to any of the disclosed categories or embodiments whereby any one or more of the recited elements, species, or embodiments, may be excluded from such categories or embodiments, for example, for use in an explicit negative limitation. The term "at least" prior to a number or series of numbers (e.g., "at least two") is understood to include the number adjacent to the term "at least," and all subsequent numbers or integers that could logically be included, as clear from context. When "at least" is present before a series of numbers or a range, it is understood that "at least" can modify each of the numbers in the series or range. As used herein, the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof, are intended to be inclusive similar to the term “comprising.” The terms “comprises,” “comprising,” and the like can have the meaning ascribed to them in U.S. Patent Law and can mean “includes,” “including” and the like. As used herein, “including” or “includes” or the like means including, without limitation. The terms “polypeptide” and “protein” refer to a polymer of amino acid residues and are not limited to a minimum length. Thus, peptides, oligopeptides, dimers, multimers, and the like, are included within the definition. Both full length proteins and fragments thereof are encompassed by the definition. The terms also include post-expression modifications of the polypeptide, for example, glycosylation, acetylation, phosphorylation, hydroxylation, and the like. Furthermore, for purposes of the present invention, a “polypeptide” refers to a protein which includes modifications, such as deletions, additions and substitutions to the native sequence, so long as the protein maintains the desired activity. These modifications may be deliberate, as through site directed mutagenesis, or may be accidental, such as through mutations of hosts which produce the proteins or errors due to PCR amplification. The term “BcsA synthase” as used herein encompasses BcsA encoded cellulose synthases from any bacterial species, and also includes biologically active fragments, variants, analogs and derivatives thereof that retain BcsA synthesis activity (i.e., catalyze synthesis of cellulose by polymerizing UDP-activated glucose). Of particular interest is the NTD of BcsA, ZIMMER-PETN (03018-02) / / 1036.375WO1 which recruits three copies of BcsG to its location so at to modify the cellulose. Provided herein the NTD of BcsA comprises about the first 140 amino acids of BcsA, see, for example, the first 140 amino acids of BcsA of SEQ ID NOS: 1-6 (FIG.8), note that this sequence (NTD) is absent in SEQ ID NO: 7 (RS). Escherichia coli MSILTRWLLIPPVNARLIGRYRDYRRHGASAFSATLGCFWMIL-AWIFIPLEHPRWQRIRAEHKNLYPHINA SRP-RPLDPVRYLIQTCWLLIGASRKETPKPR------------RRAFSGLQNIRGRYHQ----WMNELPER VSHKTQHLDEKKELGHLSAGARRLILGIIVTFSLILALICVTQPFNPLAQ-FIFLMLLWGVALIVRRMPGRF SALMLIVLSLTVSCRYIWWRYTSTLNWDD-PVSLVCGLILLFAETYAWIVLVLGYFQVVWPLNRQPVPLPKD MSLWPSVDIFVPTYNEDLNVVKNTIYASLGIDWPKDKLNIWILDDGG-------------------REEFRQ FAQNVGVKYIARTTHEHAKAGNINNALKYAKGEFVSIFDCDHVPTRSFLQMTMGWFLKEKQLAMMQTPHHFF SPDPFERNLGRFRKTPNEGTLFYGLVQDGNDMWDATFFCGSCAVIRRKPLDEIGGIAVETVTEDAHTSLRLH RRGYTSAYMRIPQAAGLATESLSAHIGQRIRWARGMVQIFRLDNPLTGKGLKFAQRLCYVNAMFHFLSGIPR LIFLTAPLAFLLLHAYIIYAPALMIALFVLPHMIHASLTNSKIQGKYRHSFWSEIYETVLAWYIAPPTLVAL INPHKGKFNVTAKGGLVEEEYVDWVISRPYIFLVLLNLVGVAVGIWRYFYGPPTEMLTVVVSMVWVFYNLIV LGGAVAVSVESKQVRRSHRVEMTMPAAIARED--GHLFSCTVQDFSDGGLGIKI------NGQAQILEGQKV NL---LLKRGQQEYVFPTQV--ARVMGNE--VGLKLM---PLTTQQHIDFVQCTFARADTWALWQDSYPEDK PLESLLDILKLGFRGYRHLAEFAPSSVKGIFRVLTS----------LVSWVVSFIPRRPERSET-AQPSDQA LAQQ-------- (SEQ ID NO: 1) Salmonella enterica subsp. enterica MSALSRWLLIPPVSARLSERYQGYRRHGASPFSAALGCLWMIL-AWIVFPLEHPRWQRIRDGHKALYPHINA ARP-RPLDPARYLIQTLWLVMISSTKERHEPR--------WR----SFARLKDVRGRYHQ----WMDTLPER VRQKTTHLEKEKELGHLSNGARRFILGVIVTFSLILALICITQPFNPLSQ-FIFLLLLWGVALLVRRMPGRF SALMLIVLSLTVSCRYIWWRYTSTLNWDD-PVSLVCGLILLFAETYAWIVLVLGYFQVVWPLNRQPVPLPKE MSQWPTVDIFVPTYNEDLNVVKNTIYASLGIDWPKDKLNIWILDDGG-------------------RESFRQ FARHVGVHYIARATHEHAKAGNINNALKHAKGEFVAIFDCDHVPTRSFLQMTMGWFLKEKQLAMMQTPHHFF SPDPFERNLGRFRKTPNEGTLFYGLVQDGNDMWDATFFCGSCAVIRRKPLDEIGGIAVETVTEDAHTSLRLH RRGYTSAYMRIPQAAGLATESLSAHIGQRIRWARGMVQIFRLDNPLFGKGLKLAQRLCYLNAMFHFLSGIPR LIFLTAPLAFLLLHAYIIYAPALMIALFVIPHMVHASLTNSKIQGKYRHSFWSEIYETVLAWYIAPPTLVAL INPHKGKFNVTAKGGLVEEKYVDWVISRPYIFLVLLNLLGVAAGVWRYYYGPENETLTVIVSLVWVFYNLVI LGGAVAVSVESKQVRRAHRVEIAMPGAIARED--GHLFSCTVHDFSDGGLGIKI------NGQAQVLEGQKV NL---LLKRGQQEYVFPTQV--VRVTGNE--VGLQLM---PLTTKQHIDFVQCTFARADTWALWQDSFPEDK PLESLLDILKLGFRGYRHLAEFAPPSVKVIFRSLTA----------LIAWIVSFIPRRPERQAA-IQPSDRV MAQAQQ------ (SEQ ID NO: 2) Citrobacter farmeri WP_195619268.1 / 2-874 MSVLSRWLLIPPVSARLSERYQGYRRHGASPFSAALGCLWVIL-AWVFIPLEHPRWQRIRAERKALYPHINA ARP-RPLDPARYAIQTFWLMASSPRQETSQPR--------WQ----TFSRMQGLIGRYHQ----WMDALPGR VTQNTTHLDKQKELGHLSPGARRFIIGIIVTFSLILALICVTQPFNPLAQ-FIFLMLLWGVALIVRRMPGRF ZIMMER-PETN (03018-02) / / 1036.375WO1 SALMLIVLSLTVSCRYIWWRYTSTLNWDD-PVSLVCGLILLFAETYAWIVLVLGYFQVVWPLNRQPVPLPKD MSLWPTVDIFVPTYNEDLHVVKNTIYASLGIDWPKDKLNIWILDDGG-------------------REEFRQ FAQTVGVKYIARTTHEHAKAGNINNALKYAKGEFVSIFDCDHVPTRSFLQMTMGWFLKDKNLAMMQTPHHFF SPDPFERNLGRFRKTPNEGTLFYGLVQDGNDMWDATFFCGSCAVIRRKPLDEIGGIAVETVTEDAHTSLRLH RRGYTSAYMRIPQAAGLATESLSAHIGQRIRWARGMVQIFRLDNPLTGKGLKFAQRLCYVNAMFHFLSGVPR LIFLTAPLAFLLLHAYIIYAPALMIALFVLPHMIHASLTNSKIQGKYRHSFWSEIYETVLAWYIAPPTLVAL INPHKGKFNVTAKGGLVEEEYVDWVISRPYIYLVLINLLGVAVGVWRYFYGPENEMLTVIVSLVWVFYNLII LGGAVAVSVESKQVRRSHRVEISMPAAIARED--GHLFSCTVHDFSDGGLGIKI------NGQAQVLEGQKV NL---LLKRGQQEYVFPTQV--ARVWGEQ--VGLQLM---PLTTKQHIDFVQCTFARADTWALWQDSYPEDK PLESLLDILKLGFRGYRHLAEFSPSSVKIIFRSLTS----------LVSWIVSFIPRRPEQDEVAVQQPDQV MAQQ-------- (SEQ ID NO: 3) Flavobacterium sp. MXW15 MCW4454233.1 ----MGWRYYP--------RREWRQERASGPVAAFVLWAFQAL-AWAFLRLEGPAWERWFRENTGSYPQFHD QRAWRAGDPLRFLVQSLWLMVVRA---EPLPRRPLDLSPLLHPLRLLRGGFERGRRRYVG----ALEDAPEA VAGSETMARLRGRERQLSPRARRWLTLALMLGTALLAMVCITQPFDYFAQ-FVFVVALWLTAMVIRRVPGRY ASLMLIVLSATVSCRYLWWRYTATLNWNN-NFDLACGIVLLAAETYSWLVLMLGYVQVAWPLNRKPAPLPAD PAQWPVVDVLIPTYNEDLDLVRNTVYAACGLDWPRDRLRIHLLDDGN-------------------REEFRR FAEVAGINYIARADNRHAKAGNLNNALGYIDGELVAIFDSDHMPVRSFLQVTVGWFLRDRRLALVQTPHHFF SPDPFERNLKVFRDAPNEGELFYGLVQDGNDLWNASFFCGSCAVLRREALDEIGGFAVETVTEDAHTALRLH RHGWNSAYLKIPQAAGLATGSLGAHVNQRIRWARGMVQIFRLDNPLLGKGLNVFQRLCYANAMLHFLSGIPR LVFLTAPLAFLLLHVYIIYAPALAIVLFVLPHMVHASITNARLQRAHRRPFWGEVYETVLSWYIARPTTVAL FSPHRGKFNVTAKDSLQEQATFDWRIARPYIVLAVLNLVGLGFAAWRLLHGPADEVGTVVVSSLWVVYNLVI IGGALAVAAEVHQLRRTHRVSTRLPAAVRTAS--GHCHPCTLTDYSSGGAGLEF------DEPVELERGVPI SL---LLGRGRRQFVFDGRV--ERSFDTR--LGLALE---FADERQRVDFAQCTFARADAWLDWHSGYQAKS LPHSLLGIVLLGWRGYRRVGDFAPAPLQKPARLASR----------SARWLASFAPRTPSSTRM-AADAASS GSPT-------- (SEQ ID NO: 4) Stutzerimonas stutzeri WP_125841845.1 -----MPNPLS--------YYQYIRYRRGASPIAALSWLIAYMTAWLLLRLDAPGWQWVMAERRRLYPHVAG KRP-TAGDPLRIVIQSIWLLVARAPGSAPARSLGDRLRSTWQRLRSVAQQLPRPPRLRHR----RTFDLPGL RNAEARVGKAERWLSSLSRPVRYAFNIAIGALAILLGILCITEPFGYTAQ-VTFVLLLWGLALLVRRVPGRF AVLMLIVLSAIISCRYLWWRYTATLHWDS-YFDLACGITLLVAETYSWIVLILGYIQTCWPLDRKPAPLPED SSSWPSVDLFIPTYNEDLSVVRTTVLAALGLDWPRDKLNVYICDDGR-------------------RDSFKQ FAEEVGVGYIVRPDNKHAKAGNLNHALTVTHSELIAIFDCDHIPVRSFLQVTTGWFLSDPKLALVQTPHHFF SPDPFERNLGSFRRKPNEGELFYGLVQNGNDMWNASFFCGSCAVLRRDAVESIGGFAVETVTEDAHTALRLH RAGWNSAYLGTPQAAGLATESLSAHIGQRIRWARGMAQIFRTDNPLLGPGLTIFQRLCYANAMLHFLAGLPR LIYLTAPLAFLLLHAYIIYAPALMIVLYVLPHMIHASLTNARMQGEYRHSFWGEVYETVLAWYIARPTTVAL FNPGKGKFNVTAKGGLIDHDQFDWRIARPYLVLAALNVAGLGFAVWRLFTGPVSEIGTVLVSSAWVIYNLLI IGAAVAVASEVRQIRRAHRVMAQLPASLKLAD--GHAYPCTLLDFAEGGAGLQI------PPGLKVDMDQPV SL---ILQRGDRSFMFPGQA--SRQIGER--LGIRLD---NLDLAQQIDLVQCTFARADVWLNRHQDFETDR ZIMMER-PETN (03018-02) / / 1036.375WO1 PLHSFIEVLRIGGRGYHRLYEQLPSVLSRPLRPLLR----------LAAWLLSYCPRTPASKSL-VTPS--- (SEQ ID NO: 5) Pseudomonas asiatica QOE10583.1 ----MTLTPLS--------AYAWFRTRGARQPVAWLFTLGVWL-AFLFLRLESPAWQALLAERQRLYPQLAG KRP-SLGDPLRLLIQSLWLLVRRQ----PLPREPGFARRAWSAVRAQLRGMHQVLRRYHGLLIESLQQLPAR YRASAFKQQASARLRGLSVFARWVFYSVLTLGAIGLAVLCVTEPFGYLAQ-LVFICLLLGIALLVRHMPGRF PTLMLIVLSTLISCRYLWWRYTSTLNWND-TTGLVCGLILLAAETYSWFVLILGYIQTSWPLQRKPANLPAN PAHWPTVDLMIPTYNEDLSVVRTTVLAALGLDWPRECLRIYILDDGR-------------------REAFRA FADEVGVGYIVRPDNKHAKAGNLNHALGVTDSELIAIFDCDHVPVRSFLQMTVGWFLKDSKLALVQTPHHFF SPDPFERNLGSFRRRPNEGELFYGLIQDGNDMWNAAFFCGSCAVLRRTALESIGGFAVETVTEDAHTALRMH RQGWASAYLSIPQAAGLATESLSAHIGQRIRWARGMVQIFRTDNPLFGRGLSLFQRVCYANAMLHFLAGLPR LVFLTAPLAFLLLHAYIIYAPALMILLYVLPHMIHASLTNSRMQGKYRQTFWGEVYETVLAWYIARPTTVAL FAPKKGKFNVTAKGGLMEQEQFDWRIAQPYLWLAALNVAGLGFAVWRLFTGPAAEIGTVIVSSLWVIYNLLI IGAAVAVAAEVRQVRRAHRVQMRLPAGLVLAS--GHAYPCTLVDYSDGGVGLQL------HRGLELQAGERV RL---LLNRGQREFAFQACV--TRTVGQH--VGLVFH---DLGQQQRIDLVHCTFARADAWLGWSEQHEVDR PLRSLVDVLKLGGVGYLRLVEHLPPWVRAWLRPLHA----------LANWLASYRPRTPRPVPS-LNPVDRD A----------- (SEQ ID NO: 6) R. sphaeroides (Cereibacter sphaeroides) ------------------------------------------------------------------------ ------------------------------------------------------------------------ --MTVRA-KARSPLRVVPVLLFLLWVALLVPFGLLAA-----APVAPSAQGLIALSAVVLVALLKPFADKMV PRFLLLSAASMLVMRYWFWRLFETLPPPALDASFLFALLLFAVETFSISIFFLNGFLSADPTDR-PFPRPLQ PEELPTVDILVPSYNEPADMLSVTLAAAKNMIYPARLRTVVLCDDGGTDQRCMSPDPELAQKAQERRRELQQ LCRELGVVYSTRERNEHAKAGNMSAALERLKGELVVVFDADHVPSRDFLARTVGYFVEDPDLFLVQTPHFFI NPDPIQRNLALGDRCPPENEMFYGKIHRGLDRWGGAFFCGSAAVLRRRALDEAGGFAGETITEDAETALEIH SRGWKSLYIDRAMIAGLQPETFASFIQQRGRWATGMMQMLLLKNPLFRRGLGIAQRLCYLNSMSFWFFPLVR MMFLVAPLIYLFFGIEIFVATFEEVLAYMPGYLAVSFLVQNALFARQRWPLVSEVYEVAQAPYLARAIVTTL LRPRSARFAVTAKDETLSENYIS-PIYRPLLFTFLLCLSGVLATLVRWVAFPGDRSVLLVVGG-WAVLNVLL VGFALRAVAEKQQRRAAPRVQMEVPAEAQIPAFGNRSLTATVLDASTSGVRLLVRLPGVGDPHPALEAGGLI QFQPKFPDAPQLERMVRGRIRSARREGGTVMVGVIFEAGQPIAVRETVAYL--IFGESAHWRTMREA--TMR PIGLLHGMARILWMAAASLPKTARDFMDEPARRRRRHEEPKEKQAHLLAFGTDFSTEPDWAGEL-LDPTAQV SARPNTVAWGSN (SEQ ID NO: 7) Other examples in which the NTD is absent include, but are not limited to, Paracoccaceae bacterium (accession no. MDT8856603.1), Oceaniglobus indicus (accession no. WP_099826996.1), Frigidibacter oleivorans (accession no. WP_126975478.1), Candidatus Azotimanducaceae bacterium (accession no. MFT6517335.1), Falsigemmobacter ZIMMER-PETN (03018-02) / / 1036.375WO1 faecalis (accession no. WP_124963257.1), and Paracoccaceae bacterium (accession no. TVR48340.1). Paracoccaceae bacterium (accession no. MDT8856603.1) 1 MTGAPFRRSA TANTPVLLAW FAVLVPVVLL ASAPTSTSGQ ALLGLVAVVC VAALKPFAAN 61 LGARFLLLGI SSVIVMRYWI WRLLETLPPP SLSLSFGIAV LLFAVESYSI LLFFLNGFIT 121 ADPTTRRFPA KVLPEDLPTV DILIPSYNEP PEMLSVTLAA AKNMIYPRAK RRVVLCDDGG 181 TDQRCNSSDP ELAARSRERR ATLQALCRDL GVRYSTRERN EHAKAGNMSA ALEQLDGQLV 241 VVFDADHVPS RDFLARTVGY FVDDPKLYLV QTPHFFINKD PIQRNLGLSE RCPPENEMFY 301 SFIHRGLDRW GGTFFCGSAA VLRRSALDSV GGFAGETITE DAETALEMHA AGWKSLYIDR 361 AMIAGLQPET FATFIEQRGR WATGMMQMLL LKNPLFRSGL RFAQRLCYIN SMSFWLFPLI 421 RLAFLLAPLV YLFFGIQIFV STYEDVLAYM LSYLAVGYLV QNALYARFRW PLISEIYEVA 481 QAPYLAKAVI KTFLSPRAAK FNVTAKDETL EEDYISPIHG PLTLLFGLML AGVVALVVRW 541 VMFPGDHSVL IVVGGWAVFN FLLVSIAWGA VSEKQQRRAS PRVDLRVPAT VRPQDAATEM 601 AATILDASTS GVRILVAPGS RDAAAPPIVP GQSFSFTPRF PDAPHLQTPV LGTVRSVRQG 661 PEGMIVGLML RPDQPMITRE AVAFLIFGDS DIWRTIREST RRRKGMFAGF GYALSLAFKV 721 LPSVIADFIR EPGRRARASM SETRRPKAAH LLAFGVEPEH LDLDAPDPGF DGSDGKVRPW 781 KVTA (SEQ ID NO: 14) Oceaniglobus indicus (accession no. WP_099826996.1) 1 MKAFAKKNRT GNFPILLLWT LIAIPIGVLI TVPTSTAGQA FLGIVAVAII ALLKPFAHLI 61 VPRFALLAVA TLLVMRYWVW RITETIPDIS QPLSFAVAML LLAVETYSIL VFLLNAFIVA 121 DPTERPFPAQ VSPEDLPTVD ILIPSYNEPV EMLSVTLAAA RNMIYPRSKR TVVLCDDGGT 181 DQRCNSSDPE LAAKSRQRRA ELQALCAELG VKYSTRARNE HAKAGNMSAA LEDLNGDLVV 241 VFDADHVPSR DFLARTVGYF NQDPKLFLVQ TPHFFINPDP IQRNLQLSPK CPPENEMFYS 301 FVHRGLDRWG GAFFCGSAAV LRRRALDSVG GFAGETITED AETALEIHAA GWKSLYLDRA 361 MIAGLQPETF ASFIEQRGRW ATGMMQMLLI KNPLFRRGLR LPQRLCYLNS MSFWLFPLIR 421 LTYLLVPLVY LFFGIEIFVA TFPEVMAYMM CYLAAAFMVQ NALYTRTRWP LISEIYEVAQ 481 APYLAKAVFK TVISPRSAKF NVTAKDETLA EDYISPIHWP LTIMWLAMLA GLIAFVIRWI 541 AFPGDHSVLS VVGGWAIFNF ILVSLAYRAV AEKQQRRASP RVDVETPATL WLAGDDSDTL 601 PVQLLDASTS GVRILLGRDV GSIQSSKLEG QLICFQPDFP EAPQLEVPVS ARIVSVQSTA 661 EGQILGLLLE ADQNMKAREA VAYLIFGDSE AWRRTRAASQ GRKGMLAGFG YTLYLFATGL 721 PALLRDLARE PARRAALDIL PPADEKPAHL LAFGVDLEAE ARARRAARQA LAEELAAAAE 781 DPPGWSDDGA VVGGKP (SEQ ID NO: 15) Frigidibacter oleivorans (accession no. WP_126975478.1) 1 MARMTSFLLL ATWIVLLVPI LVLISAPTSL ATQLCVSFVA LVIVVLVKPY ARRGWPRFIL 61 LAVGSVIVLR YWLWRVTSTL PDPGLSFSFT FAVLLLAIET YSIAVFFLNA FLVADPTRRF 121 VPPSITDVAD LPTVDILVPS YNEPTEMLSV TLSAAKNMIY PTAKRRVVLC DDGGTDQRCN 181 SSDLTLAARA RERRAELQAL CAELGIIYST RAKNEHAKAG NMSAALARLD GELVVVFDAD 241 HVPSRDFLAR TVGYFRQDPK LFLVQTPHFF INKDPIERNV GFSAKCPPEN EMFYGLIHRG ZIMMER-PETN (03018-02) / / 1036.375WO1 301 LDRWGGAFFC GSAAVLRRKA LDDVGGFAGE TITEDAETAL EIHSRGWKSL YLNRAMIAGL 361 QPETFSSFIQ QRGRWATGMM QMLLLKNPLM RKGLGLRQRL CYLNSMSFWF FPVIRLVYLL 421 APLVYLIFGI EIFVATIQEA AAYTGTYMIL SFLTQNAIFS RFRWPLVSEV YEVAQAPYLF 481 RAIVQTFRRP RAARFNVTAK DETLDSDFIS PVYGPLLALF GLMLAGVLLA IGRWIAFPGD 541 RAVLQVVGAW AVFNFLLVSL SLRATCEKQQ RRASPRIAMQ VPAEVTLAGS QAVQATIIDA 601 STSGAQVVVP RHLAGAGVAV GDMVSFVPRF PDARHLERPV RGVIRSTKEE GGTIALRLLF 661 PPDQPMTARE AVAFLIFGNS DVWEKMREDT RAGVGLTRGL LYIAWLAIRS LPLTFRDYVG 721 EPARRRKAAV APRQQQPAHL VAFGRDFVED TSPPVAAAAG GGAA (SEQ ID NO: 16) Candidatus Azotimanducaceae bacterium (accession no. MFT6517335.1) 1 MSASKKKGIK ALNAIKLGLW ACLAVVIALM VSIPTSTGGQ AFLSIVAVAV VAILKPFARL 61 LVARFFLLAT ASVIVLRYYL WRILDTLPDP GLTLSFIVAI MLIVVETYSI MVFFLNAFIG 121 ADPTRRPFPP TVAPEDLPTV DILVPSYNEP AEMLSVTLAA AKNMIYPSDK RTVILCDDGG 181 TDQRCNSSNA ELATSARARR AELQALCAEL GVVYATRARN EHAKAGNMSA ALENLTGDLV 241 VVFDADHVPS RDFLARTVNY FVDDSKLFLV QTPHFFINKD PIQRNLELSE LAPPENEMFY 301 SLIHRGLDRW GGAFFCGSAA VLRRKALDSV GGFAGETITE DAETALEIHA AGWKSLYIDR 361 AMIAGLQPET FASFIQQRGR WAAGMMQMLL LKQPLFRRGL RLPQRICYIN SMSFWLFPLM 421 RMFYLYVPLV YLFFGVEIFV ATLDEVMAYM FGYLAVSFMV QNALYARFRW PLVSEIYEVA 481 QAPYLARAVL RTIIKPRGAT FNVTAKDEVL DEDYISPIHW PLTILFLTML SGIVALAIRW 541 YEFPGDHGVL SVVGGWAIFN FILVSIAYRA VAEKQQRRAS PRVEMDVPGQ FWLPDDTEKA 601 VSVRIVDTST SGVRLLLQEG QILPERLAAA ELKSKDIVLR PRLVESPHLE SRIVGTVQSV 661 QQTPSGVVLG VIFKPDQPMH TRETVAYLIF GDSENWRRIR HSTRRPKGLL MGLAYVLKLF 721 LRGTPLLLID LIKALRAPRA IDQETPAASK PAHLLAFGID PEDRSKIEPE ESPIPTPVII 781 Q (SEQ ID NO: 17) Falsigemmobacter faecalis (accession no. WP_124963257.1) 1 MTGRRLIGSL AIPLWICLLV PVVLLCILPV SNAAQAMLGA SAVLLVMCLK PFVHHMAGRL 61 ALMGTASVVV LRYWSWRITS TLPDPGLNAS FILALALLLV ETYSILVFFL NAFITADPVD 121 RGLPPKLELD QMPTVDILVP SYNEPVEMLS ITLSAAKNMV YPARLRTVVL CDDGGTDQRC 181 NSSDPDLAAK ARARRAELQA LCRDLGVVYS TRAKNEHAKA GNMSAALAKL DGDLVVVFDA 241 DHVPSRDFLA RTVGYFVEDP KLFLVQTPHF FINKDPIERN LGLKCPPENE MFYGMIHRGL 301 DRWGGAFFCG SAAVLRRKAL DSVGGFAGET ITEDAETALE IHSKGWRSLY LDRAMIAGLQ 361 PETFASFIQQ RGRWASGMIQ MLMLKNPLFR RGLTPLQRLC YINSMSFWFF PLIRLVYLLA 421 PLTYLFFSVE IFVTTWTEAM AYTLSYMAIV LMVQNATFAR FRWPLISEIY EIAQAPYLAT 481 AIMRTVLRPR AAKFNVTAKD ETLVEDYISP IFGPLLLLFG LTAAGVAALI FRWIAFPGDR 541 SVLIVVGTWA VLNFLLVSLS LRAVSEKQQR RASPRVEMQA PAQVKLAADD SRAVTLPAEV 601 LDASTSGARI RLGAMPRGEP IAMPQKGDVL VLRPLFDEAP HLERDVRAVV RGVFNKAGEG 661 HTIGLQFLSD QPMDVREAVA QLVFGSSENW LAERMRGRKR KGLLAGLFYV FSLMLTSIPK 721 TLGDFLREPA RRTKSELTAH HKPKAAHLVA FGANFDTEEF REGARRSTLE DFIATEDRA (SEQ ID NO: 18) ZIMMER-PETN (03018-02) / / 1036.375WO1 Paracoccaceae bacterium (accession no. TVR48340.1) 1 MSLLIATRKL RPVEILLALV WFAMFVLLAV MASIPTSTTV QGALGLFAVI AVALLKPFTM 61 KNIVARFMLL AIASAVVMRY WAWRVTETLP PMDMPLSFAI ALALLVVETY AIGVFFISSF 121 ITADPVERKL PPRVMAADLP TVDILVPSYN EPVEMLSITL SAAKNMHYPA SKRTVVLCDD 181 GGTDQRCNSD DPDLAEKSRA RRAELEALCA ELGVVYSTRA RNEHAKAGNM SAALERLNGD 241 LVVVFDADHV PSRDFLARTV GYFVDDPKLF LVQTPHFFLN PDPVERNIGL RKDCPPENEM 301 FYHQGHRGLD RWGGAFFCGS AAVIRRAALD SVGGFAGETI TEDAETALEI HSQGWKSIYV 361 DHAMIAGLQP ESFVSFIQQR GRWAAGMMQL LRLKNPLRRK GMSLSQRLCY LNSMTFWLFP 421 LVRMTFILAP LAYLFFGLQI FVATIQEVMV YMGAYMAISF MVQNALYSRV RWPLISELYE 481 TAQAPYLSGV VLRTLFKPRG AKFNVTAKDE VLEEDFISPI YQPLLLVWAL AGLGVVAACV 541 RWVMFPGDQN ILTIVGGWAV FNFIILSAAL RAIAERQQRR EVPRVKMDVP AVAAIGRNGS 601 FSFVCAQVLD SSTSGASIQL VAGPDTDLKE MATISRGYVF YFTPEFPQSP HLENAVRVQV 661 KYVVRENGGI RVGVCYDRDQ PFKTRETVAH LIFGNSAIWE QERARKNKPM PMLKGMAYVV 721 GLAFRSVFHT MRALIAEPAR QQAADERARR REIETTQPAH LLAFGETFDP EIPNPDMPAP 781 DSFATMPGMS LNRMAPPTTG VEEQR (SEQ ID NO: 19) In one example, the NTD is SEQ ID NO: 9 (residues 1- 94; as shown on FIG.8 for Ec; MSILTRWLLI PPVNARLIGR YRDYRRHGAS AFSATLGCFW MILAWIFIPL EHPRWQRIRA EHKNLYPHIN ASRPRPLDPV RYLIQTCWLL IGAS), SEQ ID NO: 13 (residues 1-149; as shown in FIG. 8 for Ec; MSILTRWLLI PPVNARLIGR YRDYRRHGAS AFSATLGCFW MILAWIFIPL EHPRWQRIRA EHKNLYPHIN ASRPRPLDPV RYLIQTCWLL IGASRKETPK PRRRAFSGLQ NIRGRYHQWM NELPERVSHK TQHLDEKKEL GHLSAGARR) or a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto. Another sequence of interest is the C-terminal sequence of BcsA, in particular residues Ec 828-872 (as shown in FIG. 8; AEFAPSSVKGIFRVLTSLVSWVVSFIPRRPERSETAQPSDQALAQ (SEQ ID NO: 10)) or a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto. A BcsA polynucleotide, nucleic acid, oligonucleotide, protein, polypeptide, or peptide refers to a molecule derived from any source. The molecule need not be physically derived from an organism but may be synthetically or recombinantly produced. BcsA sequences from a number of bacterial species are well known in the art. Representative sequences are presented for BcsA (see FIG.8) from Escherichia coli (SEQ ID NO:1), S. enterica, (SEQ ID NO: 2), C. farmeri (SEQ ID NO: 3), Flavob. sp. (SEQ ID NO: 4), S. stutzeri (SEQ ID NO: 5), P. asiatica ZIMMER-PETN (03018-02) / / 1036.375WO1 (SEQ ID NO: 6) and R. sphaeroides (Rs; SEQ ID NO: 7) and additional representative sequences are listed in the National Center for Biotechnology Information (NCBI) database, all of which sequences (as entered by the date of filing of this application) are herein incorporated by reference. Any of these sequences or a variant thereof comprising a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto, can be used for cellulose modification, as described herein, wherein the variant retains biological activity. The term “BscB” as used herein encompasses BscB membrane-anchored protein from any bacterial species, and also includes biologically active fragments, variants, analogs and derivatives thereof that retain BcsB activity (i.e., cellulose synthesis involves the BcsA and BcsB proteins, where BcsA is the catalytic cellulose synthase and BcsB is a membrane- anchored protein for catalysis and polysaccharide translocation). A BcsB polynucleotide, nucleic acid, oligonucleotide, protein, polypeptide, or peptide refers to a molecule derived from any source. The molecule need not be physically derived from an organism but may be synthetically or recombinantly produced. BcsB sequences from a number of bacterial species are well known in the art. Representative sequences are listed in the National Center for Biotechnology Information (NCBI) database. See, for example, NCBI entries: AVZ42317.1, SQH82896.1, CAX57381.1, VUC65418.1, CCC32266.1, CAX61932.1, EBA46556.1, KAB1489763.1, XIW28436.1, WOY92771.1, WKE05123.1, XQF98892.1, XQC09001.1, XQC25973.1, XQB87122.1, and XPZ56261.1; all of which sequences (as entered by the date of filing of this application) are herein incorporated by reference. Any of these sequences or a variant thereof comprising a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto, can be used for cellulose modification, as described herein, wherein the variant retains biological activity, such as BcsB activity. In one embodiment, the BcsB is the regulator of cellulose synthase, cyclic di-GMP binding protein of Escherichia coli K-12; accession no. AKK14698.2, provided below as SEQ ID NO: 20: 1 MYRLFRAARS GAKRHNHRIR LWLNNDDNAM KRKLFWICAV AMGMSAFPSF MTQATPATQP 61 LINAEPAVAA QTEQNPQVGQ VMPGVQGADA PVVAQNGPSR DVKLTFAQIA PPPGSMVLRG 121 INPNGSIEFG MRSDEVVTKA MLNLEYTPSP SLLPVQSQLK VYLNDELMGV LPVTKEQLGK 181 KTLAQMPINP LFISDFNRVR LEFVGHYQDV CEKPASTTLW LDVGRSSGLD LTYQTLNVKN ZIMMER-PETN (03018-02) / / 1036.375WO1 241 DLSHFPVPFF DPSDNRTNTL PMVFAGAPDV GLQQASAIVA SWFGSRSGWR GQNFPVLYNQ 301 LPDRNAIVFA TNDKRPDFLR DHPAVKAPVI EMINHPQNPY VKLLVVFGRD DKDLLQAAKG 361 IAQGNILFRG ESVVVNEVKP LLPRKPYDAP NWVRTDRPVT FGELKTYEEQ LQSSGLEPAA 421 INVSLNLPPD LYLMRSTGID MDINYRYTMP PVKDSSRMDI SLNNQFLQSF NLSSKQEANR 481 LLLRIPVLQG LLDGKTDVSI PALKLGATNQ LRFDFEYMNP MPGGSVDNCI TFQPVQNHVV 541 IGDDSTIDFS KYYHFIPMPD LRAFANAGFP FSRMADLSQT ITVMPKAPNE AQMETLLNTV 601 GFIGAQTGFP AINLTVTDDG STIQGKDADI MIIGGIPDKL KDDKQIDLLV QATESWVKTP 661 MRQTPFPGIV PDESDRAAET RSTLTSSGAM AAVIGFQSPY NDQRSVIALL ADSPRGYEML 721 NDAVNDSGKR ATMFGSVAVI RESGINSLRV GDVYYVGHLP WFERVWYALA NHPILLAVLA 781 AISVILLAWV LWRLLRIISR RRLNPDNE In another embodiment, SEQ ID NO: 20 modified and the modified sequence was used in the experiments provided herein and is provided as SEQ ID NO: 21: AWSHPQFEKTPATQPLINAEPAVAAQTEQNPQVGQVMPGVQGADAPVVAQNGPSRDVKLTFAQIAPPPGSMVLRG INPNGSIEFGMRSDEVVTKAMLNLEYTPSPSLLPVQSQLKVYLNDELMGVLPVTKEQLGKKTLAQMPINPLFISD FNRVRLEFVGHYQDVCEKPASTTLWLDVGRSSGLDLTYQTLNVKNDLSHFPVPFFDPSDNRTNTLPMVFAGAPDV GLQQASAIVASWFGSRSGWRGQNFPVLYNQLPDRNAIVFATNDKRPDFLRDHPAVKAPVIEMINHPQNPYVKLLV VFGRDDKDLLQAAKGIAQGNILFRGESVVVNEVKPLLPRKPYDAPNWVRTDRPVTFGELKTYEEQLQSSGLEPAA INVSLNLPPDLYLMRSTGIDMDINYRYTMPPVKDSSRMDISLNNQFLQSFNLSSKQEANRLLLRIPVLQGLLDGK TDVSIPALKLGATNQLRFDFEYMNPMPGGSVDNCITFQPVQNHVVIGDDSTIDFSKYYHFIPMPDLRAFANAGFP FSRMADLSQTITVMPKAPNEAQMETLLNTVGFIGAQTGFPAINLTVTDDGSTIQGKDADIMIIGGIPDKLKDDKQ IDLLVQATESWVKTPMRQTPFPGIVPDESDRAAETRSTLTSSGAMAAVIGFQSPYNDQRSVIALLADSPRGYEML NDAVNDSGKRATMFGSVAVIRESGINSLRVGDVYYVGHLPWFERVWYALANHPILLAVLAAISVILLAWVLWRLL RIISRRRLNPDNE The term “BcsG phosphoethanolamine transferase” as used herein encompasses BcsG encoded phosphoethanolamine transferases from any bacterial species, and also includes biologically active fragments, variants, analogs, and derivatives thereof that retain BcsG phosphoethanolamine transferase activity (i.e., catalyze transfer of a phosphoethanolamine group to a cellulose hydroxyl group to produce a phosphoethanolamine-modified cellulose). A BcsG polynucleotide, nucleic acid, oligonucleotide, protein, polypeptide, or peptide refers to a molecule derived from any source. The molecule need not be physically derived from an organism but may be synthetically or recombinantly produced. BcsG sequences from a number of bacterial species are well known in the art. Representative sequence are presented for BcsG from Escherichia coli (SEQ ID NO:11), and additional representative sequences are listed in the National Center for Biotechnology Information (NCBI) database. See, for example, NCBI entries: WP_282597949.1, WP_005045901.1, WP_407959879.1, WP_406021287.1, ZIMMER-PETN (03018-02) / / 1036.375WO1 WP_323914166.1, WP_406021285.1, WP_337783804.1, WP_331853836.1, WP_331853835.1, WP_031491801.1, WP_305730834.1, WP_127674701.1, WP_117150469.1, WP_060082415.1, CAK1349375.1, CAK1207332.1, CAK1210766.1, CAK0732961.1, CAK0722996.1, CAK0724424.1, CAK0718659.1, CAK0713585.1, CAK0712198.1, CAK0693424.1, CAK0678636.1, ASF66530.1, KAB1489769.1, UZW42856.1, XPU99516.1, XPV03915.1, XPE34464.1, XPE12755.1, XNR53429.1, XNR44410.1 and XIR93626.1; all of which sequences (as entered by the date of filing of this application) are herein incorporated by reference. Any of these sequences or a variant thereof comprising a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto, can be used for cellulose modification, as described herein, wherein the variant retains biological activity, such as BcsG phosphoethanolamine transferase activity. In one embodiment, the BcsG is the cellulose biosynthesis protein BcsG from Escherichia coli with accession no. WP_001735853.1 (as used in the examples provided herein), provided below as SEQ ID NO: 11: 1 MTQFTQNTAM PSSLWQYWRG LSGWNFYFLV KFGLLWAGYL NFHPLLNLVF AAFLLMPLPR 61 YSLHRLRHWI ALPIGFALFW HDTWLPGPES IMSQGSQVAG FSTDYLIDLV TRFINWQMIG 121 AIFVLLVAWL FLSQWIRITV FVVAILLWLN VLTLAGPSFS LWPAGQPTTT VTTTGGNAAA 181 TVAATGGAPV VGDMPAQTAP PTTANLNAWL NNFYNAEAKR KSTFPSSLPA DAQPFELLVI 241 NICSLSWSDI EAAGLMSHPL WSHFDIEFKN FNSATSYSGP AAIRLLRASC GQTSHTNLYQ 301 PANNDCYLFD NLSKLGFTQH LMMGHNGQFG GFLKEVRENG GMQSELMDQT NLPVILLGFD 361 GSPVYDDTAV LNRWLDVTEK DKNSRSATFY NTLPLHDGNH YPGVSKTADY KARAQKFFDE 421 LDAFFTELEK SGRKVMVVVV PEHGGALKGD RMQVSGLRDI PSPSITDVPV GVKFFGMKAP 481 HQGAPIVIEQ PSSFLAISDL VVRVLDGKIF TEDNVDWKKL TSGLPQTAPV SENSNAVVIQ 541 YQDKPYLRLN GGDWVPYPQ The catalytic domain of BcsG can be exchanged for another catalytic domain to generate a chimeric / fusion protein that modifies cellulose differently than the wild type BcsG catalytic domain. For example, the C-terminal domain of BcsG from E. coli functions as a phosphoethanolamine transferase with substrate preference for cellulosic materials. Structural characterization of C-terminal domain of BcsG from E. coli revealed that it belongs to the alkaline phosphatase superfamily and contains a Zn2+ion at its active center. Such a C-terminal domain can comprise Ala164 to Gln559 of E coli. accession number P37659 (UniProt), provided herein below: 1 MTQFTQNTAM PSSLWQYWRG LSGWNFYFLV KFGLLWAGYL NFHPLLNLVF AAFLLMPLPR 61 YSLHRLRHWI ALPIGFALFW HDTWLPGPES IMSQGSQVAG FSTDYLIDLV TRFINWQMIG ZIMMER-PETN (03018-02) / / 1036.375WO1 121 AIFVLLVAWL FLSQWIRITV FVVAILLWLN VLTLAGPSFS LWPAGQPTTT VTTTGGNAAA 181 TVAATGGAPV VGDMPAQTAP PTTANLNAWL NNFYNAEAKR KSTFPSSLPA DAQPFELLVI 241 NICSLSWSDI EAAGLMSHPL WSHFDIEFKN FNSATSYSGP AAIRLLRASC GQTSHTNLYQ 301 PANNDCYLFD NLSKLGFTQH LMMGHNGQFG GFLKEVRENG GMQSELMDQT NLPVILLGFD 361 GSPVYDDTAV LNRWLDVTEK DKNSRSATFY NTLPLHDGNH YPGVSKTADY KARAQKFFDE 421 LDAFFTELEK SGRKVMVVVV PEHGGALKGD RMQVSGLRDI PSPSITDVPV GVKFFGMKAP 481 HQGAPIVIEQ PSSFLAISDL VVRVLDGKIF TEDNVDWKKL TSGLPQTAPV SENSNAVVIQ 541 YQDKPYVRLN GGDWVPYPQ (SEQ ID NO: 12) In some embodiments, fusion proteins were generated with the following N-terminus of BcsG from SEQ ID NO: 11: MTQFTQNTAMPSSLWQYWRGLSGWNFYFLVKFGLLWAGYLNFHPLLNLVFAAFLLMPLPRYSLHRLRHWI ALPIGFALFWHDTWLPGPESIMSQGSQVAGFSTDYLIDLVTRFINWQMIGAIFVLLVAWLFLSQWIRITV FVVAILLWLNVLTLAGPSFSLWPAGQPTTTVTTTGGNAAATVAATGGAPVVGDMPAQTAPP (SEQ ID NO: 23) A catalytic domain can then fused to the N-terminus of the BcsG protein. This C-terminal catalytic domain of BcsG can be exchanged with another catalytic that modifies cellulose. For example, WssI from the Gram-negative (e.g., Pseudomonas aeruginosa, Pseudomonas fluorescens and pathogenic Achromobacter species, e.g., Achromobacter insuavis or dolens) bacterial cellulose synthase is an O-acetyltransferase that acts on cello-oligomers with several acetyl donor substrates. For example, WssI from Achromobacter dolens, alginate O-acetyltransferase AlgX-related protein; Accession no. WP_175167905.1: GDLGPRVRRGCDGWLFLGDELQPHPAARENQAERARIVVSLRDALAARGIRLLVAVVPDKSRIESARLCG LHRSAGFEDRLSSWVGVLRAQGVATVDLSAALRGVPQDAYYRNDSHWTEAGAGAAARAVAEQVRASGVAL QAPQRWRVTAQPPAPRPGDLVRLAGVDWLPLAWQPRAEVVALHTYTPEAAASAGDADDLFGDSALPSLAL VGTSFSRTSEFLPQLSRDLGVAVGNFARDGGKFGGAAQAYFKSPAWKQSPPRLLIWEMDERDLGAPLAAE DRVGF (SEQ ID NO: 24) WssI from Pseudomonas fluorescens; alginate O-acetyltransferase AlgX-related protein; accession no.WP_012721729.1: DTGPRVRPGCPGWLFISDELRINRHAEANAQTKAQAVIDLQKQLGQKGIDLQVVVVPDKSRIAAAQRCGL YRPAVLDNRVRDWTAMLQAAGVSALDLTETLKPLGAEAYLRTDTHWSEIGSNAGAKAVAQRTQQRGIKAT PEQTFDITQAPLAVRPGDLVRLAGLDWLPPTLQPPGESVAASTTHETGGATSNADDLFGDAGLPNVALIG TSFSRNSNFVGFLQKALNAPVGNFSKDGGEFSGAAKAYFDSPAFKQTPPKLLIWEIPERDLQTPYDVITI GQ (SEQ ID NO: 25) ZIMMER-PETN (03018-02) / / 1036.375WO1 WssF from Pseudomonas fluorescens; SGNH / GDSL hydrolase family protein; accession no.WP_012721726.1: MPVSAIAGLTMLVLGESHMSFPDSLLNPLQDNLTKQGAVVHSIGACGAGAADWVVPKKVECGGERTPTGK AVIYGKNAMSTTPIQELIAKDKPDVVVLIIGDTMGSYTNPVFPKAWAWKSVTSLTKAITDTGTKCVWVGP PWGKVGSQYKKDDTRTKLMSSFLASNVAPCTYIDSLTFSKPGEWITTDGQHFTIDGYQKWAKAIGTALGD LPPSAYGKGNK (SEQ ID NO: 26) By “fragment” is intended a molecule consisting of only a part of the intact full-length sequence and structure. The fragment can include a C-terminal deletion an N-terminal deletion, and / or an internal deletion of the polypeptide. Active fragments of a particular protein or polypeptide will generally include at least about 5-10 contiguous amino acid residues of the full length molecule, preferably at least about 15-25 contiguous amino acid residues of the full length molecule, and most preferably at least about 20-50 or more contiguous amino acid residues of the full length molecule, or any integer between 5 amino acids and the full length sequence, provided that the fragment in question retains biological activity, such as BcsG activity or BscA recruitment of three copies of BcsG. “Substantially purified” generally refers to isolation of a substance (compound, cellulose or modified cellulose, oligosaccharide, monosaccharide, disaccharide, polysaccharide, polynucleotide, nucleic acid, protein, polypeptide, or peptide) such that the substance comprises the majority percent of the sample in which it resides. Typically, in a sample, a substantially purified component comprises 50%, including 80%-85%, including 90- 95% of the sample. Techniques for purifying cellulose, saccharides, polynucleotides, and polypeptides of interest are well-known in the art and include, for example, ion-exchange chromatography, affinity chromatography and sedimentation according to density. By “isolated” is meant, when referring to a cellulose or modified cellulose, oligosaccharide, monosaccharide, disaccharide, polysaccharide, or polypeptide, that the indicated molecule is separate and discrete from the whole organism with which the molecule is found in nature or is present in the substantial absence of other biological macromolecules of the same type. The term “isolated” with respect to a polynucleotide is a nucleic acid molecule devoid, in whole or part, of sequences normally associated with it in nature; or a sequence, as it exists in nature, but having heterologous sequences in association therewith; or a molecule disassociated from the chromosome. ZIMMER-PETN (03018-02) / / 1036.375WO1 The terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecule” are used herein to include a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, the term includes triple-, double- and single-stranded DNA, as well as triple-, double- and single-stranded RNA. It also includes modifications, such as by methylation and / or by capping, and unmodified forms of the polynucleotide. More particularly, the terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecule” include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D- ribose), any other type of polynucleotide which is an N- or C-glycoside of a purine or pyrimidine base, and other polymers containing nonnucleotidic backbones, for example, polyamide (e.g., peptide nucleic acids (PNAs)) and polymorpholino (commercially available from the Anti-Virals, Inc., Corvallis, Oreg., as Neugene) polymers, and other synthetic sequence-specific nucleic acid polymers providing that the polymers contain nucleobases in a configuration which allows for base pairing and base stacking, such as is found in DNA and RNA. There is no intended distinction in length between the terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecule,” and these terms will be used interchangeably. Thus, these terms include, for example, 3′-deoxy-2′,5′-DNA, oligodeoxyribonucleotide N3′ P5′ phosphoramidates, 2′-O-alkyl-substituted RNA, double- and single-stranded DNA, as well as double- and single-stranded RNA, microRNA, DNA:RNA hybrids, and hybrids between PNAs and DNA or RNA, and also include known types of modifications, for example, labels which are known in the art, methylation, “caps,” substitution of one or more of the naturally occurring nucleotides with an analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, C5-propynylcytidine, C5- propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7- deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), with negatively charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), and with positively charged linkages (e g , aminoalklyphosphoramidates, aminoalkylphosphotriesters), those containing pendant moieties, such as, for example, proteins (including nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelators (e.g., metals, radioactive metals, boron, oxidative metals, etc.), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids, etc.), as well as unmodified forms of the polynucleotide or oligonucleotide. The ZIMMER-PETN (03018-02) / / 1036.375WO1 term also includes locked nucleic acids (e.g., comprising a ribonucleotide that has a methylene bridge between the 2′-oxygen atom and the 4′-carbon atom). See, for example, Kurreck et al. (2002) Nucleic Acids Res.30: 1911-1918; Elayadi et al. (2001) Curr. Opinion Invest. Drugs 2: 558-561; Orum et al. (2001) Curr. Opinion Mol. Ther. 3: 239-243; Koshkin et al. (1998) Tetrahedron 54: 3607-3630; Obika et al. (1998) Tetrahedron Lett.39: 5401-5404. The phrase "substantial identity" or "substantially identical," used in the context of two nucleic acids or polypeptides, refers to a sequence that has at least 60% sequence identity with a reference sequence. Alternatively, percent identity can be any integer from 70% to l 00%. In some embodiments, a sequence is substantially identical to a reference sequence if the sequence has at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the reference sequence as determined using the methods described herein, such as BLAST using standard parameters. For sequence comparison, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters. A "comparison window," as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of from 20 to 600, usually 30 about 50 to about 200, more usually about 100 to about 150 in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith & Waterman, Adv. Appl. Math.2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J: Mol. Biol.48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85: 2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by manual alignment and visual inspection. Algorithms that are suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms. which are described in Altschul et al. (1990) J. Mol. Biol.215: 403-410 and Altschul et al. (1977) Nucleic Acids Res.25: 3389-3402, ZIMMER-PETN (03018-02) / / 1036.375WO1 respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (NCBI) web site. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues: always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=1, N=2, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1989)). The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin & Altschul, Proc. Natl. Acad. Sci. USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.01, including less than about 10-5, including less than about 10-20. “Recombinant” as used herein to describe a nucleic acid molecule means a polynucleotide of genomic, cDNA, viral, semisynthetic, or synthetic origin which, by virtue of its origin or manipulation, is not associated with all or a portion of the polynucleotide with which it is associated in nature. The term “recombinant” as used with respect to a protein or polypeptide means a polypeptide produced by expression of a recombinant polynucleotide. In general, the gene of interest is cloned and then expressed in transformed organisms, as ZIMMER-PETN (03018-02) / / 1036.375WO1 described further below. The host organism expresses the foreign gene to produce the protein under expression conditions. The term “transformation” refers to the insertion of an exogenous polynucleotide into a host cell, irrespective of the method used for the insertion. For example, direct uptake, transduction or f-mating are included. The exogenous polynucleotide may be maintained as a non-integrated vector, for example, a plasmid, or alternatively, may be integrated into the host genome. “Recombinant host cells”, “host cells,” “cells”, “cell lines,” “cell cultures,” and other such terms denoting microorganisms or higher eukaryotic cell lines cultured as unicellular entities refer to cells which can be, or have been, used as recipients for recombinant vector or other transferred DNA, and include the original progeny of the original cell which has been transfected. A “coding sequence” or a sequence which “encodes” a selected polypeptide, is a nucleic acid molecule which is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo when placed under the control of appropriate regulatory sequences (or “control elements”). The boundaries of the coding sequence can be determined by a start codon at the 5′ (amino) terminus and a translation stop codon at the 3′ (carboxy) terminus. A coding sequence can include, but is not limited to, cDNA from viral, prokaryotic or eukaryotic mRNA, genomic DNA sequences from viral or prokaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence may be located 3′ to the coding sequence. Typical “control elements,” include, but are not limited to, transcription promoters, transcription enhancer elements, transcription termination signals, polyadenylation sequences (located 3′ to the translation stop codon), sequences for optimization of initiation of translation (located 5′ to the coding sequence), and translation termination sequences. “Operably linked” refers to an arrangement of elements wherein the components so described are configured so as to perform their usual function. Thus, a given promoter operably linked to a coding sequence is capable of effecting the expression of the coding sequence when the proper enzymes are present. The promoter need not be contiguous with the coding sequence, so long as it functions to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between the promoter sequence and the coding sequence, and the promoter sequence can still be considered “operably linked” to the coding sequence. “Expression cassette” or “expression construct” refers to an assembly which is capable of directing the expression of the sequence(s) or gene(s) of interest. An expression cassette ZIMMER-PETN (03018-02) / / 1036.375WO1 generally includes control elements, as described above, such as a promoter which is operably linked to (so as to direct transcription of) the sequence(s) or gene(s) of interest, and often includes a polyadenylation sequence as well. Within certain embodiments of the invention, the expression cassette described herein may be contained within a donor polynucleotide, plasmid, or viral vector construct. In addition to the components of the expression cassette, the construct may also include, one or more selectable markers, a signal which allows the construct to exist as single stranded DNA (e.g., a M13 origin of replication), at least one multiple cloning site, and a “mammalian” origin of replication (e.g., a SV40 or adenovirus origin of replication). “Purified polynucleotide” refers to a polynucleotide of interest or fragment thereof which is essentially free, e.g., contains less than about 50%, preferably less than about 70%, and more preferably less than about at least 90%, of the protein with which the polynucleotide is naturally associated. Techniques for purifying polynucleotides of interest are well-known in the art and include, for example, disruption of the cell containing the polynucleotide with a chaotropic agent and separation of the polynucleotide(s) and proteins by ion-exchange chromatography, affinity chromatography and sedimentation according to density. The term “transfection” is used to refer to the uptake of foreign DNA by a cell. A cell has been “transfected” when exogenous DNA has been introduced inside the cell membrane. A number of transfection techniques are generally known in the art. See, e.g., Graham et al. (1973) Virology, 52:456, Sambrook et al. (2001) Molecular Cloning, a laboratory manual, 3rd edition, Cold Spring Harbor Laboratories, N.Y., Davis et al. (1995) Basic Methods in Molecular Biology, 2nd edition, McGraw-Hill, and Chu et al. (1981) Gene 13:197. Such techniques can be used to introduce one or more exogenous DNA moieties into suitable host cells. The term refers to both stable and transient uptake of the genetic material and includes uptake of peptide- or antibody-linked DNAs. A “vector” is capable of transferring nucleic acid sequences to target cells (e.g., viral vectors, non-viral vectors, particulate carriers, and liposomes). Typically, “vector construct,” “expression vector,” and “gene transfer vector,” mean any nucleic acid construct capable of directing the expression of a nucleic acid of interest and which can transfer nucleic acid sequences to target cells. Thus, the term includes cloning and expression vehicles, as well as plasmid and viral vectors. The terms “variant,” “analog” and “mutein” refer to biologically active derivatives of the reference molecule that retain desired activity. In general, the terms “variant” and “analog” refer to compounds having a native polypeptide sequence and structure with one or more amino acid additions, substitutions (generally conservative in nature) and / or deletions, relative to the ZIMMER-PETN (03018-02) / / 1036.375WO1 native molecule, so long as the modifications do not destroy biological activity, and which are “substantially homologous” to the reference molecule as defined below. In general, the amino acid sequences of such analogs will have a high degree of sequence homology to the reference sequence, e.g., amino acid sequence homology of more than 50%, generally more than 60%- 70%, even more particularly 80%-85% or more, such as at least 90%-95% or more, when the two sequences are aligned. Often, the analogs will include the same number of amino acids but will include substitutions, as explained herein. The term “mutein” further includes polypeptides having one or more amino acid-like molecules including but not limited to compounds comprising only amino and / or imino molecules, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids, etc.), polypeptides with substituted linkages, as well as other modifications known in the art, both naturally occurring and non-naturally occurring (e.g., synthetic), cyclized, branched molecules and the like. The term also includes molecules comprising one or more N-substituted glycine residues (a “peptoid”) and other synthetic amino acids or peptides. (See, e.g., U.S. Pat. Nos. 5,831,005; 5,877,278; and 5,977,301; Nguyen et al., Chem. Biol. (2000) 7:463-473; and Simon et al., Proc. Natl. Acad. Sci. USA (1992) 89:9367-9371 for descriptions of peptoids). Methods for making polypeptide analogs and muteins are known in the art and are described further below. As explained above, analogs generally include substitutions that are conservative in nature, i.e., those substitutions that take place within a family of amino acids that are related in their side chains. Specifically, amino acids are generally divided into four families: (1) acidic— aspartate and glutamate; (2) basic—lysine, arginine, histidine; (3) non-polar—alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan; and (4) uncharged polar— glycine, asparagine, glutamine, cysteine, serine threonine, and tyrosine. Phenylalanine, tryptophan, and tyrosine are sometimes classified as aromatic amino acids. For example, it is reasonably predictable that an isolated replacement of leucine with isoleucine or valine, an aspartate with a glutamate, a threonine with a serine, or a similar conservative replacement of an amino acid with a structurally related amino acid, will not have a major effect on the biological activity. For example, the polypeptide of interest may include up to about 5-10 conservative or non-conservative amino acid substitutions, or even up to about 15-25 conservative or non-conservative amino acid substitutions, or any integer between 5-25, so long as the desired function of the molecule remains intact. One of skill in the art may readily determine regions of the molecule of interest that can tolerate change by reference to Hopp / Woods and Kyte-Doolittle plots, well known in the art. ZIMMER-PETN (03018-02) / / 1036.375WO1 “Gene transfer” or “gene delivery” refers to methods or systems for reliably inserting DNA or RNA of interest into a host cell. Such methods can result in transient expression of non-integrated transferred DNA, extrachromosomal replication and expression of transferred replicons (e.g., episomes), or integration of transferred genetic material into the genomic DNA of host cells. Gene delivery expression vectors include, but are not limited to, vectors derived from bacterial plasmid vectors, viral vectors, non-viral vectors, adenoviruses, retroviruses, alphaviruses, pox viruses, and vaccinia viruses. The term “derived from” is used herein to identify the original source of a molecule but is not meant to limit the method by which the molecule is made which can be, for example, by chemical synthesis or recombinant means. A polynucleotide “derived from” a designated sequence refers to a polynucleotide sequence which comprises a contiguous sequence of approximately at least about 6 nucleotides, such as at least about 8 nucleotides, including at least about 10-12 nucleotides, or at least about 15-20 nucleotides corresponding, i.e., identical or complementary to, a region of the designated nucleotide sequence. The derived polynucleotide will not necessarily be derived physically from the nucleotide sequence of interest, but may be generated in any manner, including, but not limited to, chemical synthesis, replication, reverse transcription or transcription, which is based on the information provided by the sequence of bases in the region(s) from which the polynucleotide is derived. As such, it may represent either a sense or an antisense orientation of the original polynucleotide. Modification of Cellulose Phosphoethanolamine cellulose, or other modifications of cellulose, can be prepared in any suitable manner (e.g., biosynthetically, purification from cell culture, or chemical synthesis, etc.). In one embodiment, phosphoethanolamine cellulose, or other modifications of cellulose, is produced biosynthetically by expression of BcsG phosphoethanolamine transferase, or with another catalytic domain, in a cellulose producing host, wherein the expressed BcsG catalyzes, for example, phosphoethanolamine transfer to hydroxyl groups on cellulose produced by the host cell. In some embodiments, BcsG is co-expressed with BcsA (NTD and / or CTD) and BcsB to increase amount of cellulose production in the host. The modified cellulose, such as the phosphoethanolamine cellulose, produced by the methods described herein, can be recovered from host cells and further purified if desired. Suitable hosts for production of modified cellulose include bacteria, plants, and algae, or any other type of organism or cell capable of producing cellulose. In some ZIMMER-PETN (03018-02) / / 1036.375WO1 embodiments, modified cellulose is produced biosynthetically by, for example, Gram-negative bacteria, such as, but not limited to, bacteria of the Acetobacter (e.g., Acetobacter xylinum), Agrobacterium, Escherichia (e.g., Escherichia coli), or Salmonella (e.g., Salmonella enterica) genus. Any BcsG phosphoethanolamine transferase from any bacterial species, or a biologically active fragment, variant, analog, or derivative thereof that retains BcsG phosphoethanolamine transferase activity (i.e., catalyzes transfer of a phosphoethanolamine group to a cellulose hydroxyl group) may be used to produce a phosphoethanolamine-modified cellulose. The BcsG phosphoethanolamine transferase need not be physically derived from an organism but may be synthetically or recombinantly produced. Representative sequences are presented for BcsG from Escherichia coli (SEQ ID NO:11), and additional representative sequences are listed in the National Center for Biotechnology Information (NCBI) database. See, for example, NCBI entries: WP_282597949.1, WP_005045901.1, WP_407959879.1, WP_406021287.1, WP_323914166.1, WP_406021285.1, WP_337783804.1, WP_331853836.1, WP_331853835.1, WP_031491801.1, WP_305730834.1, WP_127674701.1, WP_117150469.1, WP_060082415.1, CAK1349375.1, CAK1207332.1, CAK1210766.1, CAK0732961.1, CAK0722996.1, CAK0724424.1, CAK0718659.1, CAK0713585.1, CAK0712198.1, CAK0693424.1, CAK0678636.1, ASF66530.1, KAB1489769.1, UZW42856.1, XPU99516.1, XPV03915.1, XPE34464.1, XPE12755.1, XNR53429.1, XNR44410.1 and XIR93626.1; all of which sequences (as entered by the date of filing of this application) are herein incorporated by reference. Any of these sequences or a variant thereof comprising a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto, can be used for cellulose modification, as described herein, wherein the variant retains biological activity, such as BcsG phosphoethanolamine transferase activity. The BcsG phosphoethanolamine transferase (or chimeric in which the catalytic domain of BscG is exchanged for another catalytic domain that may result in different modifications of cellulose), alone or in combination with the BcsA (or at least the NTD and / or CTD of BcsA) and BcsB-encoded proteins, can be used. Nucleic acids comprising the BcsG, BcsA (or portion thereof, such as NTD and / or CTD), or BcsB genes can be inserted into an expression vector to create an expression cassette capable of producing the encoded proteins in a suitable host cell. Numerous vectors are known in the art including, but not limited to, linear polynucleotides, ZIMMER-PETN (03018-02) / / 1036.375WO1 polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. Thus, the term “vector” includes an autonomously replicating plasmid or a virus. For purposes of this application, the terms “expression construct,” “expression vector,” and “vector,” are used interchangeably to demonstrate the application of the invention in a general, illustrative sense, and are not intended to limit the invention. The BcsG, BcsA (or portion thereof, such as NTD and / or CTD), or BcsB genes may be provided by a single vector or separate vectors. In one embodiment, the vector comprises a bcsABG operon. In certain embodiments, the nucleic acid encoding a polypeptide of interest (e.g., BcsG, BcsA (or portion thereof, such as NTD and / or CTD), or BcsB-encoded polypeptide) is under transcriptional control of a promoter. A “promoter” refers to a DNA sequence recognized by the synthetic machinery of the cell, or introduced synthetic machinery, required to initiate the specific transcription of a gene. The term promoter will be used here to refer to a group of transcriptional control modules that are clustered around the initiation site for a bacterial RNA polymerase or eukaryotic RNA polymerase (e.g., RNA polymerase I, II, or III). Typical promoters for bacterial expression include the Tac, RecA, LacZ, pBAD, OXB1-20, OXB1, ctc, gsiB, Pspv, and T7 promoters (see, e.g., Goldstein et al. (1995) Biotechnol. Annu. Rev.1:105- 128). Examples of promoters for expression in plants include the CaMV 35S, Xa27, FMV, opine promoters, plant ubiquitin promoter (Ubi), rice actin 1 promoter (Act-1), maize alcohol dehydrogenase 1 promoter (Adh-1), and various other plant pathogen, synthetic, and native promoters (see, e.g., Liu et al. (2016) Curr. Opin. Biotechnol.37:36-44, Dey et al. (2015) Planta 242(5):1077-1094, Jeong et al. (2015) J. Integr. Plant Biol.57(11):913-924, Hernandez-Garcia et al. (2014) Plant Sci. 217-218:109-119). These and other promoters can be obtained from commercially available vectors, using techniques well known in the art. See, e.g., Sambrook et al., supra. Enhancer elements may be used in association with a promoter to increase expression levels of the constructs. An expression vector for expressing BcsG, BscA (or portion thereof, such as NTD and / or CTD) or BcsB comprises a promoter “operably linked” to a polynucleotide comprising a BcsG, BscA (or portion thereof, such as NTD and / or CTD) or BcsB gene sequence. The phrase “operably linked” or “under transcriptional control” as used herein means that the promoter is in the correct location and orientation in relation to a polynucleotide to control the initiation of transcription by RNA polymerase and expression of the polynucleotide. Typically, transcription terminator / polyadenylation signals may also be present in the expression construct. Bacterial terminator sequences may include Rho-independent or Rho- dependent transcription terminator sequences. Examples of eukaryotic terminator sequences ZIMMER-PETN (03018-02) / / 1036.375WO1 include, but are not limited to, those derived from SV40, as described in Sambrook et al., supra, bovine growth hormone terminator sequence (see, e.g., U.S. Pat. No. 5,122,458), and plant terminator sequences such as the Agrobacterium nopaline synthase (NOS) terminator (see, e.g., International Patent Application Publication No. WO 2013 / 012729, Chung et al. (2005) Trends Plant Sci. 10(8):357-361). Additionally, 5′- UTR sequences can be placed adjacent to the coding sequence in order to enhance expression of the same. Such sequences may include UTRs comprising an internal ribosome entry site (IRES). Inclusion of an IRES permits the translation of one or more open reading frames from a vector. The IRES element attracts a eukaryotic ribosomal translation initiation complex and promotes translation initiation. See, e.g., Kaufman et al., Nuc. Acids Res. (1991) 19:4485-4490; Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229:295-298; Rees et al., BioTechniques (1996) 20:102-110; Kobayashi et al., BioTechniques (1996) 21:399-402; and Mosser et al., BioTechniques (199722150-161. A multitude of IRES sequences are known and include sequences derived from a wide variety of viruses, such as from leader sequences of picornaviruses such as the encephalomyocarditis virus (EMCV) UTR (fang et al. J. Virol. (1989) 63:1651-1660), the polio leader sequence, the hepatitis A virus leader, the hepatitis C virus IRES, human rhinovirus type 2 IRES (Dobrikova et al., Proc. Natl. Acad. Sci. (2003) 100(25):15125-15130), an IRES element from the foot and mouth disease virus (Ramesh et al., Nucl. Acid Res. (1996) 24:2697- 2700), a giardiavirus IRES (Garlapati et al., J. Biol. Chem. (2004) 279(5):3389-3397), and the like. A variety of nonviral IRES sequences will also find use herein, including, but not limited to IRES sequences from yeast, as well as the human angiotensin II type 1 receptor IRES (Martin et al., Mol. Cell Endocrinol. (2003) 212:51-61), fibroblast growth factor IRESs (FGF- 1 IRES and FGF-2 IRES, Martineau et al. (2004) Mol. Cell. Biol.24(17):7622-7635), vascular endothelial growth factor IRES (Baranick et al. (2008) Proc. Natl. Acad. Sci. U.S.A. 105(12):4733-4738, Stein et al. (1998) Mol. Cell. Biol.18(6):3112-3119, Bert et al. (2006) RNA 12(6):1074-1083), and insulin-like growth factor 2 IRES (Pedersen et al. (2002) Biochem. J.363(Pt 1):37-44). These elements are readily commercially available in plasmids sold, e.g., by Clontech (Mountain View, Calif.), Invivogen (San Diego, Calif.), Addgene (Cambridge, Mass.) and GeneCopoeia (Rockville, Md.). See also IRESite: The database of experimentally verified IRES structures (iresite.org). An IRES sequence may be included in a vector, for example, to express BcsG in combination with BcsE and BcsF from an expression cassette. Alternatively, a polynucleotide encoding a viral T2A peptide can be used to allow production of multiple protein products (e.g., BcsA, in combination with BscA (or portion ZIMMER-PETN (03018-02) / / 1036.375WO1 thereof, such as NTD and / or CTD), BcsB) from a single vector.2A linker peptides are inserted between the coding sequences in the multicistronic construct. The 2A peptide, which is self- cleaving, allows co-expressed proteins from the multicistronic construct to be produced at equimolar levels.2A peptides from various viruses may be used, including, but not limited to 2A peptides derived from the foot-and-mouth disease virus, equine rhinitis A virus, Thosea asigna virus and porcine teschovirus-1. See, e.g., Kim et al. (2011) PLoS One 6(4):e18556, Trichas et al. (2008) BMC Biol.6:40, Provost et al. (2007) Genesis 45(10):625-629, Furler et al. (2001) Gene Ther.8(11):864-873; herein incorporated by reference in their entireties. One of skill in the art can readily determine BcsG, BscA (or portion thereof, such as NTD and / or CTD) and BcsB nucleotide sequences using standard methodology and the teachings herein. Oligonucleotide probes can be devised based on the known sequences and used to probe genomic or cDNA libraries. The sequences can then be further isolated using standard techniques and, e.g., restriction enzymes employed to truncate the gene at desired portions of the full-length sequence. Similarly, sequences of interest can be isolated directly from cells containing the same, using known techniques, such as phenol extraction and the sequence further manipulated to produce the desired truncations. See, e.g., Sambrook et al., supra, for a description of techniques used to obtain and isolate DNA. The BcsG, BscA (or portion thereof, such as NTD and / or CTD) and BcsB sequences can also be produced synthetically, for example, based on their known sequences. The nucleotide sequence can be designed with the appropriate codons for the particular amino acid sequence desired. The complete sequence is generally assembled from overlapping oligonucleotides prepared by standard methods and assembled into a complete coding sequence. See, e.g., Edge (1981) Nature 292:756; Nambair et al. (1984) Science 223:1299; Jay et al. (1984) J. Biol. Chem.259:6311; Stemmer et al. (1995) Gene 164:49-53. Once coding sequences have been isolated and / or synthesized, they can be cloned into any suitable vector or replicon for expression. Numerous expression vectors are known to those of skill in the art, and the selection of an appropriate expression vector is a matter of choice. For example, a bacterial plasmid expression vector may be used to transform a bacterial host. Bacterial expression vectors include, but are not limited to, pACYC177, pASK75, pBAD, pBADM, pBAT, pCal, pET, pETM, pGAT, pGEX, pHAT, pKK223, pMal, pProEx, pQE, and pZA31 vectors. See, e.g., Sambrook et al., supra. Alternatively, plant expression systems can also be used to produce modified cellulose as described herein. Generally, such systems use virus-based vectors to transfect plant cells with heterologous genes. Exemplary plant viruses include the tobacco mosaic virus (TMV), ZIMMER-PETN (03018-02) / / 1036.375WO1 potato virus X, and cowpea mosaic virus. A number of plant expression systems use the Ti plasmid of Agrobacterium tumefaciens. For a description of plant expression systems, see, e.g., Zaidi et al. (2017) Front. Plant Sci. 8:539; Hefferon (2014) Biomed. Res. Int. 2014:785382; Porta et al. (1996) Mol. Biotech.5:209-221; and Hackland et al. (1994) Arch. Virol.139:1-22. In addition, algae expression systems are available for Chlamydomonas reinhardtii and Synechococcus elongatus. See, e.g., Doron et al. (2016) Front. Plant Sci.7:505 and Griesbeck et al. (2006) Mol. Biotechnol.34(2):213-223. A gene can be placed under the control of a promoter, ribosome binding site (for bacterial expression) and, optionally, an operator (collectively referred to herein as “control” elements), so that the DNA sequence encoding the desired polypeptide is transcribed into RNA in the host cell transformed by a vector containing this expression construction. The coding sequence may or may not contain a signal peptide or leader sequence. With the present invention, both the naturally occurring signal peptides and heterologous sequences can be used. Leader sequences can be removed by the host in post-translational processing. See, e.g., U.S. Pat. Nos.4,431,739; 4,425,437; 4,338,397. Such sequences include, but are not limited to, the TPA leader, as well as the honeybee mellitin signal sequence. Other regulatory sequences may also be desirable which allow for regulation of expression of the protein sequences relative to the growth of the host cell. Such regulatory sequences are known to those of skill in the art, and examples include those which cause the expression of a gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. Other types of regulatory elements may also be present in the vector, for example, enhancer sequences. The control sequences and other regulatory sequences may be ligated to the coding sequence prior to insertion into a vector. Alternatively, the coding sequence can be cloned directly into an expression vector that already contains the control sequences and an appropriate restriction site. In some cases, it may be necessary to modify the coding sequence so that it may be attached to the control sequences with the appropriate orientation, i.e., to maintain the proper reading frame. Mutants or analogs may be prepared by the deletion of a portion of the sequence encoding the protein, by insertion of a sequence, and / or by substitution of one or more nucleotides within the sequence. Techniques for modifying nucleotide sequences, such as site- directed mutagenesis, are well known to those skilled in the art. See, e.g., Sambrook et al., supra; DNA Cloning, Vols. I and II, supra; Nucleic Acid Hybridization, supra. The expression vector is then used to transform an appropriate cellulose-producing host cell. Depending on the expression system and host selected, the modified cellulose is produced ZIMMER-PETN (03018-02) / / 1036.375WO1 by growing host cells transformed by an expression vector described above under conditions whereby the BcsG phosphoethanolamine transferase is expressed with BscA (or portion thereof, such as NTD and / or CTD) and optionally BcsB The BcsG phosphoethanolamine transferase catalyzes phosphoethanolamine transfer to hydroxyl groups (or other group if catalytic domain has been exchanged with another cellulose modifying enzyme) of the cellulose produce by the host. The selection of the appropriate growth conditions is within the skill of the art. Phosphoethanolamine cellulose can be produced in bacteria, for example, by culturing bacteria in media containing a suitable carbon source. Exemplary carbon sources include monosaccharides (e.g., glucose and fructose), disaccharides (e.g., sucrose, maltose, and lactose), oligosaccharides, polysaccharides (e.g., starch hydrolysates), mannitol, ethanol, acetic acid, citric acid, glycerol, beet molasses (B-Mol), and biodiesel fuel by-product (BDF-B). One or more carbon sources may be used. Media can be supplied manually or automatically with a continuous, batch, or semi-batch fed culture system. Uses of Modified Celluloses Phosphoethanolamine cellulose is used in a wide variety of industrial, nutritional, electronic, scientific, and medical applications. For example, phosphoethanolamine cellulose can be used in various applications in which other forms of cellulose are currently used, such as in production of paper, textile, biofuels, food, pharmaceutical fillers, cellulose composites for electronic devices, nanocellulosic materials, and liquid filtration and chromatographic media. The phosphoethanolamine group enhances solubility of the cellulose and facilitates the conversion of the polymer to shorter polysaccharides and oligosaccharides or monosaccharides such as glucose through physical, chemical (e.g., acid hydrolysis), or enzymatic (e.g., cellulase catalyzed hydrolysis) methods. For example, phosphoethanolamine cellulose can be hydrolyzed by contacting the phosphoethanolamine cellulose with one or more cellulases. In certain embodiments, one or more endocellulases, exocellulases, beta-glucosidases, oxidative cellulases, cellulose phosphorylases, or a combination thereof are used in hydrolysis of phosphoethanolamine cellulose. The enhanced conversion of phosphoethanolamine cellulose to glucose provides an attractive route for production of ethanol from cellulosic biomass either in bacteria (e.g., Escherichia coli, Acetobacter xylinum, or other bacterial cellulose producer) or through bioengineering of other organisms to express BcsG (e.g., either by itself or in combination with the bcsE and bcsF genes), for example, in Miscanthus or other plants or algae. ZIMMER-PETN (03018-02) / / 1036.375WO1 Phosphoethanolamine (pEtN) modification of cellulose can endow biofilms with enhanced properties, such as the addition of reactive groups to cellulose that can be further modified. The phosphoethanolamine cellulose may also be further modified due to its reactive amine group functionality to generate a wide number of other cellulosic materials. For example, the amine group is readily alkylated, acylated, or sulfonated. In particular, the amine group can be conjugated to various agents such as peptides, antibodies, enzymes, nucleic acids, dyes, ligands, or drugs. Methods for conjugating amines are well known in the art. For example, conjugation may be performed with amine-reactive succinimidyl esters or click chemistry. For a description of various conjugation techniques, see, e.g., Bioconjugation Protocols: Strategies and Methods (S.S. Mark ed., Humana Press, 2016), G. T. Hermanson Bioconjugate Techniques (Academic Press, 3rd edition, 2013), Click Chemistry for Biotechnology and Materials Science (J. Lahann ed., Wiley, 2009); herein incorporated by reference in their entireties. In particular, phosphoethanolamine cellulose can be chemically modified to produce useful cellulose ester and ether derivatives. For example, cellulose ester derivatives can be formed by esterification of cellulose hydroxyl groups with an organic acid, acid anhydride, or acid chloride, or an inorganic acid. Exemplary organic acids that can be used in esterification include acetic acid, propanoic acid, and butyric acid. Alternatively, the corresponding acid anhydrides (e.g., acetic anhydride, propionic anhydride, and butyric anhydride) or acid chlorides (e.g., acetyl chloride) can be used. Exemplary inorganic acids that can be used in esterification include nitric acid and sulfuric acid. For a description of methods of synthesizing cellulose esters, see, e.g., Edgar et al. (2001) Progress in Polymer Science 26:1605-1688, Cao et al. (2013) J. Agric. Food Chem.61:2489-2495, Heinze et al. (2003) Cellulose 10:283-296, Krassig (1993) Cellulose (Polymer Monographs) Volume 11, CRC Press, Liebert et al. (2005) Biomacromolecules 6:333-340, El-Sakhawy et al. (2014) J. Drug Deliv. 2014:575969, U.S. Pat. No.9,624,311, U.S. Pat. No.9,458,248, U.S. Pat. No.9,217,043, U.S. Pat. No.9,708,415, U.S. Pat. No. 8,273,872, U.S. Pat. No. 6,184,373, U.S. Pat. No. 5,750,677, U.S. Pat. No. 2,651,629, and U.S. Pat. No.3,097,051; herein incorporated by reference. Cellulose esters are commonly used, for example, as binders, coating additives, and film formers or modifiers, and may find use in automotive, wood, plastic, paper, apparel, photography, and leather coatings applications. Alternatively, the cellulose hydroxyl groups can be chemically modified to produce a cellulose ether, such as an alkyl ether (e.g., methylcellulose, ethylcellulose), hydroxyalkyl ether (e.g., hydroxyethylcellulose, hydroxylpropyl cellulose), or carboxyalkyl ether (e.g., carboxymethylcellulose). Cellulose ethers are commonly prepared using an alkali metal ZIMMER-PETN (03018-02) / / 1036.375WO1 hydroxide to deprotonate cellulose hydroxyl groups, which are reacted with an etherifying agent such as an alkyl halide, alkyl sulfate, alkylene oxide, or chlorohydrin. For a description of methods of synthesizing cellulose ethers, see, e.g., Kristin Schumann et al. (2009) Macromolecular Symposia 280:86-94, Goncalves et al. (2015) Carbohydrate Polymers 116:51- 59, Lorand (1939) Ind. Eng. Chem. 31:891-897, U.S. Pat. No. 2,512,338, U.S. Pat. No. 8,541,571, and U.S. Pat. No.9,580,516; herein incorporated by reference. Cellulose ethers are commonly used, for example, as thickeners, binders, film formers, water-retention agents, suspension aids, surfactants, lubricants, and protective colloids and emulsifiers, and may find use in construction, ceramics, paints, foods, cosmetics, and pharmaceuticals. Cellulose esters and ethers may find use in pharmaceuticals for sustained and controlled release formulations, osmotic drug delivery systems, bioadhesives and mucoadhesives, compressibility enhancers in tablets, liquid dosage forms as thickening agents and stabilizers, binders in tablets, semisolid preparations as gelling agents, and various other applications. The disclosure can be better understood by reference to the following examples which are offered by way of illustration. The disclosure is not limited to the examples given herein. EXAMPLES Example I Introduction Bacterial biofilms have major impacts on our healthcare, industrial, and infrastructure systems. In a biofilm community, bacteria are encased in and protected by a complex biopolymer meshwork including protein polymers, polysaccharides, and nucleic acids (1,2). Cellulose is a common bacterial biofilm component produced by a variety of prokaryotes, such as Escherichia coli (Ec), Salmonella enterica, and Gluconacetobacter / Komagataeibacter xylinus (Gx) (3). These Gram-negative bacteria synthesize cellulose and secrete it across the cell envelope via the Bacterial cellulose synthase (Bcs) complex. The Bcs complex consists of components located in the inner and the outer membrane (IM and OM, respectively), as well as the periplasm (4). Its core machinery contains the catalytically active cellulose synthase BcsA enzyme that forms a functional complex with the periplasmic but IM-anchored BcsB subunit (5,6). A third component, BcsC, likely creates a cellulose channel in the OM (7,8). In the periplasm, the conserved cellulase BcsZ has been shown to play a role in cellulose production in vivo (9-12). BcsA is a processive family-2 glycosyltransferase (GT) that synthesizes cellulose, a linear glucose polymer, from UDP-activated glucosyl units. Additionally, the enzyme also ZIMMER-PETN (03018-02) / / 1036.375WO1 secretes the nascent cellulose chain across the IM through a channel formed by its own transmembrane (TM) segment (5). Thereby, cellulose elongation is directly coupled to membrane translocation. The periplasmic BcsB subunit is needed for BcsA’s catalytic activity (6). It contains two structural repeats consisting of a region resembling a carbohydrate-binding domain (CBD) that is fused to a flavodoxin-like domain (FD) (5). Lastly, the C-terminal segment of BcsC forms a 16-stranded β-barrel in the OM (8). This pore is preceded by 19 predicted periplasmic tetratricopeptide repeats (TPR), which consist of a pair of anti-parallel α-helices connected by a short loop. BcsC’s TPRs likely form an α-helical solenoid structure bridging the periplasm (13). E. coli and other Enterobacteriaceae can modify cellulose with lipid-derived phospho- ethanolamine (pEtN) at the sugar’s C6 position (1, 14). This modification promotes biofilm cohesion, is needed for hallmark biofilm macrocolony wrinkling phenotypes, and enhances cell-association of curli amyloid fibers and adhesion to host tissues (14-17). Transfer of pEtN is catalyzed by the membrane-bound pEtN transferase BcsG (15). Recent cryogenic electron microscopy (cryo-EM) studies of the Ec Bcs complex provided the first insights into the organization of the pEtN cellulose biosynthesis complex, Fig.1a and b (18-20). A single BcsA subunit associates with six BcsB copies that form a semicircle at the periplasmic water-lipid interface. BcsA sits at one end of the semicircle where it steers the secreted cellulose polymer towards the circle’s center. The open half of the semicircle has been proposed to accommodate multiple BcsG copies (20). Provided herein are aspects of pEtN cellulose formation and translocation across the periplasm and the OM. It is demonstrated that BcsA recruits three copies of BcsG via its N- and C-terminal domains. These domains are sufficient for BcsG binding, allowing the introduction of the pEtN modification in orthologous cellulose biosynthetic systems wherein BcsG modifies cellulose using pEtN groups derived from lipids. Further, it is shown that the OM BcsC subunit binds the sixth subunit of the BcsB semicircle via its extreme N-terminus to establish an envelope-spanning Bcs complex. The cryo-EM structure of BcsC reveals that its periplasmic domain is flexibly attached to the β-barrel. Further, BcsC binds a putative cello- oligosaccharide at its TPR solenoid, likely to facilitate translocation across the periplasm. Lastly, in vivo cellulose secretion assays reveal that the cellulase activity of BcsZ is needed for cellulose secretion. While cellulose export is dramatically reduced in the absence of BcsZ, the conserved enzyme can be functionally replaced with an off-the-shelf cellulase. The data demonstrate periplasmic translocation of cellulose along the BcsC solenoid and trimming of mis-localized polymers by BcsZ to restore translocation to the cell surface. ZIMMER-PETN (03018-02) / / 1036.375WO1 Materials and Methods Construct design and mutagenesis BcsA-NTD – The N-terminal region of BcsA (NTD; residues 1- 94; as shown on FIG. 8; SEQ ID NO: 9) was amplified from the existing pETDuet_Ec_Bcs_A-12His_nSS-Strep- B_AdrA-6His plasmid (pETDuet_Ec_Bcs_AB-AdrA) (20) with a primer-encoded C-terminal Strep-tag II and cloned into a pETDuet-1 vector using NcoI and XhoI restriction sites. BcsB has its native signal sequence (nSS) in this plasmid. The BcsG gene was amplified from an existing pACYCDuet_Ec_Bcs_PelB-8His-C-FLAG_Z_F_G-FLAG construct (pACYCDuet_Ec_Bcs_CZFG) (20) with a C-terminal dodeca-histidine tag and cloned into the vector pACYCDuet-1 using NcoI and XhoI restriction sites. BcsA chimera – The Rs-Ec (R / E) BcsA chimera was engineered by fusing the N- terminus of Ec-BcsA (1-149; as shown in FIG.8; SEQ ID NO: 13) with Rs-BcsA (17-728, as depicted in FIG. 8) and replacing the Rs-BcsA C-terminus (729-788, as depicted in FIG. 8) with the Ec-BcsA extended C-terminus (828-872, as shown in FIG.8; SEQ ID: 10). The BcsA chimera was cloned into the pETDuet-1 vector using NcoI and HindIII restriction sites, which inserted an additional Gly residue after the N-terminal Met. Next, Rs-BcsB was inserted into the second cloning site of the pETDuet-1 vector using NdeI and KpnI restriction sites, generating plasmid pETDuet_Rs_Ec_Bcs_A-12His_Rs_Bcs_B. AdrA – For the purification of in vivo synthesized cellulose, c-di-GMP generating enzyme, AdrA was amplified from the pETDuet_Ec_Bcs_ABAdrA (20) and inserted into the empty pCDFDuet-1 vector using restriction enzyme digestion and ligation cloning (pCDFDuet_AdrA). BcsB-BcsC fusion – For generating the BcsB-BcsC fusion construct, the pETDuet_Ec_Bcs_AB-AdrA plasmid was used. The N-terminal four TPRs of BcsC (residues 21-179) followed by a GSGSGSG linker were inserted by polymerase incomplete primer extension (PIPE) cloning after the N-terminal signal sequence and Strep-tag II of BcsB, followed by residues 56-779 of BcsB. This generated plasmid pETDuet_Ec_Bcs_A- 12His_nSS-Strep-TPR1-4_(GS)3G_B_AdrA. BcsC – BcsC ‘s periplasmic domain (TPR #1-18; residues 24-709) was amplified from the existing full length BcsC plasmid (8) and cloned into a pET20b vector between NcoI and XhoI restriction sites. The full-length construct contained an N-terminal and the periplasmic domain a C-terminal deca His-tag respectively. BcsZ –Mutagenesis of BcsZ was done in pACYCDuet_Ec_Bcs_C_ZFG construct using the QuikChange approach. Deletion of BcsZ was also done using the same ZIMMER-PETN (03018-02) / / 1036.375WO1 pACYCDuet_Ec_Bcs_CZFG construct by PIPE cloning. To replace Ec BcsZ with other cellulases, the Gx BcsZ (CMCax) and C. cellulolyticum Cel9M genes were synthesized with the DsbA and wild type Ec BcsZ signal sequences, respectively. The genes were inserted into the previously described pACYCDuet_Ec_Bcs_CZFG vector using PIPE cloning. This resulted in the generation of pACYCDuet_Ec_Bcs_ C_SS-Cel9M / CMCAx_FG plasmid respectively. BcsG – The catalytically inactive S278A mutant of BcsG was also generated by QuikChange mutagenesis using the pACYCDuet_Ec_Bcs_CZFG construct. Protein Expression and Purification The Strep-tagged BcsA NTD and His-tagged BcsG were co-expressed in Ec C43 (DE3) cells in Terrific Broth-M-80155 media (TB-AD) (20) containing 100 μg / mL ampicillin and 35 μg / mL chloramphenicol. 11 liter of bacterial cell cultures were grown at 37 ºC until the cell density reached OD600of 0.8 at which point the temperature was lowered to 20 ºC and growth was continued for another 18 h. Cells were harvested by centrifugation for 20 min at 5,000 rpm and 4 ºC. Pelleted cells were resuspended in ice-cold Buffer A (25 mM Tris pH 8.0, 300 mM NaCl, 5% glycerol) containing 100 mM PMSF and 1x protein inhibitor cocktail (0.8 mM Aprotinin, 5 mM E64, 10 mM Leupeptin, 15 mM Bestatin-HCl, 100 mM AEBSF-HCl, 2 mM Benzamidine-HCl and 2.9 mM Pepstatin A). The cells were lysed using a gas powered microfluidizer (25 kpsi, 3 passes) and the lysates were spun at 12,500 rpm for 20 min at 4 ºC to remove cell debris. The membrane containing supernatant was collected and subjected to ultracentrifugation at 138,000 x g for 2 hours at 4 ºC. Membrane pellets were flash frozen in liquid nitrogen and stored in -80 ºC until used. Tandem purification of BcsA NTD and BcsG To test the interaction between BcsA NTD and BcsG, a tandem affinity purification (TAP) scheme was followed. To this end, membranes were first solubilized in Buffer A containing Detergents A (1% lauryl maltose neopentyl glycol (LMNG, Anatrace), 0.2% decyl maltose neopentyl glycol (DMNG, Anatrace) and 0.2% cholesteryl hemisuccinate (CHS, Anatrace)), 40 mM imidazole, 100 mM PMSF and 1x protein inhibitor cocktail. After incubation for 1 hour at 4 ºC with mild agitation, non-solubilized material was removed by centrifugation at 138,000 x g for 40 min at 4 ºC. During this time, 7 mL of His-Pur Ni-NTA resin (Thermo Scientific) was equilibrated in Buffer A containing 40 mM imidazole. The membrane extract was applied to these pre-equilibrated beads and allowed to gently rock for 1 hour at 4 ºC. The resin was transferred to a gravity flow column and washed two times with Buffer A containing Detergents B (0.01% LMNG, 0.002% DMNG, 0.002% CHS) ZIMMER-PETN (03018-02) / / 1036.375WO1 supplemented with 40- and 50-mM imidazole respectively. The protein was eluted with 400 mM imidazole and the eluent was immediately diluted with Buffer A containing Detergents B to dilute the imidazole to 200 mM. During this Ni-NTA chromatography, Strep-Tactin resin (IBA) was also equilibrated with Buffer A containing Detergents B. The diluted eluent was passed over the Strep-Tactin resin twice at room temperature, followed by 5 column volumes of wash with Buffer A containing Detergents B and eluted with Buffer A containing 3 mM desthiobiotin and Detergents B. The eluent was concentrated and loaded onto Superose 6 increase 10 / 300 GL column (GE Healthcare) equilibrated with Buffer B (25 mM Tris pH 8.0, 100 mM NaCl) containing Detergents C (0.003% LMNG, 0.0006% DMNG, 0.0006% CHS). Initially the gel filtration fractions were run on SDS-PAGE for Coomassie staining and finally the interactions between these two components were confirmed by Western blotting and tandem mass spectrometry fingerprinting. BcsA chimera - To express the BcsA chimera protein, freshly transformed Ec C43 (DE3) cells containing the pETDuet_Rs_Ec_Bcs_A-12His_Rs_Bcs_B along with the pACYCDuet_Ec_Bcs_FG-FLAG construct (pACYCDuet_Ec_Bcs_FG) (20) were grown in 4- times 1 L of TB-AD media using the above-described protocol. After harvesting the cells, pellets were resuspended to a final volume of 250 mL in Buffer C (25 mM HEPES pH 8.0, 300 mM NaCl, 5 mM cellobiose, 5% glycerol and 5 mM MgCl2) using a glass dounce homogenizer and lysed using a microfluidizer (25 kpsi, 3 passes) in the presence of 100 mM PMSF and protein inhibitor cocktail (as described above). After pre-clearing the lysates by a low-speed centrifugation step, membranes were collected and stored as described above. Membranes were resuspended in solubilization buffer containing Buffer C, mix of Detergents (Detergents A), protease inhibitor cocktail, and 100 mM PMSF for 1 hour at 4 ºC on a rotating shaker. Insoluble material was removed by ultracentrifugation at 138,000 x g for 40 min and the supernatant was incubated with gentle rocking for 1 hour at 4 ºC with 5 mL of His-Pur Ni-NTA resin equilibrated in buffer C containing 40 mM imidazole. After batch binding, the resin was packed into a gravity flow column and washed with Buffer C containing Detergents B and 40 mM imidazole. Following this, three more washes were done with Buffer C containing Detergents B supplemented with 50 mM imidazole, 700 mM NaCl or 55 mM imidazole, respectively. Protein was eluted with Buffer C containing Detergents B and 400 mM imidazole. The protein eluent was concentrated to 500 µL and subjected to size-exclusion chromatography on a Superdex 200 increase 10 / 300 GL column (GE Healthcare) equilibrated with Buffer D (25 mM HEPES pH 8.0, 150 mM NaCl, 5 mM MgCl2, 0.5 mM cellobiose) containing Detergents C. The sample quality was evaluated by peak shape, SDS-PAGE and ZIMMER-PETN (03018-02) / / 1036.375WO1 negative stain EM. The same protocol was used to co-express and purify BcsG with the wild type Rs BcsAB (pETDuet_Rs_Bcs_A-12His_B, (6)) complex. For Congo red binding and fluorescence assays - For CR assays, the Ec Bcs total membrane complex (TMC; including inner and outer membrane components) constituting the functional cellulose synthase machinery was expressed using all three plasmids (pETDuet_Ec_Bcs_AB-AdrA, pACYCDuet_Ec_Bcs_CZFG and pCDFDuet_Ec_Bcs_R_Q_E-HA and the Ec Bcs_inner membrane complex (IMC) was expressed using the pACYCDuet_Ec_Bcs_FG construct instead of pACYCDuet_Ec_Bcs_CZFG, along with the other two constructs (pETDuet_Ec_Bcs_AB- AdrA and pCDFDuet_Ec_Bcs_R_Q_E-HA). Different versions of BcsZ or the BcsG mutant in the pACYCDuet_Ec_Bcs_CZFG plasmid generated in this study (as described above in construct design) were also expressed like the TMC complex using the two other pET and pCDFDuet_Ec_Bcs plasmids. Bcs complex purification - The Ec Bcs complex was expressed and purified as described previously (20). Briefly, the three plasmids used to express TMC complex were transformed into Ec C43 (DE3) cells followed by expression in 4 x 1 L of TB-AD media containing the necessary antibiotics. After the membrane preparation, the complex was purified using Ni-NTA affinity chromatography followed by Strep-Tactin resin purification. The strep eluent was loaded onto a Superose 6 increase 10 / 300 GL column and the inner membrane complex (IMC) eluted in a sharp peak roughly at 13 mL elution volume. The fractions containing the Ec Bcs complex were pooled and used for reconstitution into nanodiscs at a 1:4:160 molar ratio of IMC: MSP2N2: Ec total lipid extract (solubilized in 100 mM sodium cholate). Gel filtration buffer (50 mM HEPES pH 8.0, 150 mM NaCl, 5 mM MgCl2, 0.5 mM cellobiose containing Detergents C), lipid, sodium cholate (15 mM final concentration) and detergent solubilized IMC was combined and incubated for 1h at 4 ºC with gentle rocking to form the mixed micelles. MSP2N2 was added and incubated for another 30 min at 4 ºC. BioBeads were added stepwise (three times) in equal mass to remove detergent; first after the MSP2N2 addition, followed by a second addition after 1 hour and the third addition 12 hours later. Nanodisc-reconstituted IMC complex was purified on a Superose 6 increase 10 / 300 GL column equilibrated in gel equilibration buffer with no Detergents. The reconstituted fractions were screened using SDS- PAGE and negative-stain EM for the presence of MSP and IMC. For the Ec complex with the engineered BcsB-BcsC fusion, Ec C43 cells were co- transformed by electroporation with pETDuet_Ec_Bcs_A-12His_nSS-Strep-TPR1- 4_(GS)3G_B_AdrA along with two other constructs, namely pACYCDuet_Ec_Bcs_CZFG and ZIMMER-PETN (03018-02) / / 1036.375WO1 pCDFDuet_Ec_Bcs_R_Q_E-HA plasmids. The complex was expressed and purified as described for the wild type Ec Bcs complex without the nanodisc formation. BcsC TPR purification - To purify the periplasmic domain of BcsC, freshly transformed Ec Rosetta 2 cells were grown in autoinducing TB-AD media for 25 h at 28 ºC. The cells were harvested by centrifugation at 5,000 rpm for 20 min. The periplasmic extract was prepared in Tris / EDTA / Sucrose (TES) buffer as described (45). The protein was purified from the periplasmic extract via immobilized metal affinity chromatography on Ni-NTA agarose resin. The isolated extract was dialyzed overnight against buffer E (25 mM Tris pH=7.5 and 100 mM NaCl) and loaded onto Ni-NTA beads equilibrated with the buffer E. After batch binding for 1h, the resin was washed with Buffer E and Buffer E supplemented with 30mM imidazole, 500 mM NaCl or 40 mM imidazole. The protein was eluted with 400 mM imidazole and concentrated and applied to Superdex 200 increase 10 / 300 GL column (GE Healthcare) equilibrated in 0.2 M sodium bicarbonate buffer (pH 7.5), 500 mM NaCl. The peak fraction containing the protein was concentrated and aliquots were flash frozen for future use. BcsC purification – For full-length BcsC, membranes were prepared, and purification was carried out as described previously for the Ec BcsC2 porin construct (8) with some modifications. Briefly, BcsC membranes were solubilized in Buffer F containing 25 mM Tris pH 8.5, 300 mM NaCl, 5% glycerol, 35 mM imidazole, 30 mM LDAO (lauryl-dimethylamine N-oxide) and 3 mM DDM (dodecyl-β-D-maltopyranoside) for 1h. After removal of the insoluble aggregates by centrifugation at 42,000 rpm for 30 min, the supernatant was combined with 5mL Ni-NTA beads equilibrated in Buffer F containing no detergents and allowed to rock for 1h at 4 ºC. Resin was collected in a gravity flow column and washed with buffer F without LDAO but substituted with 1 mM DDM and 40 mM imidazole, 1M NaCl, or 50 mM imidazole. Protein was eluted using buffer G containing 25 mM Tris pH 8.5, 100 mM NaCl, 300 mM imidazole and 0.6% C8E4 (tetraethylene glycol monooctyl ether) and concentrated before loading onto a Superdex 200 increase 10 / 300 GL column (GE Healthcare) equilibrated with buffer G containing no imidazole. The fractions containing BcsC were concentrated and reconstituted into MSPE3D1 nanodiscs with Ec total lipids (solubilized in sodium cholate) at a final molar ratio of 1:4:100, as described above for the Ec Bcs complex, in the absence or presence of 5 mM cellotetraose. After detergent removal, the sample was purified over the Superdex 200 increase size exclusion column equilibrated in 25 mM Tris pH 8.5, 100 mM NaCl. Peak fractions corresponding to BcsC in MSPE3D1 were collected and the sample quality was evaluated by SDS-PAGE and negative stain EM. ZIMMER-PETN (03018-02) / / 1036.375WO1 BcsZ purification – BcsZ was expressed and purified as described previously using an existing plasmid (pET20b_SS_BcsZ_6His) (39). The protein was purified using Ni-NTA and size exclusion chromatography. Cellulose synthase enzyme assays Cellulose biosynthesis assays were performed as described previously (6). This assay measures the incorporation of UDP-[3H]-glucose into insoluble glucan chains. Briefly, 20 μL reaction was performed in the respective gel filtration buffer by incubating the enzyme in the presence of 20 mM MgCl2, 5 mM UDP-glucose (UDP-Glc), 0.25 μCi UDP-Glc[6-3H], 30 μM c-di-GMP at 30 ºC for the chimeric BcsAB-BcsG complex for 1 hour at an enzyme concentration of 0.5-1 mg / mL. For IMVs, assays were carried out in a similar manner but incubated for 16 hours at 37 ºC. Following biosynthesis, the reaction mixture was spotted onto the origin of a descending Whatman-2MM chromatography paper, which was developed with 60% ethanol. The high molecular weight polymer retained at the origin was quantified on a liquid scintillation counter (Beckman). Control reactions were set up by replacing the c-di- GMP with ddH2O. To confirm the formation of authentic cellulose, another set of reaction was set up wherein 5U of endo-(1,4)-β-glucanase (E-CELTR; Megazyme) was added at the beginning of the synthesis reaction for the enzymatic degradations of the in vitro synthesized glucan. Each condition was performed in triplicate and error bars represent deviations from the means. Cellulase assay In order to examine the cellulase activity of the BcsZ mutants as well as the new cellulases (Cel9M and CMCax), a cellulase activity assay was performed using Carboxymethyl cellulose agar plates, as described previously (39). CMC-agar plates were prepared by dissolving 2% CMC and 1.5% agar in LB medium, followed by autoclaving. The solution was cooled and supplemented with 0.5 mM isopropyl β-D-thiogalactopyranoside (IPTG) and respective antibiotics prior to pouring the plates. To test for cellulase activity of Cel9M and CMCax, these proteins were co-expressed along with the Ec Bcs TMC complex in C43 cells using the pACYCDuet_Ec_Bcs_C_SS-Cel9M / CMCax_FG plasmid instead of pACYCDuet_Ec_Bcs_CZFG plasmid and periplasmic fractions were extracted as described above for the purification of BcsC’s periplasmic domain. Negative and positive controls were performed by spotting purified bovine serum albumin (BSA) or Aspergillus niger cellulase (Sigma) onto the plates. CMC-agar plates were incubated at 37 ºC for 48 hours and stained with 2% CR solution for 1 hour at room temperature, followed by destaining in 1M NaCl for 2 hours. ZIMMER-PETN (03018-02) / / 1036.375WO1 To probe for BcsZ cellulase activity in the BcsZ mutants, the pACYCDuet_Ec_Bcs_C_Zwt / mutantsFG plasmid was transformed in C43 cells and plated on LB- agar plates containing 25 μg / mL chloramphenicol. After an overnight incubation of the plates at 37 ºC, a single colony was picked from each plate to inoculate 5 mL LB broth and grown overnight at 37 ºC. Next day, all cultures were normalized based on OD600 absorbance and 5 µL from each culture was spotted onto the CMC-agar plates. After incubating the agar plates at 37 ºC for 48 hours, colonies were removed from the plates prior to staining with CR as described above. All cellulase plate assays were performed in triplicates. Congo red binding and fluorescence assays Starter cultures of Ec complex transformed cells were grown overnight at 37 ºC in a shaking incubator in LB medium. The overnight cultures were normalized to an optical density at 600 nm (OD600) of 1 with sterile fresh LB medium and 5 µL of this diluted culture was spotted onto the LB agar plates lacking NaCl but containing 25 μg / mL CR, 250 μM IPTG and the antibiotics ampicillin, chloramphenicol and streptomycin. The agar plates were kept at room temperature for 48-56 hours and the bacterial cells on top of the agar were visualized using G:Box Chemi-XX6 (Syngene, Cambridge, UK). Images were acquired with GeneSys software (Syngene, version 1.8.5.0) based on the excitation and emission wavelength of the CR (497 / 610 nm). All experiments were performed in triplicate. Purification of in vivo synthesized pEtN cellulose For the purification of in vivo synthesized cellulose, Ec C43 cells were co-transformed with pETDuet_Rs_Ec_Bcs_A-12His_Rs_Bcs_B together with the pAYCDuet_Ec_Bcs_FG- FLAG and pCDFDuet_AdrA constructs. Periplasmic cellulose produced by the BcsA chimera was obtained through cell lysis; digestion of DNA, RNA, and protein; and precipitation and purification of polysaccharide. First, four 1-L cell cultures were grown in TB-AD media with slow shaking for 25 h at 28 ºC. Cells were collected by centrifugation at 5,000 g for 20 min, flash frozen in liquid nitrogen, and stored at -80 ºC until processed. To isolate pEtN cellulose, cells were thawed, resuspended in lysis buffer (10 mM Tris pH 7.4, 0.1 M NaCl and 0.5% SDS), sonicated briefly, and treated with lysozyme (final concentration of 5 mg / mL) with rocking at room temperature (RT) for 30 min. This suspension was then subjected to boiling with constant stirring for 1 hour and then cooled to room temperature. The suspension was then treated with DNase and RNase (each at final concentrations of 100 μg / mL) and incubated at RT for 1 hour. Trypsin and chymotrypsin were then added (each at final concentration of 50 μg / mL), followed by rocking at RT for 4 h. Lastly, Proteinase-K (final conc 100 μg / mL) was added, and this suspension was incubated overnight with moderate shaking at 60 ºC. The ZIMMER-PETN (03018-02) / / 1036.375WO1 solution was then cooled to RT and diluted with Milli-Q (MQ) water to reduce the final SDS concentration to less than 0.2%. This solution was subjected to dialysis against MQ water for 24 hours using 100 kDa cut-off cellulose ester dialysis membranes. Then the solution was subjected to one freeze-thaw cycle. After thawing, CR (25 μg / mL) and NaCl (170 mM) were added while stirring to facilitate purification and precipitation of pEtN cellulose from the largely clarified lysate. The solution was transferred to 50 mL falcon tubes and insoluble material was pelleted via centrifugation at 13,000 g for 1 hour to collect the enriched pEtN cellulose. The pellet was further washed with 4% SDS and 10 mM Tris pH 7.4 followed by brief sonication. The solution was allowed to rock at RT overnight, followed by centrifugation for 2 min at 13,000 g to pellet the cellulosic material. The final pellet was washed with MQ water 3-5 times to remove the SDS with pelleting by centrifugation following each resuspension. The final sample was frozen and lyophilized and used for NMR analysis. Solid-state NMR analysis 13C cross-polarization magic-angle spinning (CPMAS) solid-state NMR was performed at ambient temperature in an 89 mm bore 11.7 T magnet (Agilent Technologies, Danbury, CT) using an HCN Agilent probe with a DD2 console (Agilent Technologies) (46). Samples were spun at 7143 Hz in 36 μL capacity 3.2 mm zirconia rotors. CP was performed with a field strength of 50 kHz for 13C and with a 10% linearly ramped field strength centered at 57kHz for1H.1H decoupling was performed with two pulse phase modulation (TPPM) at 83 kHz (47). The experimental recycle time was 2 s.13C chemical shift referencing was performed by setting the high frequency adamantane peak to 38.5 ppm (48). The enriched pEtN cellulose sample from the Rs / Ec chimera was 3 mg and the13C CPMAS spectrum is the result of 40,960 scans. Preparation of Inverted Membrane Vesicles IMV preparation was carried out as described previously (49) either for the wild type Rs BcsAB complex alone, or the BcsA chimeric-BcsB or wild type BcsAB complex co- expressed with Ec BcsG and BcsF. Briefly, the complex along with AdrA was overexpressed in Ec C43 cells in TB media. When the cell density reached 0.8, expression was induced by addition of 0.6 mM IPTG. After 4 h of incubation at 37 ºC, cells were harvested and resuspended in buffer H containing 20 mM phosphate buffer (pH 7.5) and 100 mM NaCl. The cells were lysed in a microfluidizer, and cell debris was removed by low-speed centrifugation (12,500 rpm). The supernatant (approx. 22 ml) was layered over a 2 M sucrose cushion and centrifuged at 42,000 rpm for 2 hours. Following this, the dark brown ring formed at the sucrose interface was carefully collected (approx. 10ml) and diluted to 65 mL in Buffer H and the membrane vesicles were sedimented by centrifugation at 42,000 rpm for 90 min. The pellet ZIMMER-PETN (03018-02) / / 1036.375WO1 was then rinsed and resuspended in 1 mL of Buffer H, homogenized using a 2 mL dounce, and stored in aliquots at -80 ºC. The expression of BcsA and BcsG was detected by Western blotting using Anti-His and Flag antibodies, respectively. Analysis of Phosphoethanolamine modification by PACE IMVs were used for synthesizing cellulose in vitro (49). 500 μL of reactions were set up by incubating IMVs in the presence of 20 mM MgCl2, 5 mM UDP-Glc, and 30 μM c-di- GMP at 37 ºC for 16 hours. The amount of IMVs used was standardized based on radiometric cellulose quantification (as described above). Following the incubation, the reaction was terminated with 2% SDS and the insoluble polymer was pelleted by centrifugation at 21,200 g at room temperature. The obtained pellet was washed four times with water to remove the SDS. The resulting pellet was used for labelling with Alexa Fluor 647 NHS ester (succinimidyl ester, Invitrogen) in 100 mM sodium bicarbonate buffer pH 8.3. The dye was dissolved in 100% DMSO as per the manufacturer instructions at a concentration of 10 mg / mL. The labelling reaction (100 μL) was carried out for 2 hours at room temperature with continuous agitation. After the incubation, excess dye was removed by washing 3-4 times in sodium bicarbonate buffer. The resulting pellet was stored in 4 ºC and next day, one washing was done with MQ. The resulting pellet was digested with 1 mg / mL of purified Ec BcsZ cellulase (endo-β-1,4- glucanase, GH-8) in 250 μL reaction volume at 37 ºC for 4 hours in 20 mM sodium phosphate buffer pH 7.2. Control reactions were performed by omitting BcsZ. The digested samples were centrifuged at 21,200g for 20 min at room temperature and the soluble oligosaccharides were collected and dried using a centrifugal evaporator at low heat settings. The dried sample was resuspended in 15 µL urea and then analyzed by PACE as described previously (31). Briefly, from each sample, 2.5 μL was loaded onto the 240 X 180 X 0.75 mm polyacrylamide gel comprising a stacking gel with 10% polyacrylamide and a resolving gel with 20% acrylamide, both containing 0.1 M Tris-borate pH 8.2. Gels were run in 0.1 M Tris-borate buffer in a Hoefer SE660 electrophoresis tank (Hoefer Inc, Holliston, MA, USA) at 200 V for 30 min and then 1000 V for 2 hours before imaging using a G-Box Chemi-XX6 (Syngene, Cambridge, UK). Images were acquired with the GeneSys software (Syngene, version 1.8.5.0) based on the excitation and emission wavelength of the fluorophore (651 / 672 nm). A mixture of glucose and cello-oligosaccharides with DP 2-6 (each with 20 μM final concentration; (Glc)1-6ladder / Standard) was also loaded onto the gel as an internal mobility marker. For visualization, the marker was labeled with 8-aminonapthalene-1,3,6-trisulfonic acid (ANTS) as described previously (31). The ANTS labelled standard was imaged using the same G-Box equipped with a long-wavelength (365 nm) UV tube, and short pass detection filter (500-600 nm). Additional ZIMMER-PETN (03018-02) / / 1036.375WO1 standard reactions were setup initially to design these experiments wherein pure Ec pEtN cellulose (kindly provided by Lynette Cegelski, University of Stanford) or the unmodified phosphoric acid swollen cellulose was labelled with Alexa Fluor 647 NHS ester and digested with BcsZ. Western Blotting Following SDS-PAGE, the protein was transferred to a nitrocellulose membrane using a BioRad Transfer system. After blocking the membranes with 5% nonfat milk in Tris-buffered saline and 0.1% Tween 20 buffer (TBST), the membranes were washed with TBST and incubated with primary antibody at 4 ºC overnight. The membranes were then washed three times by incubating with fresh TBST for 10 min before incubation with anti-mouse IgG conjugated to a DyLight 800 fluorescent marker (Rockland) for 1 hour at room temperature. The membranes were washed three times with fresh TBST and visualized using an Odyssey Light scanner at wavelengths 700 and 800 nm. Mass spectrometry Protein identification of BcsA-NTD by mass spectrometry was performed on Coomassie stained and excised SDS-PAGE gel band. For this purpose, purified Ec NTD-BcsG complex from the gel filtration elution peak was loaded onto a 17.5% SDS-PAGE gel. After staining and destaining and rinsing in MQ, the desired band migrating close to the Ec BcsA- NTD molecular weight was excised and submitted to Biomolecular Analysis Facility at the University of Virginia. The gel pieces were digested in 20 ng / µL trypsin in 50 mM ammonium bicarbonate on ice for 30 min. Any excess enzyme solution was removed and 20 µL 50 mM ammonium bicarbonate added. The sample was digested overnight at 37 ºC and the released peptides were extracted from the polyacrylamide in a 100 µL aliquot of 50% acetonitrile / 5% formic acid. The samples were purified using C18 tips. This extract was evaporated to 20 µL for MS analysis. The liquid chromatography-mass spectrometry (LC-MS) system consisted of a Thermo Electron Orbitrap Exploris 480 mass spectrometer system with an Easy Spray ion source connected to a Thermo 75 µm x 15 cm C18 Easy Spray column. 5 µL of the extract was injected and the peptides eluted from the column by an acetonitrile / 0.1 M formic acid gradient at a flow rate of 0.3 µL / min over 2.0 hours. The nanospray ion source was operated at 1.9 kV. The digest was analysed using the rapid switching capability of the instrument acquiring a full scan mass spectrum to determine peptide molecular weights followed by product ion spectra (Top10 HCD) to determine amino acid sequence in sequential scans. This mode of analysis produces approximately 25000 MS / MS spectra of ions ranging in abundance over several ZIMMER-PETN (03018-02) / / 1036.375WO1 orders of magnitude. The data were analysed by database searching using the Sequest search algorithm against the Ec BcsA protein and Uniprot Ec. Electron microscopy Grid preparation and data acquisition Cryo-EM grids for the nanodisc reconstituted sample or the detergent solubilized samples were prepared in a similar manner on Quantifoil R1.2 / 1.3 Cu 300 mesh or C-Flat 1.2 / 1.3-4 Cu grids respectively. Grid preparation was optimized for each protein individually in regard to the protein concentration such that cryo-EM grids was prepared with purified chimeric BcsAB-BcsG complex at a concentration of 1.6 mg / mL; Ec Bcs complex with BcsB- BcsC fusion at a concentration of about 2.2 mg / mL, and BcsC reconstituted in nanodiscs at a concentration of about 2.5 mg / mL. To explore the interactions between the IMC and the outer membrane porin BcsC, Ec Bcs complex reconstituted in nanodisc was incubated with the purified periplasmic domain of BcsC (TPR #1-18) at a molar ratio of 1:5 for 1 hour on ice before cryo grid preparation. Likewise for the data collection with BcsZ, BcsC was incubated with the purified BcsZ at a molar 1:1.5. For data collection of BcsC nanodiscs incubated with cellotetraose (CTE), in addition to the 5 mM CTE present during the BcsC nanodisc reconstitution mixture, an additional 10 mM CTE was added to the final sample before grid preparation. Grids were glow-discharged in the presence of amylamine and 2-2.5 µl sample was applied to each grid and blotted and plunge-frozen in liquid ethane using a Vitrobot Mark IV plunge-freezing robot operated at 4 ºC and 100% humidity. Blotting force and time was optimized individually for each sample: force of 4 and time 6 sec for both the chimeric BcsAB- BcsG complex and the BcsB-BcsC fusion Ec complex; force 4 and blotting time 5 sec for the Ec Bcs complex in nanodiscs; and force 4 and 4 sec blotting time for BcsC. Grids were loaded into a Titan Krios electron microscope equipped with a K3 / GIF detector (Gatan) at the Molecular Electron Microscopy Core (University of Virginia School of Medicine). Data were collected using EPU in counting mode at a magnification of 81K, pixel size of 1.08 Å and a total dose of 51 e- / Å2, with a target defocus varying from -1 to -2.4 µm. Data processing Data was processed using cryoSPARC (28) or Relion (50), following similar processing pipelines. Initially, the movies were imported, and the images were first normalized by gain reference. Beam induced motion correction was performed using patch motion correction and contrast transfer function (CTF) parameters were estimated using patch CTF estimation. Micrographs were manually curated and those with outliers in defocus value, astigmatism, total ZIMMER-PETN (03018-02) / / 1036.375WO1 full frame motion, ice thickness and low resolution (below 4.5 Å) were removed. From these high-quality micrographs, particles were selected using Blob picker, extracted, and used to generate 2D classes for template-based particle picking. After extraction of these particles and 2D classification, ab initio models were generated and subjected to several rounds of heterogenous refinement. Selected particles and volumes were subjected to non-uniform refinement. For all datasets except BcsZ, this was followed by local refinement and further 3D classification using model or solvent based masks as described in the workflows. The best 3D class was subjected to local refinement again. Model building was performed in Coot starting with AlphaFold2 predicted models of the complex of BcsB and BcsC’s N-terminal TPR #1-4 aas well as for BcsC. The cryo-EM structure of the Ec BcsB hexamer (PDB: 7L2Z) and the crystal structure of BcsZ (PDB 3QXQ) were used as the corresponding initial models. The coordinates were rigid body docked into the corresponding volumes using Chimera. Models were manually refined in Coot (51) or ISOLDE (52) and real space refined in PHENIX (53). The model of BcsA in association with the BcsG trimer was predicted by AlphaFold2 and docked into the corresponding map based on BcsA’s location. The individual BcsG subunits and BcsA’s NTD were then rigid body docked into their densities. The model of the BcsA- BcsG3 complex containing the BcsB hexamer was obtained after rigid body refinement against a composite map generated from the best BcsB and BcsA-BcsG3 volumes. Figure representations were generated using Chimera (54), ChimeraX (55) and PyMOL (56). Results The Ec pEtN cellulose biosynthesis machinery catalyzes the synthesis, secretion, and pEtN modification of cellulose (Fig. 1a). The mechanism of cellulose synthesis and translocation across the IM has previously been addressed using the BcsA and BcsB components from Rhodobacter sphaeroides (Rs), producing unmodified cellulose (21,22). BcsA associates with three copies of the pEtN transferase BcsG BcsG is a membrane-bound pEtN transferase containing five N-terminal TM helices, followed by a periplasmic catalytic domain, FIGs. 1C and 7A (23-25). The homologous enzymes EptA and EptC catalyze pEtN modification of lipid A and N-glycans, respectively (26,27). Mechanistically, BcsG likely employs a catalytic triad involving Ser, His and Glu residues. The conserved Ser residue (Ser278) is assumed to serve as the nucleophile to attack the electrophilic phosphorous of a phosphatidylethanolamine (PE) lipid to form a covalent reaction intermediate with pEtN and releasing diacylglycerol, Fig.7B (25). The pEtN group is then subject to attack by the glucose C6 hydroxyl oxygen, resulting in transfer of the pEtN group to cellulose and release of the phospho-enzyme intermediate. This transfer reaction ZIMMER-PETN (03018-02) / / 1036.375WO1 requires the reorientation of BcsG’s catalytic pocket away from the membrane towards the translocating cellulose polymer. Approximately half of cellulose’s glucosyl units are modified by pEtN in Ec and Salmonella species under normal growth conditions (14). Previous intermediate-resolution cryo-EM analyses of the Ec Bcs macro-complex suggested the presence of two BcsG subunits associated with BcsA (20). Although only the TM regions were resolved, the subunits were located ‘in front’ of BcsA near the periplasmic exit of its cellulose translocation channel. Taking advantage of recent improvements in 3- dimensional variability analyses and classifications implemented in the cryoSPARC data processing workflows (28), the previously published cryo-EM data (EMD-23267) (20) was reprocessed, focusing on the BcsA-BcsG complex. Although still lacking well resolved periplasmic domains, this analysis generated improved maps for BcsA together with a trimeric complex of BcsG at approximately 5.7 Å resolution, Fig.1D. Each BcsG subunit contains five TM and two periplasmic interface helices that connect TM helices 3 and 4, Fig. 1d and e and 7a. The TM and interface helices form a ring-shaped periplasmic corral with an opening towards the phospholipid head groups. An AlphaFold2 prediction of full-length BcsG places its periplasmic domain on top of the corral, with the catalytic Ser278 facing the membrane beneath, Fig. 1C and Fig. 7A. The predicted BcsG structure suggests an arched acidic tunnel from the membrane surface, through the corral, and towards Ser278 that can position the PE headgroup for nucleophilic attack, Fig.1C. In the Bcs complex, the three BcsG subunits are arranged along a slightly curved line, Fig.1D and E. Interprotomer contacts within the BcsG trimer are mediated by TM helix 4 of one subunit and TM helices 2 and 5 of a neighboring subunit, suggesting that BcsG is prone to self-polymerize. Except for BcsA, no additional density is observed within the TM region that could correspond to other Bcs subunits, such as the single spanning subunit BcsF that has been shown to interact with BcsG and BcsE in vivo (14, 18). It is possible that this subunit is present but not resolved in the current cryo-EM map, as previously discussed (20). BcsA’s N-terminus recruits the BcsG trimer Compared to the homologous enzymes from Rs and Gx, BcsAs encoded by pEtN cellulose producers contain about 140 additional N-terminal residues of unknown function (referred to as the NTD) as well as a diverging C-terminus, Fig.8A. To test whether the NTD mediates interactions with BcsG, AlphaFold2 (29, 30) complex predictions were performed using either the full-length BcsA sequence or its NTD alone, together with three copies of BcsG’s TM region. The predictions reproducibly suggest the coordination of a BcsG trimer via ZIMMER-PETN (03018-02) / / 1036.375WO1 BcsA’s NTD, Fig.1E and Fig.8B, which is predicted to fold into five short helices (α1-α5). In addition, helices α4 and α5 are stabilized by BcsA’s extreme C-terminus, Fig.1E. The refined cryo-EM map experimentally validates the AlphaFold2-generated BcsA- BcsG3 model, Fig. 1d and e. The model was docked into the experimental map based on the location of BcsA. Its NTD as well as the three BcsG subunits were fit into the corresponding densities as rigid bodies to account for a different angle of the NTD-BcsG trimer relative to BcsA, Fig.8C. All BcsA NTD helices proposed to interact with the BcsG trimer are resolved in this map. Further, BcsA’s C-terminal helix is observed in contact with the NTD’s α4 and α5 helices, as predicted. In contrast to an earlier model that assumed a segment of the NTD to form a proper TM helix (20), the NTD resides on the cytosolic water-lipid interface with all of its helices running approximately parallel to the membrane surface, Fig.1D. Although only resolved at the backbone level in the cryo-EM map, the BcsG subunits interact with the NTD via their cytoplasmic loops connecting TM helices 4 and 5 (TM4 / 5-loop, residues 135-138), Fig. 1E. Starting with the BcsG protomer farthest away from BcsA, AlphaFold2 predicts that the TM4 / 5-loop forms backbone interactions with NTD residues 9- 11 connecting its α1 and α2 helices. Similarly, the TM4 / 5-loop of the next BcsG subunit interacts with residues 48-50 between NTD’s α3 and α4 helices, and the protomer closest to BcsA contacts NTD residues 93-95 following α5 via its TM4 / 5-loop. Past this interface, the NTD is connected to a predicted amphipathic interface helix (α6) that is only partially resolved in the cryo-EM map, Fig. 1D and E. This helix leads into BcsA’s first TM helix. No direct interactions are observed or predicted between BcsG and BcsA TM segments. The BcsG TM helices fit seamlessly into the opening of the BcsB semicircle, between the first and sixth subunit, Fig.1D. The C-terminal helical region of Ec BcsA runs roughly in the opposite direction as the preceding amphipathic helix, which is present in BcsA from Rs and other species. This segment is resolved in the cryo-EM map and interacts with helix α5 of the NTD via an SxxPRxP motif (residues 850-856), Fig.1D and e and Fig.8A. The motif is predicted to interact with residues 82 through 89 (α5) of the NTD, with possible hydrogen bonds between Arg82 and Ser850 as well as Gln86 and Arg854, Fig.1E. BcsA’s NTD is sufficient for BcsG recruitment To test whether BcsA’s NTD is sufficient to recruit BcsG, the Strep-tagged NTD (residues 1-94) was co-expressed with poly-histidine tagged BcsG. A tandem affinity purification using Ni-NTA resin followed by Strep-Tactin beads and size exclusion ZIMMER-PETN (03018-02) / / 1036.375WO1 chromatography indeed isolated an NTD-BcsG complex, Fig. 1F. The identities of the co- purified protein components were confirmed by Western blotting and tandem MS sequencing. Except for the NTD and the C-terminal region, Rs BcsA is structurally homologous to Ec BcsA. R. sphaeroides does not encode a BcsG homolog or the other pEtN-cellulose specific Bcs components BcsE and BcsF. It was sought to evaluate whether the pEtN cellulose modification could be introduced in the Rs cellulose biosynthetic system via BcsA engineering and inclusion of the pEtN transferase BcsG. To this end, a ‘BcsA chimera’ was generated containing the Rs BcsA sequence N-terminally extended with the Ec NTD (residues 1-149). In addition, the C-terminal Rs BcsA region (residues 729 to 788) was replaced with the corresponding Ec BcsA sequence (residues 828 to 872, see FIG. 8A), followed by a poly- histidine tag for purification, Fig.2a (see Methods). The BcsA chimera was co-expressed with Rs BcsB as well as Ec BcsG and BcsF. Metal affinity and size exclusion chromatography purification of the expressed complex demonstrated the co-purification of the BcsA chimera with BcsB (the native binding partner of Rs BcsA) as well as Ec BcsG, Fig. 2b. The presence of co-purified BcsA and BcsG was confirmed by Western blotting, while BcsF was not detected in the purified sample. BcsG did not co-purify with the wild type (WT) Rs BcsAB complex, suggesting that the interaction is indeed mediated by the introduced Ec BcsA NTD. The purified chimeric BcsAB-BcsG complex is catalytically active in vitro. Similar to previous reports on the wild type Rs BcsAB complex (6), the BcsA chimera synthesizes cellulose in vitro from UDP-glucose in a cyclic-di-GMP (ci-di-GMP) dependent reaction, Fig. 2C. The obtained product is readily degraded by a cellulase, as expected for a cellulose substrate. Cryo-EM analysis of the purified complex revealed the association of the BcsA chimera with BcsG in a curved micelle, Fig. 2D and E. Refinements of either the chimeric BcsAB complex alone or in association with BcsG resulted in maps of approximately 4.6 and 6 Å resolution, respectively, Fig. 2E. The maps resolve the BcsAB complex associated with a nascent cellulose polymer in a conformation similar to the previously reported crystal structure (5), as well as the TM domains of three BcsG subunits, Fig.2E. The BcsG trimer is assembled as observed in the Ec Bcs complex described above, Fig.1D, and interacts with short interface helices corresponding to the engineered Ec BcsA NTD. The also introduced Ec C-terminal extension is flexible and insufficiently resolved, likely resulting in the tilting of the BcsG trimer relative to BcsAB in the detergent micelle (Fig.2D and E). ZIMMER-PETN (03018-02) / / 1036.375WO1 To test whether the engineered chimeric cellulose synthase complex produces pEtN cellulose in vivo, cellulosic material was isolated from the periplasm of the expression host (see Methods), where it accumulates in the absence of the OM subunit BcsC (this subunit has not been identified in Rs yet). The production of pEtN-cellulose by the chimeric complex was validated through direct detection of material isolated from cell lysates using13C cross- polarization magic-angle spinning (CPMAS) solid-state NMR spectroscopy. Compared to a reference sample of pure pEtN cellulose from Ec (14), the cellulosic material produced by the BcsA chimera is highly enriched in pEtN cellulose, Fig.2F. As an additional pEtN cellulose detection assay, the synthesized product was analyzed by polysaccharide carbohydrate gel electrophoresis (PACE) (31). To this end, inverted membrane vesicles (IMVs) were prepared from cells expressing either the wild type Rs BcsAB complex alone, or the chimeric or wild type BcsAB complex together with Ec BcsG and BcsF. The presence of BcsA and BcsG (when applicable) in the IMVs was confirmed by Western blotting, Fig.2G. In vitro cellulose biosynthesis reactions were then performed with the IMVs with the expectation that BcsG would modify cellulose using PE lipids as pEtN donors. Following synthesis, the water-insoluble material was isolated after SDS denaturation and any pEtN units were subjected to modification at the amino nitrogen with N-hydroxysuccinimide (NHS)-conjugated Alexa Fluor 647. This reaction was followed by digestion with cellulase to release water-soluble cello-oligosaccharides, and the released material was analyzed by PACE and imaged (see Methods). Control reactions with either unmodified phosphoric acid swollen cellulose or purified pEtN cellulose from Ec only show the release of fluorescently labeled cello-oligosaccharides by cellulase from pEtN cellulose. This confirms the reliable detection of pEtN cello-oligosaccharides by PACE. As shown in Fig.2H, fluorescently labeled cello-oligosaccharides are readily released by cellulase from the reaction product of the chimeric BcsAB complex in the presence of BcsG. No labeled oligosaccharides are obtained from products produced by wild type Rs BcsAB alone. Minor weaker bands are detected when the wild type Rs BcsAB complex is co-expressed with BcsG and BcsF. This likely results from modification of cellulose accumulating or precipitating on the membrane surface by the abundantly expressed BcsG. While all three IMV samples produce cellulose in vitro based on3H-glucose incorporation (Fig. 2I) (6, 32), the expression level of the chimeric BcsAB complex is higher, giving rise to approximately twofold greater cellulose yields, Fig. 2G. To account for limitations in detection levels by PACE, loading approximately twice the amount of the product obtained from the wild type BcsAB complex co-expressed with BcsG and BcsF did not result in stronger PACE signals. ZIMMER-PETN (03018-02) / / 1036.375WO1 Combined, the NMR and PACE analyses suggest that the close association of BcsG with the BcsA chimera greatly facilitates pEtN cellulose formation. The BcsB semicircle tethers a single BcsC subunit Upon pEtN modification by BcsG, cellulose crosses the periplasm and the OM. This step requires BcsC, a ~130 kDa protein consisting of a C-terminal β-barrel domain preceded by 19 predicted TPR motifs. To investigate the interactions of the IM-associated Bcs complex (IMC) with BcsC by cryo-EM, the purified IMC was incubated with the separately purified N- terminal 18 TPRs of BcsC for 1h prior to cryo grid preparation (see Methods). Cryo-EM analysis of this Bcs complex revealed the previously observed BcsB hexamer architecture associated with one BcsA subunit, Fig. 3a. Viewed from the periplasm and counting clockwise, the first BcsB subunit interacts with BcsA, while the sixth sits at the opposite end of the semicircle (Fig. 1B and 3A). At low contour levels, additional density at the membrane distal tip of the sixth BcsB copy is evident, roughly extending along the perimeter of the semicircle towards the first BcsB protomer, Fig.3A. The extra density, likely belonging to BcsC’s periplasmic domain, is located between BcsB’s N-terminal CBD (CBD- 1) and the following FD region (FD-1) (referred to the ‘TPR binding groove’) (Fig.3A and B). At this contouring, fragmented density at the open side of the semicircle likely representing BcsG’s periplasmic domain is also evident near the newly identified BcsC density, Fig.3A. Similar to the approach employed for interrogating the BcsA-BcsG interaction, AlphaFold2 wasutilized to predict the interactions of BcsB with BcsC’s N-terminal TPRs. Using a single copy of BcsB and BcsC’s N-terminal TPRs #1-4, AlphaFold2 positions BcsC’s TPR#1 with high confidence into BcsB’s TPR binding groove. The predicted model is in excellent agreement with the cryo-EM map. AlphaFold2 did not generate consistent models for a truncated BcsC construct lacking TPR#1, suggesting that the interaction with BcsB indeed depends on the N-terminal region. It was concluded that the apical tip of BcsB establishes the interaction with BcsC’s TPR#1. To improve the resolution of the BcsB-BcsC complex map, BcsC’s TPR #1-4 were fused to the N-terminus of BcsB, after its N-terminal signal sequence and separated by a linker (see Methods). This fusion construct was co-expressed with all other Bcs components, followed by purification and cryo-EM analyses as described for the wild-type Bcs complex. As also observed for the wild type Bcs complex, cryo-EM analysis identified a fully assembled Bcs complex together with (likely dissociated) subcomplexes of the BcsB semicircle, containing three to five protomers. Although additional BcsC density similar to the one described above was also observed in a fully assembled complex, the highest quality map ZIMMER-PETN (03018-02) / / 1036.375WO1 revealing the BcsB-BcsC interaction was obtained for a tetrameric BcsB complex. Non- uniform and focused local refinements generated a map of about 3.2 Å resolution that delineates the specific interactions of BcsB and BcsC, Fig.3A (see Methods). The BcsC density associated with BcsB accommodates four α-helices, corresponding to BcsC’s N-terminal two TPRs, Fig. 3B. TPR#1 mediates all interactions with BcsB and is best resolved. TPR#2 is rotated by about 45 degrees relative to TPR#1 and extends away from BcsB. In complex with BcsB, TPR#1 is oriented with its interhelical loop pointing towards the center of the BcsB semicircle. It interacts extensively with one side of CBD-1’s jelly roll, as well as the region connecting the jelly roll with FD-1 (residues 208 – 220), Fig.3B. This interface contains hydrophobic, polar and charged residues. In particular, TPR#1’s N-terminal helix (helix-1) rests on a short helical segment of a CBD-1 loop, such that its Gln30, Gln34 and Leu37 stack on top of BcsB’s Leu162, Phe163, and Ile164. The side chains of Gln30 and Gln34 form hydrogen bonds with the backbone carbonyls of Val108 and Ile164, respectively. The following C-terminal helix of TRP#1 (helix-2) contacts BcsB primarily via polar interactions, including Gln49 and Arg53 that interact with Ser165 and Asp166 of CBD-1, respectively. The helix’s C-terminal end fits into a hydrophobic pocket formed by the CBD-1 / FD-1 connection. Here, Leu56 and Ile57 stack against BcsB’s Leu207 and Val209, with a backbone hydrogen bond between Leu56 and Lys210 (Fig.3B). The following TPR#2 extends from TPR#1 towards the opening of the BcsB semicircle. Due to this arrangement, the observed BcsB-BcsC interaction is only possible with the last subunit of the BcsB hexamer. Modeling a similar complex with any other subunit of the BcsB semicircle creates substantial clashes between TRP#2 and the FD domains of the neighboring subunit. This explains why, although present, the fused TPRs of the other BcsB subunits are not resolved. BcsC is an outer membrane porin with a periplasmic cellulose-binding solenoid extension The crystal structure of a C-terminal BcsC fragment containing the β-barrel, a linker, and TPR#19 revealed the porin architecture and its connection with the periplasmic TPR solenoid (8). Further, crystallographic analysis of the N-terminal six BcsC TPRs from Enterobacter CJF-002 resolved their solenoid organization (13). Lastly, the AlphaFold2- predicted model of full-length BcsC supports the solenoid arrangement of its TRP#7-19, forming roughly two helical turns that extend by about 130 Å into the periplasm, Fig.4A. ZIMMER-PETN (03018-02) / / 1036.375WO1 To gain experimental insights into the architecture of full-length BcsC and its interaction with cellulose by cryo-EM, the protein was reconstituted into a lipid nanodisc in the absence and the presence of cellotetraose. The obtained cryo-EM maps underscore the high flexibility of BcsC’s periplasmic domain. At lower contour levels and for the sample devoid of cellotetraose, the cryo-EM map confirms the solenoid architecture of TPR #9-19, while the N- terminal eight TPRs are insufficiently resolved or absent in the experimental map, Fig. 4A. Under both conditions, high resolution refinements (to about 3.2 Å) were only possible for a C-terminal portion of BcsC, beginning with TPR#15 and #16 for the ligand-bound and apo BcsC datasets, respectively, Fig.4A-C. The refined BcsC structure is consistent with the AlphaFold2-model, with only minor rigid body translations of TPR#15 towards the solenoid axis, Fig.4A. Close inspection of the cryo-EM map obtained in the presence of cellotetraose revealed additional elongated density close to the solenoid axis, contacting TPR #16-19, Fig.4A-B. Although the density cannot be identified unequivocally as a cello-oligosaccharide at the current resolution, its shape and interaction with BcsC are consistent with it representing a cellotetraose molecule, perhaps bound in different binding poses. Supporting this interpretation, no additional density at this site or elsewhere is observed in the absence of cellotetraose in a map of similar quality, Fig. 4B. Therefore, the observed molecule is referred to as a ‘putative cellulose ligand’. The most prominent interaction of the putative cellulose ligand with BcsC is mediated by Trp766 at the N-terminus of the linker region, Fig. 4B. The ligand stacks against this aromatic side chain, similar to the cellulose coordination by cellulose synthases and hydrolases (5, 33, 34). From here, additional sugar units extend towards the solenoid axis and contact TPR#18 and #16. Although the putative cello-oligosaccharide is not aligned with the center of the porin channel, Fig.4A and C, bending of the polymer or a different orientation relative to Trp766 could direct it into the OM channel, as described further below. Cellulase activity is necessary for cellulose secretion The cellulase BcsZ is a conserved subunit of Gram-negative cellulose biosynthetic systems (3, 11). Similarly, plant and tunicate cellulose synthases are also associated with cellulases for unknown reasons (35, 36). Deleting BcsZ in Gx substantially reduces cellulose production in vivo (10, 11), while BcsZ has been proposed to reduce biofilm phenotypes on Salmonella enterica typhimurium (37). Accordingly, it was determined whether BcsZ is of similar importance to pEtN cellulose production in Ec. To this end, pEtN cellulose secretion from the transformed Ec cells was monitored based on Congo red (CR) fluorescence of cells grown on nutrient agar plates, as ZIMMER-PETN (03018-02) / / 1036.375WO1 previously described (20). This assay takes advantage of substantially enhanced CR fluorescence in the presence of pEtN cellulose, compared to unmodified cellulose (38). When grown on CR agar plates, Ec C43 cells expressing the complete Ec Bcs system together with the cyclic-di-GMP producing diguanylate cyclase AdrA, give rise to strong fluorescence, indicative of pEtN cellulose secretion, Fig.5A. Cells expressing the IMC only, however, reveal background staining, similar to cells producing unmodified cellulose due to the Ser278 to Ala substitution in BcsG (25), Fig.5A. Consistent with previous observations in Gx, CR staining of Ec lacking BcsZ indicates substantially reduced pEtN cellulose secretion in the absence of the cellulase, similar to control cells expressing the IMC only, Fig.5A. It was next investigated whether BcsZ is only required for its cellulose degrading activity or whether it could be a structural component of the pEtN cellulose secretion system. In the latter case a catalytically inactive enzyme may still facilitate pEtN cellulose export. To this end, two inactive BcsZ mutants were generated by substituting the catalytic residues Glu55 and Asp243 with Gln and Ala, respectively (39). To confirm that the generated BcsZ variants are indeed catalytically inactive, we employed an agar plate-based carboxymethylcellulose digestion assay, as previously described (39). Here, cellulose digestion by cellulase secreted by plated cells is detected upon CR staining. As expected, Ec cells expressing the wild type or the generated BcsZ mutants only show cellulase activity for the wildtype enzyme, confirming that the generated BcsZ mutants are inactive within the sensitivity limits of the assay. Accordingly, analyzing pEtN cellulose secretion by these cells based on CR fluorescence shows much reduced fluorescence in the presence of the BcsZ:D243A mutant in comparison to the cells expressing Ec complex with wild type BcsZ, while no CR staining above background is observed in the presence of the E55Q or double BcsZ mutant, Fig.5A. Lack of cellulose secretion in the absence of cellulase activity indicates that BcsZ plays a modulating rather than structural role during cellulose export. Accordingly, it was investigated whether unrelated bacterial cellulases could functionally replace BcsZ. To this end, the Ec BcsZ enzyme was replaced in the Bcs expression system with either its Gx homolog CMCax (Gx produces unmodified fibrillar cellulose) (12), or the Cel9M cellulase domain from the Clostridium cellulolyticum cellulosome (40). Co-expressing the cellulases with the remaining Ec Bcs components resulted in detectable cellulase activity in the periplasmic fraction, suggesting that the cellulases were functionally expressed and translocated into the periplasm, Fig. 5b. Monitoring pEtN cellulose secretion by these cells based on CR fluorescence revealed that both cross-species complementations restored secretion by the Ec ZIMMER-PETN (03018-02) / / 1036.375WO1 Bcs complex, comparable to wild type levels, Fig. 5C. The results indicate that cellulase activity is needed for cellulose export. BcsZ assembles into a tetramer Direct interactions of BcsZ and BcsC could not be detected biochemically or by cryo- EM analysis. However, homo-oligomerization of the cellulase was discovered. Single particle cryo-EM analysis at a resolution of about 2.7 Å revealed the presence of BcsZ tetramers, in addition to small monomeric particles, Fig.5D-F. The tetramers are dimers of homodimers in which the protomers are rotated by about 180 degrees relative to each other, Fig. 5E. The homodimer interface is formed by helices 11 and 12 of the glycosylhydrolase-8 fold. It is rich in ionic interactions, including Asp344 and Arg318 of one protomer and the equivalent residues in the symmetry-related subunit. Similarly, Asp312 interacts with Arg349 across protomers and so does the Asp323 and Arg347 pair. In addition, Gln319 in helix #11 of one protomer hydrogen bonds to the backbone carbonyl oxygens of Gln345 and His346 following helix #12 of the symmetry related subunit. Two homodimers interact via the N-terminal regions of opposing BcsZ protomers, involving the loop connecting helix 1 and 2 (residues 37-53) as well as a β-strand hairpin connecting helices 3 and 4 (residues 86-114), Fig.5E and F. This interface also contains several ionic and polar interactions. In particular, Ser104 and Lys105 of the helix3 / 4 loop are in hydrogen bonding distance to Asn82 in helix 3 of the opposing subunit. Arg41 is juxtaposed to Glu39 while Gln38 hydrogen bonds to Lys50 across the dimer interface. Combined, these interactions create a square-shaped tetrameric assembly with the cellulose binding clefts of BcsZ subunits at opposing corners facing in the same direction. However, because BcsZ binds cellulose in a defined orientation (39), the 2-fold symmetry related subunits bind their polymeric substrate in opposing directions, Fig. 5E. Of note, the same tetrameric complex was previously observed but not functionally interpreted in apo and cellopentaose-bound BcsZ crystal structures, where the described tetramer either represents the crystallographic asymmetric unit or is generated by symmetry mates (39). Discussion Cellulose is a versatile biomaterial with countless biological and industrial applications. In Gram-negative bacteria, the polysaccharide is secreted across the cell envelope in a single process. Considering cellulose’s amphipathic properties, its periplasmic secretion likely requires a shielded translocation path to prevent nonspecific interactions. The cryo-EM analysis of the IM-associated Bcs complex reveals a single BcsA cellulose synthase associated with a trimer of the pEtN transferase BcsG. Although present as ZIMMER-PETN (03018-02) / / 1036.375WO1 the full-length enzyme in the analyzed sample, only its membrane-embedded region is resolved in the cryo-EM map. This suggests that the periplasmic catalytic domain is flexibly attached to the TM helices. BcsG’s catalytic domain is connected to the preceding TM helices via a linker of about 45 residues. In an extended conformation, the linker would allow the TM and periplasmic regions to separate by more than 80 Å, about the height of the BcsB semicircle. However, AlphaFold2 predictions of full-length BcsG reproducibly pack the catalytic domain against the periplasmic corral formed by its interface helices. In this conformation, the catalytic Ser278 points toward the membrane surface where it could receive a pEtN group. The corral may help to position a PE lipid to facilitate this reaction. The in vitro reconstituted pEtN cellulose biosynthesis indeed confirms that Ec lipids can serve as the pEtN donor. It was hypothesized that following formation of the phosphor-enzyme intermediate with Ser278- bound to pEtN, the catalytic domain disengages from the TM region to access the translocating cellulose polymer. Accordingly, pEtN transfer to cellulose requires an approximately 90-degree rotation and substantial translation of BcsG’s catalytic domain away from the membrane and towards the nascent polysaccharide. It is possible that the catalytic domains of the BcsG trimer operate independently, perhaps resulting in stochastic cellulose modification. Previously, it was proposed that two BcsG subunits are necessary to result in pEtN modification of cellulose’s C6 hydroxyl groups located on opposing sides of the cellulose polymer (20). Yet, the instant cryo-EM maps reveal that three copies are present. A higher copy number could compensate for limiting transfer efficiency. It is also noted that the three BcsG periplasmic domains, albeit flexible, essentially close the semicircle formed by the BcsB hexamer. This may steer the translocating cellulose chain towards BcsC in the OM. It is anticipated that the ability to associate a BcsG trimer with an unrelated cellulose synthase via grafting of BcsA’s NTD will unleash the untapped potential to modify cellulose in different systems. Because the Bcs complex produces only one cellulose polymer at a time, a single copy of the OM porin BcsC suffices to guide cellulose across the periplasm and the OM. Indeed, the arrangement of BcsC’s N-terminal TPRs and steric constraints within a BcsB hexamer ensure that only the terminal BcsB subunit interacts with BcsC. Based on the AlphaFold2-predicted model of full-length BcsC, its 19 TPRs extend by approximately 150 Å into the periplasm. BcsB’s periplasmic domain is about 70 Å tall. Combined, BcsB and BcsC suffice to span the periplasm and the OM. ZIMMER-PETN (03018-02) / / 1036.375WO1 The requirement of cellulase activity for efficient cellulose secretion suggests a model in which BcsZ prevents or reverses mislocalization of cellulose to the periplasm to ensure its translocation across the OM, Fig. 6. Thisr putative BcsC-cellotetraose complex indicates cellulose translocation along the solenoid, consistent with recent insights into poly N- acetylglucosamine interactions with the TPR of PgaA (41). Cellulose migrating away from the solenoid helix would likely be irreversibly mislocalized, thereby stalling cellulose biosynthesis. In this case, hydrolytic trimming of polymers accessible from the periplasm could reset the translocation process, Fig.6. The resulting cello-oligosaccharides may remain in the periplasm or be imported for degradation. Oligomerization of BcsZ likely increases its catalytic efficiency, as frequently observed for carbohydrate active enzymes (42). Taken together, the analyses demonstrate the synthesis and translocation of a single cellulose polymer by the Ec Bcs complex. The positioning of the first BcsB subunit ‘above’ the cellulose secretion channel steers the nascent chain towards the center of the BcsB semicircle, past the BcsG trimer. Flexibility of BcsG’s catalytic domains enable contacts with membrane lipids to recognize and attack a pEtN group as well as accessing the translocating cellulose polymer at different distances from BcsA. The modified cellulose polymer likely threads through the center of the TPR solenoid until reaching the OM channel. Polymers ‘escaping’ into the periplasm would be trimmed back by BcsZ to re-enter the TPR solenoid, Fig.6. This model explains why BcsZ’s hydrolytic activity is necessary to facilitate cellulose secretion and why the enzyme can be replaced with off-the-shelf cellulases. Hydrolytic clearance of roadblocks may also assist cellulose microfibril formation in other kingdoms of life. Bibliography 1 McCrate, O. A., Zhou, X., Reichhardt, C. & Cegelski, L. Sum of the Parts: Composition and Architecture of the Bacterial Extracellular Matrix. J. Mol. Biol.425, 4286-4294, (2013). 2 Hall-Stoodley, L., Costerton, J. W. & Stoodley, P. Bacterial biofilms: From the natural environment to infectious diseases. Nature Rev Microbiol 2, 95-108, (2004). 3 Romling, U. & Galperin, M. Y. Bacterial cellulose biosynthesis: diversity of operons, subunits, products, and functions. Trends Microbiol 23, 545-557, (2015). 4 McNamara, J. T., Morgan, J. L. W. & Zimmer, J. A molecular description of cellulose biosynthesis. Annu Rev Biochem 84, 17.11-17.27, (2015). 5 Morgan, J., Strumillo, J. & Zimmer, J. Crystallographic snapshot of cellulose synthesis and membrane translocation. Nature 493, 181-186, (2013). ZIMMER-PETN (03018-02) / / 1036.375WO1 6 Omadjela, O. et al. BcsA and BcsB form the catalytically active core of bacterial cellulose synthase sufficient for in vitro cellulose synthesis. Proc Natl Acad Sci U S A 110, 17856-17861, (2013). 7 Wong, H. C. et al. Genetic organization of the cellulose synthase operon in Acetobacter xylinum. Proc Natl Acad Sci U S A 87, 8130-8134, (1990). 8 Acheson, J. F., Derewenda, Z. S. & Zimmer, J. Architecture of the Cellulose Synthase Outer Membrane Channel and Its Association with the Periplasmic TPR Domain. Structure 27, 1855-1861 e1853, (2019). 9 Kawano, S. et al. Effects of endogenous endo-beta-1,4-glucanase on cellulose biosynthesis in Acetobacter xylinum ATCC23769. J Biosci Bioeng 94, 275-281, (2002). 10 Koo, H. M., Song, S. H., Pyun, Y. R. & Kim, Y. S. Evidence that a beta-1,4- endoglucanase secreted by Acetobacter xylinum plays an essential role for the formation of cellulose fiber. Biosci Biotechnol Biochem 62, 2257-2259, (1998). 11 Standal, R. et al. A new gene required for cellulose production and a gene encoding cellulolytic activity in Acetobacter xylinum are colocalized with the bcs operon. J Bacteriol 176, 665-672, (1994). 12 Yasutake, Y. et al. Structural characterization of the Acetobacter xylinum endo-beta- 1,4-glucanase CMCax required for cellulose biosynthesis. Proteins 64, 1069-1077, (2006). 13 Nojima, S. et al. Crystal structure of the flexible tandem repeat domain of bacterial cellulose synthesis subunit C. Sci Rep 7, (2017). 14 Thongsomboon, W. et al. Phosphoethanolamine cellulose: A naturally produced chemically modified cellulose. Science 359, 334-338, (2018). 15 Hollenbeck, E. C. et al. Phosphoethanolamine cellulose enhances curli-mediated adhesion of uropathogenic Escherichia coli to bladder epithelial cells. Proc Natl Acad Sci U S A 115, 10106-10111, (2018). 16 Jeffries, J. et al. Variation in the ratio of curli and phosphoethanolamine cellulose associated with biofilm architecture and properties. Biopolymers 112, e23395, (2021). 17 Serra, D. O., Klauck, G. & Hengge, R. Vertical stratification of matrix production is essential for physical integrity and architecture of macrocolony biofilms of. Environ Microbiol 17, 5073-5088, (2015). 18 Krasteva, P. V. et al. Insights into the structure and assembly of a bacterial cellulose secretion system. Nat Commun 8, 2065, (2017). 19 Abidi, W., Zouhir, S., Caleechurn, M., Roche, S. & Krasteva, P. V. Architecture and regulation of an enterobacterial cellulose secretion system. Sci Adv 7, (2021). ZIMMER-PETN (03018-02) / / 1036.375WO1 20 Acheson, J. F., Ho, R., Goulart, N., Cegelski, L. & Zimmer, J. Molecular organization of the E. coli cellulose synthase macrocomplex. Nat Struct Mol Biol 28, 310-318, (2021). 21 Morgan, J. L. et al. Observing cellulose biosynthesis and membrane translocation in crystallo. Nature 531, 329-334, (2016). 22 Morgan, J. L. W., McNamara, J. T. & Zimmer, J. Mechanism of activation of bacterial cellulose synthase by cyclic di-GMP. Nature Struct Mol Biol 21, 489-496, (2014). 23 Sun, L. et al. Structural and Functional Characterization of the BcsG Subunit of the Cellulose Synthase in Salmonella typhimurium. J Mol Biol 430, 3170-3189, (2018). 24 Anderson, A. C., Burnett, A. J. N., Hiscock, L., Maly, K. E. & Weadge, J. T. The Escherichia coli cellulose synthase subunit G (BcsG) is a Zn(2+)-dependent phosphoethanolamine transferase. J Biol Chem 295, 6225-6235, (2020). 25 Anderson, A. C. et al. A Mechanistic Basis for Phosphoethanolamine Modification of the Cellulose Biofilm Matrix in Escherichia coli. Biochemistry 60, 3659-3669, (2021). 26 Anandan, A. et al. Structure of a lipid A phosphoethanolamine transferase suggests how conformational changes govern substrate binding. Proc Natl Acad Sci USA 114, 2218-2223, (2017). 27 Scott, N. E. et al. Modification of the Campylobacter jejuni N-linked glycan by EptC protein-mediated addition of phosphoethanolamine. J Biol Chem 287, 29384-29396, (2012). 28 Punjani, A., Rubinstein, J. L., Fleet, D. J. & Brubaker, M. A. cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nat Methods 14, 290-296, (2017). 29 Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589, (2021). 30 Varadi, M. et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res 50, D439-D444, (2022). 31 Goubet, F., Dupree, P. & Johansen, K. S. Carbohydrate gel electrophoresis. Methods Mol Biol 715, 81-92, (2011). 32 Purushotham, P. et al. A single heterologously expressed plant cellulose synthase isoform is sufficient for cellulose microfibril formation in vitro. Proc Natl Acad Sci U S A 113, 11360-11365, (2016). 33 Knott, B. C., Crowley, M. F., Himmel, M. E., Stahlberg, J. & Beckham, G. T. Carbohydrate-protein interactions that drive processive polysaccharide translocation in enzymes revealed from a computational study of cellobiohydrolase processivity. J Am Chem Soc 136, 8810-8819, (2014). ZIMMER-PETN (03018-02) / / 1036.375WO1 34 Spiwok, V. CH / pi Interactions in Carbohydrate Recognition. Molecules 22, (2017). 35 Vain, T. et al. The Cellulase KORRIGAN Is Part of the Cellulose Synthase Complex. Plant Physiol 165, 1521-1532, (2014). 36 Matthysse, A. G. et al. A functional cellulose synthase from ascidian epidermis. Proc Natl Acad Sci U S A 101, 986-991, (2004). 37 Ahmad, I. et al. BcsZ inhibits biofilm phenotypes and promotes virulence by blocking cellulose production in serovar Typhimurium. Microbial Cell Factories 15, (2016). 38 Thongsomboon, W., Werby, S. H. & Cegelski, L. Evaluation of Phosphoethanolamine Cellulose Production among Bacterial Communities Using Congo Red Fluorescence. J Bacteriol 202, (2020). 39 Mazur, O. & Zimmer, J. Apo- and Cellopentaose-bound Structures of the Bacterial Cellulose Synthase Subunit BcsZ. J Biol Chem 286, 17601-17606, (2011). 40 Belaich, A. et al. Cel9M, a new family 9 cellulase of the Clostridium cellulolyticum cellulosome. J Bacteriol 184, 1378-1384, (2002). 41 Pfoh, R. et al. The TPR domain of PgaA is a multifunctional scaffold that binds PNAG and modulates PgaB-dependent polymer processing. PLoS Pathog 18, (2022). 42 Caveney, N. A., Li, F. K. & Strynadka, N. C. Enzyme structures of the bacterial peptidoglycan and wall teichoic acid biogenesis pathways. Curr Opin Struct Biol 53, 45-58, (2018). 43 Edgar, R. C. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucl Acid Res 32, 1792-1797, (2004). 44 Waterhouse, A. M., Procter, J. B., Martin, D. M., Clamp, M. & Barton, G. J. Jalview Version 2--a multiple sequence alignment editor and analysis workbench. Bioinform 25, 1189- 1191, (2009). 45 Pardon, E. et al. A general protocol for the generation of Nanobodies for structural biology. Nat Protoc 9, 674-693, (2014). 46 Schaefer, J. & Stejskal, E. O. Carbon-13 Nuclear Magnetic Resonance of Polymers Spinning at the Magic Angle. J Am Chem Soc 98, 1031-1032, (1976). 47 Bennett, A. E., Rienstra, C. M., Auger, M., Lakshmi, K. V. & Griffin, R. G. Heteronuclear Decoupling in Rotating Solids. J Chem Phys 103, 6951-6958, (1995). 48 Morcombe, C. R. & Zilm, K. W. Chemical shift referencing in MAS solid state NMR. J Magn Res 162, 479-486, (2003). 49 Verma, P., Kwansa, A. L., Ho, R. Y., Yingling, Y. G. & Zimmer, J. Insights into substrate coordination and glycosyl transfer of poplar cellulose synthase-8. Structure 31, 1166-+, (2023). ZIMMER-PETN (03018-02) / / 1036.375WO1 50 Scheres, S. H. W. RELION: Implementation of a Bayesian approach to cryo-EM structure determination. J Struct Biol 180, 519-530, (2012). 51 Emsley, P. & Cowtan, K. Coot: model-building tools for molecular graphics. Acta Crystallogr D Biol Crystallogr 60, 2126-2132, (2004). 52 Croll, T. I.: a physically realistic environment for model building into low-resolution electron-density maps. Acta Crystallogr D Biol Crystallogr 74, 519-530, (2018). 53 Adams, P. et al. PHENIX: a comprehensive Python-based system for macromolecular structure solution. Acta Crystallogr D Biol Crystallogr 66, 213-221, (2010). 54 Pettersen, E. F. et al. UCSF Chimera--a visualization system for exploratory research and analysis. J Comput Chem 25, 1605-1612, (2004). 55 Pettersen, E. F. et al. UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci 30, 70-82, (2021). 56 Pymol. Vol. Version 2.0 (Schroedinger, LLC). Example II Generation of an acetylated cellulose BcsG chimera It was estimated that the BcsG catalytic domain could be replaced with a different functional group to enable cellulose modifications other than pEtN transfer. To test this hypothesis, the BcsG catalytic domain is replaced with the acetyltransferase domain of Wssf, an enzyme recently shown to modify cello-oligosaccharides in vitro. The BcsG-Wssf chimera contains BcsG’s N-terminal transmembrane domain that is recognized by BcsA, followed by a flexible linker and the WssF catalytic domain. For in vitro studies, the N-terminal domain of E. coli BcsA is transferred onto Rhodobacter sphaeroides (now Cereibacter sphaeroides) BcsA, which is co-expressed with Rhodobacter BcsB as well as the engineered BcsG-Wssf chimera. Using metal affinity chromatography purification, a complex of Rs BcsA-BcsB and the engineered BcsG-Wssf chimera are isolated. This demonstrates that BcsA interacts with the engineered enzyme chimera. In vivo, Wssf likely uses acetyl-coenzyme A as a substrate. In vitro, acetyl-transfer can be achieved either with acetyl-CoA or acetyl-para nitro-phenol (pNP) as acetyl-group donors. To test the ability of the enzyme chimera to acetylate de novo synthesized cellulose, cellulose is produced by BcsA in vitro upon addition of UDP-glucose. At the same time, either acetyl- CoA or acetyl-pNP is added for acetylation. Following biosynthesis, cellulose was precipitated with 80% ethanol and digested with a cellulase, prior to mass spectrometry analysis. For both ZIMMER-PETN (03018-02) / / 1036.375WO1 acetyl-group donors, mass spectrometry identified acetylated cello-oligosaccharides ranging in length from one to six glucosyl units (Fig. 9). This demonstrates that, similar to pEtN modification by wild type BcsG, positioning cellulose acetyltransferase activity close to the BcsA cellulose synthase via BcsG’s N-terminal domain, can be employed to modify cellulose with acetyl groups. While this work with bacterial cellulose synthase is a proof-of-principle, similar efforts with other bacterial or plant cellulose synthases and engineered modifying enzymes can be used to generate novel cellulosic biomaterials. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Exemplary methods and materials are described herein, although methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention. Those skilled in the art will recognize, or be able to ascertain, using no more than routine experimentation, numerous equivalents to the specific substances and procedures described herein. Such equivalents are considered to be within the scope of this invention. All publications, patents, and patent applications, Genbank sequences, websites and other published materials referred to throughout the disclosure herein are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application, Genbank sequences, websites and other published materials was specifically and individually indicated to be incorporated by reference. In the event that the definition of a term incorporated by reference conflicts with a term defined herein, this specification shall control.

Claims

ZIMMER-PETN (03018-02) / / 1036.375WO1 WHAT IS CLAIMED IS:

1. A chimeric BcsA enzyme comprising an N-terminal domain (NTD) comprising SEQ ID NO: 13 or 95% identity thereto operably linked to a heterologous cellulose synthase catalytic domain.

2. The chimeric enzyme of claim 1, wherein the chimeric BcsA enzyme further comprising a C-terminal domain comprising SEQ ID NO: 10 or 95% identity thereto.

3. A chimeric BcsG enzyme comprising a heterologous catalytic domain, wherein the chimeric enzyme does not have a phosphoethanolamine catalytic domain.

4. The chimeric enzyme of claim 3, wherein heterologous catalytic domain is an acetyltransferase domain.

5. A nucleic acid sequence coding for the chimeric enzyme of any one of claims 1 to 4.

6. A method to modify cellulose, comprising: providing a host cell capable of synthesizing cellulose; introducing into the host cell an expression construct encoding a chimeric BcsA enzyme comprising an N-terminal domain SEQ ID NO: 13 or 95% identity thereto optionally operably linked to a heterologous cellulose synthase catalytic domain, a BcsB protein, and a BcsG enzyme; culturing the host cell under conditions sufficient to express the chimeric BcsA enzyme, BcsB protein, and BcsG enzyme, whereby the chimeric BcsA enzyme recruits a plurality of BcsG enzymes to facilitate synthesis and translocation of a cellulose polymer; and optionally recovering the cellulose polymer that has been modified.

7. The method of claim 6, wherein the chimeric BcsA enzyme further comprises a C- terminal domain comprising SEQ ID NO: 10 or 95% identity thereto.

8. The method of claim 6, wherein a nucleic acid sequence coding for SEQ ID NO: 13, SEQ ID NO: 10 or 95% identity thereto are operably integrated into a native cellulose synthase protein of the host cell.ZIMMER-PETN (03018-02) / / 1036.375WO1 9. The method of any one of claims 6 to 8, wherein the modification is phosphoethanolamine groups at C6 hydroxyl positions of the cellulose polymer.

10. The method of any one of claims 6 to 8, wherein a phosphoethanolamine transferase catalytic domain of the BscG protein is replaced with a heterologous catalytic domain.

11. The method of claim 10, wherein the heterologous catalytic domain is an acetyltransferase domain.

12. The method of any one of claims 6 to 9, wherein the expression construct is vector.

13. The method of claim 12, wherein the BcsA enzyme, BcsB protein, and BcsG enzyme are coded for by one or more expression constructs.

14. The method of any one of claims 6 to 13, wherein the host cell is a cellulose producing bacterial cell, a plant cell, an algae cell, or an animal cell.

15. A method to modify cellulose, comprising: providing a host cell capable of synthesizing cellulose; introducing into the host cell an expression construct encoding a chimeric BcsG enzyme comprising a heterologous catalytic domain, wherein the chimeric enzyme does not have a phosphoethanolamine catalytic domain, and optionally a BcsA enzyme and / or BcsB protein, culturing the host cell under conditions sufficient to express the chimeric BcsG enzyme, and optionally BcsA enzyme, and / or BcsB protein; and optionally recovering the cellulose polymer that has been modified.

16. The method of claim 15, wherein the BcsA enzyme is a chimeric enzyme comprising an N-terminal domain SEQ ID NO: 13 or 95% identity thereto operably linked to a heterologous cellulose synthase catalytic domain.

17. The method of claim 15 or 16, wherein the BcsA enzyme further comprises a C- terminal domain comprising SEQ ID NO: 10 or 95% identity thereto.ZIMMER-PETN (03018-02) / / 1036.375WO1 18. The method of claim 15, wherein a nucleic acid sequence coding for SEQ ID NO: 13, SEQ ID NO: 10 or 95% identity thereto are operably integrated into a native cellulose synthase protein of the host cell.

19. The method of any one of claims 15 to 18, wherein the heterologous catalytic domain of BcsG is an acetyltransferase domain.

20. The method of any one of claims 15 to 19, wherein the expression construct is vector.

21. The method of claim 20, wherein the BcsA enzyme, BcsB protein, and BcsG enzyme are coded for by one more expression constructs.

22. The method of any one of claims 15 to 21, wherein the host cell is a bacterial cell, a plant cell, an algae cell, or an animal cell.