Expression of nitrogenase polypeptides in plant cells
Patent Information
- Application Number
- JP2024117844
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-02-06
- Filing Date
- 2024-07-23
- Publication Date
- 2025-06-09
AI Technical Summary
The production of nitrogenase polypeptides in plant cells is challenging due to technical difficulties such as oxygen sensitivity, requirement of specific biochemical conditions, and the need for large amounts of ATP, reductant, and cofactors, making it difficult to achieve functional nitrogenase complexes for biological nitrogen fixation.
The invention involves expressing nitrogenase components as fusion proteins in plant mitochondria, using mitochondrial targeting peptides (MTP) with specific C-termini, oligopeptide linkers, and N-termini of NifD, NifK, NifE, and NifN polypeptides, ensuring equal stoichiometric ratios and functional assembly through matrix processing proteases.
This approach allows for the precise targeting and functional expression of nitrogenase components in plant mitochondria, enhancing nitrogen fixation capabilities and reducing dependence on industrially produced nitrogen fertilizers.
Smart Images

Figure 00000134_0000 
Figure 00000134_0001 
Figure 00000135_0000
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The present invention relates to methods and means for producing nitrogenase polypeptides in the mitochondria of plant cells. [Background technology]
[0002] Background of the Invention Nitrogen-fixing bacteria produce ammonia from N gas through biological nitrogen fixation (BNF), catalyzed by the enzyme complex nitrogenase. Nevertheless, the demands of modern agriculture far exceed this fixed nitrogen source, resulting in the extensive use of industrially produced nitrogen fertilizers in agriculture (Smil, 2002). However, both fertilizer production and application are sources of pollution (Good and Beatty, 2011) and are considered unsustainable (Rockstrom et al., 2009). The majority of fertilizer applied worldwide is not absorbed by crops (Cui et al., 2013; de Bruijn, 2015), leading to fertilizer runoff, weed growth, and eutrophication of waterways (Good and Beatty, 2011). The resulting algal blooms reduce oxygen levels and cause local and reef-wide offshore environmental damage (De'ath et al., 2012; Glibert et al., 2014; Sutton et al., 2008). Furthermore, over-fertilization is a problem in many developed countries, and its availability limits crop yields in certain regions (Mueller et al., 2012). Production of fertilizer itself requires significant energy input, costing an estimated 100 USD billion per year.
[0003] Clearly, strategies to reduce dependence on industrially produced nitrogen are needed. To this end, the concept of genetically engineered plants capable of biological nitrogen fixation has long attracted considerable attention (Merrick and Dixon, 1984) and has been the focus of recent reviews (e Bruijn, 2015; Oldroyd and Dixon, 2014). Promising approaches include i) the extension of diazotrophic symbiosis from legumes to cereals (Santi et al., 2013), ii) reengineering endosymbiotic microorganisms to enable nitrogen fixation (Geddes et al., 2015), and iii) genetic manipulation of nitrogenase within plant cells (Curatti and Rubio, 2014). All of these approaches remain ambitious and speculative due to technical challenges.
[0004] Nitrogenase, the enzyme complex that enables biological nitrogen fixation in diazotrophic bacteria, requires a multigene assembly pathway for its biosynthesis and function and has been extensively reviewed (Hu and Ribbe, 2013; Rubio and Ludden, 2008; Seefeldt et al., 2009). Components of a canonical iron-molybdenum nitrogenase include the catalytic proteins, designated NifD and NifK, and the electron donor NifH. Approximately 12 other proteins, specifically NifM, NifS, NifU, NifE, NifN, NifX, NifV, NifJ, NifY, NifF, NifZ, and NifQ, are involved in nitrogenase assembly in diazotrophic bacteria, including complex maturation, scaffolding, and cofactor insertion. Genetic lesions, complementation assays between diazotrophs and nondiazotrophs, and developmental analyses (Dos Santos et al., 2012; Temme et al., 2012; Wang et al., 2013) have led to the identification of a subset of Nif proteins (NifD, NifK, NifB, NifE, and NifN) as core components, while other components are considered accessory because they appear to be required for optimized activity. Specific biochemical conditions are also required for nitrogenase assembly and function. Foremost among these is that nitrogenase is highly sensitive to oxygen (Robson and Postgate, 1980). Furthermore, large amounts of ATP, reductants, readily available Fe, Mo, S-adenosylmethionine, and homocitrate are required for the biogenesis and function of the metalloprotein catalytic center (Hu and Ribbe, 2013; Rubio and Ludden, 2008). All of these factors contribute to the technical difficulty of producing a functional nitrogenase complex in plant cells. Summary of the Invention
[0005] Given the difficulties observed in producing NifD in plant cells, the inventors investigated the importance of having equal amounts of NifD and NifK in plant cells. Detailed analysis of the polypeptides showed that they can be expressed as fusion proteins and retain function, as well as the importance of the NifK component having a wild-type C-terminus. Thus, in one aspect, the present invention provides a method for the production of NifD in plant cells by combining mitochondria with the following: (i) a mitochondrial targeting peptide (MTP) having a C-terminus; (ii) a NifD polypeptide (ND) having an N-terminus and a C-terminus; (iii) an oligopeptide linker, and (iv) a NifK polypeptide having an N-terminus (NK); Including, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the ND, and wherein the linker is translationally fused to the C-terminus of the ND and the N-terminus of the NK; and a plant cell comprising the fusion polypeptide.
[0006] In one embodiment, the C-terminus of the fusion polypeptide is the C-terminus of NK. In a preferred embodiment, the C-terminus of NK is the same as the C-terminus of the wild-type NifK polypeptide, i.e., NK does not have any artificially added C-terminal extension. Having a wild-type C-terminus is important for maintaining NifK activity.
[0007] In another embodiment, the linker is of sufficient length to allow binding of the ND and NK in a functional form in a plant or bacterial cell. In one embodiment, the linker is 8 to 50 amino acids in length. Preferably, the linker is at least about 20 amino acids, at least about 25 amino acids, or at least about 30 amino acids in length. More preferably, the linker is 25 to 35 amino acids in length. Most preferably, the linker is about 30 amino acids in length. In this context, "about 30" means 27, 28, 29, 30, 31, 32, or 33 amino acids.
[0008] In certain embodiments, the NifD and NifK polypeptides in the fusion polypeptide have the same function as the NifD and NifK polypeptides when present as two separate polypeptides. Preferably, the biochemical activity of the fusion polypeptide is the same as that of the wild-type NifD and wild-type NifK polypeptides.
[0009] The present inventors have also determined the importance of having equal amounts of NifE and NifN in plant cells. Detailed analysis of the polypeptides showed that they can be expressed as fusion proteins and retain function, and that the presence of a linker is required for the Nif polypeptide to be functional. Thus, in a further aspect, the present invention provides a method for the production of NifE and NifN in plant cells by combining mitochondria with the following: (i) a mitochondrial targeting peptide (MTP) having a C-terminus; (ii) a NifE polypeptide (NE) having an N-terminus and a C-terminus; (iii) an oligopeptide linker, and (iv) a NifN polypeptide having an N-terminus (NN) Including, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NE, and wherein the linker is translationally fused to the C-terminus of the NE and the N-terminus of the NN. and a plant cell comprising the fusion polypeptide.
[0010] In some embodiments of the above aspects, the linker is of sufficient length to allow the NE and NN to bind in a functional form in a plant or bacterial cell. For example, in some embodiments, the linker is at least about 70 Å, at least about 100 Å, about 70 Å to about 150 Å, or about 100 Å to about 120 Å, or about 100 Å, or about 104 Å in length. In some embodiments, the oligopeptide linker is at least about 20 amino acids, at least about 30 amino acids, at least about 40 amino acids, or about 20 amino acids to about 70 amino acids, about 30 amino acids to about 70 amino acids, about 30 amino acids to about 60 amino acids, about 30 amino acids to about 50 amino acids, or about 25 amino acids, about 30 amino acids, about 35 amino acids, about 40 amino acids, about 45 amino acids, about 46 amino acids, about 50 amino acids, or about 55 amino acids in length.
[0011] In certain embodiments, the NifE and NifN polypeptides in the fusion polypeptide have the same function as the NifE and NifN polypeptides when present as two separate polypeptides. Preferably, the biochemical activity of the fusion polypeptide is the same as that of the wild-type NifE and NifN polypeptides.
[0012] In one embodiment, MTP comprises a protease cleavage site for a matrix processing protease (MPP) such that the fusion polypeptide can be cleaved by MPP to generate a processed fusion polypeptide (CF) comprising an N-terminal cleavage peptide and (ii)-(iv). In one embodiment, the CF comprises some, but not all, of the C-terminal amino acids of MTP, preferably 5-45 of the C-terminal amino acids of MTP. In one embodiment, cleavage by MPP removes 5-50 amino acids from the fusion polypeptide to generate the CF. In a preferred embodiment, the CF comprises about 5 to about 11 amino acid residues from the C-terminus of MTP, e.g., 6 or 7, 7 or 8, 8 or 9, 9 or 10, or 11 or 12 amino acids from the C-terminus of MTP.
[0013] In some embodiments of the above two aspects, the plant cell further comprises one or more NF fusion polypeptides (NFs), each NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, and wherein each MTP is independently the same or different, and each NP is independently the same or different, and wherein the mitochondria comprise one or more NFs and / or their processed products (CFs), wherein each CF, if present, is generated by cleavage of the corresponding NF within its MTP. Preferably, at least one of the NF polypeptides is NifH. In preferred embodiments, each CF independently comprises about 5 to about 11 amino acid residues from the C-terminus of the MTP, e.g., about 6 or 7, 7 or 8, 8 or 9, 9 or 10, or 11 or 12 amino acids from the C-terminus of the MTP.
[0014] In certain embodiments, the exogenous polynucleotide(s) encoding the fusion polypeptide(s) are integrated into the genome of the cell.
[0015] We also demonstrated that all 16 biosynthetic and functional nitrogenase (Nif) proteins of Klebsiella pneumoniae could be individually expressed as mitochondrial targeting peptide (MTP)-Nif fusions in plant cells, such as Nicotiana benthamiana leaves. We demonstrated that these fusions were precisely targeted to the mitochondrial matrix (MM), a subcellular location that potentially supports biochemical and genetic characterization of nitrogenase function. NifJ, NifH, NifD, NifK, NifY, NifE, NifN, NifX, NifU, NifS, NifV, NifM, NifF, NifB, and NifQ polypeptides were detectable by Western blot analysis. However, NifD (the core component involved in catalysis) was the least abundant. The crystal structure of the nitrogenase NifD-NifK heterodimer was used to design NifD-NifK translational fusion proteins with improved expression levels. Finally, four Nif fusion polypeptides (NifB, NifS, NifH, and NifY) were successfully coexpressed, demonstrating that multiple components of nitrogenase can be targeted to mitochondria. These results establish the feasibility of reconstituting nitrogenase components within the intracellular environment to support the reduction of nitrogen gas to ammonia.
[0016] Of all Nif proteins, NifD, an essential component of the nitrogenase catalytic complex, was the most difficult to express. Low levels of NifD protein contrasted with high levels of NifD RNA, suggesting that translation rate or protein stability limited NifD protein abundance. Given the critical importance of NifD in catalysis—its requirement that it be highly expressed in bacteria, preferably at an equimolar ratio with NifK (Poza-Carrion et al., 2014)—we fused these two crucial components together and found that NifD abundance could be increased through this strategy. This NifD-NifK fusion also had the advantage of being linked in a single expression cassette and allowing translation of both subunits in an ideal 1:1 ratio, mimicking the stoichiometry of the native heterotetramer. Furthermore, the linker itself was designed to allow sufficient flexibility for the two subunits to form the precise α2β2 heterotetrameric structure required for catalysis. We predicted that the NifD-NifK fusion polypeptide would functionally replace at least the individual NifD and NifK expression with greater efficacy than previously demonstrated.
[0017] Thus, in one aspect, the present invention provides a plant cell comprising a mitochondrion and an exogenous polynucleotide encoding a NifD fusion polypeptide (NDF), wherein the NDF comprises (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a NifD polypeptide (ND) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the ND, wherein the mitochondrion comprises the NDF and / or its processed NifD product (CDF), and wherein the CDF, if present, is produced by cleavage of the NDF within the MTP.
[0018] In one embodiment, the MTP contains a protease cleavage site for a matrix processing protease (MPP) such that the NDF can be cleaved by the MPP to generate an N-terminal cleavage peptide and a CDF. In certain embodiments, the CDF contains about 5 to about 11 amino acid residues from the C-terminus of the MTP, e.g., 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or about 11 or 12 amino acids from the C-terminus of the MTP.
[0019] In one aspect, the invention provides a plant cell comprising one or more NF fusion polypeptides (NFs), each NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the NPs are selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, wherein each MTP is independently the same or different, each NP is independently the same or different, and wherein the mitochondria comprise one or more NFs and / or their processed products (CFs), and wherein each CF, if present, is generated by cleavage of the corresponding NF within the MTP. In certain embodiments, the NF polypeptides further comprise one or more, or preferably all, of the Nif polypeptides selected from the group consisting of NifD, NifH, and NifK. In a preferred embodiment, the C-terminus of NifK and / or NifN is the same as the C-terminus of the wild-type NifK polypeptide or wild-type NifN polypeptide, respectively, i.e., NifK and / or NifN preferably both do not have any artificially added C-terminal extensions.
[0020] In a preferred embodiment, each CF independently comprises about 5 to about 11 amino acid residues from the C-terminus of MTP, for example, 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or 11 or 12 amino acids from the C-terminus of MTP.
[0021] In another aspect, the present invention provides a plant cell comprising a mitochondrion and a first exogenous polynucleotide encoding a first Nif fusion polypeptide (NF), and a second exogenous polynucleotide encoding a second NF, wherein each NF comprises (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, wherein each MTP is independently the same or different, and each NP is independently the same or different, and wherein the mitochondrion comprises (a) the first NF and / or its processed product (first CF) and (b) the second NF and / or its processed product (second CF), and wherein each CF, if present, is produced by cleavage of the corresponding NF within the MTP. In a preferred embodiment, the first and second NFs are selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY and NifZ. In a preferred embodiment, the C-terminus of NifK (if present) and / or NifN is the same as the C-terminus of a wild-type NifK polypeptide or a wild-type NifN polypeptide, respectively, i.e., NifK and / or NifN are preferably both free of any artificially added C-terminal extensions.
[0022] In a preferred embodiment, each CF independently comprises about 5 to about 11 amino acid residues from the C-terminus of MTP, for example, 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or 11 or 12 amino acids from the C-terminus of MTP.
[0023] In one embodiment, the plant cell further comprises one or more exogenous polynucleotides encoding one or more NFs, each NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, wherein each MTP is independently the same or different, and each NP is independently the same or different, and wherein the mitochondria comprise one or more NFs and / or their processed products (CFs), and wherein each CF, if present, is generated by cleavage of the corresponding NF within the MTP. In a preferred embodiment, the one or more NFs are selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ. In a preferred embodiment, the C-terminus of NifK (if present) and / or NifN is the same as the C-terminus of the wild-type NifK polypeptide or wild-type NifN polypeptide, respectively, i.e., NifK and / or NifN preferably both do not have any artificially added C-terminal extensions.
[0024] In one or further embodiments, at least one MTP, two or more MTPs, or each MTP comprises a protease cleavage site for a matrix processing protease (MPP) such that each NF comprising a protease cleavage site for the MPP can be cleaved by the MPP to generate an N-terminal truncated peptide and a corresponding CF. In preferred embodiments, each CF independently comprises about 5 to about 11 amino acid residues from the C-terminus of the MTP, e.g., 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or about 11 or 12 amino acids from the C-terminus of the MTP.
[0025] In a further aspect, the present invention provides a method for treating a mitochondria comprising: (a) a first exogenous polynucleotide encoding a NifD fusion polypeptide (NDF), the NDF comprising (i) a first mitochondrial targeting peptide (MTP1) having a C-terminus and (ii) a NifD polypeptide (ND) having an N-terminus, wherein the C-terminus of the MTP1 is translationally fused to the N-terminus of the ND; (b) a second exogenous polynucleotide encoding a NifH fusion polypeptide (NHF), the NHF comprising (i) a second mitochondrial targeting peptide (MTP2) having a C-terminus and (ii) a NifH polypeptide (NH) having an N-terminus, wherein the C-terminus of the MTP2 is translationally fused to the N-terminus of the NH; and (c) a third exogenous polynucleotide encoding a NifK fusion polypeptide (NKF), the NKF comprising (i) a third mitochondrial targeting peptide (MTP3) having a C-terminus and (ii) a NifK polypeptide (NK) having an N-terminus, wherein the C-terminus of the third MTP is translationally fused to the N-terminus of the NK; and wherein each of MTP1, MTP2, and MTP3 is independently the same or different, and
[0010] Provided herein is a plant cell comprising a first exogenous polynucleotide and a level of an NDF and / or its processed product (CDF) that is superior to the level of the NDF and / or CDF in a corresponding plant cell that does not have the second and third exogenous polynucleotides, and wherein the CDF, if present, is produced by cleavage of the NDF within MTP1. In a preferred embodiment, the plant cell further comprises one or more exogenous polynucleotides encoding one or more NFs selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ. In a preferred embodiment, the C-terminus of NifK and / or NifN (if present) is the same as the C-terminus of the wild-type NifK or NifN polypeptide, respectively, i.e., NifK and / or NifN preferably both do not have any artificially added C-terminal extensions.
[0026] In one embodiment, the CDF is present in mitochondria. In further embodiments, the mitochondria further comprise processed products of one or both of NHF (CHF) and NKF (CKF), where the CHF and CKF, if present, are produced by cleavage of NHF and NKF within MTP2 and MTP3, respectively. In preferred embodiments, CHF and / or CKF independently comprise about 5 to about 11 amino acid residues from the C-terminus of MTP, e.g., 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or about 11 or 12 amino acids from the C-terminus of MTP.
[0027] In further embodiments, the plant cell is characterized by one, more, or all of the following: (i) MTP1 contains a protease cleavage site for MPP such that NDF can be cleaved by MPP to create an N-terminal truncated peptide and CDF; (ii) MTP2 contains a protease cleavage site for MPP such that NHF can be cleaved by MPP to create an N-terminal truncated peptide and a processed NifH product (CHF); and (iii) MTP3 contains a protease cleavage site for MPP such that NKF can be cleaved by MPP to create an N-terminal truncated peptide and a processed NifK product (CKF); and wherein the mitochondria contain one, more, or all of CDF, CHF, and CKF. In a preferred embodiment, each of CDF, CHF and / or CKF independently comprises about 5 to about 11 amino acid residues from the C-terminus of MTP, for example, 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or 11 or 12 amino acids from the C-terminus of MTP.
[0028] In a further aspect, the present invention provides a plant cell comprising mitochondria and a first exogenous polynucleotide encoding a NifD polypeptide (ND) and a second exogenous polynucleotide encoding a NifK polypeptide (NK), wherein either or both of ND and NK are translationally fused at their N-terminus to a C-terminal mitochondrial targeting peptide (MTP), wherein when both ND and NK are translationally fused to the MTP, each MTP is independently the same or different, and wherein the second exogenous polynucleotide is covalently or not covalently linked to the first exogenous polynucleotide, and wherein the plant cell comprises levels of ND and NK that are approximately the same.
[0029] In a preferred embodiment, the C-terminus of NifK is the same as the C-terminus of the wild-type NifK polypeptide, i.e., NifK does not have any artificially added C-terminal extensions. In a preferred embodiment, cleavage of ND and / or NK results in a cleavage product polypeptide having about 5 to about 11 amino acid residues from the C-terminus of MTP, e.g., 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or even 11 or 12 amino acids from the C-terminus of MTP.
[0030] In one embodiment, the levels of ND and NK comprise the MTP-Nif fusion and its processed products.
[0031] In a further aspect, the present invention provides a plant cell comprising a mitochondrion and an exogenous polynucleotide encoding a NifH fusion polypeptide (NHF), wherein the NHF comprises (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a NifH polypeptide (NH) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NH, wherein the mitochondrion comprises the NHF and / or its processed NifH product (CHF), and wherein the CHF, if present, is produced by cleavage of the NHF within the MTP.
[0032] In one embodiment, MTP contains a protease cleavage site for a matrix processing protease (MPP) such that NHF can be cleaved by MPP to generate an N-terminal truncated peptide and CHF. In certain embodiments, CHF contains about 5 to about 11 amino acid residues from the C-terminus of MTP, e.g., 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or about 11 or 12 amino acids from the C-terminus of MTP.
[0033] In a further aspect, the present invention provides a plant cell comprising a mitochondrion and an exogenous polynucleotide encoding a Nif fusion polypeptide (NF), wherein the NF comprises (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, wherein the mitochondrion comprises the NF and / or its processed NP product (CF), and wherein the CF, if present, is produced by cleavage of the NF within the MTP, and wherein the NF and / or CF, preferably both the NF and CF, have the native C-terminus associated with the corresponding wild-type Nif polypeptide.
[0034] In a preferred embodiment, the NP is selected from the group consisting of a NifD polypeptide, a NifK polypeptide, and a NifN polypeptide, preferably both NifK and NifN.
[0035] In a further aspect, the invention provides a method for the preparation of a mitochondrion comprising: combining a mitochondrion with a first exogenous polynucleotide encoding a NifD fusion polypeptide (NDF), the NDF comprising (i) a first mitochondrial targeting peptide (MTP1) having a C-terminus and (ii) a NifD polypeptide (ND) having an N-terminus, wherein the C-terminus of the MTP1 is translationally fused to the N-terminus of the ND; and
[0036] (a) a second exogenous polynucleotide encoding a NifH fusion polypeptide (NHF), the NHF comprising (i) a second mitochondrial targeting peptide (MTP2) having a C-terminus and (ii) a NifH polypeptide (NH) having an N-terminus, wherein the C-terminus of the MTP2 is translationally fused to the N-terminus of the NH; and (b) a third exogenous polynucleotide encoding a NifK fusion polypeptide (NKF), the NKF comprising (i) a third mitochondrial targeting peptide (MTP3) having a C-terminus and (ii) a NifK polypeptide (NK) having an N-terminus, wherein the C-terminus of MTP3 is translationally fused to the N-terminus of the NK, preferably wherein the NKF has a native C-terminus associated with a wild-type NifK polypeptide; and (c) a fourth polynucleotide encoding a Nif fusion polypeptide (NF), excluding NDF, NHF, and NKF, the NF comprising (i) a fourth mitochondrial targeting peptide (MTP4) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP4 is translationally fused to the N-terminus of the NP; A plant cell comprising one, more, or all of:
[0037]
[0010] Provided herein is a plant cell wherein each of MTP1, MTP2, MTP3, and MTP4 is independently the same or different. In one embodiment, the fourth exogenous polynucleotide encodes one or more NPs selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ. In a preferred embodiment, the C-terminus of NifK and / or NifN (if present) is the same as the C-terminus of the wild-type NifK or NifN polypeptide, respectively, i.e., NifK and / or NifN preferably both do not have any artificially added C-terminal extensions.
[0038] In one embodiment, each NF is capable of being imported into a mitochondrion. In one embodiment of any of the above aspects, the mitochondrion comprises one, at least two, at least three, at least four, or all of the Nif polypeptides selected from the group consisting of: (i) NifD, NifH, NifK, NifB, NifE and NifN, or (ii) NifD, NifH, NifK or NifS.
[0039] In a further embodiment, one, more, or all Nif fusion polypeptides are cleaved within a matrix processing protease (MPP) in the plant cell, preferably wherein at least the NifD fusion polypeptide or the NifH fusion polypeptide is cleaved by the MPP.
[0040] In one or further embodiments, one or more polypeptides selected from the group consisting of ND, NDF, CDF, NH, NHF, CHF, NK, NKF, CKF, NF, CFNB, NE, NN and NS are capable of associating with at least one, preferably at least two, at least three, at least four, at least five, at least six or at least seven other Nif polypeptides to form a Nif protein complex (NPC), preferably such that the NPC has nitrogenase activity.
[0041] In one or further embodiments, CDF, CKF, CHF, or CF is larger in size than ND, NK, NH, or NP, respectively, preferably by 5 to 50 amino acid residues, and / or wherein CDF, CKF, CHF, or CF is smaller in size than NDF, NKF, NHF, or NF, respectively, preferably by 5 to 45 amino acid residues. In preferred embodiments, each CDF, CKF, CHF, or CF independently comprises about 5 to about 11 amino acid residues from the C-terminus of MTP, e.g., 6 or 7 amino acids, 7 or 8 amino acids, 8 or 9 amino acids, 9 or 10 amino acids, or 11 or 12 amino acids from the C-terminus of MTP.
[0042] In one or further embodiments, the MTP comprises at least 10 amino acids, preferably 10 to 80 amino acids. In one or further embodiments, at least one, more than one, or all of the MTPs comprise a mitochondrial protein precursor MTP, or a variant thereof, preferably a plant MTP. In one or further embodiments, the exogenous polynucleotide(s) are integrated into the genome of the cell.
[0043] In one or further embodiments, the cell is not a protoplast. In one or further embodiments, the cell is a cell other than an Arabidopsis thaliana protoplast. In a further aspect, the present invention provides a transgenic plant comprising a cell according to the present invention, wherein said transgenic plant is transgenic for exogenous polynucleotide(s) and / or said transgenic plant is transgenic for exogenous polynucleotide(s) encoding fusion polypeptide(s).
[0044] In one embodiment, one, more, or all of the exogenous polynucleotides are expressed in the roots of the plant, and preferably at greater levels in the roots of the plant than in the leaves of the plant. In one or further embodiments, the transgenic plant has an altered phenotype relative to a corresponding wild-type plant, which is increased yield, biomass, growth rate, vitality, nitrogen acquisition from biological nitrogen fixation, nitrogen use efficiency, abiotic stress tolerance, and / or tolerance to nutrient deficiency relative to a corresponding wild-type plant.
[0045] In an alternative embodiment, the transgenic plants have the same growth rate and / or phenotype as the corresponding wild-type plants. In one or further embodiments, the transgenic plant is a cereal plant such as, for example, wheat, rice, corn, triticale, oat, or barley, preferably wheat.
[0046] In one embodiment, the transgenic plant is homozygous or heterozygous for the exogenous polynucleotide(s). In one or further embodiments, the transgenic plants are grown in a field. In a further aspect, the present invention provides a population of at least 100 plants according to the present invention growing in a field.
[0047] In a further aspect, the present invention provides a NifD fusion polypeptide (NDF), which comprises (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a NifD polypeptide (ND) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the ND.
[0048] In a further aspect, the present invention provides a NifH fusion polypeptide (NHF), the NHF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a NifH polypeptide (NH) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NH.
[0049] In a further aspect, the present invention provides a Nif fusion polypeptide (NF), the NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, and wherein the NF has a native C-terminus associated with the corresponding wild-type Nif polypeptide.
[0050] In a preferred embodiment, the NP is selected from the group consisting of a NifD polypeptide, a NifK polypeptide, and a NifN polypeptide.
[0051] In a further aspect, the present invention provides a combination of Nif fusion polypeptides (NFs) comprising a NifK fusion polypeptide (NKF) and a NifN fusion polypeptide (NNF), each NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, wherein the MTPs of each NF are independently the same or different, and wherein each of the NKFs and NNFs has a native C-terminus associated with the corresponding wild-type Nif polypeptide.
[0052] In a further aspect, the present invention provides a method for treating a cancer cell comprising: (i) a mitochondrial targeting peptide (MTP) having a C-terminus; (ii) a NifD polypeptide (ND) having an N-terminus and a C-terminus; (iii) an oligopeptide linker, and (iv) a NifK polypeptide having an N-terminus (NK); Including, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the ND, and wherein the linker is translationally fused to the C-terminus of the ND and the N-terminus of the NK; Fusion polypeptides are provided.
[0053] In one embodiment, the linker is of sufficient length to allow binding between the ND and NK in a functional form in a plant or bacterial cell. In some embodiments, the linker is 8 to 50 amino acids in length. Preferably, the linker is at least about 20 amino acids, at least about 25 amino acids, or at least about 30 amino acids in length.
[0054] In one embodiment, the C-terminus of the fusion polypeptide is the C-terminus of NK. In another aspect, the present invention provides a method for producing a composition comprising the steps of: (i) a mitochondrial targeting peptide (MTP) having a C-terminus; (ii) a NifE polypeptide (NE) having an N-terminus and a C-terminus; (iii) an oligopeptide linker, and (iv) a NifN polypeptide having an N-terminus (NN) Including, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NE, and wherein the linker is translationally fused to the C-terminus of the NE and the N-terminus of the NN. Fusion polypeptides are provided.
[0055] In one embodiment, the linker is of sufficient length to allow the NE and NN to bind in a functional form in a plant or bacterial cell. For example, in certain embodiments, the linker is at least about 70 Å, at least about 100 Å, about 70 Å to about 150 Å, or about 100 Å to about 120 Å, or about 100 Å, or about 104 Å in length. In further embodiments of the above aspect, the oligopeptide linker is at least about 20 amino acids, at least about 30 amino acids, at least about 40 amino acids, or about 20 amino acids to about 70 amino acids, about 30 amino acids to about 70 amino acids, about 30 amino acids to about 60 amino acids, about 30 amino acids to about 50 amino acids, or about 25 amino acids, about 30 amino acids, about 35 amino acids, about 40 amino acids, about 45 amino acids, about 46 amino acids, about 50 amino acids, or about 55 amino acids in length.
[0056] In one embodiment, the fusion polypeptide(s) of any of the above aspects can be cleaved by a matrix processing protease (MPP) to generate one or more processed Nif polypeptide products.
[0057] In one further embodiment, the fusion polypeptide(s) are present in a plant or bacterial cell, preferably in the mitochondria of a plant cell.
[0058] In a further aspect, the present invention provides a processed Nif polypeptide produced from a fusion polypeptide(s) according to the invention by cleavage within the MTP of said fusion polypeptide(s).
[0059] In one embodiment, the processed Nif polypeptide was produced by cleavage of the fusion polypeptide by MPP in the mitochondrial matrix of the plant cell.
[0060] In one or further embodiments, the processed Nif polypeptide is present in the plant or bacterial cell, preferably in the plant mitochondria, more preferably in the mitochondrial matrix (MM) of the plant cell mitochondria.
[0061] In one or further embodiments, the fusion polypeptide or processed Nif polypeptide has the same biochemical activity as the corresponding wild-type Nif polypeptide, preferably having about the same level of biochemical activity as the corresponding wild-type Nif polypeptide, optionally as measured in a bacterial cell.
[0062] In one or further embodiments, the fusion polypeptide(s) or processed Nif polypeptide can associate with at least one, preferably at least two, at least three, at least four, at least five, at least six, or at least seven other Nif polypeptides to form a Nif protein complex (NPC), preferably such that the NPC has nitrogenase activity.
[0063] In a further aspect, the present invention provides polynucleotides encoding one, more or all of the fusion polypeptides or processed Nif polypeptides according to the invention.
[0064] In one embodiment, the NF is a NifD fusion polypeptide (NDF) and the polynucleotide is: (i) a NifH fusion polypeptide (NHF), the NHF comprising a second mitochondrial targeting peptide (MTP2) having a C-terminus and a NifH polypeptide (NH) having an N-terminus, such that the C-terminus of the MTP2 is translationally fused to the N-terminus of the NH; (ii) a NifK fusion polypeptide (NKF), the NKF comprising a third mitochondrial targeting peptide (MTP3) having a C-terminus and a NifK polypeptide (NK) having an N-terminus, such that the C-terminus of the MTP3 is translationally fused to the N-terminus of the NK; and
[0065] (iii) a Nif fusion polypeptide (NF) other than NDF, NHF, and NKF, the NF comprising a fourth mitochondrial targeting peptide (MTP4) having a C-terminus and a Nif polypeptide (NP) having an N-terminus, such that the C-terminus of the MTP4 is translationally fused to the N-terminus of the NP; and further comprising one or more nucleotide sequences encoding one, more, or all of: wherein each of MTP1, MTP2, MTP3 and MTP4 is independently the same or different.
[0066] In one or further embodiments, the polynucleotide is codon-modified for expression in a plant cell. In one or further embodiments, the polynucleotide comprises a promoter operably linked to the polynucleotide or each sequence therein encoding a fusion polypeptide, and / or a translational regulatory element operably linked to the polynucleotide.
[0067] In one embodiment, the promoter directs expression of the polynucleotide in the roots, leaves and / or stems of the plant, preferably the promoter directs expression of the polynucleotide in one, more, or all of the roots, leaves, or stems of the plant associated with the seed of the plant.
[0068] In one or further embodiments, the polynucleotide is present in a plant cell or a bacterial cell, and is preferably integrated into the genome of the plant cell. In a further aspect, the present invention provides a nucleic acid construct wherein a first Nif comprises a first polynucleotide encoding a fusion polypeptide (NF) and a second polynucleotide encoding a second NF, each NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, and wherein each MTP is independently the same or different.
[0069] In one embodiment, the nucleic acid construct further comprises one or more exogenous polynucleotides encoding one or more NFs, each NF comprising (i) a mitochondrial targeting peptide (MTP) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP is translationally fused to the N-terminus of the NP, and wherein each MTP is independently the same or different, and each NP is independently the same or different.
[0070] In a further aspect, the invention provides a method for the preparation of a NifD fusion polypeptide (NDF) comprising: (i) a first mitochondrial targeting peptide (MTP1) having a C-terminus; and (ii) a NifD polypeptide (ND) having an N-terminus, wherein the C-terminus of the MTP1 is translationally fused to the N-terminus of the ND;
[0071] (a) a second polynucleotide encoding a NifH fusion polypeptide (NHF), the NHF comprising (i) a second mitochondrial targeting peptide (MTP2) having a C-terminus and (ii) a NifH polypeptide (NH) having an N-terminus, wherein the C-terminus of the MTP2 is translationally fused to the N-terminus of the NH; and (b) a third polynucleotide encoding a NifK fusion polypeptide (NKF), the NKF comprising (i) a third mitochondrial targeting peptide (MTP3) having a C-terminus and (ii) a NifK polypeptide (NK) having an N-terminus, wherein the C-terminus of MTP3 is translationally fused to the N-terminus of the NK, preferably wherein the NKF has a native C-terminus relative to a wild-type NifK polypeptide; and
[0072] (c) a fourth polynucleotide encoding a Nif fusion polypeptide (NF), excluding NDF, NHF, and NKF, the NF comprising (i) a fourth mitochondrial targeting peptide (MTP4) having a C-terminus and (ii) a Nif polypeptide (NP) having an N-terminus, wherein the C-terminus of the MTP4 is translationally fused to the N-terminus of the NP; A nucleic acid construct comprising one, more, or all of:
[0073] Here, there is provided a nucleic acid construct in which each of MTP1, MTP2, MTP3 and MTP4 is independently the same or different.
[0074] In one embodiment, each NF can be transferred into the mitochondria of a plant cell. In a further aspect, the present invention provides a chimeric vector comprising or encoding a polynucleotide according to the invention or a nucleic acid construct according to the invention. In one embodiment, the polynucleotide, or each sequence therein encoding the fusion polypeptide, is operably linked to a promoter and optionally, a transcription termination sequence.
[0075] In one embodiment, the promoter directs expression of one, more, or all of the polynucleotides in the roots, leaves, and / or stems of the plant, and preferably the polynucleotides are preferentially expressed in one, more, or all of the roots, leaves, or stems of the plant associated with the seed of the plant.
[0076] In a further aspect, the present invention provides a cell comprising one, more or all of the fusion polypeptides or processed Nif polypeptides according to the invention, the polynucleotide according to the invention, the nucleic acid construct according to the invention, and / or the vector according to the invention.
[0077] In one embodiment, the cell is a plant cell or a bacterial cell. In a further embodiment, the plant cell is a cereal plant cell, such as, for example, a wheat cell, a rice cell, a corn cell, a triticale cell, an oat cell, or a barley cell, preferably a wheat cell. In a further aspect, the present invention provides a method for producing a fusion polypeptide(s) or a processed Nif polypeptide(s) according to the invention, said method comprising expressing a polynucleotide according to the invention in a cell. In a further aspect, the present invention provides a method of producing a cell according to the method, the method comprising the step of introducing into a cell a polynucleotide according to the invention, a nucleic acid construct according to the invention, and / or a vector according to the invention.
[0078] In a further aspect, the present invention provides a method for producing a transgenic plant according to the present invention, the method comprising the steps of: i) introducing a polynucleotide according to the invention, a nucleic acid construct according to the invention, and / or a vector according to the invention into a plant cell, ii) regenerating a transgenic plant from the cell, and iii) optionally collecting seeds from the plant; and / or iv) optionally producing one or more progeny plants from the transgenic plant, thereby producing a transgenic plant; Includes:
[0079] In a further aspect, the present invention provides a method for producing a plant having integrated into its genome a polynucleotide according to the present invention or a polynucleotide defined in any one of the nucleic acid constructs according to the present invention, said method comprising the steps of: i) crossing two parent plants, wherein at least one plant contains the polynucleotide(s); ii) screening one or more progeny plants from the cross for the presence or absence of said polynucleotide(s); iii) selecting progeny plants containing the polynucleotide(s), thereby producing the plant; Includes:
[0080] In one embodiment, the polynucleotide(s) encode a polypeptide(s) that confers nitrogenase activity in a plant. In one or further embodiments, at least one parent plant is a tetraploid or hexaploid wheat plant. In one or further embodiments, step ii) comprises analyzing a sample comprising DNA from one or more of the progeny plants for said polynucleotide(s).
[0081] In one or further embodiment, step iii) comprises: i) selecting progeny plants that are homozygous for the polynucleotide(s), and / or ii) analyzing one or more progeny plants for the presence and / or expression of said polynucleotide(s) or for an altered phenotype as defined above; Includes:
[0082] In one or further embodiment, the method comprises the steps of: iv) backcrossing the progeny of the cross of step i) with plants of the same genotype as the first parent that do not have the polynucleotide(s) a sufficient number of times to produce plants that have most of the genotype of the original parent but that contain the polynucleotide; and iv) selecting progeny plants which contain the polynucleotide and / or have the altered phenotype defined above; Further includes:
[0083] In one or further embodiments, the method further comprises analyzing the plant or progeny plants for at least one other genetic marker. In a further aspect, the present invention provides a plant produced using a method according to the present invention.
[0084] In a further aspect, the present invention provides the use of a polynucleotide according to the invention, a nucleic acid construct according to the invention, and / or a vector according to the invention for producing a recombinant cell and / or a transgenic plant. In one embodiment, the transgenic plant has an altered phenotype as defined above when compared to a corresponding plant that does not harbor the exogenous polynucleotide, nucleic acid construct, and / or vector.
[0085] In a further aspect, the present invention provides a method for identifying a plant comprising a polynucleotide according to the present invention or comprising a polynucleotide as defined in any one of the nucleic acid constructs according to the present invention, said method comprising the steps of: i) obtaining a nucleic acid sample from a plant; and ii) screening the sample for the presence or absence of the polynucleotide(s); Includes:
[0086] In one embodiment, the presence of the polynucleotide(s) indicates that the plant has an altered phenotype as defined above when compared to a corresponding plant that does not have the exogenous polynucleotide(s). In one or further embodiment, the method according to the invention identifies plants. In one or further embodiments, the method further comprises producing a plant from the seed prior to step i).
[0087] In a further aspect, the present invention provides a plant part of a plant according to the present invention. In one embodiment, said plant part is a seed comprising a polynucleotide according to the present invention or comprising a polynucleotide as defined in any one of the nucleic acid constructs according to the present invention.
[0088] In a further aspect, the present invention provides a method for producing a plant part, the method comprising the steps of: a) cultivating a plant according to the invention, and b) harvesting said plant parts; Includes:
[0089] In a further aspect, the present invention provides a method for producing flour, whole grain, starch, oil, seed meal or other product obtained from seeds, the method comprising the steps of: a) obtaining seeds according to the invention, and b) extracting flour, whole grain flour, starch, oil or other products or producing seed meal; Includes:
[0090] In a further aspect, the present invention provides products produced from plants and / or plant parts according to the present invention. In one embodiment, the part is a seed. In one or further embodiments, the product is a food or beverage product.
[0091] Preferably, i) the food product is selected from the group consisting of: foods containing flour, starch, fats and oils, leavened or unleavened bread, pasta, noodles, animal feed, breakfast cereals, snack foods, cakes, malt, pastries, and flour-based sauces, or ii) the beverage product is juice, beer or malt. Methods for producing such products are well known to those skilled in the art.
[0092] In alternative embodiments, the product is a non-food product. Examples of non-food products include, but are not limited to, films, coatings, adhesives, building materials, and packaging materials. Methods of making such products are well known to those skilled in the art.
[0093] In a further aspect, the present invention provides a method of preparing a food product according to the present invention, the method comprising mixing the seed, or flour, wholemeal or starch from the seed, with another food ingredient.
[0094] In a further aspect, the present invention provides a method of preparing malt comprising the step of germinating seeds according to the present invention.
[0095] In a further aspect, the present invention provides the use of a plant or part thereof according to the present invention as animal feed or for producing food for animal feed or human consumption.
[0096] In a further aspect, the present invention provides a composition comprising a fusion polypeptide(s) or a processed Nif polypeptide(s), a polynucleotide according to the invention, a nucleic acid construct according to the invention, a vector according to the invention, or a cell according to the invention, and one or more acceptable carriers.
[0097] In a further aspect, the present invention provides a method for reconstituting a nitrogenase protein complex in a plant cell, the method comprising introducing into a cell two or more polynucleotides according to the invention, two or more nucleic acid constructs according to the invention, and / or a vector according to the invention, and culturing the plant cell for a time sufficient to allow the polynucleotides or vectors to be expressed.
[0098] In a further aspect, the present invention provides a method for enhancing yield, biomass, growth rate, vitality, nitrogen acquisition from biological nitrogen fixation, nitrogen use efficiency, abiotic stress tolerance, and / or tolerance to nutrient deficiency in a plant, comprising introducing into a plant or plant cell two or more polynucleotides according to the present invention, two or more nucleic acid constructs according to the present invention, and / or a vector according to the present invention.
[0099] The present invention is not to be limited in scope by the specific embodiments described herein, which are intended for purposes of illustration only. Functionally equivalent products, compositions, and methods are clearly within the scope of the invention as described herein.
[0100] Throughout this specification, unless specifically stated otherwise or the context requires otherwise, references to a single step, composition of matter, group of steps or group of compositions of matter should be interpreted as encompassing one and more (i.e., one or more) of that step, composition of matter, group of steps or group of compositions of matter. The invention will now be described by way of the following non-limiting examples and with reference to the accompanying drawings. [Brief explanation of the drawings]
[0101] [Figure 1]Genetic map of the T-DNA of pCW440. Labels: 35SP, full-length CaMV 35S promoter, indicating the direction of transcription; T7 promoter, promoter for T7 RNA polymerase; dCoxIV MTP, dCoxIV mitochondrial targeting peptide; AscI, AscI restriction enzyme site for inserting in-frame fusions with MTP; T7 terminator, transcription termination site for T7 RNA polymerase; modified nos polyA, nos 3' polyadenylation signal. [Figure 2] Western blot analysis of NifH, NifD, NifK, and NifY fusion polypeptides produced in bacteria and N. benthamiana leaves. Each polypeptide was fused to a dCoxIV MTP, and the C-terminal extension contained an HA or FLAG epitope. Lane contents: 1, control leaf extract (p19 infiltrate only) and polypeptide molecular weight markers; filled arrow = 72 kDa; subsequent arrows (open arrows) are 55, 40, 33, 25, and 17 kDa; lanes 2–5, pCW446 dCoxIV-NifH-HA expressed in bacteria BL21-Gold (lane 2) or two separate leaf infiltrates (lanes 3 and 4); lanes 5–7, pCW447-dCoxIV-NifD-FLAG expressed in bacteria (lane 5) or two separate leaf infiltrates (lanes 6 and 7). Lanes 8–10, pCW448-dCoxIV-NifK-HA expressed in bacteria (lane 8) or two separate leaf infiltrates (lanes 9 and 10); lanes 11–13, pCW449-dCoxIV-NifY-HA expressed in bacteria (lane 11) or two separate leaf infiltrates (lanes 12 and 13); lanes 14–15, a combination of four vectors, pCW446, pCW447, pCW448, and pCW449, expressed in leaf infiltrates. Lane 15 also contains an infiltrate of vector pCW444, which expresses a cytoplasmically localized Lbh. Blot A was probed with two primary antibodies, anti-HA and anti-FLAG; blot B was probed with anti-HA alone. Longer exposure of blot A (lane 14) revealed signals for NifH, NifK, and NifY polypeptides, but not NifD. [Figure 3]Schematic diagram of the construct used for transient expression of pFAγ::GFP fusion polypeptide in N. benthamiana leaves. The wild-type pFAγ amino acid sequence is shown at the top, and the mutant amino acid sequence (mFAγ) is shown at the bottom. Arrows indicate the predicted cleavage point by MPP. 35S Pr, CaMV35S promoter; T7P, T7 RNA polymerase promoter; MTP, pFAγ or mFAγ region; GFP, GFP polypeptide; T7T, T7 RNA polymerase transcription terminator; NOS, 3' transcription terminator / polyadenylation region of the nos gene. [Figure 4] Photograph of Western blots probed with antibodies against HA (left panel) or FLAG (right panel) after SDS-PAGE of protein extracts from N. benthamiana or E. coli cells expressing constructs encoding pFAγ::NifF::HA or pFAγ::NifZ::FLAG fusion polypeptides. The + and - symbols above the lanes indicate the presence or absence, respectively, of the N. benthamiana or E. coli protein extract applied to the lane; the dilution factor of the extract used for the bacterial extract is indicated in brackets. [Figure 5] Diagram of the gene construct used to express pFAγ::NifH::HA in N. benthamiana leaf cells. Underlined residues in the nucleotide sequence indicate the site of proteolytic cleavage (carboxyl side) by trypsin. The arrow indicates the cleavage point by mitochondrial processing peptidase (MPP). Peptides ISTQVVR (SEQ ID NO: 44) and AVQGAPTMR (SEQ ID NO: 45) were detected by mass spectrometry. [Figure 6]Photograph of a Western blot probed with antibodies against HA (top left panel) or FLAG (top right panel) after SDS-PAGE of protein extracts from N. benthamiana cells expressing constructs encoding the pFAγ::Nif::HA or pFAγ::Nif::FLAG fusion polypeptides. Letters above the lanes (K, B, S, E, etc.) indicate the Nif polypeptide contained within the fusion polypeptide encoded by the gene construct. A faint band near the top of the blot for pFAγ::NifJ::FLAG is indicated with an asterisk (*). The sizes of molecular weight markers (in kDa) are indicated on the left. The lower panel shows the corresponding gel after Coomassie staining. [Figure 7] Photograph of a Western blot of proteins extracted from N. benthamiana leaves or E. coli containing the same construct encoding pFAγ::NifD::FLAG. The blot was probed with an antibody against the FLAG epitope. The molecular weight of the marker in the first lane is indicated. The dilution factor of the E. coli extract is indicated in brackets. [Figure 8] Photograph of a Western blot of protein extracts from E. coli or N. benthamiana leaves (indicated at the top of the figure) containing constructs encoding pFAγ::NifD::HA, mFAγ::NifD::HA, or pFAγ::NifK::HA and pFAγ::GFP as positive and negative controls in lanes 2 and 3, respectively. The construct used and the encoded Nif polypeptide (D or K) are indicated above each lane. The blot was probed with an antibody against the HA epitope. The molecular weight of the marker in the first lane is indicated. The bottom panel shows a Coomassie-stained gel. The major band in the Coomassie-stained gel is Rubisco from the N. benthamiana leaf sample. [Figure 9]Coexpression of NifK and NifH polypeptides increases the concentration of the NifD fusion polypeptide in plants. Photograph of a Western blot of polypeptides produced after introduction into leaf cells of constructs for coexpression of NifD, NifK, and NifH fusion polypeptides. Lane 1: pRA24 + pRA10 + pRA25; Lane 2: pRA22 + pRA10 + pRA25; Lane 3: pRA19 + pRA10 + pRA25; Lane 4: pRA19 + pRA25; Lane 5: pRA10 + pRA25; Lane 6: pRA10; Lane 7: pRA11; Lane 8: pRA25; Lane 9: pRA22; Lane 10: pRA24; Lane 11: pRA19; Lane 12: molecular weight markers; numbers to the right indicate kDa. See Table 4 for a list of constructs and encoded polypeptides, which are also shown above the blot. The asterisk in lane 1 indicates the position of the characteristic band of NifD. [Figure 10] Western blots of polypeptides produced after introduction of construct combinations into leaf cells for expression of NifD::FLAG, NifK (without C-terminal extension), NifS::HA, and NifH::HA fusion polypeptides. The upper panel was probed with anti-FLAG antibody, and the lower panel was probed with anti-HA antibody. Lanes 1 and 12 show molecular weight markers with the indicated sizes (kDa). A single asterisk in lane 2 indicates the position of the NifD-specific band, and double asterisks indicate the position of the smaller, degraded NifD::FLAG polypeptide or the internal translation initiation polypeptide. Numbers in brackets represent the quantity of the NifD-specific band relative to the background FLAG band. The positions of the NifK::HA, NifS::HA, and Nif::H fusion polypeptides are indicated in the lower panel. The construct combinations and encoded polypeptides are indicated above each lane. Abbreviations: D-FLAG (pRA07;pFAγ::NifD::FLAG), H-HA (pRA10;pFAγ::NifH::HA), K (pRA25;pFAγ::NifK), K-HA (pRA11, pFAγ::NifK::HA), S-HA (pRA16;pFAγ::NifS::HA), D-HA (pRA19, pFAγ::NifD::HA). [Figure 11] Western blots of polypeptides produced after introduction of constructs for expression of NifD::FLAG, NifK (without C-terminal extension), or NifK::HA and NifH::HA fusion polypeptides, alone or in combination, into leaf cells. The introduced constructs are indicated above each lane, and the encoded fusion polypeptides are indicated below each lane. The blots were cut and probed with either anti-HA antibody (lanes 1–7) or anti-FLAG antibody (lanes 9–12). Asterisks indicate the characteristic bands of NifD::FLAG in lanes 11 and 12. Longer exposure of the Western blot was required to observe faint NifD bands in pRA19 and pRA07 (lanes 2 and 3). [Figure 12] Model of the NifD::linker::NifK heterodimer, showing the NifD and NifK subunits as green and blue, respectively. The space-filling model shows the linker peptide in red, connecting the C-terminus of NifD to the N-terminus of NifK. [Figure 13] Homology model of the NifD::linker::NifK fusion polypeptide as a dimer complexed with two NifH polypeptides, using the K. pneumoniae amino acid sequences of NifD (green) and NifK (blue) along with the A. vinlandii amino acid sequence of the included NifH subunit (purple). The designed linker is shown in red van der Waals atom representation. The left and right images are rotated 90 degrees about the vertical axis relative to each other. [Figure 14] Upper panel: Photograph of a Western blot using an antibody detecting the HA epitope for polypeptides produced from pRA01 (lane 2, GFP), pRA11 (lane 3, pFAγ::NifK::HA), pRA19 (lane 4, pFAγ::NifD::HA), and pRA20 (lane 5, pFAγ::NifD-linker(FLAG)-NifK::HA). The size of the molecular weight marker (lane 1, kDa) is indicated on the left. The lower panel shows a Coomassie-stained gel. [Figure 15] Diagram of the structure of the multi-cassette vectors pKT100 and pKT-HC. LB: left border of T-DNA; RB: right border; P1-5 / T1-5: promoters / terminators of expression cassettes 1 to 5; PS / TS: promoters / terminators of expression cassettes containing selection genes; MAR: matrix attachment region. Restriction enzyme sites are indicated by arrows. [Figure 16] Upper panel: Photograph of a Western blot probed with a combination of anti-HA and anti-GFP antibodies after gel electrophoresis of protein extracts from N. benthamiana leaf samples that received the gene constructs. Lane 1, polypeptide molecular weight marker (kDa). Transfer vectors: lane 2, pRA01 (pFAγ::GFP); lanes 3–7, pRA01 plus vectors for expressing pFAγ::NifH::HA from cassettes 1–5. Lower panel: Ponceau staining of the same membrane. Showing the doublet of the pFAγ::NifH::HA and pFAγ::GFP bands. [Figure 17] A. Upper panel: Photograph of a Western blot probed with anti-HA antibody after gel electrophoresis of protein extracts from N. benthamiana leaf samples that received the gene construct. Lanes 1 and 3, polypeptide molecular weight markers (kDa). Transfer vectors: lane 2, pRA01 (pFAγ::GFP); lane 4, HC13. Lower panel: Ponceau staining of the same membrane. Band identities are shown. B. Same Western blot probed with anti-MYC antibody. [Figure 18]Upper panel: Photograph of a Western blot of protein extracts from N. benthamiana leaves 4 days after infiltration with a combination of pRA01 (pFAγ::GFP, lane marked G), pRA16 (pFAγ::NifS::HA, lane marked S), and pRA19 (human codon-optimized pFAγ::NifD::HA, lane marked D). The Western blot was probed with anti-HA and anti-GFP antibodies. Exposure was limited to resolve NifS and GFP without oversaturation. NifD::HA detection required longer exposure of the blot. The sizes of molecular weight markers (kDa) are indicated in lane 1. Lower panel: Photograph of a Western blot of protein extracts from N. benthamiana leaves 4 days after infiltration with combinations of pRA01 (pFAγ::GFP, lane marked with G), pRA22 (human codon-optimized mFAγ::NifD::HA, lane marked with D-mMTP), pRA19 (human codon-optimized pFAγ::NifD::HA, lane marked with D-HA), pRA26 (human codon-optimized ΔFAγ::NifD::HA, lane marked with D-stop), and pRA24 (Arabidopsis codon-optimized pFAγ::NifD::HA, lane marked with D-alt). The blot was probed with anti-HA and anti-GFP antibodies. NifD-HA detection required longer exposure of the blot. [Figure 19] Photograph of a Western blot using an anti-HA antibody after SDA-PAGE of protein extracts from N. benthamiana leaf cells after infiltration with constructs co-expressing P19 expressing NifD::HA and the SN6, SN7, or SN8 constructs (Example 20), with or without pRA25. Unpro / pro: processed or unprocessed forms of FAγ51::NifD. degr prod: a NifD-specific degradation product approximately 48 kDa in size.
[0102] Key points about the sequence table SEQ ID NO:1 - Amino acid sequence of native MTP of cytochrome c oxidase subunit IV (CoxIV MTP). SEQ ID NO:2—Derivative MTP amino acid sequence of cytochrome c oxidase subunit IV (dCoxIV). SEQ ID NO: 3 - Conserved arginine and serine residues within the motif xRxxxSSx of dCoxIV involved in MTP import and processing.
[0103] SEQ ID NO:4—Nucleotide sequence of the T-DNA region of pCW440, including the component T-DNA right border (nucleotides 1-164), the 35S promoter flanked by HindIII and XhoI sites (nucleotides 219-1564), CTCGAG(XhoI), the T7 promoter (nucleotides 1571-1587), the ATG-initiated dCoxIV coding sequence (nucleotides 1650-1742), an AscI site (nucleotides 1743-1750), the T7 terminator (nucleotides 1810-1856), the nos 3' terminator (nucleotides 1861-2084), and the T-DNA left border (nucleotides 2186-2346).
[0104] SEQ ID NO: 5—Amino acid sequence of wild-type K. pneumoniae NifH. SEQ ID NO: 6—Amino acid sequence of wild-type K. pneumoniae NifD. SEQ ID NO:7—Amino acid sequence of wild-type K. pneumoniae NifK. SEQ ID NO:8—Amino acid sequence of wild-type K. pneumoniae NifY. SEQ ID NO: 9—Amino acid sequence of wild-type K. pneumoniae NifB. SEQ ID NO: 10—Amino acid sequence of wild-type K. pneumoniae NifE. SEQ ID NO: 11—Amino acid sequence of wild-type K. pneumoniae NifN. SEQ ID NO: 12—Amino acid sequence of wild-type K. pneumoniae NifQ. SEQ ID NO: 13—Amino acid sequence of wild-type K. pneumoniae NifS. SEQ ID NO: 14—Amino acid sequence of wild-type K. pneumoniae NifU. SEQ ID NO: 15—Amino acid sequence of wild-type K. pneumoniae NifX. SEQ ID NO: 16—Amino acid sequence of wild-type K. pneumoniae NifF. SEQ ID NO: 17—Amino acid sequence of wild-type K. pneumoniae NifZ. SEQ ID NO: 18—Amino acid sequence of wild-type K. pneumoniae NifJ. SEQ ID NO: 19—Amino acid sequence of wild-type K. pneumoniae NifM. SEQ ID NO: 20—Amino acid sequence of wild-type K. pneumoniae NifV.
[0105] SEQ ID NO: 21 - Amino acid sequence of the C-terminal extension (17 aa) containing the HA epitope (amino acids 7 to 15). SEQ ID NO: 22 - Amino acid sequence of the C-terminal extension (12 aa) containing the FLAG epitope. SEQ ID NO:23—Amino acid sequence of the dCoxIV::NifH::HA fusion polypeptide encoded by pCW446. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-326 are K. pneumoniae NifH amino acids (start codon Met removed), and amino acids 327-343 contain the HA epitope. SEQ ID NO:24—Amino acid sequence of the dCoxIV::NifD::FLAG fusion polypeptide encoded by pCW447. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-515 are K. pneumoniae NifD amino acids (with the two N-terminal Met residues removed), and amino acids 516-527 contain the FLAG epitope.
[0106] SEQ ID NO:25—Amino acid sequence of the dCoxIV::NifK::HA fusion polypeptide encoded by pCW448. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-553 are K. pneumoniae NifK amino acids, and amino acids 554-570 comprise the HA epitope. SEQ ID NO:26—Amino acid sequence of the dCoxIV::NifY::HA fusion polypeptide encoded by pCW449. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-253 are K. pneumoniae NifY amino acids (without the start codon Met), and amino acids 254-270 contain the HA epitope. SEQ ID NO:27—Nucleotide sequence of the AscI fragment encoding NifH::HA, AscI-NifH-HA-AscI.
[0107] SEQ ID NO:28—Nucleotide sequence of the AscI fragment encoding NifD::FLAG, AscI-NifD-FLAG-AscI. SEQ ID NO:29—Nucleotide sequence of the AscI fragment encoding NifK::HA, AscI-NifK-HA-AscI. SEQ ID NO:30—Nucleotide sequence of the AscI fragment encoding NifY::HA, AscI-NifY-HA-AscI.
[0108] SEQ ID NO:31—Amino acid sequence of the dCoxIV::NifB::HA fusion polypeptide encoded by pCW452. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-501 are K. pneumoniae NifB amino acids (without the start codon Met), and amino acids 502-518 contain the HA epitope.
[0109] SEQ ID NO:32—Amino acid sequence of the dCoxIV::nifE::HA fusion polypeptide encoded by pCW454. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-490 are K. pneumoniae NifE amino acids (without the start codon Met), and amino acids 491-507 contain the HA epitope.
[0110] SEQ ID NO:33—Amino acid sequence of the dCoxIV::NifN::FLAG fusion polypeptide encoded by pCW455. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-493 are K. pneumoniae NifN amino acids (without the start codon Met), and amino acids 494-505 contain the FLAG epitope.
[0111] SEQ ID NO:34—Amino acid sequence of the dCoxIV::NifQ::HA fusion polypeptide encoded by pCW456. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-200 are Klebsiella sp. NifQ amino acids (without the start codon Met), and amino acids 201-217 contain the HA epitope.
[0112] SEQ ID NO:35—Amino acid sequence of the dCoxIV::NifS::HA fusion polypeptide encoded by pCW450. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-433 are K. pneumoniae NifS amino acids (without the start codon Met), and amino acids 434-450 contain the HA epitope.
[0113] SEQ ID NO:36—Amino acid sequence of the dCoxIV::NifU::FLAG fusion polypeptide encoded by pCW451. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-493 are K. pneumoniae NifU amino acids (without the start codon Met), and amino acids 494-505 contain the FLAG epitope.
[0114] SEQ ID NO:37—Amino acid sequence of the dCoxIV::NifX::FLAG fusion polypeptide encoded by pCW453. Amino acids 1-31 correspond to the dCoxIV MTP, amino acids 32-34 are the product of cloning at the AscI site, amino acids 35-189 are K. pneumoniae NifX amino acids (without the start codon Met), and amino acids 190-201 comprise the FLAG epitope.
[0115] SEQ ID NO: 38—Amino acid sequence of the N-terminal extension containing pFAγ MTP (amino acids 1-77) and the amino acid triplet GAP (78-80) added to pRA00 as a result of the cloning strategy. Cleavage by MPP occurs between amino acid residues 42 and 43.
[0116] SEQ ID NO:39 - Amino acid sequence of the modified N-terminal extension encoded by vector pRA21, containing the mFAγ sequence (amino acids 1-77) and the amino acid triplet GAP (78-80) added to pRA21 as a result of the cloning strategy. Alanine amino acid substitutions at amino acids 12-18, 24-33, and 39-45 relative to SEQ ID NO:38 were designed to abolish cleavage by MPP.
[0117] SEQ ID NO:40—Amino acid sequence of the pFAγ::NifF::HA fusion polypeptide encoded by pRA05. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning K. pneumoniae NifX at the AscI site, amino acids 81-256 are K. pneumoniae NifF amino acids (SEQ ID NO:16), and amino acids 257-267 contain the HA epitope.
[0118] SEQ ID NO:41—Amino acid sequence of the pFAγ::NifZ::FLAG fusion polypeptide encoded by pRA04. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-228 are K. pneumoniae NifZ amino acids (SEQ ID NO:17), and amino acids 229-238 contain the FLAG epitope.
[0119] SEQ ID NO:42—Amino acid sequence of the pFAγ::ifH::HA fusion polypeptide encoded by pRA10. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-372 are K. pneumoniae NifH amino acids (SEQ ID NO:5 without the initiator Met), and amino acids 373-389 contain the HA epitope.
[0120] SEQ ID NO:43—Amino acid sequence of tryptic peptide from unprocessed pFAγ. SEQ ID NO: 44—Amino acid sequence of tryptic peptide from pFAγ after treatment with MPP. SEQ ID NO: 45—Amino acid sequence of tryptic peptide from pFAγ after treatment with MPP.
[0121] SEQ ID NO:46—Amino acid sequence of the pFAγ::NifB::HA fusion polypeptide encoded by pRA03. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-547 are K. pneumoniae NifB amino acids (SEQ ID NO:9 without the initiator Met), and amino acids 548-564 contain the HA epitope.
[0122] SEQ ID NO:47—Amino acid sequence of the pFAγ::NifD::FLAG fusion polypeptide encoded by pRA07. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-561 are K. pneumoniae NifD amino acids (SEQ ID NO:6 without the initiator Met and the next Met), and amino acids 562-573 comprise the FLAG epitope.
[0123] SEQ ID NO:48—Amino acid sequence of the pFAγ::nifE::HA fusion polypeptide encoded by pRA09. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-536 are K. pneumoniae NifE amino acids (SEQ ID NO:10 without the initiator Met), and amino acids 537-553 contain the HA epitope.
[0124] SEQ ID NO:49—Amino acid sequence of the pFAγ::NifJ::FLAG fusion polypeptide encoded by pRA06. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-1251 are K. pneumoniae NifJ amino acids (SEQ ID NO:18), and amino acids 1252-1261 comprise the FLAG epitope.
[0125] SEQ ID NO:50—Amino acid sequence of the pFAγ::NifK::HA fusion polypeptide encoded by pRA11. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-599 are K. pneumoniae NifK amino acids (SEQ ID NO:7 without the initiator Met), and amino acids 600-616 contain the HA epitope.
[0126] SEQ ID NO:51—Amino acid sequence of the pFAγ::NifM::HA fusion polypeptide encoded by pRA18. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-346 are K. pneumoniae NifM amino acids (SEQ ID NO:19), and amino acids 347-357 contain the HA epitope.
[0127] SEQ ID NO:52—Amino acid sequence of the pFAγ::NifN::FLAG fusion polypeptide encoded by pRA13. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-539 are K. pneumoniae NifN amino acids (SEQ ID NO:11 without the initiator Met), and amino acids 540-551 comprise the FLAG epitope.
[0128] SEQ ID NO:53—Amino acid sequence of the pFAγ::NifQ::HA fusion polypeptide encoded by pRA08. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-246 are K. pneumoniae NifQ amino acids (SEQ ID NO:12 without the initiator Met), and amino acids 247-263 contain the HA epitope.
[0129] SEQ ID NO:54—Amino acid sequence of the pFAγ::NifS::HA fusion polypeptide encoded by pRA16. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-478 are K. pneumoniae NifS amino acids (SEQ ID NO:13 without the initiator Met), and amino acids 479-496 contain the HA epitope.
[0130] SEQ ID NO:55—Amino acid sequence of the pFAγ::NifU::FLAG fusion polypeptide encoded by pRA15. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-353 are K. pneumoniae NifU amino acids (SEQ ID NO:14 without the initiator Met), and amino acids 354-365 comprise the FLAG epitope.
[0131] SEQ ID NO:56—Amino acid sequence of the pFAγ::NifV::FLAG fusion polypeptide encoded by pRA17. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-461 are K. pneumoniae NifV amino acids (SEQ ID NO:20), and amino acids 462-471 comprise the FLAG epitope.
[0132] SEQ ID NO:57—Amino acid sequence of the pFAγ::NifX::FLAG fusion polypeptide encoded by pRA14. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-235 are K. pneumoniae NifX amino acids (SEQ ID NO:15 without the initiator Met), and amino acids 236-247 contain the FLAG epitope.
[0133] SEQ ID NO:58—Amino acid sequence of the pFAγ::NifY::HA fusion polypeptide encoded by pRA12. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-299 are K. pneumoniae NifY amino acids (SEQ ID NO:8 without the initiator Met), and amino acids 300-316 contain the HA epitope.
[0134] SEQ ID NO:59—Amino acid sequence of the pFAγ::NifD::HA fusion polypeptide encoded by pRA19. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, amino acids 81-561 are K. pneumoniae NifD amino acids (SEQ ID NO:6 without the initiator Met and the next Met), and amino acids 562-574 contain the HA epitope.
[0135] SEQ ID NO: 60 - Amino acid sequence of the pFAγ::NifK fusion polypeptide (pRA25) without any C-terminal extension. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning at the AscI site, and amino acids 81-599 were K. pneumoniae NifK amino acids (SEQ ID NO: 7 without the initiator Met).
[0136] SEQ ID NO:61—Amino acid sequence of an 11-residue portion from the known unstructured linker region from Hypocrea jecorina cellobiohydrolase II (Accession number AAG39980.1).
[0137] Amino acid sequence of the FLAG epitope of residues 62-8 of SEQ ID NO:6. SEQ ID NO: 63—Amino acid sequence of the linker. The linker is 30 residues in length and consists of an 11-residue portion from the known unstructured linker region from Hypocrea jecorina cellobiohydrolase II (Accession number AAG39980.1, ATPPPGSTTTR, SEQ ID NO: 61) with the final arginine replaced by alanine, followed by an 8-residue FLAG epitope (DYKDDDDK; SEQ ID NO: 62), and finally another copy of the 11-residue unstructured linker sequence with the arginine replaced by alanine.
[0138] SEQ ID NO:64—Amino acid sequence of the pFAγ::NifD-linker-NifK fusion polypeptide (pRA02) without any C-terminal extension of NifK. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning the AscI site, amino acids 81-561 are K. pneumoniae NifD amino acids (SEQ ID NO:6 without the initiator Met and the next Met), amino acids 562-592 are a 30 amino acid linker, and amino acids 593-1110 were K. pneumoniae NifK amino acids (SEQ ID NO:7 without the initiator Met).
[0139] SEQ ID NO:65—Amino acid sequence of the pFAγ::NifD-linker-NifK::HA fusion polypeptide (pRA20) without any C-terminal extension of NifK. Amino acids 1-77 correspond to pFAγ MTP, amino acids 78-80 (GAP) are the product of cloning the AscI site, amino acids 81-561 are K. pneumoniae NifD amino acids (SEQ ID NO:6 without the initiator Met and the next Met), amino acids 562-592 are a 30 amino acid linker, amino acids 593-1110 are K. pneumoniae NifK amino acids (SEQ ID NO:7 without the initiator Met), and amino acids 1111-1119 comprise the HA epitope.
[0140] SEQ ID NO: 66 - Amino acid sequence of a C-terminal extension containing a MYC epitope corresponding to amino acids 347 to 356 of SEQ ID NO: 67. SEQ ID NO:67 - Amino acid sequence of pFAγ::NifM::MYC fusion polypeptide. SEQ ID NO:68—Nucleotide sequence encoding the pFAγ::NifD::HA fusion polypeptide of pRA19, starting at the translation initiation codon. SEQ ID NO:69—Amino acid sequence of the last four amino acid residues at the C-terminus of the NifK polypeptide from K. pneumoniae.
[0141] SEQ ID NO:70—Nucleotide sequence of a DNA fragment encoding pFAγ-C polypeptide. SEQ ID NO:71 - Amino acid sequence of pFAγ-C polypeptide. SEQ ID NO: 72 - Oligonucleotide primer NifD-F. SEQ ID NO: 73 - Oligonucleotide primer NifD-R. SEQ ID NO:74—Amino acid sequence of wild-type K. pneumoniae NifW. SEQ ID NO:75-MTP-FAγ 51 The amino acid sequence of a polypeptide. SEQ ID NO:76-FAγ-scar 9 The amino acid sequence of a polypeptide. SEQ ID NO:77 - Amino acid sequence of CPN60 MTP. SEQ ID NO:78—Amino acid sequence of CPN60 / No GG linker MTP.
[0142] SEQ ID NO:79—Amino acid sequence of superoxide dismutase (SOD) MTP (At3G10920). SEQ ID NO:80—Amino acid sequence of superoxide dismutase double (2SOD) MTP (At3G10920). SEQ ID NO: 81—Amino acid sequence of superoxide dismutase-modified (SODmod) MTP (At3g10920). SEQ ID NO: 82—Amino acid sequence of double superoxide dismutase modified (2SODmod) MTP (At3g10920).
[0143] SEQ ID NO: 83-L29 Amino acid sequence of MTP (At1G07830). SEQ ID NO:84—Amino acid sequence of Neurospora crassa F0 ATPase subunit 9 (SU9) MTP. SEQ ID NO:85 - gATPase gamma subunit (FAγ 51 ) Amino acid sequence of MTP. SEQ ID NO:86 - Amino acid sequence of CoxIV twin-strep (ABM97483) MTP.
[0144] SEQ ID NO:87 - Amino acid sequence of CoxIV 10xHis (ABM97483) MTP. SEQ ID NO: 88—Amino acid sequence of predicted trace of superoxide dismutase (SOD) MTP (SEQ ID NO: 80).
[0145] SEQ ID NO:89—Amino acid sequence of predicted trace of superoxide dismutase double (2SOD) MTP (SEQ ID NO:81). SEQ ID NO:90-L29 Amino acid sequence of the predicted vestigial sequence of MTP (SEQ ID NO:84). SEQ ID NO: 91—Amino acid sequence of predicted vestigial protein of Neurospora crassa F0 ATPase subunit 9 (SU9) MTP (SEQ ID NO: 85).
[0146] SEQ ID NO:92 - gATPase gamma subunit (FAγ 51 ) Predicted vestigial amino acid sequence of MTP (SEQ ID NO: 86). SEQ ID NO:93—Amino acid sequence of predicted trace of CoxIV twin-strep MTP (SEQ ID NO:87). SEQ ID NO:94—Amino acid sequence of predicted trace of CoxIV 10xHis MTP (SEQ ID NO:88).
[0147] SEQ ID NO: 95—Amino acid sequence of mutant NifD. SEQ ID NO: 96—Amino acid sequence of mutant NifS. SEQ ID NO:97 - NifK9 amino acid C-terminal extension. SEQ ID NO: 98 - Oligonucleotide primer. SEQ ID NO: 99 - Oligonucleotide primer. SEQ ID NO: 100 - Oligonucleotide primer. SEQ ID NO: 101 - Oligonucleotide primer. SEQ ID NO: 102 - Oligonucleotide primer BO1. SEQ ID NO: 103 - Oligonucleotide primer BO2. SEQ ID NO: 104 - Amino acid sequence of linker.
[0148] SEQ ID NO: 105—Oligonucleotide primer pRA31DK-FW. SEQ ID NO: 106 - Oligonucleotide primer pRA31DK-RV. SEQ ID NO: 107 - Oligonucleotide primer D_start_RV. SEQ ID NO: 108 - Oligonucleotide primer K_end_FW. SEQ ID NO: 109 - Oligonucleotide primer. SEQ ID NO: 110 - Oligonucleotide primer. SEQ ID NO: 111 - N-terminal extension of the 9 amino acid sequence of the NifD-linker-NifK polypeptide.
[0149] SEQ ID NO: 112 - Amino acid sequence of HA epitope. SEQ ID NO:113—Amino acid sequence of NifE-NifN HA polypeptide linker. SEQ ID NO: 114—Oligonucleotide primer pRAEN-FW. SEQ ID NO: 115—Oligonucleotide primer pRAEN-RV. SEQ ID NO: 116 - Oligonucleotide primer E_start_RV. SEQ ID NO: 117 - Oligonucleotide primer N_end_FW. SEQ ID NO: 118 - Oligonucleotide primer. SEQ ID NO: 119 - Oligonucleotide primer.
[0150] SEQ ID NO:120 - Nucleotide sequence of SCSV-S4 version 1 promoter. SEQ ID NO:121 - Nucleotide sequence of SCSV-S4 version 2 promoter. SEQ ID NO: 122 - Nucleotide sequence of SCSV-S7 promoter. SEQ ID NO: 123—Nucleotide sequence of the long version promoter of CaMV-35S. SEQ ID NO: 124 - Nucleotide sequence of CaMV-2x25S promoter. SEQ ID NO:125—Amino acid sequence of P19 viral suppressor protein. SEQ ID NO: 126—Amino acid sequence of P19 peptide. SEQ ID NO: 127 - Amino acid sequence of P19 peptide.
[0151] SEQ ID NOs: 128-143 - Amino acid sequences of NifK peptides. SEQ ID NOs: 144 to 154 - Amino acid sequences of NifH peptides. SEQ ID NOs: 155-160 - Amino acid sequences of NifB peptides. SEQ ID NOs: 161-165 - Amino acid sequence of NifJ peptide. SEQ ID NOs: 166-174 - Amino acid sequences of NifS peptides. SEQ ID NOs: 175-182 - Amino acid sequences of NifX peptides. SEQ ID NOs: 183 to 186 - Amino acid sequences of NifF peptides. DETAILED DESCRIPTION OF THE INVENTION
[0152] Detailed Description of the Invention General Methods and Definitions Unless specifically defined otherwise, all technical and scientific terms used herein shall be construed to have the same meaning as commonly understood by one of ordinary skill in the art (e.g., in cell culture, molecular genetics, plant molecular biology, protein chemistry, and biochemistry).
[0153] Unless otherwise indicated, the recombinant protein, cell culture, and immunological techniques utilized in the present invention are standard procedures, well known to those skilled in the art. Such techniques are described, for example, in J. Perbal, A Practical Guide to Molecular Cloning, John Wiley and Sons (1984), J. Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press (1989), T.A. Brown (editor), Essential Molecular Biology: A Practical Approach, Volumes 1 and 2, IRL Press (1991), D.M. Glover and B.D. Hames (editors), DNA Cloning: A Practical Approach, Volumes 1-4, IRL Press (1995 and 1996), and F.M. Ausubel et al. (editors), Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience (1988, including all current editions), Ed Harlow and David Lane (editors), Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, (1988), and J.E. Coligan et al. (editors). It is described and explained throughout source literature such as Current Protocols in Immunology, John Wiley & Sons (including all current editions to date).
[0154] The term "and / or," e.g., "X and / or Y," shall be construed to mean either "X and Y" or "X or Y," and shall be construed as providing explicit support for both meanings or for either meaning.
[0155] Throughout this specification, the term "comprise" or variations such as "comprises" or "comprising" will be understood to imply the inclusion of a stated element, integer, or step, or group of elements, integers, or steps, but not the exclusion of any other element, integer, or step, or group of elements, integers, or steps.
[0156] Nitrogenase is a eubacterial and archaeal enzyme that catalyzes the reduction of the strong, triple bond of nitrogen (N) to produce ammonia (NH). Nitrogenase is usually found only in bacteria. It is a complex of two enzymes, dinitrogenase and dinitrogenase reductase, which can be purified separately. Dinitrogenase, also called component I or molybdenum-iron (MoFe) protein, is a tetramer of two NifD and two NifK polypeptides (α2β2) that also contain two "P-clusters" and two "FeMo cofactors" (FeMo-Co). Each pair of NifD-NifK subunits contains one P-cluster and one FeMo-Co. FeMo-Co is a metallocluster composed of a MoFe3-S3 cluster complexed with a homocitrate molecule, which is coordinated to a molybdenum atom and bridged to an Fe4-S3 cluster by three sulfur ligands. FeMo-Co is assembled separately within the cell and then incorporated into the apo-MoFe protein. The P-cluster is also a metallocluster, containing eight Fe atoms and seven sulfur atoms with a structure similar to, but distinct from, FeMo-Co. The P-cluster is located at the αβ subunit interface of dinitrogenase and is coordinated by cysteinyl residues from both subunits.
[0157] Dinitrogenase reductase, also called component II or the "Fe protein," is a dimer of NifH polypeptides that also contains a single Fe4-S4 cluster at the subunit interface and two Mg-ATP binding sites, one in each subunit. This enzyme is the electron donor to dinitrogenase, where electrons are transferred from the Fe4-S4 cluster to the P-cluster and then to FeMo-Co, the site of N2 reduction.
[0158] Although Mo-containing nitrogenase is the most commonly found nitrogenase in bacteria, two homologous nitrogenases exist that are genetically distinct but have similar cofactor and subunit compositions: vanadium-containing nitrogenase and Fe-only nitrogenase, encoded by the Vnf (vanadium nitrogen fixation) and Anf (alternative nitrogen fixation) genes, respectively. Some bacteria in nature possess all three types of nitrogenase, while others, e.g., Klebsiella pneumoniae, contain only Mo- and V-containing enzymes or only the Mo-containing enzyme.
[0159] Various nitrogen fixation (Nif) genes are required for the biosynthesis of FeMo-Co and the maturation of nitrogenase components into their catalytically active forms. The roles of NifB, NifE, NifH, NifN, NifQ, NifV, and NifX polypeptides in FeMo-Co synthesis have been described ( Rubio and Ludden, 2005 ).
[0160] Biological N2 fixation, catalyzed by the prokaryotic enzyme nitrogenase, is an alternative to the use of synthetic N2 fertilizers. The sensitivity of nitrogenase to oxygen is a major barrier to the genetic engineering of biological nitrogen fixation into plants, for example, into cereal crops by direct Nif gene introgression.
[0161] We hypothesized that targeting Nif polypeptides to the mitochondrial matrix (MM) of plant cells would overcome the oxygen sensitivity issue. MM possesses oxygen-consuming enzymes, allowing other enzymes containing oxygen-sensing Fe-S clusters to function. The mitochondrial Fe-S cluster assembly mechanism resembles its diazotrophic counterpart (Balk and Pilon, 2011; Lill and Muhlenhoff, 2008). Therefore, some of the requirements for nitrogenase biosynthesis are already in place within the MM, and the number of Nif genes required for reconstitution could be reduced. The cofactor homocitrate is produced as part of the TCA cycle. ATP potential and concentration, both prerequisites for nitrogenase enzyme catalysis, would also be greatly reduced (Geigenberger and Fernie, 2014; Mackenzie and McIntosh, 1999). Furthermore, the presence of glutamate synthase in mitochondria provides an entry point for all ammonium fixed by nitrogenase to participate in plant metabolism. Considering these characteristics, and the fact that mitochondria themselves are of α-proteobacterial origin, we considered this organelle to be an excellent location for attempting functional reconstitution of nitrogenase.
[0162] As a first step toward reconstituting nitrogenase within plant cell mitochondria, we needed evidence that individual Nif proteins are precisely targeted to the MM. To this end, we chose the model plant Nicotiana benthamiana as an expression platform (Wood et al., 2009) to provide transgene expression, either singly or, more importantly, in combination. Because the majority of proteins located in the MM are nuclear-encoded, we relied on recent advances in understanding intracellular signaling and transport processes (Huang et al., 2009; Murcha et al., 2014) and used previously characterized N-terminal peptide targeting signals (Lee et al., 2012).
[0163] The model bacterium, the diazotroph Klebsiella pneumoniae, uses 16 unique proteins for nitrogenase biosynthesis and catalytic function. We reengineered all 16 Nif proteins from K. pneumoniae to target them to plant MM and to evaluate their expression and processing in N. benthamiana leaves. All 16 Nif polypeptides were transiently expressed and tested for sequence-specific MM processing. We confirmed that all 16 Nif polypeptides were individually expressed as MTP:Nif fusion polypeptides in plant leaf cells. Furthermore, we provide evidence that these proteins can be targeted to the mitochondrial matrix (MM), a potentially compatible subcellular location for nitrogenase function, and cleaved by mitochondrial processing protease (MPP). This is the first practical demonstration of the feasibility of such an approach and represents an important step toward the goal of genetically manipulating endogenous nitrogen fixation in plants.
[0164] Mitochondrial targeting peptide (MTP)-Nif fusion polypeptide The present invention relates to mitochondrial targeting peptide (MTP)-Nif fusion polypeptides and their cleavage polypeptide products. When the MTP-Nif fusion polypeptides of the present invention are expressed in plant cells, the MTP-Nif fusion polypeptides and / or cleavage polypeptide products are targeted to the mitochondrial matrix (MM). Preferably, the fusion polypeptides confer nitrogenase reductase and / or nitrogenase activity to plant cells, or the same activity as that conferred by the corresponding wild-type Nif polypeptide in bacteria.
[0165] As used herein, the term "fusion polypeptide" refers to a polypeptide comprising two or more functional polypeptide domains covalently linked by peptide bonds. Typically, a fusion polypeptide is encoded as a single polypeptide chain by a polynucleotide of the invention. In one embodiment, a fusion polypeptide of the invention comprises a mitochondrial targeting peptide (MTP) and a Nif polypeptide (NP). In this embodiment, the C-terminus of MTP is translationally fused to the N-terminus of NP. In an alternative embodiment, a fusion polypeptide of the invention comprises the C-terminal portions of MTP and NP, wherein the C-terminal portion results from cleavage of MTP by MPP. In this embodiment, the C-terminus of the C-terminal portion of MTP is translationally fused to the N-terminus of NP.
[0166] As used herein, the term "translationally fused to the N-terminus" means that the C-terminus of the MTP polypeptide is covalently linked to the N-terminal amino acid of NP by a peptide bond, thereby resulting in a fusion polypeptide. In one embodiment, NP does not include its native translation initiation methionine (Met) residue associated with the corresponding wild-type NP or its two N-terminal Met residues. In an alternative embodiment, NP includes the translation initiation Met of the wild-type NP polypeptide, or one or both of the two N-terminal Met residues, such as with NifD.
[0167] Such polypeptides are typically produced by expression of a chimeric protein coding region in which the translational reading frame of nucleotides encoding MTP is joined in-frame with the reading frame of nucleotides encoding NP. Those skilled in the art will understand that the C-terminus of MTP may be translationally fused to the N-terminal amino acid of NP without a linker or via a linker of one or more amino acid residues, e.g., 1 to 5 amino acid residues. Such a linker may also be considered to be part of MTP. Expression of the protein coding region may be followed by cleavage of MTP within the MM of the plant cell, and such cleavage (if it occurs) is included in the concept of producing a fusion polypeptide of the present invention.
[0168] The fusion polypeptide or processed Nif polypeptide preferably has functional Nif activity comparable to that of the corresponding wild-type Nif polypeptide. The functional activity of the fusion polypeptide or processed Nif polypeptide may be measured by bacterial and biochemical complementation assays. In a preferred embodiment, the fusion polypeptide or processed Nif polypeptide has about 70-100% of the wild-type Nif activity.
[0169] A fusion polypeptide may contain two or more MTPs and / or two or more NPs. For example, a fusion polypeptide may contain an MTP, a NifD polypeptide, and a NifK polypeptide. The fusion polypeptide may also contain, for example, an oligopeptide linker connecting two NPs. Preferably, the linker is long enough to allow two or more functional domains, such as two NPs (e.g., NifD and NifK), to bind in a functional form in plant cells. Such a linker may be 8 to 50 amino acid residues long, preferably 25 to 35 amino acids long, and more preferably about 30 amino acid residues long. Fusion polypeptides may be obtained by conventional means, such as gene expression of a polynucleotide sequence encoding the fusion polypeptide in a suitable cell.
[0170] As used herein, a "substantially purified polypeptide" refers to a polypeptide that is substantially free from components (e.g., lipids, nucleic acids, carbohydrates) that are normally associated with the polypeptide, e.g., in a cell. Preferably, a substantially purified polypeptide is at least 90% free from such components.
[0171] The plant cells, transgenic plants, and components thereof of the present invention contain a polynucleotide encoding a polypeptide of the present invention. Because the polypeptide of the present invention does not naturally occur in plant cells, particularly not in the mitochondria of plant cells, the polynucleotide encoding the polypeptide is also referred to herein as a foreign polynucleotide, since it does not naturally occur in plant cells but has been introduced into a plant cell or progenitor cell. Thus, cells, plants, and plant parts of the present invention that produce a polypeptide of the present invention are sometimes referred to as producing a recombinant polypeptide. In the context of a polypeptide, the term "genetically recombinant" refers to a polypeptide that, when produced by a cell, is encoded by a foreign polynucleotide, and the polynucleotide is introduced into the cell or progenitor cell by recombinant DNA or RNA techniques, such as transformation. Typically, a plant cell, plant, or plant part contains a non-endogenous gene that is responsible for a significant amount of the polypeptide being produced, at least sometime during the plant cell's or plant's life cycle.
[0172] In certain embodiments, the polypeptides of the invention are not naturally occurring polypeptides. In alternative embodiments, the polypeptides of the invention are naturally occurring but are not present in, and do not naturally occur in, plant cells, preferably within the mitochondria of plant cells.
[0173] Nif polypeptides As used herein, the terms "Nif polypeptide" and "Nif protein" are used interchangeably and refer to a polypeptide related in amino acid sequence to a naturally occurring polypeptide involved in nitrogenase activity, wherein the polypeptide of the invention is selected from the group consisting of a NifD polypeptide, a NifH polypeptide, a NifK polypeptide, a NifB polypeptide, a NifE polypeptide, a NifN polypeptide, a NifF polypeptide, a NifJ polypeptide, a NifM polypeptide, a NifQ polypeptide, a NifS polypeptide, a NifU polypeptide, a NifV polypeptide, a NifW polypeptide, a NifX polypeptide, a NifY polypeptide, and a NifZ polypeptide, each of which is as defined herein. Nif polypeptides of the invention include "Nif fusion polypeptides," as used herein, which refer to polypeptide homologs of naturally occurring Nif polypeptides having additional amino acid residues attached to the N-terminus, C-terminus, or both, relative to the corresponding naturally occurring Nif polypeptide. As noted above, Nif fusion polypeptides may lack the translation initiation Met or two N-terminal Met residues associated with the corresponding wild-type Nif polypeptide. The amino acid residues of a Nif fusion polypeptide that correspond to a native Nif polypeptide, i.e., do not have additional amino acid residues attached to the N-terminus or C-terminus, or both, are also referred to herein as Nif polypeptides, which in this case are abbreviated as "NP" or NifD polypeptide ("ND"), etc. In preferred embodiments, "additional amino acid residues attached to the N-terminus or C-terminus, or both" comprise a mitochondrial targeting peptide (MTP) or processed MTP attached to the N-terminus of NP, or an epitope sequence ("tag") that is N-terminal or C-terminal, or both, to NP, or both an MTP or processed MTP and an epitope sequence.
[0174] Naturally occurring Nif polypeptides occur only in some bacteria, including nitrogen-fixing bacteria, including free-living, associative, and symbiotic nitrogen-fixing bacteria. Free-living nitrogen-fixing bacteria can fix significant levels of nitrogen without direct interaction with other organisms. Free-living nitrogen-fixing bacteria include, but are not limited to, members of the genera Azotobacter, Beijerinckia, Klebsiella, and Cyanobacteria (classified as aerobic organisms), as well as members of the genera Clostridium and Desulfovibrio, and those designated purple sulfur bacteria, purple non-sulfur bacteria, and green sulfur bacteria. Semi-symbiotic nitrogen-fixing bacteria are prokaryotes that can form close associations with some members of the Poaceae family (herbaceous plants). These bacteria fix significant amounts of nitrogen into the rhizosphere of the host plant. Members of the genus Azospirillum are representative of semi-symbiotic nitrogen-fixing bacteria. Symbiotic nitrogen-fixing bacteria are bacteria that fix nitrogen symbiotically by partnering with the host plant. The plant provides sugars from photosynthesis that are utilized by the nitrogen-fixing bacteria for the energy needed for nitrogen fixation. Members of the genus Rhizobia are representative of semi-symbiotic nitrogen-fixing bacteria.
[0175] The Nif polypeptide or Nif fusion polypeptide of the invention is selected from the group consisting of NifH, NifD, NifK, NifB, NifE, NifN, NifF, NifJ, NifM, NifQ, NifS, NifU, NifV, NifW, NifX, NifY and NifZ polypeptides.
[0176] A polypeptide or class of polypeptides may be defined by the degree of identity (% identity) of its amino acid sequence to a reference amino acid sequence, or by having a higher % identity to one reference amino acid sequence than to another. A type of polypeptide or polypeptide may also be defined by having the same biological activity as a native Nif polypeptide, in addition to the degree of sequence identity.
[0177] The percent identity of polypeptides is measured by GAP (Needleman and Wunsch, 1970) analysis (GCG program) with a gap creation penalty of 5 and a gap extension penalty of 0.3, or by Blastp version 2.5 or its updated version (Altschul et al., 1997), in each case aligning two sequences including the reference sequence over the entire length of the reference sequence. As used herein, reference sequences include those provided for native Nif polypeptides from K. pneumoniae, SEQ ID NOS: 1-20.
[0178] In the definitions below, the degree of identity of an amino acid sequence to a reference sequence provided as a SEQ ID NO is measured by Blastp, version 2.5 or updated versions thereof (Altschul et al, 1997), using default parameters except for the maximum number of target sequences, which is set to 10,000, and is measured along the full length of the reference amino acid sequence.
[0179] The native bacterial NifH polypeptide is a structural component of the nitrogenase complex and is often referred to as the iron (Fe) protein. It forms a homodimer with an Fe4S4 cluster bond between the subunit and two ATP-binding domains. NifH is the obligate electron donor to the MoFe protein (NifD / NifK heterotetramer), thus functioning as a nitrogenase reductase (EC 1.18.6.1). NifH is also involved in FeMo-Co biosynthesis and apo-MoFe protein maturation. As used herein, "NifH polypeptide" refers to a polypeptide whose sequence is at least 41% identical to the amino acid sequence provided as SEQ ID NO: 5 and contains one or more of the amino acid domains TIGR01287, PRK13236, PRK13233, and cd02040. The TIGR01287 domain is present in each of molybdenum-iron nitrogenase reductase (NifH), vanadium-iron nitrogenase reductase (VnfH), and iron-iron nitrogenase reductase (AnfH), but excludes homologous proteins from light-independent protochlorophyllide reductase. As used herein, therefore, NifH polypeptides include subclasses of iron-binding polypeptides, VnfH iron-binding polypeptides, and AnfH iron-binding polypeptides whose sequences contain amino acids at least 41% identical to SEQ ID NO:5. Naturally occurring NifH polypeptides are typically 260-300 amino acids in length, and the naturally occurring monomer has a molecular weight of approximately 30 kDa. Numerous NifH polypeptides have been identified, and numerous sequences are available in publicly available databases.For example, the NifH polypeptide is a member of the NifH family of bacteria found in Klebsiella michiganensis (Accession number WP_049123239. 1, 99% identical to SEQ ID NO: 5), Brenneria goodwinii (WP_048638817.1, 93% identical), Sideroxydans lithotrophicus (WP_013029017.1, 84% identical), Denitrovibrio acetiphilus (WP_013010353.1, 80% identical), Desulfovibrio africanus (WP_014258951.1, 72% identical), Chlorobium phaeobacteroides, and the like. phaeobacteroides (WP_011744626.1, 69% identical), Methanosaeta concilii (WP_013718497.1, 64% identical), Rhodobacter (WP_009565928.1, 61% identical), Methanocaldococcus infernus (WP_013099472.1, 42% identical), and Desulfosporosinus youngiae (WP_007781874.1, 41% identical). The NifH polypeptide has been described and reviewed in Thiel et al., (1997), Pratte et al., (2006), Boison et al., (2006) and Staples et al., (2007).
[0180] As used herein, a functional NifH polypeptide is a NifH polypeptide that is capable of forming a functional nitrogenase protein complex together with other required subunits, e.g., NifD and NifK, and FeMo or other cofactors.
[0181] As used herein, "NifD polypeptide" means a polypeptide comprising amino acids whose sequence is at least 33% identical to the amino acid sequence provided as SEQ ID NO:6, and comprising (i) one or both of the domains TIGR01282 and COG2710, both of which are found within iron-molybdenum binding polypeptides comprising polypeptides having the amino acid sequence set forth in SEQ ID NO:6, or (ii) the iron-vanadium binding domain TIGR01860, in which case the NifD polypeptide is within the subclass of VnfD polypeptides, or (iii) the iron-iron binding domain TIGR1861, in which case the NifD polypeptide is within the subclass of AnfD polypeptides.
[0182] As used herein, NifD polypeptides include the VnfD iron-vanadium polypeptides and AnfD polypeptides, a subclass of iron-molybdenum (FeMo-Co) binding polypeptides comprising amino acids whose sequences are at least 33% identical to SEQ ID NO: 6. Naturally occurring NifD polypeptides are typically 470-540 amino acids in length. Numerous NifD polypeptides have been identified, and many sequences are available in publicly available databases.For example, the NifD polypeptide is found in Raoultella ornithinolytica (Accession number WP_044347161.1, 96% identical to SEQ ID NO: 6), Kluyvera intermediate (WP_047370273.1, 93% identical), Dickeya dadantii (WP_038902190.1, 89% identical), Tolumonas sp. BRL6-1 (WP_024872642.1, 81% identical), Magnetospirillum gryphiswaldense (WP_024078601.1, 68% identical), Thermoanaerobacterium thermosaccharolyticum thermosaccharolyticum (WP_013298320.1, 42% identical), Methanothermobacter thermautotrophicus (WP_010877172.1, 38% identical), Desulfovibrio africanus (WP_014258953.1, 37% identical), Desulfotomaculum sp. LMa1 (WP_066665786.1, 37% identical), Desulfomicrobium baculatum (WP_015773055.1, 36% identical), Fischerella muscicola The VnfD polypeptide from A. muscicola (WP_016867598.1, 34% identity) and the AnfD polypeptide from the Opitutaceae bacterium TAV5 (WP_009512873.1, 33% identity) have been reported.
[0183] The NifD polypeptide has been described and reviewed in Lawson and Smith (2002), Kim and Rees (1994), Eady (1996), Robson et al., (1989), Dilworth et al., (1988), Dilworth et al., (1993), Miller and Eady (1988), Chiu et al., (2001), Mayer et al., (1999), and Tezcan et al., (2005).
[0184] The NifD polypeptide of the iron-molybdenum subclass is a key subunit of the nitrogenase complex and is the α subunit of the α2β2MoFe protein complex at the core of nitrogenase and the site of substrate reduction by the FeMo cofactor.
[0185] As used herein, a functional NifD polypeptide is a NifD polypeptide that is capable of forming a functional nitrogenase protein complex together with other required subunits, e.g., NifH and NifK, and FeMo or other cofactors.
[0186] As used herein, "NifK polypeptide" refers to a polypeptide whose sequence is at least 31% identical to the amino acid sequence provided as SEQ ID NO:7, and which comprises amino acids comprising one or more of the conserved domains cd01974, TIGR01286, or cd01973, where the NifK polypeptide is within the subclass of VnfK polypeptides. As used herein, NifK polypeptides include VnfK polypeptides from iron-vanadium nitrogenases. Naturally occurring NifK polypeptides are typically 430-530 amino acids in length. Numerous NifK polypeptides have been identified, and many sequences are available in publicly available databases. For example, the NifK polypeptide is found in Klebsiella michiganensis (Accession number WP_049080161.1, 99% identical to SEQ ID NO: 7), Raoultella ornithinolytica (WP_044347163.1, 96% identical), Klebsiella variicola (SBM87811.1, 94% identical), Kluyvera intermediate (WP_047370272.1, 89% identical), Rahnella aquatilis (WP_014333919.1, 82% identical), Tolumonas auensis (WP_014333919.1, 82% identical), and auensis (WP_012728880.1, 75% identical), Pseudomonas stutzeri (WP_011912506.1, 68% identical), Vibrio natriegens (WP_065303473.1, 65% identical), Azoarcus toluclasticus (WP_018989051.1, 54% identical), Frankia sp. (prf||2106319A, 50% identical), and Methanosarcina acetivorans (WP_011021239.1, 31% identical).There are several examples of polypeptides in the database annotated as "NifK" that have less than 31% identity to SEQ ID NO: 7 but do not contain any of the domains listed above and are therefore not included in the NifK polypeptide herein. NifK polypeptides have been described and reviewed in Kim and Rees (1994), Eady (1996), Robson et al., (1989), Dilworth et al., (1988), Dilworth et al., (1993), Miller and Eady (1988), Igarashi and Seefeldt (2003), Fani et al., (2000), and Rubio and Ludden (2005).
[0187] NifK polypeptides of the iron-molybdenum subclass are essential subunits of the nitrogenase complex and are the β subunit of the αβMoFe protein complex at the nitrogenase core. As used herein, a functional NifK polypeptide is one that can form a functional nitrogenase protein complex together with other necessary subunits, such as NifD and NifH, and FeMo or other cofactors. In a preferred embodiment, when aligned with SEQ ID NO:7, the amino acid sequence of a NifK polypeptide of the invention has the amino acid sequence DLVR (residues 517-520 of SEQ ID NO:7) at its C-terminus, and the arginine is the C-terminal amino acid. That is, the NifK polypeptides and NifK fusion polypeptides of the invention preferably have the same C-terminus as native NifK polypeptides, i.e., they lack an artificial addition to the C-terminus. Such preferred NifK polypeptides are better able to form functional nitrogenase complexes with NifD and NifH polypeptides, as shown in the Examples.
[0188] The native bacterial NifB polypeptide is a protein involved in FeMo-Co synthesis and converts the [4Fe-4S] cluster into NifB-co, a higher-nuclearity Fe-S cluster with a central C atom that serves as the precursor of FeMo-Co (Guo et al., 2016). Thus, NifB catalyzes the first committed step in the FeMo-Co synthesis pathway. The NifB-co product of NifB can bind to the NifE-NifN complex and be shuttled from NifB to NifE-NifN by the metallocluster carrier protein NifX.
[0189] As used herein, "NifB polypeptide" refers to a polypeptide whose amino acid sequence comprises amino acids at least 27% identical to the amino acid sequence provided as SEQ ID NO:9. Most NifB polypeptides contain one or more of the conserved domains TIGR01290, NifB conserved domain cd00852, NifX-NifB superfamily conserved domain cl00252, and Radical_SAM conserved domain cd01335. As used herein, NifB polypeptides include naturally occurring polypeptides annotated as having NifB function but which do not contain one of these domains. Naturally occurring NifB polypeptides are typically 440-500 amino acids in length, and the naturally occurring monomer has a molecular weight of approximately 50 kDa. Numerous NifB polypeptides have been identified, and many sequences are available in publicly available databases.For example, the NifB polypeptide is found in Raoultella ornithinolytica (accession number WP_041145602.1, 91% identical to SEQ ID NO: 9), Kosakonia radicincitans (WP_043953592.1, 80% identical), Dickeya chrysanthemi (WP_040003311.1, 76% identical), Pectobacterium atrosepticum (WP_011094468.1, 70% identical), Brenneria gordwinii (WP_011094468.1, 70% identical), and other strains of bacteria. goodwinii (WP_048638849.1, 63% identical), Halorhodospira halophila (WP_011813098.1, 59% identical), Methanosarcina barkeri (WP_048108879.1, 50% identical), Clostridium purinilyticum (WP_050355163.1, 40% identical), Geofilum rubicundum (GAO28552.1, 35% identical), and Desulfovibrio salexigens (WP_015850328.1, 27% identical). As used herein, a "functional NifB polypeptide" is a NifB polypeptide that can form NifB-co from a [4Fe-4S] cluster. NifB polypeptides are described and reviewed in Curatti et al. (2006) and Allen et al. (1995).
[0190] The NifEN complex is a scaffolding complex required for the correct assembly of dinitrogenase and is structurally similar to dinitrogenase (Fay et al., 2016). The NifEN complex consists of two subunits, NifE and NifN, and forms a heterotetramer, referred to herein as ENα2β2. The native bacterial NifE polypeptide is the α subunit of the ENα2β2 tetramer along with the NifN polypeptide. This ENα2β2 tetramer is required for FeMo-Co synthesis and is proposed to function as a scaffold on which FeMo-Co is synthesized.
[0191] As used herein, "NifE polypeptide" refers to a polypeptide whose sequence is at least 32% identical to the amino acid sequence provided as SEQ ID NO: 10 and which comprises amino acids comprising one or both of the domains TIGR01283 and PRK14478. Members of the TIGR01283 domain protein family are also members of the superfamily cl02775. Naturally occurring NifE polypeptides are typically 440-490 amino acids in length, and the native monomer has a molecular weight of approximately 50 kDa. Numerous NifE polypeptides have been identified, and many sequences are available in publicly available databases.For example, the NifE polypeptide is a member of the NifE family of bacteria found in Klebsiella michiganensis (accession number WP_049114606.1, 99% identical to SEQ ID NO: 10), Klebsiella variicola (SBM87755.1, 92% identical), Dickeya paradisiaca (WP_012764127.1, 89% identical), Tolumonas auensis (WP_012728883.1, 75% identical), Pseudomonas stutzeri (WP_003297989.1, 69% identical), Azotobacter vinelandii, and the like. vinelandii (WP_012698965.1, 62% identical), Trichormus azollae (WP_013190624.1, 55% identical), Paenibacillus durus (WP_025698318.1, 50% identical), Sulfuricurvum kujiense (WP_013460149.1, 44% identical), Methanobacterium formicicum (AIS31022.1, 39% identical), Anaeromusa acidaminophila (WP_018701501.1, 35% identical), and Megasphaera cerevisiae cerevisiae) (WP_048514099.1, 32% identity). As used herein, a "functional NifE polypeptide" is a NifE polypeptide that can form a functional tetramer together with NifN such that the complex can synthesize FeMo-Co. NifE polypeptides have been described and reviewed in Fay et al. (2016), Hu et al. (2005), Hu et al. (2006), and Hu et al. (2008).
[0192] The naturally occurring bacterial NifF polypeptide is a flavodoxin, an electron donor to NifH. As used herein, "NifF polypeptide" refers to a polypeptide whose sequence contains amino acids at least 34% identical to the amino acid sequence provided as SEQ ID NO: 16 and includes one or both of the flavodoxin long domain, domain TIGR01752, and the flavodoxin FLDA domain found in Nif proteins from the Azobacter and other bacterial genera PRK09267. NifF polypeptides include flavodoxins associated with pyruvate formate lyase activation and cobalamin-dependent methionine synthase activity in non-nitrogen-fixing bacteria, but exclude other flavodoxins involved in broader functions. Naturally occurring NifF polypeptides are typically 160-200 amino acids in length, and the naturally occurring monomer has a molecular weight of approximately 19 kDa. Numerous NifF polypeptides have been identified, and numerous sequences are available in publicly available databases.For example, the NifF polypeptide is a member of the NifF family of bacteria found in Klebsiella michiganensis (accession number WP_004122417.1, 99% identical to SEQ ID NO: 16), Klebsiella variicola (WP_040968713.1, 85% identical), Kosakonia radicincitans (WP_035885760.1, 76% identical), Dickeya chrysanthemi (WP_039999438.1, 72% identical), Brenneria goodwinii (WP_048638838.1, 62% identical), Methylomonas methanica, and the like. It has been reported in Azotobacter methanica (WP_064006977.1, 56% identical), Azotobacter vinelandii (WP_012698862.1, 50% identical), Chlorobaculum tepidum (WP_010933399.1, 39% identical), Campylobacter showae (WP_002949173.1, 37% identical), and Azotobacter chromococcum (WP_039801725.1, 34% identical). As used herein, a "functional NifF polypeptide" is a NifF polypeptide that is an electron donor to a NifH polypeptide. The NifF polypeptide was described and reviewed in Drummond (1985).
[0193] The native bacterial NifJ polypeptide is a pyruvate:flavodoxin (ferredoxin) oxidoreductase, an electron donor to NifH. As used herein, "NifJ polypeptide" refers to a polypeptide whose sequence is at least 40% identical to the amino acid sequence provided as SEQ ID NO: 18 and contains amino acids that comprise the conserved domain TIGR02176. Native NifJ polypeptides are typically 1100-1200 amino acids in length, and the native monomer has a molecular weight of approximately 128 kDa. Numerous NifJ polypeptides have been identified, and many sequences are available in publicly available databases. For example, the NifJ polypeptide is found in Klebsiella michiganensis (Accession number WP_024360006.1, 99% identical to SEQ ID NO: 18), Raoultella ornithinolytica (WP_044347157.1, 95% identical), Klebsiella quasipneumoniae (WP_050533844.1, 92% identical), Kosakonia oryzae (WP_064566543.1, 82% identical), Dickeya solani (WP_057084649.1, 78% identical), Rahnella aquatilis (WP_057084649.1, 78% identical), and Rhahnella aquatilis (WP_014683040.1, 72% identical), Thermoanaerobacter mathranii (WP_013149847.1, 64% identical), Clostridium botulinum (WP_053341220.1, 60% identical), Spirochaeta africana (WP_014454638.1, 52% identical), and Vibrio cholerae (CSA83023.1, 40% identical). As used herein, a "functional NifJ polypeptide" is a NifJ polypeptide that is capable of being an electron donor to a NifH polypeptide.The NifJ polypeptide was described and reviewed in Schmitz et al., (2001).
[0194] The native bacterial NifM polypeptide is a polypeptide required for the maturation of NifH. In the absence of NifM, NifH was present at only low levels in E. coli and yeast when heterologously expressed and unable to donate electrons to NifD-NifK. As used herein, "NifM polypeptide" refers to a polypeptide whose sequence is at least 26% identical to the amino acid sequence provided as SEQ ID NO: 19 and contains the amino acid domain TIGR02933. NifM polypeptides are homologous to peptidyl-prolyl cis-trans isomerases and are thought to be accessory proteins of NifH. Native NifM polypeptides are typically 240-300 amino acids in length, and the native monomer has a molecular weight of approximately 30 kDa. Numerous NifM polypeptides have been identified, and many sequences are available in publicly available databases.For example, the NifM polypeptide is a member of the NifM family of proteins found in Klebsiella oxytoca (accession number WP_064342940.1, 99% identical to SEQ ID NO: 19), Klebsiella michiganensis (WP_004122413.1, 97% identical), Raoultella ornithinolytica (WP_044347181.1, 85% identical), Klebsiella variicola (WP_063105800.1, 75% identical), Kosakonia radicinsitans, and other organisms. radicincitans (WP_035885759.1, 59% identical), Pectobacterium atrosepticum (WP_011094472.1, 42% identical), Brenneria goodwinii (WP_048638837.1, 33% identical), Pseudomonas aeruginosa PAO1 (CAA75544.1, 28% identical), Marinobacterium sp. AK27 (WP_051692859.1, 27% identical), and Teredinibacter turnerae (WP_018415157.1, 26% identical). As used herein, a "functional NifM polypeptide" is a NifM polypeptide that can complex with a NifH polypeptide for maturation of the NifH polypeptide. NifM polypeptides are described and reviewed in Petrova et al. (2000).
[0195] The native bacterial NifN polypeptide is the β subunit of an ENα2β2 tetramer along with the NifE polypeptide, and the ENα2β2 tetramer is proposed to be required for FeMo-Co synthesis and to function as a scaffold on which FeMo-Co is synthesized. As used herein, "NifN polypeptide" means (i) a polypeptide whose sequence comprises amino acids at least 76% identical to the sequence provided as SEQ ID NO: 11 and / or (ii) a polypeptide whose sequence comprises amino acids at least 34% identical to the sequence provided as SEQ ID NO: 11 and includes one or more of the conserved domains TIGR01285, cd01966, and PRK14476. NifN is related in structure to the molybdenum-iron protein β chain NifK. Polypeptides containing the conserved TIGR01285 domain encompass most examples of NifN polypeptides, but exclude some NifN polypeptides, such as the putative NifN from Chlorobium tepizumab. Therefore, the definition of NifN is not limited to polypeptides containing the conserved TIGR01285 domain. Members of the PRK14476 domain protein family are also members of the superfamily cl02775. Native NifN polypeptides are typically 410-470 amino acids long, but when naturally fused to NifB, they may contain approximately 900 amino acid residues, and the native monomer has a molecular weight of approximately 50 kDa. Numerous NifN polypeptides have been identified, and numerous sequences are available in publicly available databases.For example, the NifN polypeptide is a member of the NifN family of proteins found in Klebsiella oxytoca (accession number WP_064391778.1, 97% identical to SEQ ID NO: 11), Kluyvera intermediate (WP_047370268.1, 80% identical), Rahnella aquatilis (WP_014683026.1, 70% identical), Brenneria goodwinii (WP_048638830.1, 65% identical), Methylobacter tundripaludum (WP_027147663.1, 46% identical), and Calothrix parietina. It has been reported in NifN, Zymomonas mobilis (WP_023593609.1, 37% identical), Paenibacillus massiliensis (WP_025677480.1, 35% identical), and Desulfitobacterium hafniense (WP_018306265.1, 34% identical). As used herein, a "functional NifN polypeptide" is a NifN polypeptide that can form a functional tetramer together with NifE such that the complex can synthesize FeMo-Co. The NifN polypeptide was described and reviewed in Fay et al., (2016), Brigle et al., (1987), Fani et al., (2000), and Hu et al., (2005).
[0196] The native bacterial NifQ polypeptide is involved in FeMo-Co synthesis and possibly the initial MoO4 2-It is a polypeptide involved in processing. A conserved C-terminal cysteine residue may be involved in metal binding. As used herein, "NifQ polypeptide" refers to a polypeptide whose sequence comprises amino acids at least 34% identical to the amino acid sequence provided as SEQ ID NO: 12 and is a member of the CL04826 domain protein family and the pfam04891 domain protein family. Naturally occurring NifQ polypeptides are typically 160-250 amino acids in length, or they may be as long as 350 amino acid residues, and the natural monomer has a molecular weight of approximately 20 kDa. Numerous NifQ polypeptides have been identified, and many sequences are available in publicly available databases. For example, NifQ polypeptides are known to bind to Klebsiella oxytoca (accession number WP_064391765.1, 95% identical to SEQ ID NO: 12), Klebsiella variicola (CTQ06350.1, 75% identical), Kluyvera intermediate (WP_047370257.1, 63% identical), Pectobacterium atrosepticum (WP_043878077.1, 59% identical), Mesorhizobium metallidurans (WP_008878174.1, 46% identical), Rhodopseudomonas palustris, and the like. palustris (WP_011501504.1, 42% identical), Paraburkholderia sprentiae (WP_027196569.1, 41% identical), Burkholderia stabilis (GAU06296.1, 39% identical), and Cupriavidus oxalaticus (WP_063239464.1, 34% identical). As used herein, a "functional NifQ polypeptide" refers to a polypeptide that is a member of the NifQ family of polypeptides. 2-The NifQ polypeptide is a NifQ polypeptide that enables processing of the NifQ polypeptide. The NifQ polypeptide was described and reviewed in Allen et al., (1995) and Siddavattam et al., (1993).
[0197] The naturally occurring bacterial NifS polypeptide is a cysteine desulfurase involved in iron-sulfur (FeS) cluster biosynthesis, e.g., mobilizing sulfur for Fe-S cluster synthesis and repair. As used herein, "NifS polypeptide" refers to (i) a polypeptide whose sequence comprises amino acids at least 90% identical to the amino acid sequence provided as SEQ ID NO: 13 and / or (ii) a polypeptide whose sequence comprises amino acids at least 36% identical to the sequence provided as SEQ ID NO: 13 and comprises one or both of the conserved domains TIGR03402 and COG1104. In addition to the branch almost always found in extended nitrogen-fixing systems, the TIGR03402 domain protein family includes a second branch that is more closely related to the first than to IscS and is also part of NifS-like / NifU-like systems. The TIGR03402 domain protein family, also referred to in the literature as NifS, is instead constructed in TIGR03403 and does not extend to more distant branches found in epsilonproteobacteria such as Helicobacter pylori. The COG1104 domain protein family contains cysteine sulfinate desulfinase / cysteine desulfurase or related enzymes. Some NifS polypeptides contain the aspartate aminotransferase domain cl18945. Native NifS polypeptides are typically 370–440 amino acids in length, and the native monomer has a molecular weight of approximately 43 kDa. Numerous NifS polypeptides have been identified, and numerous sequences are available in publicly available databases.For example, the NifS polypeptide is found in Klebsiella michiganensis (accession number WP_004138780.1, 99% identical to SEQ ID NO: 13), Raoultella terrigena (WP_045858151.1, 89% identical), Kluyvera intermediate (WP_047370265.1, 80% identical), Rahnella aquatilis (WP_014333911.1, 73% identical), Agarivorans gilvus (WP_055731597.1, 64% identical), Azospirillum brasilens (WP_055731597.1, 64% identical), and Azospirillum spp. brasilense (WP_014239770.1, 60% identical), Desulfosarcina cetonica (WP_054691765.1, 55% identical), Clostridium intestinale (WP_021802294.1, 47% identical), Clostridiisalibacter paucivorans (WP_026894054.1, 36% identical), and Bacillus coagulans (WP_061575621.1, 42% identical and present in COG1104). As used herein, a "functional NifS polypeptide" is a NifS polypeptide that can function in iron-sulfur (FeS) cluster biosynthesis and / or repair. NifS polypeptides have been described and reviewed in Clausen et al. (2000), Johnson et al. (2005), Olson et al. (2000), and Yuvaniyama et al. (2000).
[0198] The native bacterial NifU polypeptide is a molecular scaffold polypeptide involved in iron-sulfur (FeS) cluster biosynthesis for nitrogenase components. As used herein, "NifU polypeptide" refers to a polypeptide whose sequence comprises amino acids at least 31% identical to the sequence provided as SEQ ID NO: 14 and contains the domain TIGR02000. Members of the TIGR02000 domain protein family are specifically involved in nitrogenase maturation. NifU contains an N-terminal domain (pfam01592) and a C-terminal domain (pfam01106). Three distinct but partially homologous Fe-S cluster assembly systems have been described: Isc, Suf, and Nif. The Nif system (of which NifU is a part) is involved in the donation of Fe-S clusters to nitrogenase in many nitrogen-fixing species. Isc and Suf homologs with domain structures equivalent to those in Helicobacter and Campylobacter are excluded from the definition of NifU herein. Therefore, NifU is specific to the NifU polypeptide involved in nitrogenase maturation. Members of the family of related TIGR01999 domain proteins, including IscU proteins (e.g., from Escherichia coli, Saccharomyces cerevisiae, and Homo sapiens) that contain homologs of the N-terminal region of NifU, are also excluded from the definition of NifU herein. Native NifU polypeptides are typically 260–310 amino acids in length, and the native monomer has a molecular weight of approximately 29 kDa. Numerous NifU polypeptides have been identified, and numerous sequences are available in publicly available databases.For example, the NifU polypeptide is a member of the NifU family of proteins found in Klebsiella michiganensis (accession number WP_049136164.1, 97% identical to SEQ ID NO: 14), Klebsiella variicola (WP_050887862.1, 90% identical), Dickeya solani (WP_057084657.1, 80% identical), Brenneria goodwinii (WP_048638833.1, 73% identical), Tolumonas auensis (WP_012728889.1, 66% identical), Agarivorans gillbus, and the like. NifU has been reported in Bacillus gilvus (WP_055731596.1, 58% identical), Desulfocurvus vexinensis (WP_028587630.1, 54% identical), Rhodopseudomonas palustris (WP_044417303.1, 49% identical), Helicobacter pylori (WP_001051984.1, 31% identical), and Sulfurovum sp. PC08-66 (KIM05011.1, 31% identical). As used herein, a "functional NifU polypeptide" is a NifU polypeptide that can function as a molecular scaffold polypeptide involved in iron-sulfur (FeS) cluster biosynthesis. The NifU polypeptide was described and reviewed in Hwang et al., (1996), Muhlenhoff et al., (2003) and Ouzounis et al., (1994).
[0199] The naturally occurring bacterial NifV polypeptide is a homocitrate synthase (EC 2.3.3.14) and produces homocitrate by the transfer of an acetyl group from acetyl-coenzyme A (acetyl-CoA) to 2-oxoglutarate. Homocitrate is then used in the synthesis of FeMo-Co. As used herein, "NifV polypeptide" means a polypeptide whose sequence comprises amino acids at least 39% identical to the amino acid sequence provided as SEQ ID NO: 20 and comprises one or both of the domains TIGR02660 and DRE_TIM. Members of the TIGR02660 domain protein family are homologous to enzymes involved in processes other than nitrogen fixation, including 2-isopropylmalate synthase, (R)-citramalate synthase, and homocitrate synthase. The cd07939 domain protein family also includes the NifV proteins of Heliobacterium chlorum and Gluconacetobacter diazotrophicus, which appear to be orthologous to FrbC. This family belongs to the DRE-TIM metallolyase superfamily, which includes 2-isopropylmalate synthase (IPMS), α-isopropylmalate synthase (LeuA), 3-hydroxy-3-methylglutaryl-CoA lyase, homocitrate synthase, citramalate synthase, 4-hydroxy-2-oxovalerate aldolase, re-citrate synthase, transcarboxylase 5S, pyruvate carboxylase, AksA, and FrbC. All of these members share a conserved triose-phosphate isomerase (TIM) barrel domain consisting of a core β(8)-α(8) motif with eight parallel β-strands forming a closed barrel surrounded by eight α-helices, harboring a catalytic center containing a divalent cation-binding site formed by a cluster of invariant residues that line the barrel core.Additionally, the catalytic site contains three invariant residues—aspartate (D), arginine (R), and glutamic acid (E)—that are the basis for the domain name "DRE-TIM." Natural NifV polypeptides are typically 360-390 amino acids in length, although some members are approximately 490 amino acid residues in length, and the natural monomer has a molecular weight of approximately 41 kDa. Numerous NifV polypeptides have been identified, and many sequences are available in publicly available databases. For example, NifV polypeptides are found in Klebsiella michiganensis (Accession number WP_049083341.1, 95% identical to SEQ ID NO: 20), Raoultella ornithinolytica (WP_045858154.1, 86% identical), Kluyvera intermediate (WP_047370264.1, 81% identical), Dickeya dadantii (WP_038912041.1, 70% identical), Brenneria goodwinii (WP_048638835.1, 59% identical), Magnetococcus marinus, and other pathogenic bacteria. marinus (WP_011712856.1, 46% identical), Sphingomonas wittichii (WP_037528703.1, 43% identical), Frankia sp. EI5c (OAA29062.1, 41% identical), and Clostridium sp. Maddingley MBC34-26 (EKQ56006.1, 39% identical). As used herein, a "functional NifV polypeptide" is a NifV polypeptide that can function as a homocitrate synthase. NifV polypeptides have been described and reviewed in Hu et al. (2008), Lee et al. (2000), Masukawa et al. (2007), and Zheng et al. (1997).
[0200] Natural bacterial NifX polypeptides are polypeptides involved in FeMo-Co synthesis, at least in the transfer of the FeMo-Co precursor from NifB to NifE-NifN. As used herein, "NifX polypeptide" refers to a polypeptide whose sequence contains amino acids at least 29% identical to the amino acid sequence provided as SEQ ID NO: 15 and contains one or both of the conserved domains TIGR02663 and cd00853. NifX is included within a larger family of iron-molybdenum cluster-binding proteins, including NifB and NifY; i.e., NifX, NafY, and the C-terminal domain of NifB all contain the pfam02579 domain and are each involved in FeMo-Co synthesis. Therefore, some NifX polypeptides are annotated in databases as NifY, and vice versa. Natural NifX polypeptides are typically 110-160 amino acids in length, and the natural monomer has a molecular weight of approximately 15 kDa. Numerous NifX polypeptides have been identified and many sequences are available in publicly available databases.For example, the NifX polypeptide is a member of the NifX family of proteins found in Klebsiella michiganensis (accession number WP_049070199.1, 97% identical to SEQ ID NO: 15), Klebsiella oxytoca (WP_064342937.1, 97% identical), Raoultella ornithinolytica (WP_044347173.1, 91% identical), Klebsiella variicola (WP_044612922.1, 83% identical), Kosakonia radicincitans (WP_043953583.1, 75% identical), Dickeya chrysanthemi (WP_044612922.1, 83% identical), and Klebsiella michiganensis (WP_044612922.1, 83% identical). chrysanthemi (WP_039999416.1, 68% identical), Rahnella aquatilis (WP_047608097.1, 58% identical), Azotobacter chroococcum (WP_039800848.1, 34% identical), Beggiatoa leptomitiformis (WP_062149047.1, 33% identical), and Methyloversatilis discipulorum (WP_020165972.1, 29% identical). As used herein, a "functional NifX polypeptide" is a NifX polypeptide that can transfer the FeMo-Co precursor from NifB to NifE-NifN. NifX polypeptides have been described and reviewed in Allen et al., (1994) and Shah et al., (1999).
[0201] The naturally occurring bacterial NifY polypeptide is a polypeptide involved in FeMo-Co synthesis, at least in assisting in the transfer of the FeMo-Co precursor from NifB to NifE-NifN. As used herein, "NifY polypeptide" refers to a polypeptide whose sequence comprises amino acids at least 34% identical to the amino acid sequence provided as SEQ ID NO:8 and contains one or both of the conserved domains TIGR02663 and cd00853. NifY is within a larger family of iron-molybdenum cluster-binding proteins that includes NifB and NifX; i.e., NifX and NafY, as well as the C-terminal domain of NifB, all contain the pfam02579 domain and each is involved in the synthesis of FeMo-Co. Numerous NifY polypeptides have been identified, and many sequences are available in publicly available databases. For example, the NifY polypeptide is found in Klebsiella michiganensis (Accession number WP_049089500.1, 99% identical to SEQ ID NO: 8), Klebsiella oxytoca (WP_064342935.1, 98% identical), Klebsiella quasipneumoniae (WP_044524054.1, 90% identical), Klebsiella variicola (WP_049010739.1, 81% identical), Kluyvera intermediate (WP_047370270.1, 69% identical), Dickeya chrysanthemi chrysanthemi (WP_039999411.1, 62% identical), Serratia sp. ATCC39006 (WP_037382461.1, 57% identical), Rahnella aquatilis (WP_014683024.1, 47% identical), Pseudomonas putida (AEX25784.1, 37% identical), and Azotobacter vinelandii (WP_012698835.1, 34% identical).As used herein, a "functional NifY polypeptide" is a NifY polypeptide that is capable of transferring the FeMo-Co precursor from NifB to NifE-NifN.
[0202] Naturally occurring bacterial NifZ polypeptides are polypeptides involved in Fe-S cluster synthesis, specifically required for coupling of the second Fe4S4 pair. As used herein, "NifZ polypeptide" refers to a polypeptide whose sequence contains amino acids at least 28% identical to the sequence provided as SEQ ID NO: 17 and contains the conserved domain pfam04319. This domain of approximately 75 amino acid residues is found in some isolated members and at the amino-terminal half of longer NifZ proteins. Naturally occurring NifZ polypeptides are typically 70 to 150 amino acids in length, and naturally occurring monomers have molecular weights of about 9 to about 16 kDa. Numerous NifZ polypeptides have been identified, and many sequences are available in publicly available databases. For example, the NifZ polypeptide is found in Klebsiella michiganensis (accession number WP_057173223.1, 93% identical to SEQ ID NO: 17), Klebsiella oxytoca (WP_064342939.1, 95% identical), Klebsiella variicola (WP_043875005.1, 77% identical), Kosakonia radicincitans (WP_043953588.1, 67% identical), Kosakonia sacchari (WP_065368553.1, 58% identical), Ferriphaselus amnicola (WP_065368553.1, 58% identical), and the like. amnicola (WP_062627625.1, 47% identical), Paraburkholderia xenovorans (WP_011491838.1, 41% identical), Acidithiobacillus ferrivorans (WP_014029050.1, 35% identical), and Bradyrhizobium oligotrophicum (WP_015665422.1, 28% identical).As used herein, a "functional NifZ polypeptide" is a NifZ polypeptide that can ligate an Fe4S4 cluster in Fe-S cluster synthesis. NifZ polypeptides are described and reviewed in Cotton (2009) and Hu et al., (2004).
[0203] Native bacterial NifW polypeptides are polypeptides that bind to NifZ polypeptides to form higher-order complexes (Lee et al., 1998) and are involved in MoFe protein (NifD-NifK) synthesis or activity. NifW and NifZ appear to be involved in the formation or accumulation of MoFe proteins (Paul and Merrick, 1989). As used herein, "NifW polypeptide" refers to a polypeptide whose amino acid sequence comprises amino acids (which sequence is at least 28% identical to the amino acid sequence provided as SEQ ID NO: 74) and contains the conserved NifW superfamily protein domain, architecture ID number 10505077, and is present within Pfamily PF03206. Numerous NifW polypeptides have been identified, and many sequences are available in publicly available databases.For example, the NifW polypeptide is a member of the NifW family of proteins found in Klebsiella oxytoca (accession number WP_064342938.1, 98% identical to SEQ ID NO: 74), Klebsiella michiganensis (WP_049080155.1, 94% identical), Enterobacter sp. 10-1 (WP_095103586.1, 90% identical), Klebsiella quasipneumoniae (WP_065877373.1, 81% identical), Pectobacterium polaris (WP_095699971.1, 69% identical), Dickeya paradisiaca, and other strains of Enterobacter sp. paradisiaca (WP_012764136.1, 58% identical), Brenneria goodwinii (WP_053085547.1, 36% identical), Aquaspirillum sp. LM1 (WP_077299824.1, 44% identical), Candidatus Muproteobacteria RBG_16_64_10 (OGI40729, 34% identical), Azotobacter vinelandii (ACO76430.1, 32% identical), and Methylocaldum marinum (BBA37427.1, 28% identical). As used herein, a "functional NifW polypeptide" is a NifW polypeptide that promotes or facilitates one or more of the formation, accumulation, or activity of the MoFe protein. A functional NifW may interact with NifZ and / or play a role in oxygen protection of the MoFe protein (Gavini et al., (1998)).
[0204] It will be understood that with respect to a defined polypeptide or enzyme, forms of higher % identity than those provided above encompass preferred embodiments. Thus, where applicable, in terms of minimum % identity, it is preferred that the polypeptide comprises an amino acid sequence that is at least 30%, more preferably at least 35%, more preferably at least 40%, more preferably at least 45%, more preferably at least 50%, more preferably at least 55%, more preferably at least 60%, more preferably at least 65%, more preferably at least 70%, more preferably at least 75%, more preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, and even more preferably at least 99.9% identical to the relevant recited SEQ ID NO.
[0205] Amino acid sequence mutants of the polypeptides defined herein can be prepared by introducing appropriate nucleotide changes into the nucleic acids defined herein or by in vitro synthesis of the desired polypeptide. Such mutants include, for example, deletion, insertion, or substitution of one or more amino acids. Combinations of deletion, insertion, and substitution mutations can be made to arrive at the final construct, provided that the final peptide product possesses the desired enzymatic activity.
[0206] Mutant (altered) polypeptides can be prepared using any technique known in the art, for example, using directed evolution or rational design strategies (see below). Products resulting from mutant / altered DNA can be readily screened using the techniques described herein to determine if their expression in a plant alters its phenotype relative to the corresponding wild-type plant, for example, if their expression results in increased yield, biomass, growth rate, vitality, nitrogen acquisition from biological nitrogen fixation, nitrogen use efficiency, abiotic stress tolerance, and / or tolerance to nutrient deficiency relative to the corresponding wild-type plant.
[0207] In designing mutants of an amino acid sequence, the location of the mutation points and the nature of the mutation will depend on the characteristic(s) to be altered. The mutation locations can be varied individually or sequentially, for example, by (1) first substituting conservative amino acid selection followed by innovative selection depending on the results achieved, (2) deleting the target residue, or (3) inserting other residues adjacent to the location to be located.
[0208] The amino acid sequence generally deletes a range of about 1 to 15 residues, more preferably about 1 to 10 residues and typically about 1 to 5 contiguous residues.
[0209] Substitution mutants have at least one amino acid residue in the polypeptide molecule removed and a different residue inserted in its place. If it is desired to maintain a particular activity, conservative substitutions are preferably absent or only at amino acid positions that are highly conserved across related protein families. Examples of conservative substitutions are shown in Table 1 under the heading "Representative Substitutions."
[0210] In preferred embodiments, mutant / variant polypeptides have only one, or no more than one, two, three, or four conservative amino acid changes when compared to a naturally occurring polypeptide. Details of conservative amino acid changes are provided in Table 1. In preferred embodiments, the changes do not occur in one or more motifs or domains that are highly conserved among different polypeptides of the present invention. As one of skill in the art will recognize, such minor changes can reasonably be expected not to alter the activity of the polypeptide when expressed in a genetically modified cell.
[0211] The primary amino acid sequence of a polypeptide of the present invention can be used to design variants / mutants based on comparison with closely related polypeptides. As one skilled in the art will appreciate, residues that are highly conserved between closely related proteins are less likely to alter the retained activity than less conserved residues, especially with non-conservative substitutions (see above). A more stringent test for identifying conserved amino acid residues is the more distantly related polypeptides with the same function are aligned. Highly conserved residues must be maintained to retain function, while non-conserved residues are amenable to substitution or deletion while maintaining function.
[0212] Also included within the scope of the invention are polypeptides of the invention that are differentially modified during or after intracellular synthesis, for example, by glycosylation, acetylation, phosphorylation, or proteolytic cleavage.
[0213] Table 1. Representative substitutions [Table 1]
[0214] Rational Design Proteins can be rationally designed based on known information about protein structure and folding. This can be achieved by design from scratch (de novo design) or by redesign based on natural scaffolds (see, e.g., Hellinga, 1997; and Lu and Berry, Protein Structure Design and Engineering, Handbook of Proteins 2, 1153-1157 (2007)). See, e.g., Example 10 herein. Protein design typically involves identifying a sequence that folds into a predetermined or target structure and is accomplished using a computer model. Computational protein design algorithms search sequence conformational space for sequences that have low energy when folded into the target structure. Computational protein design algorithms use models of protein energetics to evaluate how mutations affect protein structure and function. These energy functions typically include a combination of molecular mechanics, statistics (i.e., knowledge-based), and other empirical terms. Suitable available software includes IPRO (Interative Protein Redesign and Optimization), EGAD (A Genetic Algorithm for Protein Design), Rosetta Design, Sharpen, and Abalone.
[0215] Mitochondrial protein import in plants Nearly all mitochondrial proteins are encoded in the nucleus and imported into the cytosol, thus requiring their translocation into mitochondria. Signal sequences in polypeptides direct their import into four different mitochondrial locations: the outer membrane (OM), the intermembrane space (IS), the inner membrane (IM), or the matrix (MM). These signal sequences are distinguished by their biochemical properties and guide transport through at least four different import pathways that direct polypeptides to one or more of the four locations (Chacinska et al., 2009). These four pathways are as follows: (1) the general import pathway, also known as the "classical" presequence pathway, which targets polypeptides to the MM, IS, or IM; (2) the carrier import pathway, which is used for import into the IM; (3) the mitochondrial intermembrane space (MIA) assembly pathway; and (4) the sorting and assembly machinery (SAM) pathway, which is used for import of polypeptides into the OM. The general import pathway also imports polypeptides with cleavable presequences, also known as signal sequences. These polypeptides may also have hydrophobic sorting signals (HSSs). The carrier import pathway imports polypeptides with internal presequences, such as signals or hydrophobic regions. The MIA pathway imports polypeptides with twin cysteine residues. The SAM pathway imports polypeptides containing a β signal and a putative TOM20 signal. All of these pathways use translocases in the outer membrane (TOM), and the first and second pathways also use the TIM23 translocase in the intermembrane complex. Only the first pathway uses a matrix processing peptidase (matrix processing protease, MPP).
[0216] A common feature of all mitochondrial-targeting polypeptides is the presence of at least one domain within the polypeptide that directs translocation to a precise location. The most well-studied of these is the "classical" N-terminal presequence domain, which is cleaved by MPP within the matrix (Murcha et al., 2004). Although approximately 70% of plant and animal mitochondrial proteins possess a cleavable presequence, both internal and C-terminal signal sequences have also been found (reviewed in Pfanner and Geissler (2001) and Schleiff and Soll (2000)). In Arabidopsis, these presequences range in length from 11 to 109 amino acid residues, with an average length of 50 amino acid residues. While there is no consensus sequence that completely defines presequences for the first pathway, they tend to contain a high proportion of hydrophobic and positively charged amino acids. An additional characteristic is their ability to form amphipathic α-helices, usually beginning within the first 10 amino acid residues (Roise et al., 1986). These domains are rich in hydrophobic (Ala, Leu, Phe, Val), hydroxylated (Ser, Thr), and electropositive (Arg, Lys) amino acid residues and lack acidic amino acids. For many mitochondrial proteins, serine (16–17%) and alanine (12–13%) are overrepresented in mitochondrial signal peptides, and arginine is abundant (12%). The MPP cleavage point is defined for most presequences by the presence of a conserved arginine residue, usually at position P2 (-2 aa from the uniformly cleaved bond) or, in most other cases, at P3 (Huang et al., 2009).
[0217] Mitochondrial presequences interact with the Tom20 receptor through hydrophobic residues. Studies have shown that the hydrophobic surface of the α-helix facilitates peptide recognition by the TOM20 component of the TOM import complex, while the positive charge is recognized by the TOM22 subunit (Abe et al., 2000). Finally, most presequences direct the translocation of polypeptides bound to Hsp70; therefore, almost all plant presequences contain at least one binding motif for the Hsp70 molecular chaperone (Zhang and Glaser, 2002). The chaperone Hsp70 is involved in protein folding, prevention of protein aggregation, and functioning as a molecular motor, pulling precursors across the mitochondrial membrane. The membrane potential (ΔΨ) across the inner membrane (~100 mV, negative inside) also drives the translocation of positively charged presequences via an electrophoretic effect.
[0218] The majority of proteins with cleavable presequences are targeted to the mitochondrial matrix via a common import pathway, which utilizes the transporter of the outer membrane (TOM) complex and the transporter of the inner membrane 23 complex (TIM23). However, some proteins with cleavable presequences can assemble in the inner membrane (Murcha et al., 2005) or intermembrane space if they also contain a hydrophobic sorting signal (HSS) (Glick et al., 1992). Only a few examples exist of matrix-localized proteins that do not have their presequences cleaved. In Arabidopsis, only glutamate dehydrogenase, with its full-length unprocessed presequence, has been found in the matrix (Huang et al., 2009).
[0219] Non-matrix-targeted proteins employ a variety of endogenous non-cleavable localization signals. These are typically associated with specific transport pathways and are further tailored to specific types of proteins. In plants, no studies have yet determined the precise composition of internal signal sequences in intermembrane space proteins. However, motifs with twin cysteine residues appear to be associated with translocation through the mitochondrial intermembrane space assembly pathway (MIA) (Carrie et al., 2010; Darshi et al., 2012). Finally, non-cleavable internal sequences are also utilized by proteins targeted to the inner membrane via carrier pathways, which utilize the TOM and TIM22 machinery to insert proteins with multiple transmembrane domains (Kerscher et al., 1997; Sirrenberg et al., 1996). These sequences typically contain a hydrophobic region followed by an internal-like presequence, making them similar to N-terminal presequences but distinguishable by their internal location within their cognate proteins.
[0220] In photosynthetic organisms, the nuclear coding for mitochondrial proteins necessitated a distinction between chloroplast and mitochondrial transport, despite many similarities between these two organelles and their proteomes. The α-helices that appear mostly in mitochondrial presequences are typically absent in chloroplast presequences (Zhang and Glaser, 2002), which tend to be more unstructured and therefore exhibit a high degree of β-sheet structure (Bruce, 2001).
[0221] In plants, MPP is anchored in the inner membrane-bound Cytbc1 complex, but the active MPP site is positioned facing the matrix, so the functions of the two proteins are independent (Glaser and Dessi, 1999).
[0222] Mitochondrial Targeting Peptides As used herein, the term "mitochondrial targeting peptide" or "MTP" refers to an amino acid sequence of at least 10 amino acids in length, and preferably 10 to about 80 amino acid residues in length, that targets a target protein to mitochondria and can be used non-homologously in MTP-target protein translational fusions to target a selected target protein, such as a Nif polypeptide, Gus, or GFP, to mitochondria.
[0223] MTPs typically contain a methionine at their N-terminus, which is the translation initiator of the polypeptide from which they are derived. MTPs are translationally fused to Nif polypeptides or "target proteins" by a peptide bond to a Met residue corresponding to the initiator Met of the target protein; alternatively, the Met residue may be omitted and the peptide bond fused directly to an amino acid residue that is the second amino acid of the wild-type target protein. MTPs are typically rich in basic and hydroxylated amino acids and usually lack acidic amino acids or extended hydrophobic stretches. MTPs may form amphipathic helices.
[0224] Without wishing to be limited by theory, MTPs typically contain an import targeting sequence that binds to a receptor on the outer membrane of mitochondria. Upon binding to the outer membrane, the fusion polypeptide preferably undergoes membrane translocation, transporting the channel protein, and crosses the mitochondrial double membrane to the mitochondrial matrix (MM). The import targeting sequence is then typically cleaved, and the mature fusion protein is folded.
[0225] The MTP may subsequently contain additional signals that target the protein to different regions of the mitochondria, such as the mitochondrial matrix (MM). In certain embodiments, the uptake targeting sequence is a matrix targeting sequence.
[0226] MTP, when translationally fused to a Nif polypeptide, may be cleavable or non-cleavable. In certain embodiments, at least 50% of the MTP-Nif fusion polypeptide produced in a cell is cleaved intracellularly. In alternative embodiments, less than 50% of the MTP-Nif fusion polypeptide is cleaved intracellularly, e.g., MTP is not cleaved. In certain embodiments, MTP does not contain a cleavage site for MPP. MTP may contain a cleavage site (within the MTP sequence). Upon cleavage, the N-terminal portion of the resulting processed product (i.e., mature NP) may contain one or more C-terminal amino acids of MTP. Alternatively, the cleavage site may be located within the fusion polypeptide such that the entire MTP sequence is cleaved; for example, the linker may contain the cleavage sequence.
[0227] A unique mitochondrial targeting peptide is localized at the N-terminus of the precursor protein, and the N-terminal portion is typically cleaved off during or after import into mitochondria. Cleavage is typically catalyzed by a general matrix processing protease (MPP), which in plants is integrated into the bc1 complex of the respiratory chain. This protease recognizes cleavage sites in approximately 1,000 precursor proteins with a wide range of amino acid sequences that show little conservation. In some embodiments, the MTP contains a protease cleavage site for the MPP. In further embodiments, cleavage of the fusion protein within the MTP by the MPP results in a processed product.
[0228] In some embodiments, the MTP is not cleaved. We have demonstrated that incorporation of the MTP does not always lead to complete processing of the Nif protein. In some cases (NifX-FLAG, NifD-HA), opt1Both processed and unprocessed Nif proteins were observed (NifDK-HA and NifDK-HA). Given that there is no general consensus sequence for MTPs and that internal protein sequences can influence mitochondrial targeting (Becker et al., 2012), it may not be surprising that we found differences in processing efficiency among Nif proteins.
[0229] Suitable MTPs for use in the context of the present invention include, but are not limited to, peptides having the general structure defined by von Heijne (1986) or Roise and Schatz (1988). Non-limiting examples of MTPs are the mitochondrial targeting peptides defined in Table I of von Heijne (1986).
[0230] In one embodiment, the MTP is the FlATPase gamma-subunit (pFAγ). An example of a suitable pFAγMTP is that derived from A. thaliana (Lee et al., 2012). In one embodiment, pFAγMTP is 77 amino acids in length, and its cleavage by an MMP leaves 38 MTP residues at the N-terminus of the fusion polypeptide. In a preferred embodiment, pFAγMTP is less than 77 amino acids in length. For example, pFAγMTP may be approximately 51 amino acids in length, and its cleavage by an MMP leaves 9 MTP residues at the N-terminus of the fusion polypeptide.
[0231] Those skilled in the art will appreciate that software exists to predict mitochondrial proteins and their targeting sequences, for example, MitoProtII, PSORT, TargetP, NNPSL.
[0232] MitoProtII is a program that predicts the mitochondrial localization of a sequence based on several physiochemical parameters (e.g., the amino acid composition of the N-terminal portion or the highest overall hydrophobicity over a 17-residue window). PSORT is a program that predicts subcellular location based on various sequence-derived features, such as the presence of sequence motifs and amino acid composition. TargetP predicts the subcellular location of eukaryotic proteins based on the predicted presence of any N-terminal presequence: chloroplast transit peptide, mitochondrial targeting peptide, or secretory pathway signal peptide. TargetP requires the N-terminal sequence as input into a two-layer artificial neural network (ANN) and utilizes the initial binary predictors SignalP and ChloroP. For sequences predicted to contain an N-terminal presequence, potential cleavage sites can also be predicted. NNPSL is another ANN-based method that uses amino acid composition to assign one of four subcellular localizations (cytosolic, extracellular, nuclear, and mitochondrial) to a query sequence.
[0233] Those skilled in the art can easily determine whether a selected MTP targets a fusion polypeptide to the mitochondrial matrix based on routine methods and those disclosed herein. We chose a relatively long targeting peptide because it was previously demonstrated to be capable of transporting GFP in Arabidopsis protoplasts (Lee et al., 2012) and to aid in the detection of processed proteins. As shown in the Examples herein, the selected MTP targeted all of the selected nitrogenase proteins to the MM. This conclusion is based on several lines of evidence. First, the observed sizes of N. benthamiana-expressed Nif polypeptides were consistent with the expected sizes resulting from MM peptidase processing. This was also reflected by the size differences observed between bacterial (full-length, unprocessed) and small Nifs (NifF and NifZ) expressed in plant mitochondria. Furthermore, mutations in MTP that render it unable to be processed by the mitochondrial import machinery, resulting in larger bands for both NifD and GFP fusions, were consistent with the size difference between the processed and unprocessed proteins. Finally, mass spectrometry of a representative fusion polypeptide determined that MTP-NifH was cleaved between residues 42 and 43 of MTP, as predicted for specific processing within the matrix.
[0234] In some embodiments of the present invention, it may be useful to use multiple tandem copies of a selected MTP. The coding sequence for a dual or multiple targeting peptide may be obtained by genetic engineering from an existing MTP. The amount of MTP can be measured by cell fractionation, followed by, for example, quantitative immunoblot analysis. Thus, in the present invention, the term "mitochondrial targeting peptide" or "MTP" encompasses one or more copies of an amino acid peptide that targets Nif proteins to mitochondria. In a preferred embodiment, the MTP contains two copies of the selected MTP. In another embodiment, the MTP contains three copies of the selected MTP. In another embodiment, the MTP contains four or more copies of the selected MTP.
[0235] Those skilled in the art will recognize that the MTP sequence is not limited to the native MTP sequence, but may include amino acid substitutions, deletions and / or insertions relative to the native MTP, provided that the sequence variant is still functional for mitochondrial targeting.
[0236] Those skilled in the art will understand that an MTP may be flanked at its N- or C-terminus by amino acids as a result of a cloning strategy and may function as linkers. These additional amino acids may be considered to form part of the MTP.
[0237] Those skilled in the art will also understand that the MTP may be fused at the N- or C-terminus to an oligopeptide linker and / or a tag, such as, for example, an epitope tag. In a preferred embodiment, one or more, or all, of the Nif fusion polypeptides of the invention produced in plant cells lack an additional epitope tag relative to the corresponding wild-type Nif polypeptide.
[0238] Linker As used herein in the context of polypeptides, the term "linker" or "oligopeptide linker" refers to one or more amino acids that covalently link two or more functional domains, e.g., MTP and NP, two NPs, or NP and tag. The amino acids are covalently linked through peptide bonds both within the linker and between the linker and the functional domains. The linker can provide freedom of movement of one functional domain relative to another without causing substantial adverse effects on the function of the two or more domains. The linker can help promote proper folding and function of one or both functional domains. Those skilled in the art will understand that the size of the linker can be determined empirically or modeled based on protein folding information.
[0239] The linker may include a cleavage site for a protease, such as MPP. Such a linker may also be considered to be part of MTP.
[0240] Those skilled in the art will recognize that the C-terminus of MTP can be translationally fused to the N-terminal amino acid of NP without a linker or via a linker of one or more amino acid residues, e.g., 1 to 5 amino acid residues, such that the linker can also be considered to be part of MTP.
[0241] In embodiments, the linker comprises at least 1 amino acid, at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, at least 10 amino acids, at least 12 amino acids, at least 14 amino acids, at least 16 amino acids, at least 18 amino acids, at least 20 amino acids, at least 25 amino acids, at least 30 amino acids, at least 35 amino acids, at least 40 amino acids, a minimum of 45 amino acids, at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, or about 100 amino acids. In embodiments, the maximum size of the linker is 100 amino acids, preferably 60 amino acids, and more preferably 40 amino acids.
[0242] In some embodiments, the linker allows movement of one functional domain relative to the other to enhance the stability of the fusion polypeptide. If desired, the linker includes polyglycine repeats or a combination of glycine, proline, and alanine residues. Linkers for connecting two Nif polypeptides, such as NifD-linker-NifK and NifE-linker-NifN, are preferably selected with regard to the number and sequence of amino acids within the linker based on several criteria: the absence of cysteine residues to avoid undesired disulfide bond formation; few charged residues (Glu, Asp, Arg, Lys), preferably the absence of charged residues, to reduce the possibility of undesired surface salt-bridge interactions; few hydrophobic residues (Phe, Trp, Tyr, Met, Val, Ile, Leu), or the absence of hydrophobic residues, if such residues may promote the tendency of the polypeptide to infiltrate the surface; and the absence of amino acids that may be post-translationally modified.
[0243] In this context, "few charged residues" means less than 10% of the amino acid residues in the linker, and "few hydrophobic residues" means less than 15% of the amino acid residues in the linker. In certain embodiments, the linker does not contain any cysteine residues. In certain embodiments, the linker contains 4, 3, 2, 1, or no charged residues. Preferably, the linker contains a total of 4, 3, 2, 1, or no glutamic acid, aspartic acid, arginine, and lysine residues.
[0244] In certain embodiments, the linker contains 4, 3, 2, 1, or no hydrophobic residues, preferably 4, 3, 2, 1, or no total of phenylalanine, tryptophan, tyrosine, methionine, valine, isoleucine, and leucine residues. In some embodiments, at least 70%, at least 80%, or at least 90% of the linker comprises residues selected from threonine, serine, glycine, and alanine. The use of oligopeptide linkers in modifying polypeptides is reviewed in Chen et al., (2013) and Zhang et al., (2009).
[0245] tag In certain embodiments, the fusion polypeptide contains at least one tag suitable for detection or purification of the fusion polypeptide or its processed products. The tag is typically attached to the C-terminus or N-terminal domain of the fusion polypeptide. In a preferred embodiment, the tag is attached to the C-terminus of the Nif polypeptide. The tag is generally a peptide or amino acid sequence capable of binding with high affinity to one or more ligands, e.g., one or more ligands of an affinity matrix, such as a chromatographic support or bead, or an antibody. Those skilled in the art will understand that the tag is preferably positioned within the fusion protein at a position that does not result in removal of the tag from the NP when the MTP is cleaved off after import into mitochondria. Furthermore, the tag should not interfere with the mitochondrial import machinery. In a preferred embodiment, the polynucleotide of the present invention encodes a fusion polypeptide comprising, from N- to C-terminus, an N-terminal MTP, a Nif polypeptide, and a detection / purification tag. In an alternative embodiment, the fusion polypeptide comprises, from N- to C-terminus, an N-terminal MTP, a detection / purification tag, and a Nif polypeptide.
[0246] Additional illustrative, non-limiting examples of tags useful for detecting, isolating, or purifying a fusion polypeptide or its processed product include human influenza hemagglutinin (HA) tags, such as histidine tags containing 6 or 8 histidine residues, fluorescent tags, such as fluorescein, resorufin, and its derivatives, Arg tags, FLAG tags, Strep tags, antibody-recognizable epitopes, such as c-myc tags (recognized by anti-c-myc antibodies), SBP tags, S tags, calmodulin-binding peptides, cellulose-binding domains, chitin-binding domains, glutathione S-transferase tags, maltose-binding proteins, NusA, TrxA, DsbA, Avi tags, and the like.
[0247] Translational fusions involving Nif polypeptides Translational fusions were performed on several Nif polypeptides, as reported in the scientific literature. These are summarized in Table 2 and in the review by Buren and Rubio (2018). Most of these involved the artificial addition of epitopes or binding domains, such as histidine or Strep tags, to the proteins for detection and purification purposes, and only a few were expressed in plant cells. There are several reports of natural fusions between Nif polypeptides in bacteria. For assays in bacterial hosts, His tags of different lengths (7–10 histidines) were added to NifD (Christiansson et al., 1998), NifE (Goodwin et al., 1998), NifM (Gavini et al., 2006), and both full-length and truncated versions of NifB (Fairy et al., 2015). In each case, Nif function was retained in the modified Nif polypeptides, as demonstrated by bacterial or in vitro nitrogenase reconstitution assays.
[0248] Thiel et al. (1995) identified a natural deletion of 29 nucleotides in the intergenic region between the NifE and NifN genes of the cyanobacterium Anabaena variabilis, resulting in the deletion of 9 amino acids and the NifE stop codon. The deletion resulted in a NifE-NifN polypeptide fusion that retained at least some of the nitrogenase function of the NifE and NifN polypeptides. The NifE-NifN fusion polypeptide also contained another 19 amino acid substitutions in the fusion junction region, which may affect Nif function in an unknown manner. The fusion gene was expressed, but only under strictly anaerobic conditions. It was not reported whether there was a reduction in activity relative to the non-fused gene.
[0249] Suh et al. (1996) created an artificial junction between the NifD and NifK genes on the A. vinelandii chromosome by deleting the NifD stop codon and the NifK translation initiation codon (ATG), forming a vector designated pBG1404. The deletion resulted in a net deletion of three amino acids and a substitution of seven amino acids in amino acids 2 to 10 of the NifK polypeptide. A. vinelandii host cells containing pBG1404 were impaired in their growth in low-nitrogen medium relative to the corresponding wild-type bacteria.
[0250] Wiig et al. (2011) used a natural translational fusion between the NifN and NifB genes found in Clostridium pasteurianum and determined that it was functional in terms of NifN and NifB activity in bacterial and biochemical complementation assays. The fusion was direct, without any peptide linker; i.e., the C-terminus of NifN was directly and covalently linked to the N-terminus of NifB.
[0251] In yeast and plant cells, translational fusions have been used to translocate nuclear-encoded proteins to the mitochondrial matrix. Yeast expression assays showed that translational fusions (MTPs) of mitochondrial targeting peptides with several Nif polypeptides (NifH, NifM, NifS, and NifU) were functional when grown under aerobic conditions (Lopez-Torrejon et al., 2016). While these fusions were intended for localization within the yeast cytoplasm and were functional only when the yeast was grown under anaerobic conditions, epitope fusions (FLAG and HIS) were also shown to be functional when fused to NifH, NifM, NifS, and NifU. Buren et al. (2017) showed that a soluble mutant, mitochondrial matrix-targeted version of NifB was functional in in vitro complementation assays when reisolated from yeast mitochondria. This version of NifB contains an N-terminal MTP, a truncated mutant of NifB (without the NifX-like domain), and a C-terminal 10xHis epitope tag. Numerous MTP-Nif fusions have also been generated in yeast expression assays. However, this large collection of co-expressed proteins showed no activity in yeast (Buren et al., 2017).
[0252] The MTP from the CPN-60 gene was fused to the N-terminus of NifH, NifM, NifS, and NifU and shown to be functional by in vitro complementation assays when FeProtein was reisolated from plants grown under hypoxic conditions at 10% oxygen (US2016 / 0304842).
[0253] Table 2. Summary of gene fusions of Nif polypeptides reported in the literature [Table 2]
[0254] Polynucleotides The terms "polynucleotide" and "nucleic acid" are used interchangeably herein. They refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. A polynucleotide, as defined herein, may be of either origin or manipulation: (1) unrelated to all or a portion of the polynucleotides with which it is naturally associated (e.g., a Nif polynucleotide that does not contain a native promoter-coding sequence); (2) linked to a polynucleotide other than that with which it is naturally associated (e.g., a Nif polynucleotide linked to an MTP-encoding nucleotide sequence and / or a non-native promoter-coding sequence); or (3) non-naturally occurring (e.g., a polynucleotide encoding an MTP-Nif fusion polypeptide of the invention), of genomic, cDNA, semisynthetic, or synthetic origin, single-stranded, preferably double-stranded. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, locus(s) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, chimeric DNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise modified nucleotides, such as, for example, methylated nucleotides or nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The nucleotide sequence may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
[0255] An "isolated polynucleotide" is generally substantially free from components (e.g., regulatory sequences) with which it is linked or associated. Thus, an isolated polynucleotide is substantially free of other cellular material or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Preferably, an isolated polynucleotide is at least 60% free, more preferably at least 75% free, and more preferably at least 90% free from such components.
[0256] The term "gene" as used herein is intended to be interpreted in its broadest context and includes the deoxyribonucleotide sequences that comprise the transcribed region of a structural gene and, when translated, the protein-coding region, located adjacent to the coding region on both the 5' and 3' ends over a distance of at least about 2 kb on either end, and that include sequences involved in gene expression. In this regard, a gene may contain regulatory signals, such as promoters, enhancers, translation and transcription termination, and / or polyadenylation signals, that are naturally associated with a given gene, or may contain heterologous regulatory signals, in which case the gene is referred to as a "chimeric gene." Sequences located 5' of the protein-coding region and present on the mRNA are referred to as 5' non-translated sequences. Sequences located 3' or downstream of the protein-coding region and present on the mRNA are referred to as 3' non-translated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. Genomic forms or clones of a gene contain the coding region, which may be interrupted by non-coding sequences called "introns," "intervening regions," or "intervening sequences." Introns are segments of a gene that are transcribed into nuclear RNA (nRNA). Introns can contain regulatory elements such as enhancers. Introns are removed or "spliced out" from the nuclear or primary transcription product; therefore, introns are absent in the mRNA transcription product. During translation, the mRNA functions to specify the sequence or order of amino acids in a nascent polypeptide. The term "gene" includes synthetic or fusion molecules encoding all or part of the proteins of the invention described herein, as well as nucleotide sequences complementary to any one of the above.
[0257] As used herein, "chimeric DNA," also referred to herein as a "DNA construct," refers to any DNA molecule not naturally found in nature but artificially combining two DNA segments into a single molecule, each of which is found in nature but not the whole. For example, the DNA construct encodes an MTP-Nif fusion polypeptide of the present invention. Typically, chimeric DNA contains regulatory and transcriptional or protein-coding sequences that are not actually found together in nature (e.g., a Nif polynucleotide linked to a non-native promoter-coding sequence). Thus, chimeric DNA can contain regulatory and coding sequences from different sources, or regulatory and coding sequences from the same source but arranged in a manner different from that in which they are found in nature. An open reading frame may or may not be linked to its natural upstream and downstream regulatory elements. An open reading frame may, for example, be incorporated into a plant genome in which it is not naturally found, or into a replicon or vector in which it is not naturally found, such as a bacterial plasmid or viral vector. The term "chimeric DNA" is not limited to DNA molecules that are replicable in a host, but includes DNA that can be ligated into a replicon, for example, by specific adapter sequences.
[0258] A "transgene" is a gene that has been introduced into the genome by transformation techniques. The term includes a gene of a progeny cell, plant, seed, non-human organism, or part thereof, that has been introduced into the genome of its progenitor cell. Such progeny, etc., may be at least the third or fourth generation progeny of a progenitor cell that was the original transformed cell. Progeny may be produced by sexual reproduction or vegetative reproduction, such as from potato tubers or sugarcane shoots. The term "genetically modified" and variations thereof is a general term that includes introducing a gene into a cell by transformation or transduction, mutating a gene in a cell, and genetically altering or modulating the regulation of a gene in a cell or of a gene in the progeny of any cell so modified.
[0259] As used herein, a "genomic region" refers to a region into which a transgene or group of transgenes (also referred to herein as a cluster) is inserted into a cell or its ancestor. Such a region only contains nucleotides that have been incorporated by human intervention, e.g., by the methods described herein.
[0260] A "recombinant polynucleotide" of the present invention refers to a nucleic acid molecule constructed or modified by artificial recombinant methods. A recombinant polynucleotide may be present in a cell in an altered amount or expressed at an altered rate (e.g., in the case of mRNA) compared to its native state. In one embodiment, a polynucleotide is introduced into a cell that does not naturally contain the polynucleotide. Typically, exogenous DNA is used as a template for transcription of mRNA, which is translated in the transformed cell into a contiguous sequence of amino acid residues encoding a polypeptide of the present invention. In another embodiment, the polynucleotide is endogenous to the bacterial cell, and its expression is altered by recombinant means, e.g., an exogenous regulatory sequence is introduced upstream of the endogenous gene of interest to enable the transformed cell to express the polypeptide encoded by the gene.
[0261] Recombinant polynucleotides of the present invention include polynucleotides that have not been separated from other components of the cell-based or cell-free expression system in which they are present, as well as polynucleotides that are produced in this cell-based or cell-free expression system and then purified from at least some of the other components. A polynucleotide can be a naturally occurring contiguous stretch of nucleotides (e.g., a Nif polynucleotide) or can comprise two or more contiguous stretches of nucleotides from different sources (natural and / or synthetic) joined to form a single polynucleotide (e.g., a Nif polynucleotide linked to a nucleotide sequence encoding MTP and / or a sequence encoding a non-native promoter). Typically, such chimeric polynucleotides comprise at least one open reading frame encoding a polypeptide of the present invention operably linked to a promoter suitable for driving transcription of the open reading frame in a cell of interest.
[0262] It will be understood that with respect to a specified polynucleotide, higher % identity figures than those provided above encompass preferred embodiments. Thus, when applicable, taking into account minimum % identity figures, it is preferred that a polynucleotide comprises a polynucleotide sequence that is at least 60%, more preferably at least 65%, more preferably at least 70%, more preferably at least 75%, more preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, and even more preferably at least 99.9% identical to the relevant designated SEQ ID NO.
[0263] Polynucleotides of the present invention or useful in the present invention can selectively hybridize to the polynucleotides defined herein under stringent conditions. As used herein, stringent conditions include (1) using a denaturing agent such as formamide during hybridization, for example, 50% (v / v) formamide containing 0.1% (w / v) bovine serum albumin, 0.1% Ficoll, 0.1% polyvinylpyrrolidone, 50 mM sodium phosphate buffer at pH 6.5 containing 750 mM NaCl, and 75 mM sodium citrate at 42°C, or (2) using 50% formamide in 0.2×SSC and 0.1% SDS at 42°C. amide, 5x SSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5x Denhardt's solution, sonicated salmon sperm DNA (50 g / ml), 0.1% SDS, and 10% dextran sulfate; and / or (3) using 0.015 M NaCl / 0.0015 M sodium citrate / 0.1% SDS at low ionic strength and high temperature, e.g., 50°C, for washing.
[0264] The polynucleotides of the present invention may have one or more mutations that are deletions, insertions, or substitutions of nucleotide residues when compared to naturally occurring molecules. Polynucleotides having mutations compared to a reference sequence may be naturally occurring (i.e., isolated from a natural source) or may be synthetic (e.g., by performing site-directed mutagenesis or DNA shuffling on nucleic acids, as described above).
[0265] The polynucleotides of the present invention may be codon-modified for expression in plant cells. Those skilled in the art will appreciate that the protein coding region may be codon-optimized relative to the coding region of a native polynucleotide of, for example, a nitrogen-fixing bacterium.
[0266] nucleic acid construct The present invention includes nucleic acid constructs comprising one or more polynucleotides of the invention, as well as vectors and host cells containing them, methods for their production and use, and their uses. The present invention refers to operably connected or linked elements. "Operably connected" or "operably linked," and the like, refer to the linkage of polynucleotide elements in a functional relationship. Typically, operably connected nucleic acid sequences are linked contiguously, and, where necessary to join two protein-coding regions, contiguous and in reading frame. A coding sequence is "operably linked" to another coding sequence when RNA polymerase transcribes the two coding sequences into a single RNA that, when transcribed, is then transcribed into a single peptide having amino acids from both coding sequences. Coding sequences need not be contiguous to each other, so long as the expressed sequence is ultimately processed to produce the desired protein.
[0267] As used herein, the terms "cis-acting sequence," "cis-acting element," or "cis-regulatory region" or "regulatory region," or similar terms, refer to any sequence of nucleotides that, when properly positioned and connected to an expressible gene sequence, is capable of regulating, at least in part, the expression of the gene sequence. Those skilled in the art will recognize that cis-regulatory regions activate, silence, promote, repress, or otherwise alter the expression level, and / or cell-type specificity and / or developmental specificity, of a gene sequence at the transcriptional or post-transcriptional level. In a preferred embodiment of the present invention, the cis-acting sequence is an activator sequence that promotes or stimulates expression of an expressible gene sequence.
[0268] "Operably linking" a promoter or enhancer element to a transcribable polynucleotide means placing a transcribable polynucleotide (e.g., a protein-coding polynucleotide or other transcript) under the regulatory control of the promoter, which then controls the transcription of that polynucleotide. In the construction of a heterologous promoter / structural gene combination, it is generally preferred to position the promoter, or a variant thereof, at a distance from the transcription start site of the transcribable polynucleotide that is approximately the same as the distance between the promoter and the protein-coding region it controls in its natural context: i.e., the gene from which the promoter is derived. As is known in the art, some variation in this distance can be accommodated without loss of function. Similarly, the preferred location of a regulatory sequence element (e.g., an operator, enhancer, etc.) with respect to the transcribable polynucleotide to be placed under its control is dictated by the location of that element in its natural context, i.e., in the gene from which it is derived.
[0269] As used herein, "promoter" or "promoter sequence" refers to a region of a gene, generally upstream (5') of the RNA coding region, that controls the initiation and level of transcription in a cell of interest. "Promoter" includes classical genomic gene transcription control sequences, such as TATA box and CCAAT box sequences, as well as additional regulatory elements (i.e., upstream activating sequences, enhancers, and silencers) that alter gene expression in response to developmental and / or environmental stimuli or in a tissue- or cell-type-specific manner. Promoters are usually, but not necessarily (e.g., some Pol III promoters), located upstream of the structural gene whose expression they regulate. Furthermore, regulatory elements, including promoters, are usually located within 2 kb of the transcription start site of a gene. Promoters may also contain additional specific regulatory elements located more distally from the start site to further enhance expression in the cell and / or to alter the timing or inducibility of expression of the structural gene to which they are operably linked.
[0270] A "constitutive promoter" refers to a promoter that directs the expression of an operably linked transcription sequence in many or all tissues of an organism, such as a plant. As used herein, the term "constitutive" indicates, but is not required to indicate, that a gene is expressed at the same level in all cell types; however, the gene is expressed in a wide range of cell types, although some variation in the level is often detectable. "Preferential expression," as used herein, refers to expression exclusively in a particular organ of a plant, such as the endosperm, embryo, leaf, fruit, tuber, or root. In a preferred embodiment, the promoter is selectively or preferentially expressed in the roots, leaves, and / or stems of a plant, preferably a cereal plant. Thus, selective expression can be contrasted with constitutive expression, which refers to expression in many or all tissues of a plant under most or all conditions experienced by the plant.
[0271] Selective expression can also result in compartmentalization of the gene expression product in specific plant tissues, organs, or developmental stages. For example, compartmentalization in specific subcellular locations such as plastids, cytosol, vacuoles, or apoplastic spaces may be achieved by incorporating into the structure of the gene product an appropriate signal, e.g., a signal peptide, for translocation to the required cellular compartment, or, in the case of semi-autonomous organelles (plastids and mitochondria), by inclusion of a transgene with appropriate regulatory sequences directly within the organelle genome.
[0272] A "tissue-specific promoter" or "organ-specific promoter" is a promoter that is preferentially expressed in one tissue or organ relative to many, preferably most, if not all, other tissues or organs, e.g., in plants. Typically, the promoter is expressed at a 10-fold higher level in a particular tissue or organ than in other tissues or organs.
[0273] In certain embodiments, the promoter is a stem-specific promoter, a leaf-specific promoter, or a promoter that directs gene expression in the aerial parts of the plant (at least the stems and leaves) (a green tissue-specific promoter), such as the ribulose-1,5-bisphosphate carboxylase oxygenase (RUBISCO) promoter.
[0274] Examples of stem-specific promoters include, but are not limited to, those described in US Pat. No. 5,625,136 and Bam et al. (2008). In certain embodiments, the promoter is a root-specific promoter, examples of which include, but are not limited to, the acidic promoter of the chitinase gene and certain subdomains of the CaMV35S promoter.
[0275] The promoters contemplated by the present invention may be native to the host plant to be transformed, or may be derived from alternative sources, provided that the region is functional in the host plant. Other sources include tissue-specific promoters, such as Agrobacterium T-DNA genes, such as promoters for the biosynthesis of nopaline, octapine, mannopine, or other opine promoters (see, for example, US Pat. No. 5,459,252 and WO 91 / 13992); promoters from viruses (including host-specific viruses), or partial or complete synthetic promoters. Numerous promoters functional in monocotyledonous and dicotyledonous plants are well known in the art, including various promoters isolated from plants and viruses, such as the cauliflower mosaic virus promoter (CaMV35S, 19S) (see, for example, Greve, 1983; Salomon et al., 1984; Garfinkel et al., 1983; Barker et al., 1983). Non-limiting methods for assessing promoter activity are disclosed by Medberry et al. (1992, 1993), Sambrook et al. (1989, supra) and US Pat. No. 5,164,316.
[0276] Alternatively or additionally, the promoter may be an inducible or developmentally regulated promoter capable of driving expression of the introduced polynucleotide at the appropriate developmental stage in, for example, a plant. Other cis-acting sequences that may be utilized include transcriptional and / or translational enhancers. Enhancer regions are well known to those skilled in the art and may include the ATG translation initiation codon and adjacent sequences. When included, the initiation codon, when translated, must be in phase with the reading frame of the coding sequence associated with the foreign or exogenous polynucleotide to ensure translation of the entire sequence. The translation initiation region may be provided from the source of the transcription initiation region or from the foreign or exogenous polynucleotide. The sequence may also be derived from the source of the promoter selected to drive transcription and may be specifically modified to enhance translation of mRNA.
[0277] The nucleic acid constructs of the present invention may also include a 3' untranslated sequence of approximately 50 to 1,000 nucleotide base pairs, which may include a transcription termination sequence. The 3' untranslated sequence may include a transcription termination signal, which may include a polyadenylation signal, and any other regulatory signals that can affect mRNA processing. The polyadenylation signal functions to add polyadenylic acid to the 3' end of the mRNA precursor. Polyadenylation signals are generally recognized by their homology to the standard form 5'AATAAA-3', although variations are not uncommon. Transcription termination sequences that do not include a polyadenylation signal include Pol I or Pol III RNA polymerase terminators, which contain a stretch of four or more thymidines. An example of a suitable 3' non-translated sequence is the 3' transcribed untranslated region containing the polyadenylation signal from the octopine synthase (ocs) or nopaline synthase (nos) gene of Agrobacterium tumefaciens (Bevan et al., 1983). Suitable 3' non-translated sequences may also be derived from plant genes, such as the ribulose-1,5-bisphosphate carboxylase (ssRUBISCO) gene, although other 3' elements known to those skilled in the art can also be used.
[0278] When a DNA sequence is inserted between the transcription start site and the beginning of the coding sequence, i.e., a non-translated 5' leader sequence (5'UTR), its translation and transcription may affect gene expression, and those skilled in the art may also utilize specific leader sequences. Suitable leader sequences include those selected to direct optimal expression of foreign or endogenous DNA sequences. For example, such leader sequences include the preferred consensus sequences described by Joshi (1987), which may enhance or maintain mRNA stability and prevent improper initiation of translation.
[0279] vector The present invention includes the use of vectors for the manipulation or transfer of genetic constructs. A vector is a nucleic acid molecule, preferably a DNA molecule, that can be used to artificially transport foreign genetic material into another cell where it can be replicated or expressed. A vector containing foreign DNA is also called a "genetic recombination vector." Examples of vectors include, but are not limited to, plasmids, viral vectors, cosmids, extrachromosomal elements, minichromosomes, and artificial chromosomes. A vector may also contain transposition elements.
[0280] A vector is preferably double-stranded DNA and contains one or more unique restriction sites. It may be capable of autonomous replication in a given host cell, including the target cell or tissue or its progenitor cell or tissue, or it may be capable of integration into the genome, preferably the nuclear genome, of a given host cell so that cloned sequences can be replicated. Thus, a vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity whose replication is independent of chromosomal replication, such as a linear or closed circular plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. A vector may contain any means for ensuring autonomous replication. Alternatively, a vector may be one that, upon introduction into a cell, integrates into the genome, preferably the nuclear genome, of a recipient cell and is replicated together with the chromosome(s) into which it has been integrated. A vector system may comprise a single vector or plasmid, two or more vectors or plasmids, which together contain the total DNA to be introduced into the host cell, or a transposon. The choice of vector typically depends on the compatibility of the vector with the cell into which it will be introduced. The vector may also contain a selection marker, such as, for example, an antibiotic resistance gene, a herbicide resistance gene or other gene used for selection of successful transformants. Examples of such genes are well known to those skilled in the art.
[0281] The nucleic acid constructs of the present invention may be incorporated into a vector, such as a plasmid. Plasmid vectors typically contain additional nucleic acid sequences that provide for easy selection, amplification, and transformation of the expression cassette in prokaryotic and eukaryotic cells, such as pUC-, pSK-, pGEM-, pSP-, pBS-, or binary vectors containing one or more T-DNA regions. The additional nucleic acid sequences include an origin of replication that provides for autonomous replication of the vector, a selectable marker gene that preferably encodes antibiotic or herbicide resistance, a unique multiple cloning site that provides multiple sites for inserting nucleic acid sequences or genes encoded by the nucleic acid construct, and sequences that enhance transformation of prokaryotic and eukaryotic (particularly plant) cells.
[0282] "Marker gene" refers to a gene that confers a distinctive phenotype on cells that express the marker gene, thus allowing such transformed cells to be distinguished from cells that do not possess the marker. A selectable marker gene provides a trait that can be "selected" for based on resistance to a selection agent (e.g., herbicide, antibiotic, radiation, heat, or other treatment that damages non-transformed cells). A screenable marker gene (or reporter gene) provides a trait that can be identified through observation or testing, i.e., "screening" (e.g., β-glucuronidase, luciferase, GFP, or other enzyme activity that is not present in non-transformed cells). The marker gene and the nucleotide sequence of interest do not have to be linked.
[0283] To facilitate identification of transformants, the nucleic acid construct preferably includes a selection or screening marker gene as, or in addition to, the foreign or exogenous polynucleotide. The actual choice of marker is not important, as long as it is functional (i.e., selectable) in combination with the host cell, preferably a plant host cell. The marker gene of interest and the foreign or exogenous polynucleotide do not need to be linked, as co-transformation of unlinked genes, as described, for example, in U.S. Pat. No. 4,399,216, is also an efficient step in plant transformation.
[0284] Examples of bacterial selectable markers include markers that confer antibiotic resistance, such as ampicillin, erythromycin, chloramphenicol, or tetracycline resistance, preferably kanamycin resistance. Exemplary selectable markers for selection of plant transformants include the hyg gene, which encodes hygromycin B resistance; the neomycin phosphotransferase (nptII) gene, which confers resistance to kanamycin, paromomycin, and G418; the rat liver glutathione-S-transferase gene, which confers resistance to herbicide-derived glutathione, as described, for example, in EP 256223; the glutamine synthetase gene, which, upon overexpression, confers resistance to glutamine synthetase inhibitors, such as phosphinothricin, as described, for example, in WO 87 / 05327; and the acetyltransferase gene from Streptomyces viridochromogenes, which confers resistance to the selective agent phosphinothricin, as described, for example, in EP 275957; see, for example, Hinchee et al. (1988) , conferring resistance to N-phosphonomethylglycine; the bar gene conferring resistance to bialaphos, e.g., as described in WO 91 / 02071 ; a nitrilase gene such as bxn from Ozonophorus , conferring resistance to bromoxynil ( Stalker et al., 1988 ); a dihydrofolate reductase (DHFR) gene conferring resistance to methotrexate ( Thillet et al., 1988 ); a mutant acetolactate synthase gene (ALS) conferring resistance to imidazolinones, sulfonylureas, or other ALS-inhibiting chemicals ( EP 154,204 ); a mutated anthranilate synthase gene conferring resistance to 5-methyltryptophan; or a dalapon dehalogenase gene conferring resistance to herbicides.
[0285] Preferred screen markers include, but are not limited to, the uidA gene, which encodes the β-glucuronidase (GUS) enzyme, for which various chromogenic substrates are known; the β-galactosidase gene, which encodes an enzyme for which chromogenic substrates are known; the aequorin gene, which can be used for calcium-sensitive bioluminescent detection (Prasher et al., 1985); the green fluorescent protein gene (Niedz et al., 1995) or its derivatives; the luciferase (luc) gene, which allows bioluminescent detection (Ow et al., 1986), and others known in the art. As used herein, "reporter molecule" refers to a molecule whose chemical nature provides an analytically identifiable signal that facilitates promoter determination by reference to the protein product.
[0286] Preferably, the nucleic acid construct is stably integrated into the genome of, for example, a plant. Thus, the nucleic acid contains appropriate components that allow the molecule to be integrated into the genome, or the construct is placed in a suitable vector that can be integrated into the chromosomes of the plant cell.
[0287] One embodiment of the present invention includes a recombinant vector, which contains at least one polynucleotide as defined herein and is capable of delivering the polynucleotide into a host cell. Such a vector contains heterologous nucleic acid sequences, i.e., nucleic acid sequences that are not naturally found adjacent to the nucleic acid molecules of the invention and that are preferably derived from a species other than the species from which the nucleic acid molecule(s) are derived. Vectors can be either RNA or DNA, either prokaryotic or eukaryotic, and typically are viruses or plasmids.
[0288] The recombinant vectors of the present invention contain a fusion sequence that leads to the expression of a nucleic acid molecule as a fusion protein. Genetic recombinant vectors may also include intervening and / or untranslated sequences surrounding and / or within the nucleic acid sequences of the polynucleotides defined herein. Preferably, the recombinant vector is stably integrated into the genome of the host cell, e.g., a plant cell. Thus, the recombinant vector may comprise appropriate elements that allow the vector to be integrated into the genome or into the chromosome of the cell.
[0289] Recombinant cells Another embodiment of the present invention includes recombinant cells, e.g., recombinant plant cells, which are host cells transformed with one or more polynucleotides, constructs, or vectors of the present invention, or their progeny. The term "recombinant cell" is used interchangeably herein with the term "transgenic cell."
[0290] Transformation of a nucleic acid molecule into a cell can be accomplished by any method by which a nucleic acid molecule is inserted into a cell. Transformation techniques include, but are not limited to, transfection, electroporation, microinjection, lipofection, adsorption, and protoplast fusion. Recombinant cells may remain unicellular or may be propagated in tissues, organs, or multicellular organisms. Transforming nucleic acid molecules of the present invention may remain extrachromosomal or may be integrated into one or more sites in the chromosomes of the transformed cell in a manner that retains their ability to be expressed.
[0291] Preferred host cells are plant cells, more preferably cells of cereal plants, more preferably barley or wheat cells, or even more preferably wheat cells.
[0292] The recombinant cell may be a cultured cell, in vitro, or in an organism such as, for example, a plant, or in an organ such as, for example, a root, leaf, or stem. Preferably, the cell is present in a plant, more preferably in the root, leaf, and / or stem of a plant.
[0293] In certain embodiments, expression of an active NifDK in a plant cell requires expression of NifD, NifK, NifH, NifB, NifE, NifN, and optionally NifU, NifS, NifO, NifV, NifY, NifW, and / or NifZ.
[0294] In other or further embodiments, expression of active NifH in plant cells requires expression of NifH and NifM, and optionally NifU and / or NifN.
[0295] In certain embodiments, reconstitution of nitrogenase activity in plant cells requires expression of at least NifD, NifK, NifH, NifB, NifE, NifN, and NifM.
[0296] Those skilled in the art will understand that a small subset of Nif proteins can result in functional nitrogenase reconstitution in plant cells. To our knowledge, the only report of nitrogenase gene transfer into any photosynthetic organism described the introduction of NifH in the chloroplast genome of Chlamydomonas (Cheng et al., 2005). NifH was able to complement a chlorophyll biosynthesis mutant despite the fact that the NifH biosynthetic precursor proteins, NifM, NifS, and NifU, were not coexpressed. This demonstrated that endogenous eukaryotic equivalents can functionally replace specific Nif proteins. Indeed, a recent report demonstrating that E. coli can reconstitute nitrogenase function using only eight Nif proteins (Wang et al., 2013) implies that achieving function in plants may not be as complex as expressing the entire complex of Nif proteins. Although we have not yet established the functionality of Nif proteins in planta, it is promising that a repertoire of biosynthetic and functional Nif proteins can be expressed in environments that potentially support nitrogenase function.
[0297] Transgenic plants The term "plant," when used as a noun herein, refers to an entire plant and to any member of the plant kingdom, while when used as an adjective, it refers to any material present in, derived from, or associated with a plant, such as plant organs (e.g., leaves, stems, roots, flowers), single cells (e.g., pollen), seeds, plant cells, etc. Also included within the meaning of "plant" are plantlets from which roots and shoots emerge, and germinated seeds. The term "plant part," as used herein, refers to one or more plant tissues or organs obtained from a plant and containing the plant's genomic DNA. Plant parts include vegetative structures (e.g., leaves, stems), roots, floral organs / structures, seeds (including embryos, cotyledons, and seed coats), plant tissues (e.g., vascular tissue, ground tissue, etc.), cells, and their progeny. The term "plant cell," as used herein, refers to a cell obtained from or in a plant, including protoplasts or other cells derived from a plant, gametogenic cells, and cells that regenerate into whole plants. Plant cells may also be cells in culture. "Plant tissue" refers to differentiated tissue in or obtained from a plant (an "explant"), or undifferentiated tissue derived from various aggregates of plant cells in culture, such as immature or mature embryos, seeds, roots, shoots, fruits, tubers, pollen, tumor tissue, e.g., crown gall, and callus. Exemplary plant tissues in or from seeds are cotyledons, embryos, and hypocotyls. Thus, the present invention includes plants, plant parts, and products comprising these.
[0298] As used herein, the term "seed" refers to a "mature seed" of a plant, which is either ready for harvest or has been harvested from a plant, e.g., typically commercially harvested in a field, or exists as a "developing seed" which occurs on a plant after fertilization and before reaching seed dormancy and before harvest.
[0299] The term "genetically modified plant," as used herein, refers to a plant containing a nucleic acid construct not found in a wild-type plant of the same species, variety, or cultivar. That is, a genetically modified plant (transformed plant) contains genetic material (transgene) that it did not contain prior to transformation. A transgene may include a genetic sequence obtained or derived from a plant cell, another plant cell, a non-plant source, or a synthetic sequence. Typically, a transgene is introduced into a plant by artificial manipulation, such as by transformation, although those skilled in the art will recognize that any method may be used. The genetic material is preferably stably integrated into the plant's genome, preferably the nuclear genome. The introduced genetic material may comprise a homologous native sequence, but in a rearranged order, e.g., an antisense sequence, or in a different configuration. Plants containing such sequences are included herein as "genetically modified plants."
[0300] In a preferred embodiment, the transgenic plants are homozygous for each and every gene (transgene) introduced so that their progeny do not segregate for the desired phenotype. Transgenic plants may also be heterozygous for the introduced transgene(s), preferably, for example, in the case of F1 progeny grown from hybrid seed. Such plants can provide advantages such as hybrid vigor, which are well known in the art.
[0301] A transgenic plant, as defined in the context of the present invention, includes the progeny of a plant that has been genetically modified using recombinant DNA technology, wherein the progeny contain the transgene of interest. Such progeny may be obtained by self-fertilization of the primary transgenic plant or by crossing such a plant with another plant of the same species. This will generally be for the purpose of regulating the production of at least one protein defined herein in the desired plant or plant organ. Transgenic plant parts include all parts and cells of the plant that contain the transgene, such as cultured tissue, callus, and protoplasts.
[0302] Transgenic plants can generally be produced using techniques known in the art, such as those described in A. Slater et al., Plant Biotechnology - The Genetic Manipulation of Plants, Oxford University Press (2003), and P. Christou and H. Klee, Handbook of Plant Biotechnology, John Wiley and Sons (2004).
[0303] A "non-transgenic plant" is one that has not been genetically engineered by the introduction of genetic material via recombinant DNA technology. As used herein, the term "compared to an isogenic plant," or similar phrases, refers to a plant that is identical or similar in most characteristics, preferably isogenic or near-isogenic to the transgenic plant, but lacks the transgene of interest. Preferably, the corresponding non-transgenic plant is the same cultivar or variety as the ancestor of the transgenic plant of interest, or a sibling plant line lacking the construct, often referred to as a "segregant," or a plant of the same cultivar or variety, which may be a non-transgenic plant transformed with an "empty vector" construct. As used herein, "wild-type" refers to a cell, tissue, or plant that has not been modified according to the present invention. A wild-type cell, tissue, or plant may be used as a control to compare the level of expression of an exogenous nucleic acid or the degree and nature of trait modification in cells, tissues, or plants modified as described herein.
[0304] A transgenic plant, as defined in the context of the present invention, includes the progeny of a plant that has been genetically modified using recombinant genetic techniques, wherein the progeny contain the transgene of interest. Such progeny may be obtained by self-fertilization of the primary transgenic plant or by crossing such a plant with another plant of the same species. Transgenic plant parts include all parts and cells of the plant that contain the transgene, such as cultured tissue, callus, and protoplasts.
[0305] Plants contemplated for use in the practice of the present invention include both monocotyledonous and dicotyledonous plants. Target plants include, but are not limited to, cereals (e.g., wheat, barley, rye, oats, rice, corn, sorghum, and related crops); grapes; beets (sugar beet, fodder beet); soft fruits (apple, pear, plum, peach, almond, cherry, strawberry, raspberry, blackberry); legumes (beans, lentils, peas, soybeans); oil plants (rapeseed or other Brassicas, mustard, poppy, olive, sunflower, safflower, hemp, coconut, castor bean, cocoa bean, groundnut); cucumber plants (cucumber, plants (Cucurbita pepo, cucumber, melon); fiber plants (cotton, flax, hemp, jute); citrus fruits (orange, lemon, grapefruit, mandarin); vegetables (spinach, lettuce, asparagus, cabbage, carrot, onion, tomato, potato, pepper); Lauraceae (avocado, cinnamon, camphor); or plants such as corn, tobacco, nuts, coffee, sugarcane, tea, grapes, hops, turf, banana, and natural rubber plants, and ornamental plants (flowers, shrubs, broad-leaved trees, and evergreen trees such as conifers). Preferably, the plant is a cereal plant, more preferably wheat, rice, corn, triticale, oats, or barley, and even more preferably wheat.
[0306] As used herein, the term "wheat" refers to any species in the genus Triticum, including its ancestors and its descendants resulting from hybridization with other species. Wheat includes "hexaploid wheat," which has a genome organization of 42 chromosomes (AABBDD), and "tetraploid wheat," which has a genome organization of 28 chromosomes (AABB). Hexaploid wheat includes Triticum aestivum, Triticum spelta, Triticum macha, Triticum compactum, Triticum sphaerococcum, Triticum vavilovii, and interspecific hybrids thereof. A preferred species of hexaploid wheat is Triticum aestivum ssp. aestivum (also known as "bread wheat"). Tetraploid wheats include T. durum (also referred to herein as durum wheat or Triticum turgidum ssp. durum), T. dicoccoides, T. dicoccum, T. polonicum, and interspecific hybrids thereof. Additionally, the term "wheat" includes predicted ancestors of hexaploid or tetraploid wheat species such as T. urartu (T. uartu), T. monococcum, or wild-type T. boeoticum in the A genome, Aegilops speltoides in the B genome, and T. tauschii (also known as Aegilops squarrosa or Aegilops tauschii) in the D genome. Particularly preferred progenitors are those of the A genome, and even more preferably, the A genome progenitor is T. monococcum. Wheat cultivars for use in the present invention may belong to any of the species listed above, but are not limited thereto.Also included are plants produced by conventional techniques using wheat species as parents in sexual crosses with non-wheat species, including but not limited to Triticale, such as rye (Secale cereale).
[0307] As used herein, the term "barley" refers to any species of the genus Hordeum, including its ancestors as well as its progeny by crossing with other species. Preferably, the plant is a commercially grown barley species, such as a strain or cultivar or variety of Hordeum vulgare, or is suitable for commercial grain production.
[0308] Four general methods for direct gene delivery into cells have been described: (1) chemical methods (Graham et al., 1973); (2) physical methods, such as microinjection (Capecchi, 1980), electroporation (see, e.g., WO87 / 06614, US5,472,869, US5,384,253, WO92 / 09696, and WO93 / 21335), and gene guns (see, e.g., US4,945,050 and US5,141,131); (3) viral vectors (Clapp, 1993; Lu et al., 1993; Eglitis et al., 1988); and (4) receptor-mediated mechanisms (Curiel et al., 1992; Wagner et al., 1992).
[0309] Acceleration methods that can be used include, for example, microprojectile bombardment. One example of a method for delivering transforming nucleic acid molecules into plant cells is microprojectile bombardment. This method is reviewed by Yang et al., Particle Bombardment Technology for Gene Transfer, Oxford Press, Oxford, England (1994). Non-biological microprojectiles can be coated with nucleic acids and delivered into cells by propelling force. Exemplary particles include those made of tungsten, gold, platinum, etc. In addition to being an effective means of reproducibly transforming monocotyledonous plants, a particular advantage of microprojectile bombardment is that it does not require protoplast isolation or susceptibility to Agrobacterium infection. A suitable particle delivery system for use in the present invention is the helium-accelerated PDS-1000 / He gun available from Bio-Rad Laboratories. For bombardment, immature embryos or induced target cells, such as scutellum or callus from immature embryos, can be placed on solid culture medium.
[0310] In other alternative embodiments, the plasmid can be stably transformed. Methods disclosed for plasmid transformation in higher plants include particle gun delivery of DNA containing a selection marker and targeting the DNA to the plasmid genome via homologous recombination (US 5,451,513, US 5,545,818, US 5,877,402, US 5,932479, and WO99 / 05265).
[0311] Agrobacterium-mediated transfer is a widely applicable system for gene transfer into plant cells because DNA can be introduced into whole plant tissues, thereby avoiding the need to regenerate intact plants from protoplasts. The use of Agrobacterium-mediated plant integrating vectors to introduce DNA into plant cells is well known in the art (see, e.g., US 5,177,010, US 5,104,310, US 5,004,863, US 5,159,135). Furthermore, T-DNA integration is a relatively precise process that results in little rearrangement. The DNA region to be transferred is determined by border sequences, and intervening DNA is usually inserted into the plant genome.
[0312] Agrobacterium transformation vectors can replicate in both E. coli and Agrobacterium, allowing for convenient manipulation as described (Klee et al., Plant DNA Infectious Agents, Hohn and Schell, (editors), Springer-Verlag, New York, (1985): 179-203). Furthermore, technological advances in vectors for Agrobacterium-mediated gene transfer have facilitated the construction of vectors capable of expressing various polypeptide-encoding genes by improving the arrangement of genes and restriction sites in the vectors. The described vectors, which contain convenient multilinker regions flanked by promoters and polyadenylation sites for direct expression of inserted polypeptide-encoding genes, are suitable for this purpose. Furthermore, Agrobacterium containing both armed and disarmed Ti genes can be used for transformation. For plant species in which Agrobacterium-mediated transformation is effective, this is the method of choice due to the facile and unambiguous nature of gene transfer.
[0313] Transgenic plants generated using the Agrobacterium transformation method typically contain a single gene locus on one chromosome. Such transgenic plants can be said to be hemizygous for the added gene. More preferred are transgenic plants that are homozygous for the added structural gene, i.e., contain two added genes, one gene at the same locus on each chromosome of a chromosome pair. Homozygous transgenic plants can be obtained by sexually crossing (selfing) independent, segregating transgenic plants containing a single added gene, germinating several seeds produced, and analyzing the resulting plants for the gene of interest.
[0314] It should also be understood that two different transgenic plants can be crossed to produce offspring containing two independently segregating exogenous genes. Selfing of suitable offspring can produce plants that are homozygous for both exogenous genes. Backcrossing of parent plants and outcrossing with non-transgenic plants are also contemplated, as is vegetative propagation. Descriptions of other breeding methods commonly used for different traits and crops can be found in Fehr, Breeding Methods for Cultivar Development, J. Wilcox (editor), American Society of Agronomy, Madison, Wis. (1987).
[0315] Transformation of plant protoplasts can be achieved using methods based on calcium phosphate precipitation, polyethylene glycol treatment, electroporation, and combinations of these treatments. The application of these systems to different plant species depends on the ability to regenerate this particular plant strain from protoplasts. Exemplary methods for regenerating cereals from protoplasts have been described (Fujimura et al., 1985; Toriyama et al., 1986; Abdullah et al., 1986).
[0316] Other methods of cell transformation can also be used and include, but are not limited to, direct DNA transfer into pollen, direct DNA transfer into the reproductive organs of a plant, or direct DNA transfer into a plant by direct injection of DNA into the cells of an immature embryo followed by rehydration of the desiccated embryo.
[0317] Regeneration, development, and cultivation from single plant protoplast transformants or from various transformed explants are well known in the art (Weissbach et al., Methods for Plant Molecular Biology, Academic Press, San Diego, (1988)). This regeneration and growth process typically involves a selection step of transformed cells, i.e., culturing these individualized cells through the usual stages of embryo development to the rooted plantlet stage. Transgenic embryos and seeds are similarly regenerated. The resulting rooted transgenic shoots are then planted in a suitable plant growth medium, such as soil.
[0318] The generation or regeneration of plants containing foreign exogenous genes is well known in the art. Preferably, the regenerated plants are self-pollinated to provide homozygous transgenic plants. Alternatively, pollen obtained from the regenerated plants is crossed to seed-propagated plants of agriculturally important lines. Conversely, pollen from plants of these important lines is used to pollinate the regenerated plants. The transgenic plants of the present invention containing the desired exogenous nucleic acid are cultivated using methods well known to those skilled in the art.
[0319] Methods for transforming dicotyledonous plants and obtaining transgenic plants, primarily by the use of Agrobacterium tumefaciens, have been published for cotton (US 5,004,863, US 5,159,135, US 5,518,908); soybean (US 5,569,834, US 5,416,011); Brassica (US 5,463,174); peanut (Cheng et al., 1996); and pea (Grant et al., 1995).
[0320] Methods for transforming cereal plants such as wheat and barley to introduce genetic diversity into plants by introducing exogenous nucleic acids, as well as methods for regenerating plants from protoplasts or immature plant embryos, are well known in the art. See, for example, CA 2,092,588, AU 61781 / 94, Australian Patent No. 667939, US 6,100,447, WOPCT / US97 / 10621, US 5,589,617, US 6,541,257, and other methods are described in patent specification WO 99 / 14314. Preferably, transgenic wheat or barley plants are produced by Agrobacterium tumefaciens-mediated transformation procedures. Vectors carrying the desired nucleic acid constructs can be introduced into regenerable wheat cells of appropriate plant systems, such as tissue culture plants or explants, or protoplasts. Regenerable wheat cells are preferably derived from the scutellum of an immature embryo, a mature embryo, callus derived therefrom, or meristem tissue.
[0321] To confirm the presence of the transgene in transgenic cells and plants, polymerase chain reaction (PCR) amplification or Southern blot analysis can be performed using methods known to those skilled in the art. Transgene expression products can be detected by any of a variety of methods, depending on the nature of the product, including Western blot and enzyme assays. One particularly useful method for quantifying protein expression and detecting replication in different plant tissues is to use a reporter gene such as GUS. Once a transgenic plant is obtained, it can be cultivated to produce plant tissues or parts with the desired phenotype. The plant tissues or plant parts can be harvested and / or seeds can be collected. The seeds can serve as a source for cultivating another plant with tissue or parts with the desired characteristics.
[0322] "Polymerase chain reaction" ("PCR") is a reaction that uses a "primer pair" or "primer set" consisting of an "upstream" and a "downstream" primer, and a polymerization catalyst, such as a DNA polymerase, typically a thermostable polymerase enzyme, to create replicate copies of a target polynucleotide. Methods for PCR are known in the art and are taught, for example, in "PCR" (MJ McPherson and SG Moller (editors), BIOS Scientific Publishers Ltd, Oxford, (2000)). PCR can be performed on cDNA obtained by reverse transcribing mRNA isolated from plant cells expressing the polynucleotide of the present invention. However, it will generally be easier if PCR is performed on genomic DNA isolated from plant cells.
[0323] A primer is an oligonucleotide sequence that hybridizes to a target sequence in a sequence-specific manner and can be extended during PCR. An amplicon, or PCR product, or PCR fragment, or amplification product, is an extension product containing primers and newly synthesized copies of the target sequence. A multiplex PCR system contains multiple sets of primers that result in the simultaneous generation of more than one amplicon. Primers may perfectly match the target sequence or contain internal mismatched bases that can introduce restriction enzyme or catalytic nucleic acid recognition / cleavage sites within a specific target sequence. Primers may also contain additional sequences and / or modified or labeled nucleotides to facilitate amplicon capture or detection. Repeated cycles of thermal denaturation of DNA, annealing of primers to complementary sequences, and extension of the annealed primers by polymerase result in exponential amplification of the target sequence. The terms target, target sequence, or template refer to the nucleic acid sequence to be amplified.
[0324] Methods for direct sequencing of nucleotide sequences are well known to those skilled in the art and can be found, for example, in Ausubel et al. (see above) and Sambrook et al. (see above). Sequencing can be performed by any suitable method, such as dideoxy sequencing, chemical sequencing, or variations thereof. Direct sequencing has the advantage of determining any base pair mutations in a specific sequence.
[0325] Plant / grain processing The grain / seed of the present invention, preferably the cereal grain, or other plant parts of the present invention may be processed to produce food ingredients, food or non-food products using any technique known in the art.
[0326] In one embodiment, the product is a whole grain flour, such as, for example, an ultra-finely milled whole grain flour or a flour made from about 100% of the grain. The whole grain flour includes refined flour components (refined flour or refined flour) and a coarse fraction (an ultra-finely milled coarse fraction).
[0327] The refined flour may be prepared, for example, by grinding and bolting washed grains, such as wheat or barley grains. The grain size of the refined flour is described as a powder that passes 98% or more through a cloth with openings no larger than those of a woven wire cloth designated "212 micrometers (USA Wire 70)." The coarse fraction includes at least one of bran and germ. For example, germ is the plant embryo located within the kernel of the grain. The germ contains lipids, fiber, vitamins, proteins, minerals, and phytonutrients (such as flavonoids). Bran contains several cell layers and has significant amounts of lipids, fiber, vitamins, proteins, minerals, and phytonutrients, such as flavonoids. Additionally, the coarse fraction may include the aleurone layer, which also contains lipids, fiber, vitamins, proteins, minerals, and phytonutrients, such as flavonoids. Although the aleurone layer is technically considered part of the endosperm, it is typically removed along with the bran and germ during the milling process because it exhibits many of the same characteristics as the bran. The aleurone layer contains proteins, vitamins, and phytonutrients such as ferulic acid.
[0328] Additionally, the coarse fraction can be mixed with refined flour. The coarse fraction can be mixed with refined flour to form a whole grain flour, thus providing a whole grain flour with improved nutritional value, fiber content, and antioxidant capacity compared to refined flour. For example, the coarse fraction or whole grain flour can be used in various amounts to replace refined or whole grain flour in baked products, snack products, and food products. The whole grain flour of the present invention (i.e., ultrafine-milled whole grain flour) can be sold directly to consumers for use in their homemade baked products. In an exemplary embodiment, the granulation profile of the whole grain flour is such that 98% of the particles by weight of the whole grain flour are less than 212 micrometers.
[0329] In a further embodiment, enzymes found in the bran and germ of the whole grain flour and / or coarse fraction are inactivated to stabilize the whole grain flour and / or coarse fraction. Stabilization is a process in which the enzymes found in the bran and germ layers are inactivated using steam, heat, radiation, or other treatments. The stabilized flour retains its cooking properties and has a long shelf life.
[0330] In further embodiments, the whole grain flour, coarse fraction, or refined flour may be a component (ingredient) of a food product or may be used to produce a food product, such as bagels, biscuits, bread, buns, croissants, dumplings, English muffins, muffins, pita bread, quick bread, frozen / frozen dough products, dough, baked beans, burritos, chili, tacos, tamales, tortillas, pot pies, ready to eat cereals, and the like. cereal), ready to eat meals, fillings, microwave meals, brownies, cakes, cheesecakes, coffee cakes, cookies, desserts, pastries, sweet rolls, candy bars, pie crusts, pie fillings, baby food, baking mixes, batter, breadcrumbs, gravy mixes, meat extenders, meat substitutes, seasoning mixes, soup mixes, gravy, roux, salad dressings, soups, sour cream, noodles, pasta, ramen, chow mein noodles, lo mein noodles, ice cream inclusions, ice cream bars, ice cream cones, ice cream sandwiches, crackers, croutons, donuts, egg rolls, extruded snacks, fruit and grain bars, microwaveable snack products, nutritional bars, pancakes, par-bake bakery products, pretzels, puddings, granola-based products, snack chips, snack foods, snack mixes, waffles, pizza crusts, animal foods, or pet foods.
[0331] In another embodiment, the whole grain flour, refined flour, or crude fraction can be a component of a dietary supplement. For example, a dietary supplement may be a product added to food containing one or more additional ingredients, typically including vitamins, minerals, herbs, amino acids, enzymes, antioxidants, herbs, spices, probiotics, extracts, prebiotics, and fiber. The whole grain flour, refined flour, or crude fraction of the present invention contains vitamins, minerals, amino acids, enzymes, and fiber. For example, the crude fraction contains essential nutrients, such as B vitamins, selenium, chromium, manganese, magnesium, and antioxidants, which are essential for a healthy diet, as well as a high concentration of dietary fiber. For example, 22 g of the crude fraction of the present invention provides 33% of an individual's recommended daily fiber intake. Dietary supplements may also contain known nutritional ingredients that contribute to an individual's overall health, including, but not limited to, vitamins, minerals, other fiber components, fatty acids, antioxidants, amino acids, peptides, proteins, lutein, ribose, omega-3 fatty acids, and / or other nutritional ingredients. The nutritional supplement may be provided in the following forms, but is not limited to: instant drink mixes, ready-to-drink beverages, nutritional bars, wafers, cookies, crackers, gel shots, capsules, chews, chewable tablets, and pills. One embodiment provides the fiber supplement in the form of a flavored shake or malt-type beverage, which may be particularly attractive as a fiber supplement for children.
[0332] In further embodiments, the milling process can be used to create multi-grain flours or coarse fractions of multiple grains. For example, the bran and germ from one type of grain can be milled and blended with milled endosperm or whole grain flour from another type of grain. Alternatively, the bran and germ from one type of grain can be milled and blended with milled endosperm or whole grain flour from another type of grain. The present invention is intended to encompass blending any combination of one or more bran, germ, endosperm, and whole grain flour from one or more grains. This multi-grain approach allows for custom flour creation and can utilize the qualities and nutritional content of multiple types of grains to create a single flour.
[0333] It is contemplated that the whole grain flours, coarse fractions, and / or grain products of the present invention may be produced by any milling method known in the art. An exemplary embodiment includes milling the grain in a single stream without separating the endosperm, bran, and germ of the grain into separate streams. The washed and tempered grain is conveyed to a first-pass grinder, such as a hammer mill, roller mill, pin mill, impact mill, disc mill, air attrition mill, gap mill, and the like. After milling, the grain is discharged and conveyed to a sifter. Furthermore, it is contemplated that the whole grain flours, coarse fractions, and / or grain products of the present invention may be modified or enhanced by numerous other processes, such as fermentation, instantization, extrusion, encapsulation, toasting, roasting, or the like.
[0334] malt production The malt-based beverages provided by the present invention include alcoholic beverages (including distilled beverages) and non-alcoholic beverages produced by using malt as a starting material, either in part or in whole. Examples include beer, happoshu (low-malt beer beverages), whiskey, low-alcohol malt-based beverages (e.g., malt-based beverages containing less than 1% alcohol), and non-alcoholic beverages.
[0335] Malting is the drying process of grains, such as barley and wheat, following controlled steeping and germination. This sequence of events is important for the synthesis of numerous enzymes that result in grain modifications, primarily the breakdown of dead endosperm cell walls and the mobilization of grain nutrients. During the subsequent drying process, aroma and color are produced by chemical browning reactions. While malt's primary use is for beverage production, it is also utilized in other industrial processes, such as in the bakery and confectionery industry as an enzyme source, in the food industry as a flavoring and coloring agent, as malt or malt flour, or indirectly as malt syrup.
[0336] In one embodiment, the present invention relates to a method for producing a malt composition, preferably comprising the following steps: (i) providing grain of the invention, e.g., barley or wheat grain; (ii) soaking the grain; (iii) germinating the soaked grain under predetermined conditions; and (iv) drying the germinated grains; Includes:
[0337] For example, malt can be produced by any of the methods described in Hoseney (Principles of Cereal Science and Technology, Second Edition, 1994: American Association of Cereal Chemists, St. Paul, Minn.). However, any other suitable method for producing malt (e.g., specialty malt production methods, including, but not limited to, malt roasting) can also be used with the present invention.
[0338] Malt is primarily used for beer production, but is also used to produce distilled spirits. Beer production involves wort production, primary and secondary fermentation, and post-processing. First, the malt is crushed, stirred into water, and heated. During this "mashing," enzymes activated by the malt process break down the starch in the grain into fermentable sugars. The resulting wort is clarified, yeast is added, the mixture is fermented, and post-processing is carried out.
[0339] Detection of nitrogenase complex Detection of the nitrogenase complex is carried out by any method that allows detection of the interaction between the NifDK protein complex and the NifH protein. Suitable methods for detecting the interaction between the NifDK protein complex and the NifH protein include any method known in the art for detecting protein-protein interactions, including co-immunoprecipitation, affinity blotting, pull-down, FRET, etc. Alternatively, detection of the nitrogenase complex can be carried out by measuring the activity of the resulting nitrogenase complex. Suitable methods for measuring nitrogenase activity include any method known in the art for detecting the enzymatic reduction of dinitrogen to ammonia, in which electrons are transferred from the NifH protein to the NifDK protein complex. For example, nitrogen fixation activity can be estimated by an acetylene reduction assay. Briefly, this technique is an indirect method that utilizes the ability of the nitrogenase complex to reduce a triple-bond substrate. The nitrogenase enzyme reduces acetylene (C2H2) to ethylene (C2H4). Both gases can be quantified using gas chromatography. Nitrogen fixation may also be measured by a hydrogen release assay. H2 is an essential by-product of N2 fixation. Therefore, an indirect measurement of nitrogenase activity can be obtained by quantifying the H2 concentration in a gas stream using a flow-through H2 sensor or chromatograph.
[0340] Detection of N2 fixation Nitrogen fixation can be estimated by 1) measuring the net gain of total N in the plant-soil system (N balance method), 2) separating plant N into fractions absorbed from the soil and fractions derived from N fixation (N subtraction, 15N natural abundance, 15N isotype dilution, and ureido methods), and 3) measuring nitrogenase activity (acetylene reduction and hydrogen release assays). [Example]
[0341] Example 1 Materials and Methods Gene expression in plant cells using transient expression systems Genes were expressed in plant cells using a transient expression system essentially as described by Wood et al. (2009). Binary vectors containing coding regions expressed in plant cells by the strong, constitutive 35S promoter were introduced into Agrobacterium tumefaciens strains AGL1 or GV3101. A chimeric binary vector, 35S:p19, for expression of the p19 viral silencing suppressor was separately introduced into AGL1 as described in WO 2010 / 057246. Recombinant A. tumefaciens cells were grown to stationary phase at 28°C in LB broth supplemented with 50 mg / L kanamycin and 50 mg / L rifampicin. The bacteria were then pelleted by centrifugation at 5000 g for 5 minutes at room temperature and then resuspended to an OD of 1.0 in infiltration buffer containing 10 mM MES pH 5.7, 10 mM MgCl, and 100 μM acetosyringone. The cells were then incubated with shaking at 28°C for 3 hours, after which the OD was measured and an aliquot of each culture containing the viral suppressor construct 35S:p19 was added to a new tube to reach a final concentration of OD of 0.125. The final volume was adjusted with infiltration buffer. The leaves were then infiltrated with the culture mixture. After infiltration, the plants were typically cultured for an additional 3–5 days, after which leaf discs were harvested for analysis. For combinatorial overexpression of two or more genes of interest, each additional gene was introduced separately into an A. tumefaciens strain and cultured as before. Bacterial suspensions were mixed so that each bacterial strain had a final OD600 of 0.125. A bacterial strain containing the gene encoding the viral silencing suppressor 35S:p19 was included in all mixtures at the same concentration. For example, to express four genes in a transient leaf assay, the final OD600 of the infiltration mixture containing the viral suppressor construct was 5 × 0.125 = 0.625 units. Simultaneous overexpression of at least five genes, each from separate T-DNA vectors, in plant cells in a transient assay format has been previously demonstrated (Wood et al., 2009).
[0342] Protein extraction from leaf tissue To analyze polypeptides produced in plant cells after T-DNA introduction, Nicotiana benthamiana leaf samples were collected by removing approximately 2 x 2 cm leaf pieces from the infiltrated area 5 days after infiltration (unless otherwise specified). These were immediately frozen in liquid nitrogen and ground into powder in a 2 mL Eppendorf tube. 300 μL of buffer was added to each powdered sample. The buffer contained 125 mM Tris-HCl pH 6.8, 4% sodium dodecyl sulfate (SDS), 20% glycerol, and 60 mM dithiothreitol (DTT). The samples were heated at 95°C for 3 minutes and then centrifuged at 12,000 g for 2 minutes. The supernatant containing the extracted polypeptides was removed, and 10 μL to 100 μL was used for Western blotting, depending on the expected level of the polypeptide to be detected.
[0343] Western blot analysis Polypeptides in the extracted samples were separated by SDS-polyacrylamide gel electrophoresis (SDS-PAGE) on a 4-12% NuPAGE Bis Tris gel (ThermoFisher) at 200 V for approximately 1 hour. Separated polypeptides were transferred from each gel to a PVDF membrane using a semi-dry apparatus according to the supplier's instructions (ThermoFisher). After blotting, the gel was stained with Coomassie stain for 1 hour and then rinsed with water to visualize the retained proteins and demonstrate polypeptide transfer. The membrane with bound polypeptides was blocked overnight at 4°C in TBST buffer containing 5% non-fat milk. TBST buffer is. Anti-HA and anti-FLAG antibodies were purchased from Sigma. Anti-GFP antibody was a gift from Leila Blackman (Australian National University, Canberra, Australia). The antibody was added at a 1:5000 dilution in TBST containing 5% non-fat milk, and the membrane was incubated in the solution for 2 hours. The membrane was then washed 3 x 20 minutes with TBST. The secondary antibody, Immun-Star Goat Anti-Mouse (GAM)-HRP conjugate (Biorad), was added at a 1:5000 dilution in TBST containing 5% non-fat milk, and the membrane was incubated for 1 hour, followed by washing the membrane 3 x 15 minutes with TBST. Amersham ECL reagent was used for secondary antibody detection, and the membrane was developed using either an X-ray developer or an Amersham image display device (Amersham).
[0344] Preparation of protoplasts To isolate protoplasts from leaf tissue, a protocol adapted from Breuers et al. (2012) was used. Three days after infiltration (3 dpi), a 2 cm square area of the infiltrated leaf was excised, minced, and transferred to a 5 ml syringe. Two ml of digestion solution containing 1.5% (w / v) Cellulase R-10, 0.4% (w / v) Macrozyme R-10, 0.4 M mannitol, 20 mM KCl, 20 mM MES pH 5.6, 10 mM CaCl2, and 0.1% (w / v) BSA was added, and a slight vacuum was applied manually to facilitate penetration of the solution into the intercellular spaces of the leaf tissue. The solution and leaf pieces were transferred to a 2 ml Eppendorf tube, and the mixture was incubated at room temperature for 1 hour. The resulting protoplasts were gently extracted by manually inverting the tube. Leaf debris was removed using forceps and the protoplasts were allowed to settle, after which the solution was replaced with imaging solution (0.4 M mannitol, 20 mM KCl, 20 mM MES pH 5.6, 10 mM CaCl2, 0.1% BSA).
[0345] Confocal laser scanning microscopy and mitochondrial staining Protoplasts were imaged using an upright Leica confocal laser scanning microscope equipped with a 40x water objective. GFP was excited at 488 nm, and emission was recorded between 499 and 535 nm. Mitochondria were stained for 10–20 min using a 100 nM solution of MitoTracker® Red CMXRos (ThermoFisher Scientific). MitoTracker® Red CMXRos was excited at 561 nm, and emission was recorded between 570 and 624 nm.
[0346] RNA extraction, cDNA synthesis, and analysis To extract RNA from Agrobacterium-infiltrated N. benthamiana leaf cells, leaf pieces approximately 2 × 2 cm in area were frozen in liquid nitrogen, ground to a powder, and 500 μl of Trizol buffer (Thermo Fisher Scientific) was added per sample. The Trizol supplier's instructions were then followed, with these modifications: the chloroform extraction was repeated, and the RNA was dissolved at 37°C. The extracted RNA was treated with RQ1 DNAse (Promega) to remove all extracted DNA. The RNA preparation was then further purified using Plant RNAeasy columns (Qiagen). When performed, cDNA synthesis was performed using Superscript III reverse transcriptase (Thermo Fisher Scientific) with oligo-dT primers according to the supplier's protocol. For RT-PCR analysis of each RNA sample, three separate cDNA synthesis reactions were performed. A 20-μl cDNA reaction was diluted 20-fold in nuclease-free water. qRT-PCR was performed using a Qiagen rotor gene Q real-time PCR instrument. 9.6 μl of each cDNA was added to 10 μl of 2x sensifast no ROX SYBR Taq (Bioline) and 0.4 μl of 10 μmol each of forward and reverse primers for a final reaction volume of 20 μl. All qPCR reactions (for both reference and specific genes) were performed in triplicate under the following cycling conditions: 1 cycle of 95°C / 5 min, 45 cycles of 95°C / 15 s, 60°C / 15 s, and 72°C / 20 s. Fluorescence was measured at 72°C steps. A subsequent melt cycle from 55°C to 99°C was performed. A control amplification of constitutively expressed N. benthamiana GADPH mRNA was used to normalize gene expression using the comparative quantification program in the rotor gene software package. Values for each set of three cDNAs, representing the mean of triplicate assays, were averaged, allowing for calculation of the standard error of the mean (SEM).
[0347] Tandem mass spectrometryInfiltrated N. benthamiana tissues were ground under liquid N2 and then processed using a Retsch tissue-lyser in 50 mM Tris-HCl pH 7.5, 1 mM EDTA, 150 mM NaCl, 0.2% SDS, 10% glycerol, 5 mM DTT, 0.5 mM PMSF, and 1% plant protease inhibitor cocktail (Sigma, catalog no. P9599) in a 2 mL Eppendorf tube. Protein extracts were clarified by centrifugation, and the supernatant was used directly as input for overnight incubation with a monoclonal anti-HA antibody conjugated to agarose beads (Sigma, catalog no. A2095). Unbound proteins were removed by a series of washes with 150 mM Tris-HCl pH 7.5, 5 mM EDTA, 150 mM NaCl, 0.1% Triton X-100, 5% glycerol, 5 mM DTT, 0.5 mM PMSF, and 1% protease inhibitor cocktail. Bound proteins were eluted by incubating the beads in Laemmli buffer (50 mM Tris-HCl, pH 6.8, 2% (w / v) SDS, 0.1% (w / v) bromophenol blue, 10% (v / v) glycerol, 100 mM DTT) at 95°C for 10 min. Input and immunoprecipitated protein samples were separated by SDS-PAGE, and the gel region containing the fusion polypeptide (MTP::NifH::HA) was determined by simultaneous Western analysis from replicate gels. The region containing the fusion polypeptide was excised and subjected to in-gel trypsin digestion, and the trypsin digest was analyzed by tandem mass spectrometry using an Agilent Chip Cube system coupled to an Agilent Q-TOF 6550 mass spectrometer (Campbell et al., 2014).For example, mass spectra from tryptic peptides from common contaminants such as added trypsin and keratin were determined, and the remaining mass spectral data were then used to search against a database containing all protein sequences from Nicotiana species in NCBI (database number 10 / 3 / 2015) in addition to HA sequences using SpectrumMill software (Agilent Rev. B.04.01.141SP1) with a precursor mass tolerance of 15 ppm, a product mass tolerance of 50 ppm, default Q-TOF scoring, and stringent default "autovalidation" settings. Modification of cysteine residues with acrylamide was a required modification, and oxidation of methionine was a variable modification. Initially, tryptic cleavage was required, and a maximum of two erroneous cleavages was allowed. After validation of the peptide matches, the search was repeated using the remaining unmatched spectra, allowing for non-tryptic cleavages.
[0348] Software used for molecular modeling All homology models were constructed using the MODELLER program (Sali and Blundell, 2013) implemented in Accelrys Discovery Studio 3.5. Suitable templates for building homology models were identified using BLAST searches against the Brookhaven Protein Databank. All sequence alignments were performed using the ClustalW algorithm (Sali and Blundell, 2013) implemented in Discovery Studio 3.5. All molecular dynamics simulations were performed using Amber12.
[0349] Transformation of Azotobacter vinelandii Plasmids were transformed into Azotobacter vinelandii according to the method of Dos Santos (2011). Briefly, A. vinelandii DJ1271 was spiked from a DMSO stock onto solid Burk medium lacking molybdate and subcultured for an additional period to remove the cells from DMSO contaminants. A loopful of cells was used to inoculate 50 mL of modified Burk medium lacking molybdate, and iron was added to a 125 mL Erlenmeyer flask. The culture was incubated at 28°C with shaking at 160 rpm for 20–24 hours. 1 ng of the desired plasmid DNA was added to a 50 μL aliquot of competent cells and incubated at room temperature for 20 minutes. The cell-DNA mixture was then added to 3.8 mL of modified Burk medium and allowed to recover for 24 hours at 28°C with shaking at 160 rpm. Aliquots of recovered cells were plated onto solid modified Burk medium containing 6 μg / mL kanamycin and 20 μg / mL ampicillin to select for carriers of pMMB66EH and its derivatives. Plates were incubated at 28°C for 3-5 days. Single colonies were re-passaged onto solid modified Burk medium to obtain single-colony isolates. Re-passaged isolates that retained ampicillin resistance were used to prepare DMSO stocks and inoculate solid modified Burk medium for additional testing.
[0350] Example 2. Use of MTP from the yeast CoxIV gene to target Nif polypeptides to plant mitochondria To the best of our knowledge, there have been no published reports on the production of bacterial nitrogenase (Nif) polypeptides in higher plants, including plant mitochondria. To test such production in mitochondria, we developed a plant-based transient expression system in Nicotiana benthamiana leaves to determine whether Nif polypeptides could be produced and detected in plant cells and whether fusion polypeptides based on bacterial Nif polypeptides could be targeted to mitochondria in plant cells by using a mitochondrial targeting peptide (MTP). To test mitochondrial localization, we chose MTP, derived from the N-terminal region of the yeast cytochrome c oxidase subunit IV (CoxIV) protein. CoxIV MTP has been shown to mediate mitochondrial localization and processing of fusion polypeptides with GFP in plant cells (Kohler et al., 1997). The CoxIV MTP (SEQ ID NO: 1) was only 29 amino acids long before cleavage, shorter than many other MTPs (Huang et al., 2009). When this MTP was processed in yeast cells, 17 or 25 amino acids were removed (Hurt et al., 1985), potentially leaving a minimum of four amino acid residues from the MTP attached to the N-terminus of the polypeptide. However, more recent studies in plant cells (Huang et al., 2009) predicted that processing in the mitochondrial matrix (MM) by mitochondrial matrix protease (MMP) would cleave the peptide immediately upstream of the serine-serine at amino acids 20–21 and generate a processed fusion polypeptide with an additional 10 amino acid residues from the MTP at the N-terminus of the fusion polypeptide. We reasoned that a shorter MTP sequence would be advantageous, especially if the fusion added only 10 amino acids after cleavage.
[0351] A derivative of CoxIV MTP, also referred to herein as dCoxIV (SEQ ID NO: 2), was designed. dCoxIV contains the conserved arginine and serine residues of the motif xRxxxSSx (SEQ ID NO: 3), which are involved in polypeptide import and processing in mitochondria according to Huang et al. (2009), who found from a genome-wide search for mitochondrial targeting and processing sites that the most important residues for import and processing in the plant mitochondrial matrix (MM) are an arginine (-3R or -4R) three or four residues upstream of the cleavage site and two serine residues (+1S, +2S) immediately following the cleavage site. dCoxIV had two additional amino acids inserted toward the N-terminus and also had a glutamic acid at position 28 of SEQ ID NO: 2 rather than the corresponding glutamine in the native CoxIV MTP. These changes were not expected to affect the localization or cleavage of the dCoxIV peptide.
[0352] Construction of vectors pCW440 and pCW441 For gene introduction into plant cells, we designed and constructed a universal plant / bacterial expression vector, pCW440. It is based on a binary vector to enable replication and selection in both Escherichia coli and Agrobacterium tumefaciens bacteria, and contains a T-DNA region for gene transfer from A. tumefaciens into plant cells. To create pCW440, pORE1 (Coutu et al., 2007) was engineered to remove the plant selectable marker gene, creating pORE1-null. This vector contained the 35S promoter and nos3' transcription terminator region. A DNA fragment was synthesized containing, in order, the T7 RNA polymerase promoter, 5' UTR sequence, a nucleotide sequence encoding the dCoxIV MTP initiated by the ATG start codon, a cloning site for the AscI restriction enzyme, and a transcription termination site for T7 RNA polymerase. The sequence of this fragment was partially based on the pET14b vector (Novogene). The fragment was flanked by restriction sites and ligated into the pORE-null between the 35S promoter and the nos3' region, thereby creating pCW440. The nucleotide sequence of the T-DNA region of pCW440 is provided as SEQ ID NO:4. The components of the expression cassette in the vector's T-DNA, in transcriptional order, are the CaMV35S promoter (nucleotides 219-1564 of SEQ ID NO:4), flanked by HindIII and XhoI sites for cloning purposes, to drive expression of the downstream protein-coding region in plant cells to produce a fusion polypeptide; the T7 promoter (nucleotides 1571-1587), which allows expression of the coding region in suitable E. coli cells to produce the same polypeptide; the nucleotides encoding the dCoxIV MTP gene (nucleotides 1650-1742), initiated by the ATG start codon; the AscI restriction enzyme site (nucleotides 1743-1750), which serves as a cloning site for insertion of the protein-coding region; the T7 RNA polymerase transcription termination sequence (nucleotides 1810-1856); and finally, the nos3' plant transcription termination sequence (nucleotides 1861-2084). The genetic map of pCW440 is shown diagrammatically in Figure 1.
[0353] The nucleotide sequence encoding MTP was inserted into pCW440 in a manner that allows for the insertion of any desired protein-coding region within the AscI site to provide an in-frame fusion of the encoded protein to the C-terminus of the dCoxIV amino acid in the translated polypeptide. Because the vector and its derivatives are stable in E. coli, they are used to produce proteins driven by the T7 polymerase promoter system (Studier and Moffatt, 1986). Therefore, the versatile pCW440 was designed as a base vector for expressing fusion polypeptides in bacteria using commercially available T7 ribonucleic acid polymerase cell lines, such as BL21 gold, in which the T7 promoter and terminator control fusion protein expression. Because bacteria do not possess mitochondria, processing of full-length fusion polypeptides containing MTP does not occur in E. coli. The same vector can be used in plant cells to express the same fusion polypeptide under the control of the 35S promoter and nos3' terminator, which controls gene expression using the endogenous transcription machinery in plant cells.
[0354] To generate pCW441, a DNA fragment encoding the open reading frame of GFP (Brosnan et al., 2007) without its initial methionine residue was amplified by PCR and inserted into the AscI site of pCW440, flanked by AscI sites, allowing a translational fusion between dCoxIV and GFP. Insertion at the AscI site introduced three additional amino acids at the junction of the fusion polypeptide. The DNA fragment was inserted into pCW440 to generate pCW441.
[0355] To test whether the dCoxIV region could target the GUS fusion polypeptide to plant mitochondria, A. tumefaciens cells containing pCW441(dCoxIV::GFP) were infiltrated into N. benthamiana leaves using the method described in Example 1. Because N-terminal fusions to GFP generally do not affect its fluorescent activity (Kohler et al., 1997), the dCoxIV::GFP polypeptide was expected to retain fluorescent activity when expressed. Control infiltrations were performed simultaneously with A. tumefaciens containing pUQ214, a construct for expressing cytoplasmically localized GFP (Brosnan et al., 2007). All infiltrations included A. tumefaciens cells containing a construct encoding the viral suppressor of silencing, p19, used alone as a control infiltration. Four days after infiltration, infiltrated leaf sections were examined for fluorescence by light microscopy using blue light excitation at 488 nm and a GFP filter to detect any GFP polypeptide. Fluorescence was observed in leaf cells infiltrated with pCW441. Under the microscope, numerous small intracellular structures were observed to fluoresce, a result consistent with a previous report (Kohler et al., 1997). These structures were highly mobile and tended to cluster at the cell edge, consistent with their presence in mitochondria. In contrast, the introduced cytoplasmic GFP-encoding gene fluoresced more evenly. It was concluded that dCoxIV is sufficient to target the dCoxIV::GFP polypeptide into plant mitochondria, as evidenced by the movement of small, mitochondrial-like particles within plant cells.
[0356] Design and construction of vectors pCW446, pCW447, pCW448 and pCW449 Based on the observation that dCoxIV::GFP fusion polypeptides are produced in plant cells and localize to mitochondria after transient expression, a series of vectors were designed and constructed to provide expression of dCoxIV::Nif fusion polypeptides. First, gene constructs were designed to express fusions to Klebsiella pneumoniae NifH, NifD, NifK, and NifY. The amino acid sequences of wild-type K. pneumoniae NifH, NifD, NifK, and NifY are provided as SEQ ID NOS: 5, 6, 7, and 8, respectively. In an attempt to improve translation efficiency, the protein coding regions of these polypeptides were codon-modified using human codon bias, and cryptic splice sites, potential polyadenylation signals, and internal repeat sequences were removed from the nucleotide sequences provided by a commercial supplier (Geneart). The Nif initiator methionine was removed from each fusion polypeptide. Furthermore, each open reading frame contains a nucleotide sequence encoding a C-terminal extension for each polypeptide, which contains either an HA or FLAG epitope to provide easier detection of the polypeptide using antibodies against the epitope from commercial sources (e.g., Wood et al., 2006). When the transferred polypeptides had similar sizes, various epitopes (commercially available) were added to allow the proteins to be distinguished by using various antibodies. For example, the NifD and NifK fusion polypeptides have similar sizes, so NifD was fused to the FLAG epitope and NifK was fused to the HA epitope. The amino acid sequences added to each C-terminus are provided as SEQ ID NO: 21 for the HA epitope and SEQ ID NO: 22 for the FLAG epitope. The amino acid sequences of the fusion polypeptides containing the dCoxIV MTP and epitopes are provided as SEQ ID NOs: 23, 24, 25, and 26, respectively.
[0357] The codon-optimized nucleotide sequences encoding the NifH::HA, NifD::FLAG, NifK::HA, and NifY::HA polypeptides are provided as SEQ ID NOs: 27, 28, 29, and 30, respectively. DNA fragments having these nucleotide sequences, each containing flanking AscI restriction sites, were synthesized by a commercial supplier and each was inserted into the AscI site of pCW440 to generate vectors pCW446 (pCoxIV::NifH::HA), pCW447 (pCoxIV::NifD::FLAG), pCW448 (pCoxIV::NifK::HA), and pCW449 (pCoxIV::NifY::HA).
[0358] Expression of fusion polypeptides in bacteria The versatile nature of these vectors allows for gene expression and production of fusion polypeptides in suitable bacterial cells, providing a source of polypeptides that can be used as controls in gel electrophoresis and immunodetection experiments. Therefore, these vectors were introduced into the E. coli strain BL21.1-Gold (Stratagene), which allows expression of gene constructs from the T7 promoter after induction by culturing in Overnight Express medium (EMD-Millipore) at 37°C. Bacterial cultures harboring pCW446, pCW447, pCW448, or pCW449 were centrifuged to harvest cells. The cells were lysed in half a volume of BugBuster reagent (EMD-Millipore) at room temperature. The lysates were further centrifuged to harvest inclusion bodies, which were then washed two more times with half a volume of BugBuster reagent. The inclusion bodies were dissolved in standard Laemmli buffer containing SDS and further diluted as needed to produce a clear signal in Western blots. Western blots were probed with commercially available antibodies recognizing the HA or FLAG epitope and a rabbit anti-mouse HRP secondary antibody at a 1:5000 dilution, using ChemStar chemiluminescence reagent (Amersham) as the final detection solution.
[0359] Expression of fusion polypeptides in plant cells Gene constructs for expression of dCoxIV::Nif::HA- or FLAG-tagged fusion polypeptides were separately introduced into N. benthamiana leaves using the methods described in Example 1. As previously described, all infiltrations contained A. tumefaciens cells carrying p19, a construct encoding a viral suppressor of silencing, to reduce the gene silencing response. Five days later, the infiltrated areas were harvested, and bacterially produced extracts were processed as described in Example 1 and above by Western blotting using antibodies binding the HA or FLAG epitope to detect polypeptides and assay their size and relative expression levels. In the gel electrophoresis step, aliquots of bacterial and plant extracts were applied to adjacent lanes to allow for the detection of possible small changes in polypeptide size, which would predict whether the dCoxIV MTP was cleaved. For example, the size of the full-length dCoxIV::NifY::HA polypeptide is approximately 30 kDa, whereas the size of the processed dCoxIV::NifY::HA polypeptide is predicted to be 28 kDa.
[0360] When Western blots (Fig. 2) were examined, the presence of a band corresponding to the dCoxIV::NifH::HA polypeptide was readily observed in samples transfected with the pCW446 T-DNA. Based on the band intensity, we observed that this NifH fusion polypeptide was more strongly expressed in leaf cells than the NifK and NifY fusion polypeptides. The size of the NifH fusion polypeptide produced in plant cells appeared to be approximately 40 kDa on gel electrophoresis, identical to the size of the bacterially expressed polypeptide. This size was as expected for the full-length dCoxIV::NifH::HA fusion protein. In a similar manner, a band corresponding to the dCoxIV::NifK::HA polypeptide was observed in samples transfected with the pCW448 T-DNA, and a band corresponding to dCoxIV::NifY::HA was observed in samples transfected with the pCW449 T-DNA. In each case, the plant-expressed polypeptides appeared to be the same size as the corresponding bacterially expressed polypeptides, approximately 60 kDa for the NifK polypeptide and approximately 30 kDa for the NifY polypeptide. Based on the lack of significant differences in the migration of the bacterially and plant-expressed polypeptides, we concluded that the dCoxIV MTP in each case was not cleaved to any significant extent in plant cells.
[0361] We observed completely different results for the dCoxIV::NifD::FLAG fusion polypeptide in plant cells compared with bacterial cells. When the fusion polypeptide was expressed from pCW447 in bacteria and the samples were assayed by Western blot using an antibody to detect the FLAG-tagged polypeptide, a strong band was observed at 55 kDa, as expected for the full-length fusion polypeptide (Fig. 2). However, when the T-DNA from pCW447 was introduced into plant cells, there was no detectable band for the dCoxIV::NifD::FLAG polypeptide, even after prolonged exposure of the Western blot by imaging techniques. This suggested a lack of production of the fusion polypeptide containing the NifD sequence or its very rapid turnover.
[0362] The experiment with the NifD construct was repeated several times to check the results. Again, the dCoxIV::NifD::FLAG polypeptide was not detected in extracts from plant cells, even after extended exposure of Western blots. The gene construct was checked by sequencing using NifD-specific and 35S-specific primers to confirm the correct sequence of the promoter and NifD fragment within the pCW440 backbone. The same construct was successfully expressed in an E. coli strain, demonstrating the integrity of the open reading frame.
[0363] Expression of other Nif fusion polypeptides in plant cells In each case, similar gene constructs encoding a larger set of Nif fusion polypeptides were generated using the K. pneumoniae sequence and pCW440 as the base vector, which included codon optimization of the protein coding region, omission of the natural Nif initiator methionine, and addition of a C-terminal extension containing an HA or FLAG epitope tag to each fusion polypeptide. These constructs were pCW452 (dCoxIV::NifB::HA; amino acid sequence ID No. 31); pCW454 (dCoxIV::NifE::HA; sequence ID No. 32); pCW455 (dCoxIV::NifN::FLAG; sequence ID No. 33); pCW456 (dCoxIV::NifQ::HA; sequence ID No. 34); pCW450 (dCoxIV::NifS::HA; sequence ID No. 35); pCW451 (dCoxIV::NifU::FLAG; sequence ID No. 36) and pCW453 (dCoxIV:NifX::FLAG; sequence ID No. 37).
[0364] For each of these, production of full-length fusion polypeptides in N. benthamiana cells was readily detected after introduction of the relevant T-DNA from Agrobacterium. In each case, the size of the detected full-length polypeptide was identical to that of the polypeptide produced in the corresponding bacteria, suggesting the lack of MTP cleavage by MMPs in plant cells. In some cases, multiple bands were observed on Western blots. In particular, expression of dCoxIV-translated fusion polypeptides to NifH::HA, NifS::HA, and NifN::FLAG gave rise to several smaller bands on Western blots, possibly due to translation initiated at internal initiation codons within the open reading frame, or to cleavage of the polypeptide in plant cells or during extraction, such that the C-terminal epitope was still present in the detected band. From this analysis, the NifK, NifY, NifE, NifN, NifB, and NifQ fusion polypeptides appeared to be at higher levels than the others, whereas production of the NifH, NifS, NifU, and NifX fusion polypeptides was detected at moderate to low levels. Again, production of the NifD fusion polypeptide in plant cells was not detected. We concluded that all of the Nif fusion polypeptides except the NifD polypeptide could be expressed in plant cells using this approach, or that they were expressed at different levels or accumulated at different abundances despite similar codon usage parameters and identical promoter and polyadenylation regulatory sequences. Importantly, we concluded that there was something special about the NifD construct or fusion polypeptide because it stood out clearly.
[0365] Example 3. Further attempts to detect production of NifD fusion polypeptides in plant cells Given the lack of detection of the dCoxIV::NifD::FLAG fusion polypeptide described in Example 2, we hypothesized that this Nif fusion polypeptide might be susceptible to degradation, possibly related to oxygen concentration, photosynthesis in leaf cells, or potential misfolding and thus instability of the polypeptide, possibly due to the loss of a putative chaperone protein (Ribbe and Burgess, 2001). To test these hypotheses, N. benthamiana leaves were infiltrated with A. tumefaciens harboring the vector pCW447, whose T-DNA encoded the pCoxIV::NifD::FLAG polypeptide. Infiltrated plant tissue was excised and maintained in liquid medium for 24 hours under various conditions. Combinations of each of these variations were also tested. The variations included maintaining plant tissue in the dark or light at 21%, 5%, or 1% atmospheric oxygen. Co-transfection of gene constructs expressing the GroEL polypeptide (Ribbe and Burgess, 2001) or leghemoglobin (Ott et al., 2005) (Lhb; Figures 6 and 3) was also performed in some infiltrations. This was achieved by inserting the GroEL- and Lbh-coding sequences into the pORE1-35S expression vector (Wood et al., 2009), generating the pCW-GroEL and pCW444 constructs, respectively. Both constructs were transformed into Agrobacterium and used for transient leaf expression, as previously described.
[0366] None of these variations in conditions or genes resulted in the detection of the dCoxIV::NifD::FLAG fusion polypeptide in N. benthamiana cells, even though control infiltrations resulted in the robust production of other Nif fusion polypeptides, such as the dCoxIV::NifN::FLAG polypeptide.
[0367] The inventors concluded from these experiments that expression of NifD fusion polypeptides in plant cells posed a problem that needed to be solved.
[0368] Example 4. Expression of multiple vectors expressing combinations of Nif polypeptides in plant cells We also tested whether the four Nif fusion polypeptides could be coexpressed in N. benthamiana leaf systems and whether this would improve production levels of the dCoxIV::NifD polypeptide. This was tested by mixing four A. tumefaciens cell suspensions, each containing a different gene for expression: dCoxIV fusions to NifY::HA, NifD::FLAG, NifK::HA, and NifH::HA; and A. tumefaciens containing the p19 construct alone. When extracts from infiltrated leaf tissue were assayed by Western blotting as before, this four-gene coinfiltration experiment resulted in binding of dCoxIV::NifK, dCoxIV::NifY, and dCoxIV::NifH, but at much lower levels than when each was expressed alone. However, there was no detectable band for dCoxIVNifD. Thus, once again, the NifD fusion polypeptide appeared to be distinct from other Nif polypeptides: the combinations of dCoxIV::NifK::HA, dCoxIV::NifH::HA, and dCoxIV::NifY::HA did not enhance the production of coexpressed dCoxIV::NifD::FLAG to detectable levels.
[0369] Consideration In the previously described experiments, we synthesized and used gene constructs encoding dCoxIV, a derivative of the yeast mitochondrial targeting peptide (MTP), separately fused to 10 different Nif polypeptides. These were also fused to either an HA or FLAG epitope tag at their C-terminus to detect the polypeptides when produced in bacteria or plant cells, and were shown to be expressed in leaf tissue after T-DNA transfer by A. tumefaciens. Processing of MTP in mitochondria was examined by carefully comparing the sizes of polypeptides extracted from bacteria and N. benthamiana cells. We observed that all 10 dCoxIV-Nif fusion polypeptides were readily detected in Western blot assays after production in E. coli bacteria, but only nine polypeptides were detected after transfer into N. benthamiana cells. The exception was the dCoxIV::NifD::FLAG polypeptide, which was consistently undetectable. Changes in the environmental conditions in which the plant tissue was maintained did not result in the detection of the polypeptides—reduced oxygen concentrations, maintenance of plant tissue in the dark versus the light, or the presence of co-expressed leghemoglobin or GroEL molecular chaperone proteins.
[0370] NifD is an essential component of the nitrogenase enzyme complex, and therefore the experiments in the following examples were carried out in an attempt to alleviate the deficiency of NifD fusion polypeptides.
[0371] Example 5. Validation of pFAγ MTP for targeting proteins to plant mitochondria As described in Examples 2–4, production of 10 of 11 tested Nif fusion polypeptides was demonstrated in N. benthamiana leaf cells using dCoxIV MTP, with the only exception being the NifD fusion polypeptide. In all cases, no processing of the targeting peptide was observed. Therefore, we tested whether various MTP sequences fused to NifD and other Nif polypeptides would provide detectable expression and cleavage of the fusion polypeptide in plant cells. To this end, we selected the MTP from the A. thaliana F1-ATPase γ-subunit (pFAγ). The 77-amino acid MTP was functionally validated in Arabidopsis protoplasts (Lee et al., 2012) by fusion with a GFP reporter polypeptide. Cleavage of the pFAγ MTP sequence by matrix processing protease (MPP) occurred after amino acid 42 and cleaved off the 35-amino acid C-terminal portion fused to the N-terminus of the polypeptide. Because it is not known how many of the 35 amino acids are required for mitochondrial localization and processing in N. benthamiana leaf cells, we therefore decided to use the entire 77-amino acid sequence. Therefore, the ability of pFAγ MTP to transport fusion Nif polypeptides to the MM of intact plant leaf cells was tested using the N. benthamiana transient leaf assay system.
[0372] A 970-bp DNA fragment was chemically synthesized encoding a 319-amino acid polypeptide consisting of 77 amino acids of pFAγ MTP (amino acids 1–77 of SEQ ID NO: 38) fused to GFP (GFP 65T, GenBank accession number U43284; Haas et al., 1996), separated by three amino acids gly-ala-pro (GAP) and also containing flanking NcoI and AscI restriction sites. After NcoI and AscI (partial) digestion, the fragment was inserted into the NcoI-AscI sites of pCW441 to generate vector pRA01. Digestion of pRA01 with AscI excised the GFP-coding region but left the pFAγ MTP sequence and GAP amino acids, generating vector pRA00, which was used as a base vector for cloning Nif fusion polypeptides. When Nif or other polypeptides are fused downstream of pFAγMTP and GAP amino acids by insertion of a DNA sequence at the AscI site, the resulting gene construct encodes pFAγMTP and GAP amino acids at the junction of the fusion, thereby adding an 80-amino acid N-terminal extension to Nif or other polypeptides (SEQ ID NO: 38). Mitochondrial processing of pFAγMTP is predicted to cleave 42 amino acids from the N-terminus, including the initiator methionine, and reduce the size of the expressed polypeptide to approximately 4.6 kDa. Therefore, cleavage of MTP is predicted to leave 38 amino acids from the C-terminal extension fused to the N-terminus of Nif or other polypeptides.
[0373] As a control vector, we generated a second vector encoding a disabled version of pFAγMTP, a vector of the same length that would encode an N-terminal extension but not provide mitochondrial processing by MPP. In this vector, 24 amino acid substitutions were introduced within the region of MTP required for its mitochondrial recognition and processing (Lee et al., 2012) and included amino acids at the cleavage site of pFAγ. Each substitution replaced the wild-type amino acid in pFAγMTP of pRA1 with an alanine residue (Figure 3). Therefore, this modified version of pFAγMTP, also referred to herein as mFAγ, would not be processed correctly but would instead yield a full-length fusion polypeptide with the same number of amino acid residues as the corresponding unprocessed pFAγ fusion polypeptide. Therefore, this modified vector was used as a base vector for expressing fusion polypeptides to provide an unprocessed molecular weight control in Western blot analysis. Both vector pRA00 and its mFAγ derivatives are binary vectors and provided for the transfer of their T-DNA from A. tumefaciens to plant cells. The amino acid sequence of the N-terminal extension of modified mFAγ, containing a GAP triplet at its C-terminus, is provided as SEQ ID NO: 39.
[0374] To test the ability of pFAγ MTP to localize the fusion polypeptide to the mitochondria of plant cells, pRA01 was utilized in a transient N. benthamiana leaf assay. As a control for detecting matrix processing of the pFAγ::GFP fusion polypeptide in plant cells, a corresponding vector encoding the mFAγ::GFP fusion polypeptide was also constructed and designated pRA21 (Table 3). Five days after infiltration of leaves with A. tumefaciens containing either pRA01 or pRA21, leaf samples from the infiltrated area were harvested, and protein extracts were prepared as described in Example 1. SDS-PAGE gel electrophoresis and Western blotting were performed on the protein extracts using a GFP antibody. From the introduction of pRA01 (pFAγ::GFP), a polypeptide band of the predicted size (∼30 kDa) was observed by Western blot for the GFP polypeptide cleaved at the predicted site, whereas for pRA21 (mFAγ::GFP), a larger band (∼35 kDa) was observed for the predicted size of the unprocessed fusion polypeptide (Figure 4). A fainter band at ∼28 kDa was observed for both pRA01 and pRA21, which was not observed in the negative control lacking either gene encoding the GFP polypeptide. This likely represented a degradation product of the GFP polypeptide or one resulting from alternative transcription or translation. Because the cleaved band was much more intense than the uncleaved band (Figure 4), processing of the pFAγ::GFP fusion polypeptide appeared to be efficient. We conclude that the pFAγ::GFP fusion polypeptide is processed in the MM of intact N. benthamiana leaf cells, implying that at least the MTP portion of the fusion polypeptide is transported into the MM and is accessible to the MPP.
[0375] To visualize the localization of the GFP polypeptide, protoplasts were prepared from leaf tissue containing pRA01(pFAγ::GFP) as described in Example 1 and examined by confocal microscopy. Protoplasts were imaged using an upright Leica confocal laser scanning microscope equipped with a 40x water objective. GFP was excited at 488 nm, and emission was recorded between 499 and 535 nm. As a counterstain to identify mitochondria, protoplasts were also stained for 10–20 minutes with a 100 nM solution of MitoTracker® Red CMXRos (ThermoFisher Scientific, Cat. No. M7512). MitoTracker® Red CMXRos was excited at 561 nm, and emission was recorded between 570 and 624 nm. In this way, fluorescence images of both fluorophores were sometimes overlaid. By these means, we observed that GFP fluorescence colocalized with MitoTracker® and, therefore, was localized to mitochondria. Therefore, we concluded that, at least for pFAγ::GFP, MTP both transported the fusion polypeptide to the MM of N. benthamiana leaf cells and provided for its cleavage by MPP (Fig. 4). This conclusion was based on the knowledge that processing of MTP by MPP occurs only in the MM.
[0376] Table 3. Gene constructs encoding Nif fusion polypeptides [Table 3]
[0377] Example 6 Transfer of Nif fusion polypeptides into plant mitochondria We next wanted to test whether pFAγ MTP could transfer Nif fusion polypeptides to the MM and whether it could provide for cleavage of the MTP by MPP. To initially test this, two Nif proteins (NifF and NifZ) were selected for the construction of fusion polypeptides because of their relatively small molecular weight, which allowed clear differentiation of cleaved and uncleaved polypeptides by Western blotting. The protein coding regions of Klebsiella pneumoniae NifF and NifZ polypeptides fused to either the HA or FLAG epitope as C-terminal fusions were human codon-optimized and synthesized by a commercial supplier. DNA fragments were inserted into the AscI site of pRA00 so that the Nif open reading frame was translationally fused to the N-terminus of pFAγ MTP, generating vectors pRA05 and pRA04 (Table 3). HA and FLAG epitopes were incorporated at the C-terminus of the NifF and NifZ polypeptides, respectively, to allow detection with the corresponding antibodies. To generate unprocessed versions of these pFAγ::Nif fusion polypeptides as controls, the same constructs were expressed in E. coli using T7 RNA polymerase. Given that MTPs are not processed in bacteria lacking MPPs, the size difference between plant- and bacterially expressed polypeptides allowed processing to be detected by gel electrophoresis and Western blotting. The amino acid sequences of the unprocessed pFAγ::NifF::HA and pFAγ::NifZ::FLAG fusion polypeptides are provided as SEQ ID NOs:40 and 41, respectively.
[0378] Western blots (Fig. 4) revealed that the size of the polypeptides detected in N. benthamiana leaves was smaller in each case than the corresponding polypeptides produced in E. coli. For pRA05 (pFAγ::NifF::HA) and pRA04 (pFAγ::NifZ::FLAG), the polypeptides detected in plant cells corresponded to the size predicted for cleavage of the fusion polypeptides from those MTPs, whereas the polypeptides detected in E. coli extracts were of the size predicted for the unprocessed pFAγ::Nif fusion polypeptide. We conclude from these data that pFAγ MTP can transport at least the MTP portion of the Nif fusion polypeptide to the MM and provide for cleavage of the MTP by the MPP in plant cells.
[0379] Example 7. Demonstration of MTP cleavage by mass spectrometry To verify that the predicted MTP processing site within the pFAγ portion of the fusion polypeptide was cleaved intramitochondrially by MPP, the peptide product was analyzed by mass spectrometry. To do this, we designed and generated a pRA00-based gene construct encoding the fusion polypeptide pFAγ::NifH::HA, designated pRA10 (Table 3). The amino acid sequence of this fusion polypeptide before processing is provided as SEQ ID NO:42. The NifH fusion polypeptide was chosen due to the importance of NifH as a core component of the nitrogenase enzyme complex and the previously observed high expression of the dCoxIV::NifH::HA polypeptide in plant cells (Example 2). The pRA10 construct was introduced into N. benthamiana leaf cells, and infiltrated tissue was harvested 4 days later. The tissue was pulverized under liquid nitrogen and then processed in a 2 mL Eppendorf tube using a Retsch tissue lyser in the presence of protein extraction buffer (PEB) containing 50 mM Tris-HCl, pH 7.5, 1 mM EDTA, 150 mM NaCl, 10% glycerol, 5 mM DTT, 0.5 mM PMSF, and 1% plant-specific protease inhibitor cocktail (Sigma, catalog no. P9599). To improve the solubility of the mitochondria-targeted NifH fusion polypeptide, 0.2% (w / v) SDS was added to PEB. The crude protein extract was clarified by centrifugation and then incubated in the presence of a monoclonal anti-HA antibody coupled to agarose beads (Sigma, catalog no. A2095) to immunoprecipitate HA-containing polypeptides. Unbound proteins were removed by a series of washes with 150 mM Tris HCl, pH 7.5, 5 mM EDTA, 150 mM NaCl, 0.1% Triton X-100, 5% glycerol, 5 mM DTT, 0.5 mM PMSF, and 1% plant-specific protease inhibitor cocktail. Bound proteins were eluted by incubating the beads in Laemmli buffer at 95°C for 10 min, followed by further purification by electrophoresis on denaturing SDS-PAGE.Regions of the gel determined to contain the pFAγ::NifH::HA fusion polypeptide by simultaneous Western blot analysis were excised and subjected to in-gel digestion with trypsin, followed by tandem mass spectral analysis of the resulting peptides as described in Example 1 using an Agilent Chip Cube system interfaced to an Agilent Q-TOF 6550 mass spectrometer (Campbell et al., 2014). This analysis found five fully digested tryptic peptides identical to regions in NifH and six quasi-tryptic peptides consistent with precise cleavage of the MTP between residues 42 and 43 (Figure 5). The tryptic peptide SISTQVVR (SEQ ID NO: 43), which would have been obtained from unprocessed MTP, was not observed. Instead, the most abundant N-terminal peptide detected was the quasi-tryptic ISTQVVR (SEQ ID NO: 44), confirmed by the full series of y-ions in its MS / MS spectrum.
[0380] These data conclusively demonstrated that at least the MTP portion of the pFAγ::NifH::HA polypeptide was translocated to MM and cleaved by MPP at a preselected site in the MTP within the N-terminal extension, implying that the pFAγ MTP contains all of the signals necessary for metastasis and processing in MM.
[0381] Example 8. Expression of multiple nitrogenase proteins in the plant mitochondrial matrix Given the successful mitochondrial expression and processing of NifF, NifZ, and NifH polypeptides fused to GFP and pFAγ MTP in N. benthamiana leaf cells, we attempted to express the remaining 13 Nif polypeptides as fusion polypeptides to pFAγ MTP. In the model diazotroph, Klebsiella pneumoniae, 16 Nif proteins are involved in nitrogenase biosynthesis or function, while four others are of unknown function or involved in transcriptional regulation (Oldroyd and Dixon, 2014). Given the data presented in Examples 2–4, we were particularly interested in whether pFAγ MTP could mediate the production and cleavage of the NifD fusion polypeptide.
[0382] Codon-optimized versions of DNA fragments encoding the 16 K. pneumoniae Nif polypeptides were obtained and each was separately inserted into the AscI site of pRA00 to generate a series of gene constructs (Table 3). Each gene construct encoded a fusion polypeptide with an N-terminal pFAγ MTP, followed by the Nif sequence (with or without its initiator methionine), and then a C-terminal extension containing either the HA or FLAG epitope for detection with an appropriate antibody. For plant expression, each construct contained a 35S promoter and nos3' transcription terminator region flanking the protein coding region. The amino acid sequences of the 16 fusion polypeptides are presented in SEQ ID NOs: 40-42 and 46-58.
[0383] A. tumefaciens cells containing the 16 gene constructs were separately infiltrated into N. benthamiana leaves, and after 4 days, protein extracts were prepared and analyzed by gel electrophoresis and Western blotting as before. For each of the constructs encoding the HA-tagged pFAγ-Nif polypeptide, a band of the approximate size predicted for MPP was detected by Western blot (Table 3, Figure 6). Protein abundance varied among the HA-tagged polypeptides, with the NifB, NifH, NifK, NifS, and NifY fusion polypeptides being the easiest to detect. NifF, NifE, and NifM polypeptides were present at low levels, while detection of the NifQ polypeptide required longer exposure of the blot to be visualized. Interestingly, additional, higher molecular weight bands were detected for infiltration with the NifB, NifS, NifH, and NifY constructs, size-specific to each of the individual Nif::HA constructs (Figure 6). These additional bands were approximately twice as large as the initial bands and suggested to us that these polypeptides were dimerized despite the denaturing conditions during gel electrophoresis. The NifB, NifS, and NifH proteins have been reported to function as homodimers in bacteria (Rubio and Ludden, 2008; Yuvaniyama et al., 2000).
[0384] FLAG-tagged pFAγ::Nif::FLAG fusion polypeptides were visualized by Western blot for each of the constructs containing NifJ, NifN, NifV, NifU, NifX, and NifZ sequences (Fig. 6, upper right panel). The FLAG antibody yielded higher background bands from N. benthamiana extracts than the HA antibody. Nevertheless, the results were similar to those for HA-tagged proteins. Considerable variation was observed in the signal intensity of the various Nif::FLAG fusion polypeptides. An additional, higher molecular weight band specific to the NifU construct was observed (Fig. 6). The pFAγ::NifX::FLAG construct also yielded an additional, smaller band of higher intensity than the predicted, processed molecular weight.
[0385] Surprisingly, given that 15 of the 16 gene constructs resulted in detectable polypeptide production, and despite frequent replication of infiltrations, we were unable to detect any distinctive bands for the pFAγ::NifD::FLAG construct (pRA07) in these plant assays, even though expression of the same gene construct in E. coli readily yielded visible bands of the predicted molecular weights (Figure 7). In fact, even a 1:100 dilution of the bacterially produced extract yielded bands readily observable in Western blots. The results from the bacterial extracts confirmed that the gene construct pRA07 was functional, at least with respect to the protein-coding regions. Thus, with the notable exception of the pFAγ::NifD::FLAG polypeptide, all of the complete set of essential Nif polypeptides required for nitrogenase biosynthesis and function were successfully expressed in plant leaf cells using fusions to pFAγ MTP. Furthermore, the molecular weights of the observed bands were consistent with processing of the fusion polypeptides in MM (Table 3).
[0386] Because nitrogenase activity requires the coordinated action of multiple Nif proteins, we anticipated that reconstitution of functional nitrogenase in plants would require the multiple Nif proteins utilized by diazotrophic bacteria for biosynthesis and function. Therefore, we wanted to determine whether multiple Nif proteins could be expressed in N. benthamiana MM using pFAγ MTP. To test this concept, we selected genetic constructs encoding four Nif fusion polypeptides of different sizes, namely, pFAγ::NifB::HA (pRA03), pFAγ::NifS::HA (pRA16), pFAγ::NifH::HA (pRA10), and pFAγ::NifY::HA (pRA12), whose resulting polypeptides could be identified by Western blot analysis. Four A. tumefaciens cultures transformed with these constructs were mixed in equal amounts, and the mixture was then infiltrated into N. benthamiana leaves. The accumulated polypeptide concentrations were compared when each culture was infiltrated separately. It was observed that each polypeptide was more abundant when expressed from a single construct than from the four-gene combination. Nevertheless, all four Nif fusion polypeptides were readily detected in protein extracts from the gene combination, and the molecular weights observed for each polypeptide were identical for both individual and combined infiltrates. This demonstrated that combinations of Nif fusion polypeptides can be produced in plant cells with the desired targeting to mitochondria and processing of each Nif fusion polypeptide.
[0387] Example 9. Attempts to Improve NifD Fusion Polypeptide Production in N. benthamiana Given the essential role of NifD in nitrogenase catalytic activity, we attempted to identify the reason for the lack of NifD fusion polypeptide production and tested several approaches to improve its abundance in plant assays. First, they tested whether the lack of NifD fusion polypeptide production / accumulation could be attributed to low transgene transcription or mRNA instability by measuring mRNA expression levels. To do this, we first measured mRNA levels in infiltrated N. benthamiana cells from pRA07 by qRT-PCR as described in Example 1 and compared them to the levels of mRNA transcribed from a construct encoding pFAγ::NifU::FLAG (pRA15). This second construct was used as a control because, as previously described, it provides high levels of polypeptide production in plant cells. To eliminate any bias in amplification efficiency, we used oligonucleotide primers that anneal within the pFAγ MTP region shared by both Nif fusion genes. Results from RT-PCR assays showed that pFAγ::NifU::FLAG mRNA was more than threefold higher than the level of pFAγ::NifD::FLAG mRNA. Second, cDNA was synthesized from plant-produced mRNA transcribed from pRA07, cloned, and sequenced, verifying the base integrity of the nucleotide sequence. Therefore, the lack of accumulation of the pFAγ::NifD::FLAG fusion polypeptide was not due to low mRNA expression or instability. These experiments also demonstrated that the T-DNA of pRA07 encoding pFAγ::NifD::FLAG was fully functional, and the 35S promoter of the construct was functional as well. Therefore, the failure to detect the NifD fusion polypeptide in N. benthamiana cells was not due to any impairment in mRNA expression.
[0388] Because transcription and mRNA accumulation of the gene encoding pFAγ::NifD::FLAG were not clearly limiting NifD fusion polypeptide production, several modifications were made to the gene construct in an attempt to overcome the lack of NifD fusion polypeptide accumulation. First, we considered and tested the possibility that the presence of the FLAG epitope in the C-terminal extension might be causing either lack of production or instability of the NifD fusion polypeptide. To test this possibility, we designed and generated a construct designated pRA19(pFAγ::NifD::HA), replacing the FLAG epitope with an HA epitope. The HA epitope enabled the accumulation and detection of each of the tagged Nif fusion polypeptides tested by its epitope (Table 3). The amino acid sequence of this fusion polypeptide is provided as SEQ ID NO:59. Second, the codon usage of the NifD::HA open reading frame was modified to more closely resemble codon usage in A. thaliana rather than optimized for translation in human cells in an attempt to determine whether a different mRNA sequence would improve translation efficiency. The construct was called pRA24; it encoded the same pFAγ::NifD::HA polypeptide (SEQ ID NO:59) as pRA19. Additionally, a genetic construct was generated encoding a version of the NifD fusion polypeptide with a mutant mFAγ N-terminal extension rather than pFAγ as described in Example 5 (mFAγ::NifD::HA). This construct (pRA22) was generated to test whether mitochondrial targeting and / or processing, if any, is at least partially responsible for the lack of NifD fusion polypeptide production. Three constructs designed to address these questions are listed in Table 3.
[0389] N. benthamiana leaves were infiltrated with A. tumefaciens containing these constructs, and protein extracts were prepared and analyzed from the infiltrated tissue. Surprisingly, Western blots revealed HA-containing bands of the predicted molecular weights for both matrix-processed and unprocessed NifD fusion polypeptides when either the pRA19 or pRA24(pFAγ::NifD::HA) constructs were introduced (Fig. 8). Introduction of pRA22 yielded only the larger (unprocessed) fusion polypeptide. In this experiment, introduction of pRA19 into plant cells yielded a stronger NifD::HA fusion polypeptide band on Western blots than pRA24, although the difference in intensity was not significant in subsequent replicate experiments. Matrix processing was confirmed by comparing the band positions with those produced from pRA22(mFAγ::NifD::HA) and bacterially produced polypeptides (Fig. 8). The observation of two bands, the sizes of which corresponded to the processed and unprocessed forms of pFAγ::NifD::HA, indicated that processing of this NifD fusion polypeptide was less efficient than that of the other Nif fusion polypeptides described previously. The observed levels of the pFAγ::NifD::HA polypeptide were much lower than those of the pFAγ::NifK::HA fusion polypeptide (lane 2), used as a positive control, despite the same expression construct design and expression conditions. Furthermore, for each of the three modified NifD::HA constructs, additional bands of comparable molecular weight were observed on Western blots, some of which were specific to particular pFAγ::NifD::HA versions. For example, a strong band at approximately 50 kDa was distinct despite being present in all modified NifD::HA constructs, and a strong band at approximately 40 kDa appeared unique to samples in which pRA22 was introduced. As these bands were absent in either the pFAγ::NifK::HA or GFP control, it is possible that they represent NifD::HA degradation products or possibly alternative transcription products or translation initiation signals.
[0390] We conclude that the replacement of the FLAG epitope with the HA epitope is at least partially responsible for improving the accumulation of the NifD fusion polypeptide.
[0391] Consideration In the experiments described above, we demonstrated that all 16 Nif polypeptides required for nitrogenase function in Klebsiella can be expressed from a genetic construct in plant leaf cells as MTP::Nif fusion polypeptides, and that the polypeptides accumulate in processed forms in mitochondria. Furthermore, the experiments showed that these proteins can be targeted to the MM, a subcellular location potentially corresponding to nitrogenase function. We believe that these experiments represent the first practical demonstration of the feasibility of such an approach.
[0392] The conclusion for targeting Nif fusion polypeptides to MM was based on several lines of evidence. First, the size of each of the plant-expressed Nif polypeptides was consistent with the predicted size that would result from processing by MPP. The smaller molecular weight observed for the plant-expressed Nif fusion polypeptides compared with the bacterially produced polypeptides (full-length, unprocessed) indicated processing of the Nif fusion polypeptides by MPP. Furthermore, when the MTP sequence was mutated, rendering MTP inaccessible to processing by the mitochondrial import machinery, larger polypeptides were observed for both the NifD and GFP fusion polypeptides, consistent with the size difference between the processed and unprocessed polypeptides. Finally, mass spectrometry determined that pFAγ::NifH was cleaved between residues 42–43 of MTP, as predicted for specific processing in the matrix, indicating that the MTP sequence can be cleaved when fused to a Nif polypeptide.
[0393] The presence of MTP did not always lead to complete processing of Nif proteins. For example, the NifX::FLAG construct and two NifD::HA constructs resulted in the accumulation of both processed and unprocessed fusion polypeptides. Moreover, despite the use of a strong, constitutive 35S promoter, a great deal of variability in polypeptide accumulation levels was observed with the various Nif fusion polypeptides. Additional, shorter-than-predicted polypeptides were observed for some Nif fusion polypeptides, detected by the presence of epitopes at the C-terminus of the polypeptides.
[0394] Of all the Nif fusion polypeptides, the NifD polypeptide was the most difficult to produce at detectable levels. This was not due to lack of NifD gene expression or mRNA instability, but rather poor translation and protein instability likely limited the abundance of the NifD fusion polypeptide. Given the critical importance of NifD in nitrogenase function, including its requirement for high expression in bacteria (Poza-Carrion et al., 2014), this problem must be overcome.
[0395] Example 10. Exploring Nitrogenase Polypeptide Structure and Function Using In Silico Modeling Bacterial nitrogenase is a metalloprotein complex composed of multiple distinct Nif polypeptides that must bind to f...
Claims
1. A plant cell comprising a mitochondrion and one or more exogenous polynucleotides encoding one or more Nif fusion polypeptides (NFs), each NF comprising: (i) a mitochondrial targeting peptide (MTP) having a C-terminus, and (ii) a Nif polypeptide (NP) having an N-terminus, said NP being selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ; Including, the C-terminus of the MTP is translationally fused to the N-terminus of the NP; each MTP is independently the same or different, and each NP is independently the same or different; the mitochondria contain the one or more NFs and / or their processed products (CFs); Each CF, if present, is generated by cleavage of the corresponding NF within its MTP; plant cells.
2. 2. The plant cell of claim 1, wherein the NF polypeptide further comprises one or more, or all, Nif polypeptides selected from the group consisting of NifD, NifH, and NifK.
3. A plant cell comprising a mitochondrion and a first exogenous polynucleotide encoding a first Nif fusion polypeptide (NF) and a second exogenous polynucleotide encoding a second NF, each NF comprising: (i) a mitochondrial targeting peptide (MTP) having a C-terminus, and (ii) a Nif polypeptide having an N-terminus (NP); Including, the C-terminus of the MTP is translationally fused to the N-terminus of the NP; each MTP is independently the same or different, and each NP is independently the same or different; The mitochondria are (a) the first NF and / or its processed product (first CF), and (b) the second NF and / or its processed product (second CF), Including, Each CF, if present, is generated by cleavage of the corresponding NF within its MTP; plant cells.
4. 4. The plant cell of claim 3, wherein the first and second NFs are selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ.
5. The plant cell of any one of claims 1 to 4, wherein each CF independently comprises from about 5 to about 11 amino acid residues from the C-terminus of the MTP.
6. Mitochondria and (a) a first exogenous polynucleotide encoding a NifD fusion polypeptide (NDF), said NDF comprising (i) a first mitochondrial targeting peptide (MTP1) having a C-terminus, and (ii) a NifD polypeptide (ND) having an N-terminus, wherein the C-terminus of MTP1 is translationally fused to the N-terminus of said ND; (b) a second exogenous polynucleotide encoding a NifH fusion polypeptide (NHF), said NHF comprising (i) a second mitochondrial targeting peptide (MTP2) having a C-terminus, and (ii) a NifH polypeptide (NH) having an N-terminus, wherein the C-terminus of MTP2 is translationally fused to the N-terminus of said NH; (c) a third exogenous polynucleotide encoding a NifK fusion polypeptide (NKF), said NKF comprising (i) a third mitochondrial targeting peptide (MTP3) having a C-terminus, and (ii) a NifK polypeptide (NK) having an N-terminus, wherein the C-terminus of the third MTP is translationally fused to the N-terminus of the NK; A plant cell comprising: each of MTP1, MTP2, and MTP3 is independently the same or different; the plant cell comprises a level of NDF and / or CDF that is greater than the level of NDF and / or its processed product (CDF) in a corresponding plant cell comprising the first exogenous polynucleotide but lacking the second and third exogenous polynucleotides; The CDF, if present, is generated by cleavage of the NDF in the MTP1; plant cells.
7. 7. The plant cell of claim 6, further comprising one or more exogenous polynucleotides encoding one or more NFs selected from the group consisting of NifE, NifF, NifJ, NifM, NifN, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ.
8. 8. A plant cell according to claim 2, 4 or 7, wherein the C-terminus of the NifK and / or NifN, if present, is identical to the C-terminus of a wild-type NifK or NifN polypeptide, respectively.
9. 1. A plant cell comprising a mitochondrion, a first exogenous polynucleotide encoding a NifD polypeptide (ND), and a second exogenous polynucleotide encoding a NifK polypeptide (NK), Either or both of the ND and NK are translationally fused at their N-terminus to a mitochondrial targeting peptide (MTP) having a C-terminus; When both the ND and NK are translationally fused to an MTP, each MTP is independently the same or different; said second exogenous polynucleotide is covalently or non-covalently linked to said first exogenous polynucleotide; The plant cells contain approximately the same levels of ND and NK; plant cells.
10. The plant cell according to any one of claims 1 to 9, wherein the MTP comprises at least 10 amino acids, preferably between 10 and 80 amino acids.
11. 11. The plant cell according to any one of claims 1 to 10, wherein at least one, more than one or all of the MTPs comprise a mitochondrial protein precursor MTP or a variant thereof, preferably a plant MTP.
12. The plant cell according to any one of claims 1 to 11, wherein an exogenous polynucleotide encoding the fusion polypeptide is integrated into the nuclear genome of the cell.
13. 13. A transgenic plant, or a transgenic part thereof, comprising a cell according to any one of claims 1 to 12, wherein said transgenic plant is transgenic for one or more exogenous polynucleotides encoding a fusion polypeptide.
14. 14. The transgenic plant or part thereof of claim 13, wherein one, more or all of the exogenous polynucleotides are expressed in the roots of the plant, and preferably at a greater level in the roots of the plant than in the leaves of the plant.
15. 15. A transgenic plant or part thereof according to claim 13 or claim 14, which is a cereal plant or part thereof, such as wheat, rice, maize, triticale, oat or barley, preferably wheat.
16. A population of at least 100 plants according to any one of claims 13 to 15 growing in a field.
17. 14. The transgenic plant part of claim 13, which is a seed.
18. 1. A method for producing flour, whole grain flour, starch, oil, seed meal or other product derived from seeds, comprising the steps of: a) obtaining a seed according to claim 17, and b) extracting flour, wholemeal, starch, oil or other products or producing seed meal.
19. A product produced from a plant according to any one of claims 13 to 15 and / or a seed according to claim 17.
20. 20. A method of preparing a food product comprising mixing the seed, or flour, wholemeal or starch from the seed, of claim 17 with another food ingredient.