Expression of nitrogenase polypeptides in plant cells

JP2026026077A5Pending Publication Date: 2026-05-27COMMONWEALTH SCI & IND RES ORG

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
COMMONWEALTH SCI & IND RES ORG
Filing Date
2025-10-10
Publication Date
2026-05-27

Smart Images

  • Figure 00000260_0000
    Figure 00000260_0000
  • Figure 00000260_0001
    Figure 00000260_0001
  • Figure 00000260_0002
    Figure 00000260_0002
Patent Text Reader

Abstract

The present invention provides methods and means for producing nitrogenase polypeptides in the mitochondria of plant cells. The present invention provides methods and means for producing nitrogenase polypeptides in the mitochondria of plant cells.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to methods and means for producing nitrogenase polypeptides in the mitochondria of plant cells. [Background technology]

[0002] Nitrogen-fixing bacteria produce ammonia from N2 gas through biological nitrogen fixation (BNF), catalyzed by the enzyme complex nitrogenase. However, modern agricultural demands exceed this source of fixed nitrogen, resulting in the widespread use of industrially produced nitrogenous fertilizers in agriculture (Smil, 2002). However, both fertilizer production and application are sources of pollution (Good and Beatty, 2011) and are considered unsustainable (Rockstrom et al., 2009). The vast majority of fertilizers applied worldwide are not taken up by crops (Cui et al., 2013; de Bruijn, 2015), leading to fertilizer runoff, weed growth, and eutrophication of rivers and streams (Good and Beatty, 2011). The resulting algal blooms reduce oxygen levels, causing environmental damage locally and across offshore reefs (De'ath et al., 2012; Glibert et al., 2014; Sutton et al., 2008). Furthermore, excessive fertilization is a problem in many developed countries, while its availability limits crop yields in some regions (Mueller et al., 2012). Fertilizer production itself requires significant energy inputs, costing an estimated US$100 billion per year.

[0003] Clearly, strategies to reduce dependence on industrially produced nitrogen are needed. To that end, the idea of ​​engineering plants capable of biological nitrogen fixation has long attracted considerable interest (Merrick and Dixon, 1984) and has been the focus of recent reports (de Bruijn, 2015; Oldroyd and Dixon, 2014). Potential approaches include i) extending nitrogen-fixing bacterial symbiosis from legumes to cereals (Santi et al., 2013), ii) modifying endosymbiotic microorganisms to enable nitrogen fixation (Geddes et al., 2015), and iii) genetically engineering nitrogenase into plant cells (Curatti and Rubio, 2014). All of these approaches are technically challenging and therefore ambitious and speculative.

[0004] It has been widely reported that nitrogenase, the enzyme complex that enables biological nitrogen fixation in nitrogen-fixing bacteria, requires a multigene assembly pathway for its biosynthesis and function (Hu and Ribbe, 2013; Rubio and Ludden, 2008; Seefeldt et al., 2009). Components of a canonical iron-molybdenum nitrogenase include catalytic proteins named NifD and NifK and the electron donor NifH. Approximately 12 other proteins are involved in the assembly of nitrogenase in nitrogen-fixing bacteria, including NifM, NifS, NifU, NifE, NifN, NifX, NifV, NifJ, NifY, NifF, NifZ, and NifQ, among others, in complex maturation, scaffolding, and cofactor insertion. Genetic lesions, complementation assays between nitrogen-fixing bacteria and non-nitrogen-fixing prokaryotes, and phylogenetic analyses (Dos Santos et al., 2012; Temme et al., 2012; Wang et al., 2013) have led to a subset of Nif proteins (NifD, NifK, NifB, NifE, and NifN) that are considered core components, while others are considered accessory because they are thought to be required for optimal activity. Specific biochemical conditions are also required for nitrogenase assembly and function. First among these is that nitrogenase is extremely oxygen-sensitive (Robson and Postgate, 1980). Furthermore, large amounts of ATP, reducing agents, readily available Fe, Mo, S-adenosylmethionine, and homocitrate are required for the biogenesis and function of the metalloprotein catalytic center (Hu and Ribbe, 2013; Rubio and Ludden, 2008). All of these factors contribute to the technical difficulty of producing a functional nitrogenase complex in plant cells. Summary of the Invention

[0005] In light of the difficulties observed in producing functional NifD in plant cells, the inventors of the present invention determined that it would be important to express a NifD in plant cells that is resistant to secondary cleavage / degradation.

[0006] Thus, in one aspect, the present invention provides a plant cell comprising an exogenous polynucleotide encoding a NifD polypeptide (ND) that is resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97 to 100 of SEQ ID NO:18.

[0007] In a related aspect, the invention provides a plant cell comprising an exogenous polynucleotide encoding a NifD polypeptide (ND) comprising an amino acid sequence other than RRNY (amino acid sequence 101) at positions corresponding to amino acids 97 to 100 of SEQ ID NO:18.

[0008] In a preferred embodiment, the ND is more resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO: 18 than a corresponding ND having an amino acid sequence other than RRNY (amino acid sequence 101) at a position corresponding to amino acids 97-100 of SEQ ID NO: 18.

[0009] In one embodiment of the above aspect, the ND comprises a mitochondrial targeting peptide (MTP), preferably the MTP is at the N-terminus of the ND.

[0010] In a further embodiment, the ND can be cleaved within or immediately after the MTP, such that when the foreign polynucleotide is expressed in a plant cell, a processed NifD polypeptide (CND) is produced, whereby the CND either contains an amino acid sequence from the C-terminal amino acid of the MTP (the scar sequence) at its N-terminus, or does not contain the scar sequence.

[0011] In a preferred embodiment, MTP is cleaved with at least 50% efficiency in the plant cell and / or CND is present in the plant cell at a higher level than ND, preferably at a ratio of greater than 2:1, more preferably greater than 3:1, or greater than 4:1.

[0012] In a preferred embodiment, the CND has NifD function.

[0013] In a further or alternative embodiment of the above aspect, the exogenous polynucleotide encodes ND, a fusion polypeptide (NifD-linker-NifK fusion polypeptide) comprising, in order, the amino acid sequence of NifD, a linker amino acid sequence (linker), and a NifK polypeptide (NK) amino acid sequence. The linker amino acid sequence is 8 to 50 residues in length, preferably about 30 residues, and is translationally fused to ND and NK. In a preferred embodiment, the ND further comprises a mitochondrial targeting peptide (MTP), which is translationally fused to the N-terminus of the NifD amino acid sequence. In a highly preferred embodiment, the ND can be cleaved within or immediately after the MTP, such that, when the exogenous polynucleotide is expressed in a plant cell, a processed NifD polypeptide (CND) is generated, whereby the CND either contains or does not contain the scar sequence at its N-terminus.

[0014] In one embodiment of the above aspect, the ND or CND has NifD function, or the ND (NifD-linker-NifK polypeptide) has both NifD and NifK functions. In one embodiment, the NifD polypeptide is an AnfD polypeptide and the NifK polypeptide is an AnfK polypeptide.

[0015] In one embodiment of the above aspect, the MTP comprises any MTP disclosed herein, for example, the MTP comprises about 51 amino acids in length from the F1-ATPase γ-subunit.

[0016] In one embodiment, the CND comprises a scar sequence translationally fused to the N-terminus of the NifD amino acid sequence, the scar sequence being 1 to 45 amino acids in length, preferably 1 to 20 amino acids in length, more preferably 1 to 10 or 11 to 20 amino acids in length.

[0017] In a further or alternative embodiment, one or both of the ND and CND, eg the NifD-linker-NifK polypeptide, is within the mitochondria of the plant cell, preferably within the mitochondrial matrix (MM) of the plant cell.

[0018] In a further embodiment, one or both of the ND and CND, e.g., the NifD-linker-NifK polypeptide, are predominantly soluble in plant mitochondria. Preferably, at least 60% or at least 75% of the CND in plant mitochondria is soluble. The degree of solubility is preferably determined as described in the Examples.

[0019] In a further or alternative embodiment, the ND, eg, the NifD-linker-NifK polypeptide, comprises an amino acid other than tyrosine (Y) at the position corresponding to amino acid 100 of SEQ ID NO:18.

[0020] In one embodiment, the ND, e.g., the NifD-linker-NifK polypeptide, comprises a glutamine (Q) or lysine (K) at the position corresponding to amino acid 100 of SEQ ID NO: 18, or a leucine (L) or methionine (M) or phenylalanine (F) at the position corresponding to amino acid 100 of SEQ ID NO: 18.

[0021] In another embodiment, the ND comprises a Q, K, L, or M at a position corresponding to amino acid 100 of SEQ ID NO:18.

[0022] In another embodiment, the ND comprises an L or M at a position corresponding to amino acid 100 of SEQ ID NO:18.

[0023] In another embodiment, the ND comprises a Q, K, or L at a position corresponding to amino acid 100 of SEQ ID NO:18.

[0024] In another embodiment, the ND comprises a Q, K, or M at a position corresponding to amino acid 100 of SEQ ID NO:18.

[0025] In another embodiment, the ND comprises a Q, K, or F at a position corresponding to amino acid 100 of SEQ ID NO:18.

[0026] In a further or alternative embodiment, the ND, e.g., the NifD-linker-NifK polypeptide, comprises the sequence RRNX (SEQ ID NO: 154) at positions corresponding to amino acids 97-100 of SEQ ID NO: 18, where X is any amino acid other than Y.

[0027] In one embodiment, X is Q or K; L, M, or F; L or M; Q, K, or L; Q, K, or M; Q, K, or F.

[0028] In a further embodiment, the plant cell comprises one or more exogenous polynucleotides, preferably 2 to 8 exogenous polynucleotides, encoding one or more Nif fusion polypeptides (NFs) other than ND, each NF comprising (i) an MTP at the N-terminus of the NF and (ii) a Nif polypeptide sequence (NP), wherein each MTP is independently the same or different, and each NP is independently the same or different.

[0029] In one embodiment, each NF can be cleaved within or immediately after the MTP, and when one or more exogenous polynucleotides are expressed in a plant cell, processed Nif polypeptides (CNFs) are produced, whereby each CNF either contains a scar sequence or does not contain a scar sequence at its N-terminus.

[0030] In one embodiment, at least one of the NF polypeptide sequences is a NifK polypeptide or a NifH polypeptide, or both a NifK polypeptide and a NifH polypeptide.

[0031] In a further or alternative embodiment, the plant cell comprises an NK amino acid sequence, and the C-terminus of this polypeptide is the C-terminus of wild-type NifK, i.e., NK lacks any artificially added C-terminal extension.

[0032] In a further or alternative embodiment of the above aspect, the exogenous polynucleotide encodes a NifE-linker-NifN fusion polypeptide (NifE-linker-NifN) comprising, in order, a NifE amino acid sequence (NE), a linker amino acid sequence (linker), and a NifN polypeptide (NN) amino acid sequence. The linker amino acid sequence is 20 to 70 residues in length, preferably about 46 residues, and is translationally fused to NE and NN. In a preferred embodiment, the NifE-linker-NifN polypeptide comprises a mitochondrial targeting peptide (MTP) that is translationally fused to the N-terminus of the NE amino acid sequence. In a highly preferred embodiment, the NifE-linker-NifN polypeptide can be cleaved within or immediately after the MTP, such that, when the exogenous polynucleotide is expressed in a plant cell, a processed NifD polypeptide (CNE) is generated, whereby the CNE either contains or does not contain a scar sequence at its N-terminus.

[0033] In a further or alternative embodiment, the linker of the NifE-linker-NifN polypeptide is at least about 30 amino acids in length, or at least about 40 amino acids, or from about 20 amino acids to about 60 amino acids, or from about 30 amino acids to about 70 amino acids, or from about 30 amino acids to about 60 amino acids, or from about 30 amino acids to about 50 amino acids, or about 25 amino acids, or about 30 amino acids, or about 35 amino acids, or about 40 amino acids, or about 45 amino acids, or about 46 amino acids, or about 50 amino acids, or about 55 amino acids. Most preferably, the linker is about 30 amino acids in length for NifD-linker-NifK fusion polypeptides and about 46 amino acids in length for NifE-linker-NifN fusion polypeptides. In this context, "about 30" means 27, 28, 29, 30, 31, 32, or 33 amino acids, and "about 46" means 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or 51 amino acids.

[0034] In a further or alternative embodiment, the linker is long enough to allow the ND and NK, or the NE and NN, to associate into a functional configuration in a plant or bacterial cell. In one embodiment, the linker is 8 to 50 amino acids in length. Preferably, the linker is at least about 20 amino acids, at least about 25 amino acids, or at least about 30 amino acids in length. More preferably, the linker is 25 to 35 amino acids in length for a NifD-linker-NifK fusion polypeptide.

[0035] In a further or alternative embodiment, the fusion polypeptide can be cleaved within its MTP or immediately after the MTP, such that when the exogenous polynucleotide is expressed in a plant cell, a processed polypeptide (CNK) is produced, whereby the CNK comprises, in order, an optional scar sequence, a NifD amino acid sequence, a linker amino acid sequence, and an NK amino acid sequence. If cleavage occurs immediately after the MTP, the scar sequence is not present.

[0036] In one embodiment, the plant cell comprises the fusion polypeptide, the CDK, or both.

[0037] In a further or alternative embodiment, the CDK comprises a scar sequence translationally fused to the N-terminus of the NifD amino acid sequence, the scar sequence being 1 to 45 amino acids in length, preferably 1 to 20 amino acids in length, more preferably 1 to 10 or 11 to 20 amino acids in length.

[0038] In a further or alternative embodiment, the CDK has both the functions of NifD and NifK.

[0039] In a further or alternative embodiment, the plant cell further comprises an exogenous polynucleotide encoding one or more Nif polypeptides (NF) other than ND and NK, each NF comprising (i) an MTP at the N-terminus of the NF and (ii) a Nif polypeptide sequence (NP), wherein each MTP is independently the same or different, and each NP is independently the same or different.

[0040] In a further or alternative embodiment, each NF can be cleaved within its MTP or immediately after the MTP, and when the exogenous polynucleotide is expressed in a plant cell, a processed Nif polypeptide (CNF) is produced, whereby each CNF either contains a scar sequence or does not contain a scar sequence at its N-terminus.

[0041] In one embodiment, at least one of the NF polypeptides is a NifH polypeptide.

[0042] In one embodiment of any of the above aspects, the plant cell comprises (i) an exogenous polynucleotide encoding Nif polypeptides, including NifD, NifH, NifK, NifB, NifE, and NifN polypeptides, preferably within the mitochondrial matrix of the plant cell.

[0043] In a further or alternative embodiment of any of the above aspects, each MTP comprises at least 10 amino acids and preferably has a length of 10-80 amino acids.

[0044] In a further or alternative embodiment of any of the above aspects, each MTP, or at least one MTP, or all of the MTPs, independently comprise an MTP of a mitochondrial protein precursor or a variant thereof, preferably a plant MTP.

[0045] In a further or alternative embodiment of any of the above aspects, one or more or all of the exogenous polynucleotides are integrated into the nuclear genome of the cell, preferably as a contiguous nucleic acid sequence, and / or are expressed in the nucleus of the cell.

[0046] In one embodiment of any of the above aspects, the cell is other than an Arabidopsis thaliana protoplast or a cell other than a Nicotiana benthamiana cell.

[0047] The inventors of the present invention have also generated plant cells that produce combinations of Nif polypeptides that are at least partially soluble in plant mitochondria.

[0048] Thus, in one aspect, the invention provides a plant cell comprising a mitochondrion and at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides, wherein the Nif polypeptides are selected from the group consisting of NifF, NifM, NifN, NifS, NifU, NifW, NifY, NifZ, NifV, NifH, and NifD-NifK, and wherein each of the at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides is at least partially soluble in the mitochondrion.

[0049] In one embodiment, the plant cell comprises a NifV polypeptide. Preferably, the NifV produces homocitric acid. More preferably, the NifV polypeptide is at least partially soluble in the mitochondria of the plant cell. In one embodiment, the NifV polypeptide is a NifV of the present invention.

[0050] In another embodiment, the plant cell comprises at least a NifS polypeptide or a NifU polypeptide, or both a NifS polypeptide and a NifU polypeptide, and optionally a NifV polypeptide.

[0051] In another embodiment, the plant cell comprises at least a NifH polypeptide or a NifM polypeptide, or both a NifH polypeptide and a NifM polypeptide, and optionally one or more or all of NifV, NifS, and NifU.

[0052] In another embodiment, the plant cell comprises NifF, NifH, or NifD-NifK polypeptides, or comprises NifH and NifD-NifK, or comprises NifF, NifH, and NifD-NifK, and optionally one or more or all of NifV, NifS, NifU, NifH, and NifM polypeptides.

[0053] In one embodiment, the NifD polypeptide is an AnfD polypeptide, the NifH polypeptide is an AnfH polypeptide, and the NifD-NifK polypeptide is an AnfD-AnfK polypeptide. In a preferred embodiment, the plant cell further comprises an AnfG polypeptide that is at least partially soluble in the mitochondria.

[0054] In one embodiment, each of the at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides after cleavage by MPP is independently at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% soluble in mitochondria. The Nif polypeptides are up to 80%, or up to 90%, or even completely soluble in the mitochondria of plant cells.

[0055] In one embodiment, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven of the Nif polypeptides each independently comprise a mitochondrial targeting peptide (MTP), or a C-terminal peptide resulting from cleavage of an MTP, or a combination of both MPP-processed and unprocessed forms is present, preferably an MTP is at the N-terminus of each of at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides, or the MPP-processed forms do not have a C-terminal peptide at the N-terminus of the Nif polypeptide.

[0056] In one embodiment, each MTP is independently cleaved with at least 50% efficiency in the plant cell, and / or each of the at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven processed Nif polypeptides is independently present in the plant cell at a higher level than the corresponding Nif polypeptide, preferably in a ratio of greater than 1:1, greater than 2:1, greater than 3:1, or greater than 4:1.

[0057] In one embodiment, the plant cell comprises a NifD-linker-NifK fusion polypeptide comprising, in order, a NifD amino acid sequence (ND), a linker amino acid sequence, and a NifK polypeptide (NK) amino acid sequence, wherein the linker amino acid sequence has a length of 8 to 50 residues, preferably 16 to 50 residues, more preferably about 26 or about 30 residues, or most preferably 26 or 30 residues, and is translationally fused to ND and NK.

[0058] In a further embodiment, the NifD-linker-NifK fusion polypeptide comprises a mitochondrial targeting peptide (MTP), or a C-terminal peptide resulting from cleavage of MTP, or a combination of both the processed and unprocessed forms by MTP, wherein MTP is translationally fused to the N-terminus of the NifD-NifK fusion polypeptide.

[0059] In one embodiment, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven processed Nif polypeptides each independently comprise a C-terminal peptide resulting from cleavage of MTP and consisting of 1 to 45 amino acids in length, preferably 1 to 20 amino acids in length, more preferably 1 to 10 or 11 to 20 amino acids in length, translationally fused to the N-terminus of the Nif polypeptide.

[0060] In one embodiment, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides, or at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven processed Nif polypeptides, are functional Nif polypeptides.

[0061] In one embodiment, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides, or preferably at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven processed Nif polypeptides, are within the mitochondria of the plant cell, preferably within the mitochondrial matrix (MM) of the plant cell.

[0062] In one embodiment, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides, or preferably at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven processed Nif polypeptides, or both, are independently predominantly soluble in plant mitochondria (i.e., greater than 50% soluble in mitochondria). The processed Nif polypeptides are preferably up to 80%, or even 90%, or even completely soluble in the mitochondria of plant cells. The solubility of the polypeptides can be determined as described herein.

[0063] In one embodiment, the NifD fusion polypeptide, or NifD-linker-NifK fusion polypeptide, or MPP cleavage product thereof, is present in a plant cell and (a) is resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO: 18, and / or (b) comprises an amino acid sequence other than RRNY (amino acid sequence 101) at a position corresponding to amino acids 97-100 of SEQ ID NO: 18. In one embodiment, the ND comprises an amino acid other than tyrosine (Y) at a position corresponding to amino acid 100 of SEQ ID NO: 18. In one embodiment, the ND comprises glutamine (Q) or lysine (K) at a position corresponding to amino acid 100 of SEQ ID NO: 18, or leucine (L) or methionine (M) or phenylalanine (F) at a position corresponding to amino acid 100 of SEQ ID NO: 18.

[0064] In one embodiment, the MTP is about 51 amino acids in length from the F1-ATPase gamma subunit.

[0065] In one embodiment, the plant cell comprises an NK amino acid sequence, the C-terminus of the polypeptide being the C-terminus of wild-type NifK.

[0066] In one embodiment, the linker is at least about 20 amino acids in length, or at least about 30 amino acids, or at least about 40 amino acids, or from about 20 amino acids to about 70 amino acids, or from about 30 amino acids to about 70 amino acids, or from about 30 amino acids to about 60 amino acids, or from about 30 amino acids to about 50 amino acids, or about 25 amino acids, or about 30 amino acids, or about 35 amino acids, or about 40 amino acids, or about 45 amino acids, or about 46 amino acids, or about 50 amino acids, or about 55 amino acids.

[0067] In one embodiment, the NifD-linker-NifK fusion polypeptide can be cleaved within or immediately after the MTP to generate a processed polypeptide (CNK), whereby CNK comprises, in order, the optimal C-terminal peptide resulting from cleavage of the MTP, the NifD amino acid sequence (ND), the linker amino acid sequence, and the NK amino acid sequence.

[0068] In one embodiment, the plant cell further comprises a fusion polypeptide or a CDK, or both.

[0069] In one embodiment, the CDK comprises a scar sequence translationally fused to the N-terminus of the NifD amino acid sequence and having a length of 1 to 45 amino acids, preferably 1 to 20 amino acids, more preferably 1 to 10 or 11 to 20 amino acids.

[0070] In one embodiment, the CDK has the functions of both NifD and NifK.

[0071] In one embodiment, ND is AnfD and NK is AnfK.

[0072] In one embodiment, the MTP is about 51 amino acids in length from the F1-ATPase gamma subunit.

[0073] In one embodiment, each MTP comprises at least 10 amino acids, and preferably has a length of 10-80 amino acids.

[0074] In one embodiment, the MTP, or at least one MTP, or all of the MTPs independently comprise an MTP of a mitochondrial protein precursor or a variant thereof, preferably a plant MTP.

[0075] In one embodiment, the at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides are encoded by at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven exogenous polynucleotides, of which at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven are integrated into the nuclear genome of the cell, preferably as a contiguous nucleic acid sequence, and / or are expressed in the nucleus of the plant cell.

[0076] In another embodiment of any of the above aspects, the cell is other than an Arabidopsis thaliana protoplast or other than a Nicotiana benthamiana cell.

[0077] The present inventors have also succeeded in expressing in plant mitochondria the combination of Nif polypeptides required for the minimal nitrogenase complex.

[0078] Thus, in another aspect, the present invention provides a plant cell comprising mitochondria and exogenous polynucleotides encoding at least eight or at least nine Nif fusion polypeptides, each of the exogenous polynucleotides comprising a promoter operably linked to a nucleotide sequence encoding one of the Nif fusion polypeptides and directing expression of that nucleotide sequence in the plant cell, each Nif fusion polypeptide independently comprising a mitochondrial targeting peptide (MTP), the Nif fusion polypeptides comprising: (i) NifH, NifB, NifF, NifJ, NifS, NifU, and NifV fusion polypeptides; (ii) NifD fusion polypeptides and NifK fusion polypeptides; or (iii) NifD-linker-NifK fusion polypeptides comprising a NifD sequence having a C-terminus, an oligopeptide linker, and a NifK sequence having an N-terminus, wherein the oligopeptide linker and (iii) the MPP cleavage product of the NifD-linker-NifK fusion polypeptide, when present in the plant cell, is at least partially soluble in the mitochondria of the plant cell; and wherein the NifV fusion polypeptide and / or its MPP cleavage product produces homocitric acid in the plant cell and is at least partially soluble in the mitochondria of the plant cell.

[0079] In another aspect, the invention provides a plant cell comprising mitochondria and exogenous polynucleotides encoding at least two, at least three, at least four, at least five, or at least six Nif fusion polypeptides, each of the exogenous polynucleotides comprising a promoter operably linked to a nucleotide sequence encoding one of the Nif fusion polypeptides and directing expression of that nucleotide sequence in the plant cell, each Nif fusion polypeptide independently comprising a mitochondrial targeting peptide (MTP), the Nif fusion polypeptides comprising one, two or more, or all of: (i) NifW, NifX, NifY, and NifZ fusion polypeptides; and either (ii) a NifD fusion polypeptide and a NifK fusion polypeptide; or (iii) a NifD-linker-NifK fusion polypeptide comprising a NifD sequence having a C-terminus, an oligopeptide linker, and a NifK sequence having an N-terminus, wherein the oligopeptide linker is translationally fused to the C-terminus of the NifD sequence and the N-terminus of the NifK sequence, and wherein the NifD-linker-NifK fusion polypeptide comprises at least NifW, NifX, NifY, and NifZ fusion polypeptides. The mitochondrial processing protease (MPP) cleavage products of (ii) the NifD fusion polypeptide and NifK fusion polypeptide are each at least partially soluble in the mitochondria of the plant cell, and either the MPP cleavage products of (ii) the NifD fusion polypeptide and NifK fusion polypeptide are at least partially soluble in the mitochondria of the plant cell when present in the plant cell, or the MPP cleavage products of (iii) the NifD-linker-NifK fusion polypeptide are at least partially soluble in the mitochondria of the plant cell when present in the plant cell, wherein the MPP cleavage products of the NifD fusion polypeptide and NifK fusion polypeptide of (ii) or the MPP cleavage products of the NifD-linker-NifK fusion polypeptide of (iii) are present in the plant cell in an amount that is greater than the amount of the MPP cleavage products of the NifD fusion polypeptide and NifK fusion polypeptide that is present in a corresponding plant cell lacking an exogenous polynucleotide encoding one, two or more, or all of the NifW, NifX, NifY, and NifZ fusion polypeptides of (i).

[0080] In another aspect, the invention provides a plant cell comprising a mitochondrion and exogenous polynucleotides encoding at least five, at least six, at least seven, at least eight, or at least nine Nif fusion polypeptides, each of the exogenous polynucleotides comprising a promoter operably linked to a nucleotide sequence encoding one of the Nif fusion polypeptides and directing expression of that nucleotide sequence in the plant cell, each Nif fusion polypeptide independently comprising a mitochondrial targeting peptide (MTP), the Nif fusion polypeptides comprising one or more of: (i) NifH, NifS, and NifU fusion polypeptides, and optionally a MifM polypeptide; (ii) one or more or all of NifW, NifX, NifY, and NifZ fusion polypeptides; (iii) a NifD fusion polypeptide and a NifK fusion polypeptide; or (iv) a NifD-linker-NifK fusion polypeptide comprising a NifD sequence having a C-terminus, an oligopeptide linker, and a NifK sequence having an N-terminus, wherein the oligopeptide linker (iii) a mitochondrial processing protease (MPP) cleavage product of the NifS fusion polypeptide and the NifU fusion polypeptide is at least partially soluble in the mitochondria of the plant cell; (iv) a NifD-linker-NifK fusion polypeptide is at least partially soluble in the mitochondria of the plant cell when present in the plant cell; and (v) a MPP cleavage product of the NifD fusion polypeptide and the NifK fusion polypeptide is at least partially soluble in the mitochondria of the plant cell when present in the plant cell; and (vi) a MPP cleavage product of the NifD-linker-NifK fusion polypeptide is at least partially soluble in the mitochondria of the plant cell when present in the plant cell; and (vi) a MPP cleavage product of the NifD fusion polypeptide and the NifK fusion polypeptide is present in the plant cell as a complex with a P cluster.

[0081] In one embodiment, the plant cell comprises a NifH fusion polypeptide that is an AnfH fusion polypeptide, the NifD fusion polypeptide, if present, is an AnfD fusion polypeptide, the NifK fusion polypeptide, if present, is an AnfK fusion polypeptide, and the NifD-linker-NifK fusion polypeptide, if present, is an AnfD-linker-AnfK fusion polypeptide; the plant cell comprises an exogenous polynucleotide encoding an AnfG fusion polypeptide comprising an MTP, the exogenous polynucleotide encoding the AnfG fusion polypeptide comprising a promoter operably linked to a nucleotide sequence encoding the AnfG fusion polypeptide and causing expression of the nucleotide sequence in the plant cell; and an MPP cleavage product of the AnfG fusion polypeptide is at least partially soluble in the mitochondria of the plant cell.

[0082] In one embodiment of the above three aspects, the NifD fusion polypeptide or NifD-linker-NifK fusion polypeptide is present in a plant cell and (a) is resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO: 18, and / or (b) comprises an amino acid sequence other than RRNY (amino acid sequence 101) at a position corresponding to amino acids 97-100 of SEQ ID NO: 18.

[0083] In another aspect, the present invention provides a plant cell comprising a mitochondrion and exogenous polynucleotides encoding at least two, at least three, or at least four Anf fusion polypeptides, each of the exogenous polynucleotides comprising a promoter operably linked to a nucleotide sequence encoding one of the Anf fusion polypeptides and directing expression of the nucleotide sequence in the plant cell, each Anf fusion polypeptide independently comprising a mitochondrial targeting peptide (MTP), the Anf fusion polypeptides comprising either (i) an AnfG fusion polypeptide, or an AnfG and AnfH fusion polypeptide, and (ii) an AnfD fusion polypeptide and an AnfK fusion polypeptide, or (iii) an AnfD-linker-AnfK fusion polypeptide comprising an AnfD sequence having a C-terminus, an oligopeptide linker, and an AnfK sequence having an N-terminus, the oligopeptide linker being at the C-terminus of the AnfD sequence and the AnfK sequence. A plant cell is provided in which (ii) the AnfD and AnfK fusion polypeptides are translatably fused to the N-terminus thereof, and the mitochondrial processing protease (MPP) cleavage products of at least the AnfG and AnfH fusion polypeptides are at least partially soluble in the mitochondria of the plant cell when present in the plant cell, and either (i) the MPP cleavage products of the AnfD and AnfK fusion polypeptides are at least partially soluble in the mitochondria of the plant cell when present in the plant cell, or (iii) the MPP cleavage products of the AnfD-linker-AnfK fusion polypeptide are at least partially soluble in the mitochondria of the plant cell when present in the plant cell, and the MPP cleavage products of the AnfD fusion polypeptide and AnfK fusion polypeptide of (ii) or the MPP cleavage product of the AnfD-linker-AnfK fusion polypeptide of (iii) form a protein complex in the plant cell with the MPP cleavage product of the AnfG fusion polypeptide when present in the plant cell.

[0084] In some embodiments, the plant cell further comprises one or more exogenous polynucleotides encoding one or more Nif fusion polypeptides as defined herein.

[0085] As will be appreciated by those skilled in the art, the embodiments of Nif polypeptides provided herein apply equally and specifically to the corresponding Nif polypeptides that are Anf polypeptides. For example, the NifD, NifK, and NifH polypeptide embodiments described herein according to one aspect of the invention apply equally and specifically to AnfD, AnfK, and AnfH polypeptides, respectively.

[0086] The present inventors are, to their knowledge, the first to generate plant cells comprising a NifV polypeptide that is at least partially soluble in mitochondria. Thus, in another aspect, the present invention provides a plant cell comprising a NifV polypeptide (NV), wherein the NV is at least partially soluble, preferably at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or completely soluble, in the mitochondria of the plant cell, preferably in the MM of the plant cell.

[0087] In one embodiment, the NV is capable of producing or is producing homocitrate in the cell.

[0088] In one embodiment, the NV polypeptide comprises an amino acid sequence provided as any one of SEQ ID NOs: 163, 206-209, 211, or 212, or a biologically active fragment thereof, or has an amino acid sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to the amino acid sequence provided as any one or more of SEQ ID NOs: 163, 206-209, 211, or 212, and is capable of producing homocitric acid in a cell.

[0089] In one embodiment of this aspect, the invention provides a plant cell comprising a mitochondrion and an exogenous polynucleotide encoding a NifV polypeptide (NV), wherein the exogenous polynucleotide comprises a promoter operably linked to a nucleotide sequence encoding the NV and causing expression of the nucleotide sequence in the plant cell, wherein the NV produces homocitric acid in the plant cell and is at least partially soluble in the mitochondria of the plant cell, wherein the exogenous polynucleotide is preferably integrated into the nuclear genome of the plant cell and / or expressed in the nucleus of the plant cell, and optionally the NV comprises a mitochondrial targeting peptide (MTP).

[0090] In another aspect, the present invention provides a plant cell comprising an exogenous polynucleotide encoding a NifD polypeptide that is (a) resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO: 18, and / or (b) comprises an amino acid sequence other than RRNY (amino acid sequence 101) at a position corresponding to amino acids 97-100 of SEQ ID NO: 18, wherein the exogenous polynucleotide comprises a promoter operably linked to a nucleotide sequence encoding ND and causing expression of the nucleotide sequence in the plant cell, and wherein the NifD polypeptide preferably comprises MTP.

[0091] In some embodiments, the plant cell further comprises one or more exogenous polynucleotides encoding one or more, or all, of the Nif fusion polypeptides defined herein that are present in the cell and / or cleavage products of the Nif fusion polypeptides that are present in the cell. Preferably, the plant cell comprises an exogenous polynucleotide for each Nif fusion polypeptide and / or cleavage product that is present in the cell.

[0092] In one embodiment, the plant cell comprises an exogenous polynucleotide encoding a NifK polypeptide (NK), wherein the exogenous polynucleotide encoding the NK comprises a promoter operably linked to a nucleotide sequence encoding the NK such that the nucleotide sequence is expressed in the plant cell; the ND has a C-terminus; the NK has an N-terminus; and either (i) the NK has a mitochondrial targeting peptide (MTP), or (ii) the ND and NK are translationally fused as a NifD-linker-NifK fusion polypeptide comprising an oligopeptide linker, in which the oligopeptide linker is translationally fused to the C-terminus of the ND and the N-terminus of the NK.

[0093] In one embodiment, the plant cell comprises an exogenous polynucleotide encoding a NifH polypeptide (NH), wherein the exogenous polynucleotide encoding the NH comprises a promoter operably linked to a nucleotide sequence encoding the NH such that the nucleotide sequence is expressed in the plant cell, and wherein the NH comprises a mitochondrial targeting peptide (MTP), and preferably the NH and / or its MPP cleavage product is at least partially soluble in the mitochondria of the plant cell.

[0094] In one embodiment, at least one or more, and preferably all, MPP cleavage products of the Nif fusion polypeptides are at least partially soluble in the mitochondria of the plant cell, and preferably the MPP cleavage products of each of the NifD, NifK, and NifD-linker-NifK fusion polypeptides (when present in the plant cell) and the NifH polypeptide are at least partially soluble in the mitochondria of the plant cell.

[0095] The present inventors are, to their knowledge, the first to generate plant cells comprising a NifH polypeptide that is at least partially soluble in mitochondria. Thus, in another aspect, the present invention provides a plant cell comprising a NifH polypeptide (NH), wherein the NH is at least partially soluble in mitochondria.

[0096] In one embodiment, the NH is encoded by an exogenous polynucleotide which, when present in a plant cell, is preferably integrated into the nuclear genome of the cell as a nucleic acid sequence contiguous with the exogenous polynucleotides encoding NifD, NifK, and the NifD-linker-NifK fusion polypeptide.

[0097] In another aspect, the invention provides a plant cell comprising an exogenous polynucleotide encoding a NifH fusion polypeptide (NH), wherein the exogenous polynucleotide comprises a promoter operably linked to a nucleotide sequence encoding the NH such that the nucleotide sequence is expressed in the plant cell, the NH comprising a mitochondrial targeting peptide (MTP), and an MPP cleavage product of the NH is at least partially soluble in the mitochondria of the plant cell, and optionally the exogenous polynucleotide is integrated into the nuclear genome of the plant cell and / or expressed in the nucleus of the plant cell.

[0098] In some embodiments, the plant cell further comprises one or more exogenous polynucleotides encoding one or more Nif fusion polypeptides as defined herein that are present in the cell and / or cleavage products of the Nif fusion polypeptides that are present in the cell. Preferably, the plant cell comprises an exogenous polynucleotide for each Nif fusion polypeptide and / or cleavage product that is present in the cell.

[0099] In embodiments of each of the above aspects, the plant cell further comprises an exogenous polynucleotide encoding a NifM polypeptide (NM), wherein the exogenous polynucleotide encoding NM comprises a promoter operably linked to a nucleotide sequence encoding NM such that the nucleotide sequence is expressed in the plant cell, and the NM optionally comprises a mitochondrial targeting peptide (MTP).

[0100] In embodiments of each of the above aspects, the plant cell comprises exogenous polynucleotides encoding a NifS fusion polypeptide and a NifU fusion polypeptide, each of which comprises a promoter operably linked to a nucleotide sequence encoding one of the Nif fusion polypeptides and causing expression of the nucleotide sequence in the plant cell, and each of the NifS fusion polypeptide and the NifU fusion polypeptide comprises a mitochondrial targeting peptide (MTP).

[0101] In embodiments of each of the above aspects, each Nif polypeptide is produced in the plant cell as a Nif fusion polypeptide that includes a mitochondrial targeting peptide (MTP), each MTP being independently the same or different, and preferably an MTP being at at least one, or more than one, or all, N-termini of the Nif fusion polypeptide.

[0102] In each embodiment of the above aspects, each Nif fusion polypeptide produced in the plant cell is independently either (i) cleaved by MPP within the MTP sequence to generate an MPP-cleaved Nif polypeptide, such that the MPP-cleaved Nif polypeptide contains the C-terminal peptide (scar peptide) from MTP at its N-terminus, or (ii) cleaved by MPP immediately after MTP, such that the MPP-cleaved Nif polypeptide does not contain the C-terminal peptide from MTP.

[0103] In each embodiment of the above aspects, each MTP is independently cleaved with at least 50% efficiency in the plant cell, and / or each cleaved Nif polypeptide is independently present in the plant cell at a higher level than the corresponding uncleaved Nif fusion polypeptide, preferably in a ratio greater than 1:1, 2:1, or 3:1.

[0104] In embodiments of each of the above aspects, each Nif fusion polypeptide is at least partially cleaved in the plant cell within its MTP sequence to generate MPP-truncated Nif polypeptides, and each MPP-truncated Nif polypeptide independently comprises a peptide (scar peptide) derived from the MTP sequence and having a length of 1 to 45 amino acids, preferably 1 to 20 amino acids, and more preferably 1 to 11 amino acids or 11 to 20 amino acids, translationally fused to the N-terminus of the MPP-truncated Nif polypeptide. In embodiments, one or more of the scar peptides are independently 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in length. In embodiments, one or more of the scar peptides are independently 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids in length, or 20 to 30, 20 to 40, or 20 to 50 amino acids in length, although shorter scar sequences are preferred. In these embodiments, as used herein, a scar peptide comprises an optional linker sequence (such as the Gly-Gly linker used in the Examples herein) fused to the N-terminus of the Nif sequence. In embodiments, the Nif sequence retains the Met (translation initiation Met) from the wild-type sequence at its N-terminus, and this Met is not included in the scar sequence. Alternatively, the translation initiation Met is omitted from the Nif sequence. In embodiments, additional amino acids compared to the wild-type Nif sequence can be excised from the N-terminus of the Nif sequence, provided that the excised Nif sequence retains its Nif function.

[0105] In an embodiment of each of the above aspects, the plant cell further comprises an exogenous polynucleotide encoding a ferredoxin fusion polypeptide, preferably an FdxN fusion polypeptide, wherein the exogenous polynucleotide encoding the ferredoxin fusion polypeptide comprises a promoter operably linked to a nucleotide sequence encoding the ferredoxin fusion polypeptide such that the nucleotide sequence is expressed in the plant cell, and wherein the ferredoxin fusion polypeptide comprises a mitochondrial targeting peptide (MTP).

[0106] In one embodiment, the MPP cleavage product of the ferredoxin fusion polypeptide is at least partially soluble in the mitochondria of the plant cell, and preferably the exogenous polynucleotide is integrated into the nuclear genome of the plant cell and / or expressed in the nucleus of the plant cell.

[0107] In one embodiment, the plant cell comprises a NifD-linker-NifK fusion polypeptide comprising, in order, a NifD amino acid sequence (ND), an oligopeptide linker, and a NifK polypeptide (NK) amino acid sequence, wherein the oligopeptide linker is 8 to 50 residues in length, preferably 16 to 50 residues in length, more preferably about 26 or about 30 residues in length, or most preferably 30 residues in length, translationally fused to the ND and NK.

[0108] In one embodiment, each Nif fusion polypeptide is cleaved in the plant cell to generate a Nif polypeptide that is a functional Nif polypeptide.

[0109] In one embodiment, the plant cell comprises an exogenous polynucleotide encoding a NifD fusion polypeptide (ND) or a NifD-linker-NifK fusion polypeptide, said ND or NifD-linker-NifK fusion polypeptide comprising an amino acid sequence other than RRNY (amino acid sequence 101) at positions corresponding to amino acids 97 to 100 of SEQ ID NO: 18, and said ND or NifD-linker-NifK fusion polypeptide preferably comprising an amino acid other than tyrosine (Y) at positions corresponding to amino acids 97 to 100 of SEQ ID NO: 18.

[0110] In one embodiment, the ND or NifD-linker-NifK fusion polypeptide comprises a glutamine (Q) or lysine (K) at the position corresponding to amino acid 100 of SEQ ID NO: 18, or a leucine (L) or methionine (M) or phenylalanine (F) at the position corresponding to amino acid 100 of SEQ ID NO: 18.

[0111] In one embodiment, the plant cell comprises an exogenous polynucleotide encoding a NifK fusion polypeptide or a NifD-linker-NifK fusion polypeptide, wherein the NifK fusion polypeptide or the NifD-linker-NifK fusion polypeptide has a C-terminal amino acid sequence identical to that of a wild-type NifK polypeptide, in some embodiments, at least the last two, at least the last three, or at least the last four amino acids of the sequence are identical to that of a wild-type NifK polypeptide. Suitable wild-type NifK polypeptide sequences include SEQ ID NO: 3 as well as accession numbers WP_049080161.1, WP_044347163.1, SBM87811.1, WP_047370272.1, WP_014333919.1, WP_012728880.1, WP_011912506.1, WP_065303473.1, WP_018989051.1, prf||2106319A, WP_011021239.1, and the like.

[0112] In one embodiment, the NifK fusion polypeptide or NifD-linker-NifK fusion polypeptide and the MPP cleavage products derived therefrom have an amino acid sequence in which the last four amino acids are the same as the last four amino acids of the wild-type NifK polypeptide.

[0113] In one embodiment, the amino acid sequence of a NifK polypeptide of the invention has at its C-terminus the amino acids DLVR (SEQ ID NO:58). In another embodiment, a NifK polypeptide has at its C-terminus the amino acids DLIR (SEQ ID NO:239), DVVR (SEQ ID NO:240), DIIR (SEQ ID NO:241), DLTR (SEQ ID NO:242), or INVW (SEQ ID NO:243). In one embodiment, an NifK polypeptide has at its C-terminus the amino acids LNVW (SEQ ID NO:244), LNTW (SEQ ID NO:245), LNMW (SEQ ID NO:246), LAMW (SEQ ID NO:247), or LSVW (SEQ ID NO:248).

[0114] In embodiments of the above aspects, the plant cell comprises an exogenous polynucleotide encoding an AnfD-linker-AnfK fusion polypeptide, the AnfD-linker-AnfK fusion polypeptide comprising an AnfD sequence having a C-terminus, an oligopeptide linker, and an AnfK sequence having an N-terminus, wherein the oligopeptide linker is translationally fused to the C-terminus of the AnfD sequence and the N-terminus of the AnfK sequence, and the oligopeptide linker has a length of at least about 20 amino acids, at least about 30 amino acids, at least about 40 amino acids, from about 20 amino acids to about 70 amino acids, from about 30 amino acids to about 70 amino acids, from about 30 amino acids to about 60 amino acids, from about 30 amino acids to about 50 amino acids, about 25 amino acids, about 30 amino acids, about 35 amino acids, about 40 amino acids, about 45 amino acids, about 46 amino acids, about 50 amino acids, or about 55 amino acids. That is, in these embodiments, the NifD sequence of the above embodiments is an AnfD sequence, and the NifK sequence is an AnfK sequence.

[0115] In one embodiment, at least one, or more than one, or preferably all, of the exogenous polynucleotides are integrated into the nuclear genome of the plant cell and / or are expressed in the nucleus of the plant cell.

[0116] In one embodiment, each MTP comprises at least 10 amino acids, and preferably has a length of 10-80 amino acids.

[0117] In one embodiment, at least one of the Nif fusion polypeptides comprises an MTP that is approximately 51 amino acids in length from the F1-ATPase γ-subunit.

[0118] In one embodiment, the MTP, or at least one MTP, or all of the MTPs independently comprise an MTP of a mitochondrial protein precursor or a variant thereof, preferably a plant MTP.

[0119] In one embodiment, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven Nif polypeptides are encoded by at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven exogenous polynucleotides, of which at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or at least eleven are integrated into the nuclear genome of the cell, preferably as contiguous nucleic acid sequences.

[0120] In embodiments of the above aspects, the cells are unable to give rise to progeny cells, e.g., unable to regenerate cell cultures or living plants.

[0121] In one embodiment, the plant cells of the present invention are further defined by one or more of the characteristics mentioned herein. Each possible combination of characteristics is expressly contemplated.

[0122] In a further aspect, the present invention provides a plant or plant part, organ or tissue comprising a plant cell of the invention, preferably a transgenic plant or part thereof, wherein said transgenic plant or part thereof is transgenic for one or more exogenous polynucleotides encoding a Nif polypeptide.

[0123] In one embodiment, the plant part is a seed. In one embodiment, the seed is capable of germinating, or alternatively, has been treated or processed so that it is no longer capable of germinating. The cells of the seed may not be able to be regenerated into cell cultures or living plants.

[0124] In embodiments of the above aspects, one or more of the one or more exogenous polynucleotides are expressed in the roots of the plant, preferably at a higher level in the roots than in the leaves of the plant, using a promoter sequence that provides the desired tissue-specific expression.

[0125] In one embodiment, the transgenic plant has an altered phenotype compared to a corresponding wild-type plant (e.g., increased yield, biomass, growth rate, vigor, nitrogen gain from biological nitrogen fixation, nitrogen use efficiency, abiotic stress tolerance, and / or tolerance to nutrient deficiency compared to a corresponding wild-type plant).

[0126] In an alternative embodiment, the transgenic plant has the same growth rate and / or phenotype as the corresponding wild-type plant.

[0127] In embodiments of the above aspects, the plant cell, plant, or part thereof is a cereal plant cell, cereal plant, or part thereof, such as wheat, rice, corn, triticale, oats, barley, etc., preferably wheat.

[0128] In embodiments of the above aspects, the plant cell, plant, or part thereof is homozygous or heterozygous for one or more exogenous polynucleotides, preferably homozygous for all of the exogenous polynucleotides.

[0129] In embodiments of the above aspects, the plant cell, plant, or part thereof is a monocotyledonous plant cell, monocotyledonous plant, or part thereof (e.g., a cereal plant cell, cereal plant, or part thereof, such as wheat, rice, corn, triticale, oats, barley, etc., preferably wheat), or a dicotyledonous plant cell, dicotyledonous plant, or part thereof.

[0130] In a further or alternative embodiment, the transgenic plants are grown outdoors or plant parts are taken from plants grown outdoors, or the plants are grown in a greenhouse.

[0131] In a further embodiment, the invention provides a population of at least 100 plants according to the invention growing outdoors or in a greenhouse, or plant parts taken therefrom.

[0132] In another aspect, the present invention provides an isolated or recombinant NifD polypeptide (ND) that is resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO:18.

[0133] In a further aspect, the present invention provides an isolated or recombinant NifD polypeptide (ND) comprising an amino acid sequence other than RRNY (amino acid sequence 101) at positions corresponding to amino acids 97 to 100 of SEQ ID NO:18.

[0134] The isolated or recombinant ND can be further defined by any of the above-mentioned characteristics applicable to Nif polypeptides, with all possible combinations of characteristics being considered as part of the present invention.

[0135] In a related aspect, the invention provides a NifD fusion polypeptide comprising a mitochondrial targeting peptide (MTP) translationally fused to a NifD polypeptide (ND), or a cleavage product thereof comprising the ND and, optionally, a scar peptide, wherein the NifD fusion polypeptide or cleavage product thereof (a) is resistant to protease cleavage at a site within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO: 18, and / or (b) comprises an amino acid sequence other than RRNY (amino acid sequence 101) at a position corresponding to amino acids 97-100 of SEQ ID NO: 18.

[0136] In one embodiment, the NifD fusion polypeptide comprises an oligopeptide linker and a NifK polypeptide (NK) translationally fused as a NifD-linker-NifK fusion polypeptide, wherein ND comprises the C-terminus and NK comprises the N-terminus, and the oligopeptide linker is translationally fused to the C-terminus of ND and the N-terminus of NK.

[0137] In another aspect, the invention provides a cleavage product of the NifD fusion polypeptide of the invention, the cleavage product comprising ND, an oligopeptide linker, and NK, wherein the oligopeptide linker is translationally fused to the C-terminus of .

[0138] In one embodiment, the NifD fusion polypeptide or cleavage product thereof is at least partially soluble in the mitochondria of a plant cell when the NifD polypeptide is produced in said plant cell.

[0139] In one embodiment, the NifD fusion polypeptide is an AnfK polypeptide, NK is an AnfK polypeptide, and the NifD-linker-NifK fusion polypeptide is an AnfD-linker-AnfK fusion polypeptide.

[0140] In another aspect, the present invention provides a NifD fusion polypeptide comprising a mitochondrial targeting peptide (MTP) translationally fused to a NifK polypeptide (NK), wherein the NifK polypeptide or a cleavage product thereof, when produced in a plant cell, is at least partially soluble in the mitochondria of the plant cell.

[0141] In another aspect, the present invention provides a cleavage product of the NifD fusion polypeptide of the present invention, comprising NK and optionally a scar peptide, which cleavage product, when produced in a plant cell, is at least partially soluble in the mitochondria of the plant cell.

[0142] In one embodiment, the NK is an AnfK polypeptide (AK).

[0143] In one embodiment, the NifK polypeptide has a C-terminal amino acid sequence that is the same as the C-terminal amino acid sequence of a wild-type NifK polypeptide. Suitable wild-type NifK polypeptide sequences are described herein.

[0144] In another aspect, the present invention provides an AnfD fusion polypeptide comprising a mitochondrial targeting peptide (MTP) and an AnfD polypeptide (AD), or a cleavage product thereof comprising an AD and optionally a scar peptide, preferably wherein the AnfD fusion polypeptide, or the cleavage product thereof, when produced in a plant cell, is at least partially soluble in the mitochondria of the plant cell.

[0145] In another aspect, the present invention provides an AnfH fusion polypeptide comprising a mitochondrial targeting peptide (MTP) and an AnfH polypeptide (AH), or a cleavage product thereof comprising the AH and optionally a scar peptide, wherein preferably the AnfH fusion polypeptide, or the cleavage product thereof, when produced in a plant cell, is at least partially soluble in the mitochondria of the plant cell.

[0146] In another aspect, the present invention provides an AnfG fusion polypeptide comprising a mitochondrial targeting peptide (MTP) and an AnfG polypeptide (AG), or a cleavage product thereof comprising the AG and optionally a scar peptide, wherein preferably the AnfG fusion polypeptide, or the cleavage product thereof, when produced in a plant cell is at least partially soluble in the mitochondria of the plant cell.

[0147] In another aspect, the present invention provides an AnfD-linker-AnfK fusion polypeptide or a cleavage product thereof, comprising a translatably fused AnfD polypeptide (AD), an oligopeptide linker, and (a) a polypeptide (AK), wherein the AD comprises an N-terminus and a C-terminus, the AK comprises an N-terminus, and the oligopeptide linker is translatably fused to the C-terminus of the AD and the N-terminus of the AK, and preferably the fusion polypeptide comprises a mitochondrial targeting peptide (MTP) or the cleavage product comprises a scar peptide translatably fused to the N-terminus of the AD.

[0148] In another aspect, the present invention provides a combination of Anf polypeptides, wherein the Anf is according to an embodiment described herein, preferably a combination of cleavage products of an Anf fusion polypeptide. Preferably, at least one or more, or all, of the cleavage products comprises a scar peptide, for example, fused to the N-terminus of the Anf polypeptide. Preferred combinations are AnfD and AnfK, AnfD, AnfK, and AnfG, AnfD-linker-AnfK and AnfG, or more preferably, AnfD, AnfK, AnfG, and AnfH, or AnfD-linker-AnfK, AnfG, and AnfH polypeptides. In some embodiments, the features of Nif polypeptides described herein apply to the corresponding Anf polypeptides. In preferred embodiments, the combination of Anf polypeptides, preferably cleavage products thereof, is present in a plant cell, transgenic plant, or part thereof, or product derived therefrom, as described herein.

[0149] In another aspect, the present invention provides a protein complex comprising (i) a cleavage product of a NifD fusion polypeptide, preferably an AnfD fusion polypeptide, (ii) a cleavage product of a NifK fusion polypeptide, preferably an AnfK fusion polypeptide, and, optionally, (iii) an Fe-S cluster, preferably a P cluster. Preferably, at least one or more, or all, of the cleavage products comprise a scar peptide, for example, fused at the N-terminus of the Anf polypeptide.

[0150] In another aspect, the present invention provides a protein complex comprising (i) cleavage products of an AnfD fusion polypeptide and an AnfK fusion polypeptide, and optionally, a cleavage product of an AnfG fusion polypeptide, or (ii) cleavage products of an AnfD-linker-AnfK fusion polypeptide and an AnfG fusion polypeptide, and optionally, (iii) an Fe-S cluster, preferably a P-cluster. Preferably, at least one or more, or all, of the cleavage products comprise a scar peptide, e.g., fused to the N-terminus of an Anf polypeptide.

[0151] In one embodiment, the protein complex of the invention is in a plant cell, preferably in the mitochondria of a plant cell or a transgenic plant or part thereof. In one embodiment, the plant cell, the transgenic plant or part thereof comprising an Anf polypeptide, a combination of Anf polypeptides or a protein complex of the invention is used in a method of the invention as described.

[0152] In another aspect, the present invention provides a substantially purified or recombinant NifV polypeptide (NV) that, when expressed in a plant cell, is at least partially soluble in plant mitochondria.

[0153] In a related aspect, the present invention provides an isolated or recombinant NifV polypeptide or NifV fusion polypeptide comprising a mitochondrial targeting peptide (MTP) translationally fused to a NifV polypeptide (NV), or a cleavage product thereof comprising the NV and optionally a scar peptide, wherein the NifV polypeptide and / or NifV fusion polypeptide and / or cleavage product thereof when expressed in a plant cell is at least partially soluble in the plant cell, preferably at least partially soluble in the mitochondria of the plant cell.

[0154] In one embodiment, the isolated or recombinant NifV polypeptide, or NifV fusion polypeptide, or cleavage product thereof, is capable of producing homocitric acid in a plant cell, preferably within the mitochondria of the plant cell.

[0155] In another aspect, the present invention provides a substantially purified or recombinant NifH polypeptide (NH) that is at least partially soluble in plant mitochondria when expressed in plant cells, preferably in transgenic plants.

[0156] In another aspect, the invention provides a NifH fusion polypeptide comprising a mitochondrial targeting peptide (MTP) translationally fused to a NifH polypeptide (NH), or a cleavage product thereof comprising the NH and, optionally, a scar peptide, wherein the NifH fusion polypeptide or cleavage product thereof is at least partially soluble in the mitochondria of a plant cell. In embodiments of these aspects, the NH polypeptide is at least partially cleaved within the MTP sequence in the plant cell to generate an MPP-cleaved Nif polypeptide. The MPP-cleaved NH comprises a peptide (scar peptide) derived from the MTP sequence and having a length of 1 to 45 amino acids, preferably 1 to 20 amino acids, and more preferably 1 to 11 amino acids or 11 to 20 amino acids, translationally fused to the N-terminus of the MPP-cleaved Nif polypeptide. In embodiments, one or more of the scar peptides are independently 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in length. In embodiments, one or more of the scar peptides are independently 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids in length, or 20-30, 20-40, or 20-50 amino acids in length, although shorter scar sequences are preferred.

[0157] In one embodiment of these aspects, the NH is an AnfH polypeptide.

[0158] In one embodiment, the NifH fusion polypeptide, or preferably its MPP cleavage product, binds one or two Fe-S clusters, preferably one or two Fe4-S4 clusters.

[0159] In another aspect, an isolated or exogenous polynucleotide encoding a NifV polypeptide (NV) is provided, wherein the NV is at least partially soluble in plant mitochondria when expressed in a plant cell.

[0160] In one embodiment, the NV polypeptide comprises an amino acid sequence provided as any one of SEQ ID NOs: 163, 206-209, 211, or 212, or a biologically active fragment thereof, or has an amino acid sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to the amino acid sequence provided as any one or more of SEQ ID NOs: 163, 206-209, 211, or 212.

[0161] In one embodiment, a polypeptide of the invention is an isolated or recombinant polypeptide. In another embodiment, a polypeptide of the invention (e.g., a recombinant polypeptide) is present in a cell, preferably a plant cell.

[0162] Suitable amino acid sequences for the Nif polypeptide of any one of the above embodiments are known in the art and include those provided herein.

[0163] In one embodiment, the NifH polypeptide has the following sequence: i. SEQ ID NO: 1; ii. SEQ ID NO: 218; iii. SEQ ID NO: 224; iv.Accession number WP_049123239.1; v. Accession number WP_048638817.1; vi.Accession number WP_013029017.1; vii.Accession number WP_013010353.1; viii.Accession number WP_014258951.1; ix. Accession number WP_011744626.1; x.Accession number WP_013718497.1; xi.Accession number WP_009565928.1; xii.Accession number WP_013099472.1; xiii.Accession number WP_007781874.1; xiv.Accession number WP_012703362; xv.Accession number WP_153472986; xvi.Accession number WP_015854293; xvii.Accession number WP_123927773; xviii. Accession number WP_073538802; and xix. Accession number RCV6483 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0164] In one embodiment, the NifH polypeptide comprises one or more of the amino acid sequence motifs set forth in SEQ ID NOs:225-231.

[0165] In one embodiment, the NifH polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:1.

[0166] In one embodiment, the NifH polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:218.

[0167] In one embodiment, the NifD polypeptide has the following sequence: i. SEQ ID NO:2; ii. SEQ ID NO: 18; iii. SEQ ID NO: 148; iv. SEQ ID NO: 149; v.SEQ ID NO: 150; vi. SEQ ID NO: 151; vii. SEQ ID NO: 152; viii. SEQ ID NO: 153; ix. SEQ ID NO: 216; x.Accession number WP_044347161.1; xi.Accession number WP_047370273.1; xii.Accession number WP_038902190.1; xiii.Accession number WP_024872642.1; xiv.Accession number WP_024078601.1; xv.Accession number WP_013298320.1; xvi.Accession number WP_010877172.1; xvii.Accession number WP_014258953.1; xviii.Accession number WP_066665786.1; xix.Accession number WP_015773055.1; xx.Accession number WP_016867598.1; xxi.Accession number WP_009512873.1; xxii.Accession number WP_012703361; xxiii.Accession number WP_075356167; xxiv.Accession number WP_038590013; xxv.Accession number WP_ 023922817; xxvi. Accession number WP_011021232; and xxvii. Accession number OAV73823 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0168] In one embodiment, the NifD polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:2.

[0169] In one embodiment, the NifD polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:216.

[0170] In one embodiment, the NifK polypeptide has the following sequence: i. SEQ ID NO:3; ii. SEQ ID NO: 217; iii.Accession number WP_049080161.1; iv.Accession number WP_044347163.1; v. Accession number SBM87811.1; vi.Accession number WP_047370272.1; vii.Accession number WP_014333919.1; viii.Accession number WP_012728880.1; ix. Accession number WP_011912506.1; x.Accession number WP_065303473.1; xi.Accession number WP_018989051.1; xii.Accession number prf||2106319A; xiii.Accession number WP_011021239.1; xiv.Accession number WP_012703359; xv.Accession number WP_144571040; xvi.Accession number WP_077859050; xvii. Accession number WP_122630336; and xviii. Accession number WP_088520366 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0171] In one embodiment, the NifK polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:3.

[0172] In one embodiment, the NifK polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:217.

[0173] In one embodiment, the NifB polypeptide has the following sequence: i. SEQ ID NO: 4; ii. Accession number WP_041145602.1; iii.Accession number WP_043953592.1; iv.Accession number WP_040003311.1; v. Accession number WP_011094468.1; vi.Accession number WP_048638849.1; vii.Accession number WP_011813098.1; viii.Accession number WP_048108879.1; ix. Accession number WP_050355163.1; x. Accession number WP_015850328.1; and xi. Accession number P10930 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0174] In one embodiment, the NifB polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:4.

[0175] In one embodiment, the NifE polypeptide has the following sequence: i. SEQ ID NO:5; ii. Accession number WP_049114606.1; iii. Accession number SBM87755.1; iv.Accession number WP_012764127.1; v. Accession number WP_012728883.1; vi.Accession number WP_003297989.1; vii.Accession number WP_012698965.1; viii.Accession number WP_013190624.1; ix. Accession number WP_025698318.1; x.Accession number WP_013460149.1; xi.Accession number AIS31022.1; xii. Accession number WP_018701501.1; and xiii. Accession number WP_048514099.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0176] In one embodiment, the NifE polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:5.

[0177] In one embodiment, the NifF polypeptide has the following sequence: i. SEQ ID NO: 6; ii. Accession number WP_004122417.1; iii.Accession number WP_040968713.1; iv.Accession number WP_035885760.1; v. Accession number WP_039999438.1; vi.Accession number WP_048638838.1; vii.Accession number WP_064006977.1; viii.Accession number WP_012698862.1; ix. Accession number WP_010933399.1; x. Accession number WP_002949173.1; and xi. Accession number WP_039801725.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0178] In one embodiment, the NifF polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:6.

[0179] In one embodiment, the AnfG polypeptide has the following sequence: i. SEQ ID NO: 219; ii. Accession number WP_012703360; iii.Accession number WP_144571041; iv.Accession number HBE76208; v. Accession number WP_144349445; vi. Accession number WP_112317428; and vii. Accession number WP_048515315 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0180] In one embodiment, the AnfG polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:219.

[0181] In one embodiment, the NifJ polypeptide has the following sequence: i. SEQ ID NO: 7; ii. Accession number WP_024360006.1; iii.Accession number WP_044347157.1; iv.Accession number WP_050533844.1; v. Accession number WP_064566543.1; vi.Accession number WP_057084649.1; vii.Accession number WP_014683040.1; viii.Accession number WP_013149847.1; ix. Accession number WP_053341220.1; x. Accession number WP_014454638.1; and xi. Accession number CSA83023.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0182] In one embodiment, the NifJ polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:7.

[0183] In one embodiment, the NifM polypeptide has the following sequence: i. SEQ ID NO:8; ii. Accession number WP_064342940.1; iii.Accession number WP_004122413.1; iv.Accession number WP_044347181.1; v. Accession number WP_064566543.1; vi.Accession number WP_063105800.1; vii.Accession number WP_035885759.1; viii.Accession number WP_011094472.1; ix. Accession number WP_048638837.1; x.Accession number CAA75544.1; xi. Accession number WP_051692859.1; and xii. Accession number WP_018415157.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0184] In one embodiment, the NifM polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:8.

[0185] In one embodiment, the NifN polypeptide has the following sequence: i. SEQ ID NO: 9; ii. Accession number WP_064391778.1; iii.Accession number WP_047370268.1; iv.Accession number WP_014683026.1; v. Accession number WP_048638830.1; vi.Accession number WP_027147663.1; vii.Accession number WP_015195966.1; viii.Accession number WP_023593609.1; ix. Accession number WP_025677480.1; and x.Accession number WP_018306265.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0186] In one embodiment, the NifN polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:9.

[0187] In one embodiment, the NifQ polypeptide has the following sequence: i. SEQ ID NO: 10; ii. Accession number WP_064391765.1; iii.Accession number CTQ06350.1; iv.Accession number WP_047370257.1; v. Accession number WP_043878077.1; vi.Accession number WP_008878174.1; vii.Accession number WP_011501504.1; viii.Accession number WP_027196569.1; ix. Accession number GAU06296.1; and x. Accession number WP_063239464.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0188] In one embodiment, the NifQ polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:10.

[0189] In one embodiment, the NifS polypeptide has the following sequence: i. SEQ ID NO: 11; ii. SEQ ID NO: 19; iii.Accession number WP_004138780.1; iv.Accession number WP_045858151.1; v. Accession number WP_047370265.1; vi.Accession number WP_014333911.1; vii.Accession number WP_055731597.1; viii.Accession number WP_014239770.1; ix. Accession number WP_054691765.1; x.Accession number WP_021802294.1; xi. Accession number WP_026894054.1; and xii. Accession number WP_061575621.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0190] In one embodiment, the NifS polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:11.

[0191] In one embodiment, the NifS polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:19.

[0192] In one embodiment, the NifU polypeptide has the following sequence: i. SEQ ID NO: 12; ii. Accession number WP_049136164.1; iii.WP_050887862.1; iv.WP_057084657.1; v.WP_048638833.1; vi.WP_012728889.1; vii.WP_055731596.1; viii.WP_028587630.1; ix.WP_044417303.1; x.WP_001051984.1; and xi.KIM05011.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0193] In one embodiment, the NifU polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:12.

[0194] In one embodiment, the NifV polypeptide has the following sequence: i. SEQ ID NO: 13; ii. SEQ ID NO: 163; iii. SEQ ID NO: 164; iv. SEQ ID NO: 206; v. SEQ ID NO: 207; vi. SEQ ID NO: 208; vii. SEQ ID NO: 209; viii. SEQ ID NO: 210; ix. SEQ ID NO: 211; x.SEQ ID NO:212; xi. SEQ ID NO: 213; xii. SEQ ID NO: 214; xiii. SEQ ID NO: 215; xiv.Accession number WP_049083341.1; xv.Accession number WP_045858154.1; xvi.Accession number WP_047370264.1; xvii.Accession number WP_038912041.1; xviii.Accession number WP_048638835.1; xix.Accession number WP_011712856.1; xx.Accession number WP_037528703.1; xxi. Accession number OAA29062.1; and xxii. Accession number EKQ56006.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0195] In one embodiment, the NifV polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:13.

[0196] In one embodiment, the NifX polypeptide has the following sequence: i. SEQ ID NO: 14; ii. Accession number WP_049070199.1; iii.Accession number WP_064342937.1; iv.Accession number WP_044347173.1; v. Accession number WP_044612922.1; vi.Accession number WP_043953583.1; vii.Accession number WP_039999416.1; viii.Accession number WP_047608097.1; ix. Accession number WP_039800848.1; x. Accession number WP_062149047.1; and xi. Accession number WP_020165972.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0197] In one embodiment, the NifX polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:14.

[0198] In one embodiment, the NifY polypeptide has the following sequence: i. SEQ ID NO: 15; ii. Accession number WP_049089500.1; iii.Accession number WP_064342935.1; iv.Accession number WP_044524054.1; v. Accession number WP_049010739.1; vi.Accession number WP_047370270.1; vii.Accession number WP_039999411.1; viii.Accession number WP_037382461.1; ix. Accession number WP_014683024.1; x. Accession number AEX25784.1; and xi. Accession number WP_012698835.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0199] In one embodiment, the NifY polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:15.

[0200] In one embodiment, the NifZ polypeptide has the following sequence: i. SEQ ID NO: 16; ii. Accession number WP_057173223.1; iii.Accession number WP_064342939.1; iv.Accession number WP_043875005.1; v. Accession number WP_043953588.1; vi.Accession number WP_065368553.1; vii.Accession number WP_062627625.1; viii.Accession number WP_011491838.1; ix. Accession number WP_014029050.1; and x. Accession number WP_015665422.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0201] In one embodiment, the NifZ polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:16.

[0202] In one embodiment, the NifW polypeptide has the following sequence: i. SEQ ID NO: 17; ii. Accession number WP_064342938.1; iii.Accession number WP_049080155.1; iv.Accession number WP_095103586.1; v. Accession number WP_065877373.1; vi.Accession number WP_095699971.1; vii.Accession number WP_012764136.1; viii.Accession number WP_053085547.1; ix. Accession number WP_077299824.1; x.Accession number OGI40729; xi. Accession number ACO76430.1; and xii. Accession number BBA37427.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0203] In one embodiment, the NifW polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:17.

[0204] In one embodiment, the ferredoxin polypeptide has the following sequence: i. SEQ ID NO: 232; ii. Accession number WP_012703542; iii.Accession number WP_065835964.1; iv.Accession number WP_069124666.1; v. Accession number WP_101942980; vi.Accession number WP_049076934.1; vii.Accession number WP_072048756.1; viii. Accession number WP_130674512.1; and ix. Accession number WP_103805005.1 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0205] In one embodiment, the ferredoxin polypeptide comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:232.

[0206] Suitable amino acid sequences for MTPs relating to any of the above aspects are known in the art and include those provided herein. In one embodiment, the MTP has the following sequence: i. SEQ ID NO: 36; ii. SEQ ID NO:21; iii. amino acids 1-77 of SEQ ID NO:20; iv. SEQ ID NO: 28; v. SEQ ID NO: 29; vi. SEQ ID NO: 30; vii. SEQ ID NO: 31; viii. SEQ ID NO: 32; ix. SEQ ID NO: 33; x.SEQ ID NO:34; xi. SEQ ID NO: 35; xii. SEQ ID NO: 37; and xiii. SEQ ID NO: 38 and / or a sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to any one or more of:

[0207] In one embodiment, the MTP comprises amino acids having a sequence at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, at least 99% identical, or identical to the sequence set forth in SEQ ID NO:36.

[0208] In another aspect, the invention provides polynucleotides encoding any one or more of the polypeptides of the invention.

[0209] In one embodiment, the protein coding region of the polynucleotide is codon-modified for expression in a plant cell compared to the corresponding protein coding region of the naturally occurring polynucleotide in a bacterium, hi one embodiment, most, or even all, of the protein coding region is codon-optimized for expression in a plant cell, preferably a plant cell of the invention.

[0210] In a further embodiment, each exogenous polynucleotide comprises a promoter operably linked to the polynucleotide, and / or a translational regulatory element operably linked to the polynucleotide.

[0211] In another embodiment, the promoter directs expression of the one or more polynucleotides in the roots, leaves, and / or stems of the plant, preferably the promoter directs expression of the one or more polynucleotides in one, more than one, or all of the roots, leaves, and / or stems of the plant rather than the seeds of the plant.

[0212] In another embodiment, one or more or all of the polynucleotides are present in a plant or bacterial cell and are preferably integrated into the nuclear genome of the plant cell, e.g., as a contiguous DNA sequence integrated into the chloroplast genome, or preferably, the nuclear genome of the plant cell. The plant cell can contain multiple copies of the contiguous DNA sequence integrated into the nuclear genome, e.g., as multiple T-DNAs.

[0213] In one embodiment, each polynucleotide, or each sequence therein encoding a polypeptide, is operably linked to a promoter and optionally a transcription termination sequence.

[0214] In a further or alternative embodiment, the promoter directs expression of the one or more polynucleotides in the roots, leaves, and / or stems of the plant, and preferably the one or more polynucleotides are preferentially expressed in one, more than one, or all of the roots, leaves, and / or stems of the plant relative to the seeds of the plant.

[0215] In a further aspect, there is provided a chimeric vector comprising or encoding a polynucleotide of the invention.

[0216] In another aspect, the present invention provides a vector comprising a polynucleotide of the present invention.

[0217] In one embodiment, the vector comprises a polynucleotide encoding at least three, at least four, or at least five Nif fusion polypeptides.

[0218] In another aspect, the present invention provides a vector comprising a polynucleotide encoding at least three, at least four, or at least five Nif fusion polypeptides as defined in any one of the above aspects of the invention.

[0219] In one embodiment, the vector comprises: a) a NifD fusion polypeptide and a NifK fusion polypeptide, or a NifD-linker-NifK fusion polypeptide; b) a NifH fusion polypeptide and a NifV fusion polypeptide; c) optionally, an AnfG fusion polypeptide and / or a ferredoxin fusion polypeptide The polynucleotide encoding the

[0220] In one embodiment, the vector comprises: a) NifF, NifJ, NifU, and NifB fusion polypeptides and, optionally, a NifS fusion polypeptide; and / or b) NifW, NifX, NifY, and NifZ fusion polypeptides The polynucleotide encoding the

[0221] In a further aspect, the present invention provides a cell comprising one or more polypeptides of the invention, one or more exogenous polynucleotides according to the invention, and / or a vector according to the invention.

[0222] In a related aspect, the invention provides a cell, preferably a plant cell, comprising a fusion polypeptide or cleavage product according to the invention, or a combination of two or more of said fusion polypeptides or cleavage products, a protein complex according to the invention, and / or a polynucleotide according to the invention, or a vector according to the invention. In a preferred embodiment, the cell comprises an exogenous polynucleotide for each fusion polypeptide or cleavage product present in the cell.

[0223] In one embodiment, the fusion polypeptide or cleavage product, or a combination of fusion polypeptides or cleavage products, or protein complex, is present in the mitochondria of a cell. As will be readily understood, not all of the fusion polypeptide or cleavage products need to be present in the mitochondria of a cell, provided that at least some are present in the mitochondria.

[0224] In one embodiment, the cell is a plant cell or a bacterial cell, preferably a cell of a transgenic plant, more preferably a cell in which at least one of the exogenous polynucleotides has been integrated into the nuclear genome of the cell.

[0225] In a further embodiment, the plant cell is a monocotyledonous plant cell (e.g., a cereal plant cell, such as a wheat cell, rice cell, corn cell, triticale cell, oat cell, or barley cell, preferably a wheat cell), or a dicotyledonous plant cell. The plant cell can be further characterized by a polypeptide or polynucleotide defined by any of the above-mentioned characteristics. All possible combinations of the above-mentioned characteristics are considered part of the invention in the context of plant cells and in other aspects of the invention.

[0226] In a further aspect, the present invention provides a transgenic plant or transgenic part thereof (preferably a seed) comprising one or more polypeptides according to the invention, one or more exogenous polynucleotides according to the invention, and / or a vector according to the invention.

[0227] In one embodiment, the transgenic plant is a monocotyledonous plant (e.g., a cereal plant such as wheat, rice, corn, triticale, oats, or barley, preferably wheat), or a dicotyledonous plant. The plant or part thereof can be further characterized by a polypeptide or polynucleotide defined by any of the above-mentioned characteristics. All possible combinations of the above-mentioned characteristics are considered part of the present invention in the context of the plant or part thereof and in other aspects of the invention.

[0228] In a further aspect, the invention provides a method for producing a polypeptide according to the invention, which method comprises expressing a polynucleotide according to the invention in a cell.

[0229] In a further aspect, the present invention provides a method of producing a cell according to the invention, the method comprising the step of introducing one or more polynucleotides according to the invention and / or a vector according to the invention into the cell.

[0230] In another aspect, the present invention provides a method for producing homocitric acid in a plant cell, the method comprising expressing a recombinant NifV polypeptide or NifV fusion polypeptide of the present invention in the plant cell, wherein the recombinant NifV polypeptide or NifV fusion polypeptide, and / or its cleavage products, produce homocitric acid in the plant cell.

[0231] In one embodiment, the method further comprises introducing into the plant cell a polynucleotide encoding the recombinant NifV polypeptide or NifV fusion polypeptide.

[0232] In another aspect, the present invention provides the use of a NifV polypeptide of the present invention for producing homocitric acid in a plant cell.

[0233] In another aspect, the invention provides a method for increasing the amount of NifD, NifK, or a NifD-linker-NifK fusion polypeptide in a plant cell, the method comprising expressing in the plant cell one or more or all of NifW, NifX, NifY, and NifZ fusion polypeptides (wherein each Nif fusion polypeptide independently comprises a mitochondrial targeting peptide (MTP)), thereby increasing the amount of NifD, NifK, or a NifD-linker-NifK fusion polypeptide in the plant cell relative to a corresponding plant cell that does not express one or more or all of the NifW, NifX, NifY, and NifZ fusion polypeptides.

[0234] In one embodiment, the method further comprises: i) introducing into a plant cell one or more polynucleotides encoding NifD, NifK, or a NifD-linker-NifK fusion polypeptide; ii) introducing into the plant cell one or more polynucleotides encoding one or more or all of the NifW, NifX, NifY, and NifZ fusion polypeptides.

[0235] In another aspect, the present invention provides a method for increasing the amount of a NifY polypeptide in a plant cell, the method comprising expressing in the plant cell one or more or all of NifW, NifX, and NifZ fusion polypeptides (wherein each Nif fusion polypeptide independently comprises a mitochondrial targeting peptide (MTP)), thereby increasing the amount of NifY polypeptide in the plant cell relative to a corresponding plant cell that does not express one or more or all of the NifW, NifX, and NifZ fusion polypeptides.

[0236] In one embodiment, the method further comprises: i) introducing into a plant cell a polynucleotide encoding a NifY polypeptide; ii) introducing into the plant cell one or more polynucleotides encoding one or more or all of the NifW, NifX, and NifZ fusion polypeptides.

[0237] In another aspect, the present invention provides the use of one or more polynucleotides encoding one or more or all of the NifW, NifX, and NifZ fusion polypeptides to increase the amount of NifY polypeptide in a plant cell.

[0238] In another aspect, the present invention provides the use of a polynucleotide of the present invention and / or a vector of the present invention for producing a transgenic plant.

[0239] In another aspect, the present invention provides a method for producing a transgenic plant, the method comprising: i) introducing one or more polynucleotides of the invention and / or one or more vectors of the invention into cells of a plant; ii) regenerating a transgenic plant of the invention from the cells of step i); and iii) optionally generating transgenic seeds and / or progeny plants from the transgenic plants regenerated in step ii).

[0240] In a further aspect, the present invention provides a method of producing a transgenic seed, the method comprising: i) harvesting seeds from a transgenic plant of the invention; and / or ii) harvesting seeds from one or more transgenic progeny plants produced by the methods of the present invention.

[0241] In a further aspect, the present invention provides a method for producing a plant having a polynucleotide according to the present invention integrated into its genome, the method comprising: i) crossing two parent plants, wherein at least one of the plants comprises the polynucleotide; ii) screening one or more progeny plants from this cross for the presence or absence of the polynucleotide; and iii) selecting progeny plants containing the polynucleotide; This produces the above-mentioned plant.

[0242] In a further or alternative embodiment, at least one of the parent plants is a tetraploid or hexaploid wheat plant.

[0243] In a further or alternative embodiment, step ii) comprises analyzing a sample comprising DNA from one or more of the progeny plants for said polynucleotide.

[0244] In a further or alternative embodiment, step iii) comprises: i) selecting progeny plants that are homozygous for the polynucleotide; and / or ii) analyzing the plant, or one or more of its progeny plants, for the presence and / or expression of the polynucleotide or for an altered phenotype as defined above.

[0245] In one or a further embodiment, the method further comprises: iv) backcrossing the progeny of the cross of step i) a sufficient number of times with plants of the same genotype as the first parent plant lacking said polynucleotide to produce plants having most of the genotype of the first parent but containing said polynucleotide; v) selecting progeny plants which contain the polynucleotide and / or have an altered phenotype as defined above.

[0246] In a further or alternative embodiment, the method further comprises analyzing the plant or progeny plants for at least one other genetic marker.

[0247] In a further aspect, the invention provides a plant produced using the method of the invention.

[0248] In a further aspect, the present invention provides the use of a polynucleotide according to the invention and / or a vector according to the invention for producing a recombinant cell and / or a transgenic plant.

[0249] In one embodiment, the transgenic plant has an altered phenotype as defined above when compared to a corresponding plant lacking said exogenous polynucleotide and / or said vector.

[0250] In a further aspect, the present invention provides a method for identifying a plant comprising a polynucleotide according to the present invention, the method comprising: i) obtaining a nucleic acid sample from a plant; ii) screening the sample for the presence or absence of said polynucleotide.

[0251] In one embodiment, the presence of the polynucleotide indicates that the plant has an altered phenotype as defined above when compared to a corresponding plant lacking the exogenous polynucleotide.

[0252] In a further or alternative embodiment, the method identifies a plant according to the present invention.

[0253] In a further or alternative embodiment, the method further comprises generating a plant from the seed prior to step i).

[0254] In another aspect, the invention provides a part of a transgenic plant comprising a plant cell of the invention, or a plant cell obtained from a transgenic plant of the invention.

[0255] In one embodiment, the plant part is a seed comprising a polynucleotide of the invention.

[0256] In another aspect, the present invention provides a method for producing flour, whole grain, starch, oil, seed meal, or other product derived from seeds, the method comprising: a) obtaining seeds of the invention, and / or b) Extracting grain flour, whole grain, starch, oil or other products, or producing seed meal.

[0257] In a further aspect, the present invention provides products produced from transgenic plants of the invention and / or parts of plants of the invention comprising a polypeptide of the invention and / or a polynucleotide of the invention.

[0258] In one embodiment, the plant part is a seed.

[0259] In a further or alternative embodiment, the product is a food ingredient or beverage ingredient or a foodstuff or beverage. Preferably, i) the food ingredient or product is selected from the group consisting of foods containing cereal flour, starch, fat, leavened or unleavened bread, pasta, noodles, animal feed, breakfast cereals, snacks, cakes, malt, pastries, and flour-based sauces, or ii) the beverage is juice, beer, or malt. Methods for producing such products are well known to those skilled in the art.

[0260] In an alternative embodiment, the product is a non-food product. Non-limiting examples of non-food products include films, coatings, adhesives, building materials, and packaging materials. Methods of making such products are well known to those skilled in the art.

[0261] In a further aspect, the invention provides a method of preparing a food product, the method comprising mixing the seeds of the invention, or a flour, wholemeal, starch, oil or other product therefrom, with another food ingredient, or preferably processing the seeds or flour or wholemeal by grinding, cracking, grinding, flaking, blanching, cooking or baking the seeds or a composition comprising the seeds and / or flour or wholemeal obtained from the seeds.

[0262] In a further aspect, the present invention provides a method of preparing malt, the method comprising germinating seeds according to the present invention.

[0263] In a further aspect, the present invention provides the use of a plant or part thereof according to the present invention as animal feed or for producing feed for animal consumption or food for human consumption.

[0264] In a further aspect, the present invention provides a composition comprising any one of a polypeptide according to the present invention, a polynucleotide according to the present invention, a vector according to the present invention, or a cell according to the present invention, and one or more acceptable carriers.

[0265] In a further aspect, the present invention provides a method for reconstituting a nitrogenase protein complex in a plant cell, the method comprising introducing into the cell two or more polynucleotides according to the invention, two or more nucleic acid constructs according to the invention, and / or a vector according to the invention, and culturing the plant cell for a period of time sufficient to allow expression of the polynucleotides or vectors.

[0266] In another aspect, the invention provides a plant cell comprising a mitochondrion and three, four, five, six, seven, eight, nine, ten, or eleven Nif polypeptides selected from the group consisting of NifF, NifM, NifN, NifS, NifU, NifW, NifY, NifZ, NifV, NifH, and NifD-NifK, and each of the three, four, five, six, seven, eight, nine, ten, or eleven Nif polypeptides is at least partially soluble in the mitochondrion.

[0267] In one embodiment, the plant cell contains NifV. Preferably, the NifV produces homocitric acid. In one embodiment, the NifV is a NifV of the present invention.

[0268] In another embodiment, the plant cell contains NifS, NifU, or both NifS and NifU, and optionally NifV.

[0269] In another embodiment, the plant cell comprises NifH, NifM, or both NifH and NifM, and optionally all one or more of NifV, NifS, and NifU.

[0270] In another embodiment, the plant cell comprises NifF, NifH, or NifD-NifK, or NifH and NifD-NifK, or NifF, NifH, and NifD-NifK, and optionally all one or more of NifV, NifS, NifU, NifH, and NifM.

[0271] In one embodiment, the NifH polypeptide is an AnfH polypeptide and the NifD-NifK polypeptide is an AnfD-AnfK polypeptide.

[0272] In a further embodiment, the plant comprises an AnfG polypeptide that is at least partially soluble in the mitochondria.

[0273] In one embodiment, at least 10%, at least 20%, at least 30%, or at least 40%, up to 50%, of each of the 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides are soluble in mitochondria.

[0274] In one embodiment, the 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides each independently comprise a mitochondrial targeting peptide (MTP), or a C-terminal peptide resulting from cleavage of the MTP, or both, and preferably the MTP or the C-terminal peptide or both are located at the N-terminus of each of the 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides, or there is no MTP and no C-terminal peptide at the N-terminus of the Nif polypeptide.

[0275] In a further embodiment, 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides are each independently cleaved within or immediately after the MTP to yield 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides, whereby each of the 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides either contains an MTP-derived C-terminal peptide at its N-terminus, or does not contain an MTP-derived C-terminal peptide.

[0276] In another or further embodiment, 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides are each independently cleaved within or immediately after the MTP to yield 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides, whereby each of the 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides either contains an MTP-derived C-terminal peptide at its N-terminus, or does not contain an MTP-derived C-terminal peptide.

[0277] In one embodiment, each MTP is independently cleaved in the plant cell with at least 50% efficiency, and / or each of the 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides is present in the plant cell in an amount greater than the corresponding Nif polypeptide, preferably 1:1, 2:1, or 3:1.

[0278] In one embodiment, the plant cell comprises a NifD-NifK fusion polypeptide comprising, in order, a NifD amino acid sequence (ND), a linker amino acid sequence, and a NifK polypeptide (NK) amino acid sequence, wherein the linker amino acid sequence has a length of 8 to 50 residues, preferably 16 to 50 residues, more preferably about 26 or about 30 residues, or most preferably 26 or 30 residues, and is translationally fused to ND and NK.

[0279] In a further embodiment, the NifD-NifK fusion polypeptide further comprises a mitochondrial targeting peptide (MTP), or a C-terminal peptide resulting from cleavage of MTP, or both, which are translatably fused to the N-terminus of the NifD-NifK fusion polypeptide.

[0280] In one embodiment, the 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides each independently comprise a C-terminal peptide resulting from cleavage of MTP, 1 to 45 amino acids in length, preferably 1 to 20 amino acids, translatably fused at the N-terminus of the Nif polypeptide.

[0281] In one embodiment, the 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides, or the 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides, or both, are functional Nif polypeptides.

[0282] In one embodiment, 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides, or preferably 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides, or both, are present in the mitochondria of the plant cell, preferably in the mitochondrial matrix (MM) of the plant cell.

[0283] In one embodiment, the 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides, or preferably the 3, 4, 5, 6, 7, 8, 9, 10, or 11 processed Nif polypeptides, or both, independently, are predominantly soluble in plant mitochondria (i.e., are greater than 50% soluble in mitochondria).

[0284] In one embodiment, the ND comprises an amino acid other than tyrosine (Y) at a position corresponding to amino acid 100 of SEQ ID NO:18.

[0285] In one embodiment, the ND comprises a glutamine (Q) or lysine (K) at a position corresponding to amino acid 100 of SEQ ID NO: 18, or a leucine (L) or methionine (M) or phenylalanine (F) at a position corresponding to amino acid 100 of SEQ ID NO: 18.

[0286] In one embodiment, the MTP is about 51 amino acids in length from the F1-ATPase gamma subunit.

[0287] In one embodiment, the plant cell comprises an NK amino acid sequence, the C-terminus of the polypeptide being the C-terminus of wild-type NifK.

[0288] In one embodiment, the linker is at least about 20 amino acids in length, or at least about 30 amino acids, or at least about 40 amino acids, or from about 20 amino acids to about 70 amino acids, or from about 30 amino acids to about 70 amino acids, or from about 30 amino acids to about 60 amino acids, or from about 30 amino acids to about 50 amino acids, or about 25 amino acids, or about 30 amino acids, or about 35 amino acids, or about 40 amino acids, or about 45 amino acids, or about 46 amino acids, or about 50 amino acids, or about 55 amino acids.

[0289] In one embodiment, the fusion polypeptide can be cleaved within or immediately after the MTP to produce a processed polypeptide (CNK), whereby the CNK comprises, in order, the optimal C-terminal peptide resulting from cleavage of the MTP, the NifD amino acid sequence (ND), a linker amino acid sequence, and the NK amino acid sequence.

[0290] In one embodiment, the plant cell further comprises a fusion polypeptide or a CDK, or both.

[0291] In one embodiment, the CDK comprises a scar sequence of 1 to 45 amino acids, preferably 1 to 20 amino acids in length, translationally fused to the N-terminus of the NifD amino acid sequence.

[0292] In one embodiment, the CDK has the functions of both NifD and NifK.

[0293] In one embodiment, ND is AnfD and NK is AnfK.

[0294] In one embodiment, the MTP is about 51 amino acids in length from the F1-ATPase gamma subunit.

[0295] In one embodiment, each MTP comprises at least 10 amino acids, and preferably has a length of 10-80 amino acids.

[0296] In one embodiment, the MTP, or at least one MTP, or all of the MTPs independently comprise an MTP of a mitochondrial protein precursor or a variant thereof, preferably a plant MTP.

[0297] In one embodiment, the 3, 4, 5, 6, 7, 8, 9, 10, or 11 Nif polypeptides are encoded by 3, 4, 5, 6, 7, 8, 9, 10, or 11 exogenous polynucleotide(s), which 3, 4, 5, 6, 7, 8, 9, 10, or 11 are integrated into the nuclear genome of the cell, preferably as a contiguous nucleic acid sequence.

[0298] In another embodiment of any of the above aspects, the cell is a cell other than an Arabidopsis thaliana protoplast.

[0299] The present inventors are the first to generate plant cells containing NifV polypeptides that are at least partially soluble in mitochondria. Thus, in another aspect, the present invention provides a plant cell containing a NifV polypeptide (NV), wherein the NV is at least partially soluble in mitochondria.

[0300] In one embodiment, the NV is capable of producing or is producing homocitrate in the cell.

[0301] In one embodiment, the NV polypeptide comprises an amino acid sequence provided as any one of SEQ ID NOs: 205-209, or 211, or a biologically active fragment thereof, or has an amino acid sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to the amino acid sequence provided as any one or more of SEQ ID NOs: 205-209, or 211.

[0302] The present inventors are also the first to generate plant cells containing a NifH polypeptide that is at least partially soluble in mitochondria. Thus, in another aspect, the present invention provides a plant cell containing a NifH polypeptide (NH), wherein the NH is at least partially soluble in mitochondria.

[0303] In one embodiment, the NH is encoded by an exogenous polynucleotide, which is preferably integrated into the nuclear genome of the cell as a contiguous nucleic acid sequence.

[0304] In one embodiment, the plant cell of one or both of the above two aspects is further defined by one or more of the characteristics mentioned herein.

[0305] In a further aspect, the present invention provides a transgenic plant comprising a plant cell of the present invention, wherein the transgenic plant is transgenic for one or more exogenous polynucleotide(s) encoding Nif polypeptide(s).

[0306] In one embodiment, one or more of the one or more exogenous polynucleotide(s) is expressed in the roots of the plant, preferably at a higher level in the roots of the plant than in the leaves of the plant.

[0307] In a further or alternative embodiment, the transgenic plant has an altered phenotype compared to a corresponding wild-type plant (e.g., increased yield, biomass, growth rate, vigor, nitrogen gain from biological nitrogen fixation, nitrogen use efficiency, abiotic stress tolerance, and / or tolerance to nutrient deficiency compared to a corresponding wild-type plant).

[0308] In an alternative embodiment, the transgenic plant has the same growth rate and / or phenotype as the corresponding wild-type plant.

[0309] In one embodiment, the plant is a cereal plant (such as wheat, rice, corn, triticale, oats, or barley, preferably wheat).

[0310] In one embodiment, the plant is homozygous or heterozygous for one or more exogenous polynucleotide(s), preferably homozygous for all of the exogenous polynucleotides.

[0311] In another embodiment, the transgenic plant is a monocotyledonous plant (eg, a cereal plant such as wheat, rice, corn, triticale, oats, or barley, preferably wheat), or a dicotyledonous plant.

[0312] In a further or alternative embodiment, the transgenic plants are grown outdoors.

[0313] In a further embodiment, the invention provides a population of at least 100 plants according to the invention growing outdoors.

[0314] Also provided is a substantially purified or recombinant NifV polypeptide (NV) that, when expressed in a plant cell, is at least partially soluble in plant mitochondria.

[0315] Further provided is a substantially purified or recombinant NifH polypeptide (NH) that, when expressed in a plant cell, is at least partially soluble in plant mitochondria.

[0316] In another aspect, an isolated or exogenous polynucleotide encoding a NifV polypeptide (NV) is provided, wherein the NV is at least partially soluble in plant mitochondria when expressed in a plant cell.

[0317] In one embodiment, the NV polypeptide comprises an amino acid sequence provided as any one of SEQ ID NOs: 205-209, or 211, or a biologically active fragment thereof, or has an amino acid sequence that is at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 97% identical, or at least 99% identical to the amino acid sequence provided as any one or more of SEQ ID NOs: 205-209, or 211.

[0318] In a further aspect, there is provided an isolated or exogenous polynucleotide encoding a NifH polypeptide (NH) of the invention, said NV being at least partially soluble in plant mitochondria when expressed in a plant cell.

[0319] In one embodiment, the protein coding region of the polynucleotide has been codon-modified for expression in a plant cell compared to the corresponding protein coding region of the polynucleotide naturally occurring in bacteria.

[0320] In a further embodiment, the polynucleotide further comprises a promoter operably linked to the polynucleotide, and / or a translational regulatory element operably linked to the polynucleotide.

[0321] In another embodiment, the promoter directs expression of the one or more polynucleotide(s) in the roots, leaves, and / or stems of the plant, preferably the promoter directs expression of the one or more polynucleotide(s) in one, more than one, or all of the roots, leaves, and / or stems of the plant rather than the seeds of the plant.

[0322] In another embodiment, the polynucleotide is present in a plant or bacterial cell and is preferably integrated into the nuclear genome of the plant cell, for example, as a contiguous DNA sequence integrated into the nuclear genome or chloroplast genome of the plant cell. Plant cells can contain multiple copies of a contiguous DNA sequence integrated into their nuclear genome.

[0323] In a further aspect, there is provided a chimeric vector comprising or encoding a polynucleotide of the invention.

[0324] In one embodiment, the polynucleotide, or each sequence therein that encodes a polypeptide, is operably linked to a promoter and optionally a transcription termination sequence.

[0325] In a further or alternative embodiment, the promoter causes expression of the one or more polynucleotide(s) in the roots, leaves, and / or stems of the plant, and preferably the one or more polynucleotide(s) are preferentially expressed in one, more than one, or all of the roots, leaves, and / or stems of the plant relative to the seeds of the plant.

[0326] In a further aspect, the present invention provides a cell comprising one or more polypeptides of the invention, one or more exogenous polynucleotides according to the invention, and / or a vector according to the invention.

[0327] In one embodiment, the cell is a plant cell or a bacterial cell.

[0328] In a further embodiment, the plant cell is a monocotyledonous plant cell (e.g., a cereal plant cell, such as a wheat cell, rice cell, corn cell, triticale cell, oat cell, or barley cell, preferably a wheat cell), or a dicotyledonous plant cell.

[0329] In a further aspect, the present invention provides a transgenic plant or transgenic part thereof (preferably a seed) comprising one or more polypeptides according to the invention, one or more exogenous polynucleotides according to the invention, and / or a vector according to the invention.

[0330] In one embodiment, the transgenic plant is a monocotyledonous plant (eg, a cereal plant such as wheat, rice, corn, triticale, oats, or barley, preferably wheat), or a dicotyledonous plant.

[0331] In a further aspect, the invention provides a method for producing a polypeptide according to the invention, which method comprises expressing a polynucleotide according to the invention in a cell.

[0332] In a further aspect, the present invention provides a method of producing a cell according to the invention, the method comprising the step of introducing one or more polynucleotides according to the invention and / or a vector according to the invention into the cell.

[0333] In another aspect, the present invention provides a method for producing a transgenic plant, the method comprising: i) introducing the polynucleotide of the present invention and / or the vector of the present invention into a plant cell; ii) regenerating a transgenic plant of the invention from the cells of step i); and iii) optionally generating transgenic progeny plants from one or more transgenic plants regenerated in step ii). Includes.

[0334] In a further aspect, the present invention provides a method of producing a transgenic seed, the method comprising: i) harvesting seeds from a transgenic plant of the invention; and / or ii) harvesting seeds from one or more transgenic progeny plants produced by the methods of the present invention. Includes.

[0335] In a further aspect, the present invention provides a method for producing a plant having a polynucleotide according to the present invention integrated into its genome, the method comprising: i) crossing two parent plants, wherein at least one of the plants comprises the polynucleotide; ii) screening one or more progeny plants from this cross for the presence or absence of the polynucleotide; and iii) selecting progeny plants containing the polynucleotide; This produces the above-mentioned plant.

[0336] In a further or alternative embodiment, at least one of the parent plants is a tetraploid or hexaploid wheat plant.

[0337] In a further or alternative embodiment, step ii) comprises analyzing a sample comprising DNA from one or more of the progeny plants for said polynucleotide.

[0338] In a further or alternative embodiment, step iii) comprises:

[0339] i) selecting progeny plants that are homozygous for the polynucleotide; and / or ii) analyzing the plant, or one or more of its progeny plants, for the presence and / or expression of the polynucleotide or for an altered phenotype as defined above. Includes.

[0340] In one or a further embodiment, the method further comprises: iv) backcrossing the progeny of the cross of step i) a sufficient number of times with plants of the same genotype as the first parent plant lacking said polynucleotide to produce plants having most of the genotype of the first parent but containing said polynucleotide; iv) selecting progeny plants which contain the polynucleotide and / or have an altered phenotype as defined above. Includes.

[0341] In a further or alternative embodiment, the method further comprises analyzing the plant or progeny plants for at least one other genetic marker.

[0342] In a further aspect, the invention provides a plant produced using the method of the invention.

[0343] In a further aspect, the present invention provides the use of a polynucleotide according to the invention and / or a vector according to the invention for producing a recombinant cell and / or a transgenic plant.

[0344] In one embodiment, the transgenic plant has an altered phenotype as defined above when compared to a corresponding plant lacking said exogenous polynucleotide and / or said vector.

[0345] In a further aspect, the present invention provides a method for identifying a plant comprising a polynucleotide according to the present invention, the method comprising: i) obtaining a nucleic acid sample from a plant; ii) screening the sample for the presence or absence of said polynucleotide. Includes.

[0346] In one embodiment, the presence of the polynucleotide indicates that the plant has an altered phenotype as defined above when compared to a corresponding plant lacking the exogenous polynucleotide.

[0347] In a further or alternative embodiment, the method identifies a plant according to the present invention.

[0348] In a further or alternative embodiment, the method further comprises generating a plant from the seed prior to step i).

[0349] In another aspect, the invention provides a part of a transgenic plant comprising a plant cell of the invention, or a plant cell obtained from a transgenic plant of the invention.

[0350] In one embodiment, the plant part is a seed comprising a polynucleotide of the invention.

[0351] In another aspect, the present invention provides a method for producing flour, whole grain, starch, oil, seed meal, or other product derived from seeds, the method comprising: a) obtaining the seeds of the present invention; and b) Extracting cereal flours, whole grains, starches, oils or other products, or producing seed meals Includes.

[0352] In a further aspect, the present invention provides products produced from transgenic plants of the invention and / or parts of plants of the invention comprising a polypeptide of the invention and / or a polynucleotide of the invention.

[0353] In one embodiment, the plant part is a seed.

[0354] In a further or alternative embodiment, the product is a food ingredient or beverage ingredient or a foodstuff or beverage. Preferably, i) the food ingredient or product is selected from the group consisting of foods containing cereal flour, starch, fat, leavened or unleavened bread, pasta, noodles, animal feed, breakfast cereals, snacks, cakes, malt, pastries, and flour-based sauces, or ii) the beverage is juice, beer, or malt. Methods for producing such products are well known to those skilled in the art.

[0355] In an alternative embodiment, the product is a non-food product. Non-limiting examples of non-food products include films, coatings, adhesives, building materials, and packaging materials. Methods of making such products are well known to those skilled in the art.

[0356] In a further aspect, the invention provides a method of preparing a foodstuff, the method comprising mixing the seed of the invention, or a flour, whole grain, starch, oil or other product therefrom, with another food ingredient.

[0357] In a further aspect, the present invention provides a method of preparing malt, the method comprising germinating seeds according to the present invention.

[0358] In a further aspect, the present invention provides the use of a plant or part thereof according to the present invention as animal feed or for producing feed for animal consumption or food for human consumption.

[0359] In a further aspect, the present invention provides a composition comprising any one of a polypeptide according to the present invention, a polynucleotide according to the present invention, a vector according to the present invention, or a cell according to the present invention, and one or more acceptable carriers.

[0360] Any embodiment herein is intended to apply mutatis mutandis to any other embodiment unless otherwise indicated. For example, those skilled in the art will appreciate that the examples of Nif polypeptides outlined above with respect to one aspect of the invention apply equally to other aspects of the invention.

[0361] The scope of the present invention is not limited by the specific embodiments described herein, which are intended for purposes of example only. Functionally equivalent products, compositions, and methods are clearly within the scope of the invention as described herein.

[0362] Throughout this specification, unless otherwise stated or the context requires otherwise, a single step, composition of matter, group of steps, or composition of matter is intended to encompass one and more than one (i.e., one or more) of that step, composition of matter, group of steps, or group of compositions of matter.

[0363] The invention will now be described by the following non-limiting examples, with reference to the accompanying drawings. [Brief explanation of the drawings]

[0364] [Figure 1] Western blot analysis using anti-HA antibody to detect unprocessed and MPP-processed individual pFAγ51::Nif::HA or 6×HIS::Nif::HA polypeptides after transient expression in Nicotiana benthamiana leaves. C, cytoplasmic expression (6×His); M, mitochondrial targeting. [Figure 2] Western blots of protein extracts from N. benthamiana leaves after transfection with MTP:Nif gene constructs. The first and last lanes on each blot show molecular weight markers in kDa from the Invitrogen Prestained BenchMark ladder. The gene construct used in each sample is indicated above each lane, and the Nif polypeptide contained in each fusion polypeptide is indicated below the lane. Paired infiltrations of constructs SN26–SN32 were performed with or without co-infiltration of pRA25, encoding the MTP-FAγ77::NifK fusion polypeptide (WO 2018 / 141030). Western blots were probed with an HA antibody. [Figure 3] Western blot analysis of individual MTP-FAγ51::Nif::HA polypeptides (except MTP-FAγ51::HA::NifK) and their MPP-processed products after expression in Nicotiana benthamiana leaf cells using anti-HA antibody. T, total protein; I, insoluble fraction; S, soluble fraction. [Figure 4] The upper panel shows a schematic representation of the genetic constructs tested to generate secondary cleavage products from wild-type NifD fusion polypeptides. MTP represents the FAγ51 or L29 sequence, NifD represents the wild-type K. oxytoca sequence, and HA represents the HA epitope. The lower panel shows a Western blot of protein extracts after transfection of the genetic constructs into N. benthamiana leaf cells. The Western blot was probed with an HA antibody. Lane 1 shows molecular weight markers using the Prestained Benchmark ladder. Paired lanes indicate the absence (-) or presence (+) of the NifK construct pRA25. Band 1 represents the unprocessed MTP::NifD fusion polypeptide; band 2 represents the MPP-processed fusion polypeptide; band 3 represents the approximately 48 kDa degradation product. [Figure 5] Western blot of protein extracts after introduction of the MTP:NifD gene construct into N. benthamiana leaf cells. Lane 1 shows molecular weight markers in kDa using the ThermoFisher Prestained Benchmark ladder. The gene construct used in each sample is indicated above each lane. pRA24 encodes the MTP-FAγ::NifD::HA polypeptide, within which the NifD coding region was codon-optimized for Arabidopsis (WO 2018 / 141030). Each construct was cotransfected into plant cells with pRA25 (MTP-FAγ77::NifK) to enhance accumulation of the NifD fusion polypeptide. Western blots were probed with an HA antibody. The arrow indicates the position of the approximately 48-kDa secondary cleavage product from NifD. The arrow indicates the position of the approximately 48-kDa secondary cleavage polypeptide from NifD. [Figure 6]Western blot of protein extracts after introduction of the MTP:NifD gene construct into N. benthamiana leaf cells. Lane 1 shows molecular weight markers in kDa using the ThermoFisher Prestained Benchmark ladder. The gene construct used in each sample is indicated above each lane. SN64 encodes the mMTP-CPN60::NifD polypeptide, in which the mMTP-CPN60 amino acid sequence was altered by alanine substitution, making it resistant to MPP cleavage. pRA24 encodes the MTP-FAγ::NifD::HA polypeptide, in which the NifD coding region was codon-optimized for Arabidopsis (WO2018 / 141030). Western blots were probed with an HA antibody. [Figure 7] Alignment of the mutated mMTP-FA γ51 amino acid sequence in SN66 (SEQ ID NO:59) with the unmodified MTP-FA γ51 sequence in SN10 (SEQ ID NO:122) (SEQ ID NO:21). A region of five consecutive amino acid residues and a region of eight consecutive amino acid residues were substituted with alanine, rendering MPP processing inactive. [Figure 8] Western blots of protein extracts from plant or yeast cells transfected with MTP:Nif gene constructs were probed with an HA antibody, demonstrating secondary NifD cleavage / degradation in yeast cells and reduced cleavage with the Y100Q amino acid substitution (SN114, SNY114). Protein extracts from N. benthamiana leaf cells (SN10, SN196, SN114) or yeast cells (SNY10, SNY196, SNY114) were electrophoresed in the indicated lanes. Lanes 1 and 8 show molecular weight markers in kDa using the ThermoFisher Prestained Benchmark ladder. The band at approximately 64 kDa represents the unprocessed MTP::NifD::HA fusion polypeptide, and the band at approximately 58 kDa represents the MPP-processed polypeptide. The arrow indicates the approximately 48 kDa C-terminal polypeptide generated by secondary cleavage. [Figure 9] Western blot of protein extracts from N. benthamiana leaf cells after transfection with gene constructs encoding MTP::NifD::HA amino acid substitution variants, along with SN46 (MTP-Su9::NifK). Lane 12 shows molecular weight markers in kDa using the ThermoFisher Prestained Benchmark ladder. The most intense band at approximately 58 kDa in lanes 5–11 was MTP-FAγ51::NifD processed by MPP. Lanes 2 and 3 show the 48 kDa polypeptide generated by secondary cleavage. Note the absence of the 48 kDa polypeptide in lanes 5–11. [Figure 10] Amino acid sequence alignment of the region corresponding to amino acids 49-108 of K. oxytoca NifD (SEQ ID NO: 18) with the wild-type NifD polypeptide. One representative sequence was selected from each cluster containing at least 10 members in the sequence similarity network. The number of members in each cluster of NifD sequences is indicated in parentheses. Completely conserved amino acids are indicated above the alignment. [Figure 11] Location of the proposed secondary cleavage site shown in the crystal structure of the NifD polypeptide from K. oxytoca (PDB: 1QGU). The cofactor FeMoco is shown as multiple spheres on the right. NifK-Ser515, NifK-Asp517, the C-terminus, and the structures to the upper left are from the NifK polypeptide. Arg97, Arg98, Asn99, Tyr100, Tyr101, Thr102, and the structures to the lower right, excluding FeMoco, are from NifD. Dotted lines indicate possible hydrogen bonds between Tyr100 and Ser515, and between Asp517 and the hydroxyl of Arg98. [Figure 12]Western blot analysis showing mitochondrial processing of NifD fusion polypeptides from six different bacteria. Three constructs were analyzed for each NifD sequence in adjacent lanes: the mMTP-FAγ51::NifD::HA fusion polypeptide that was not cleaved by MPP at the canonical MPP cleavage site (lane labeled A), the MTP-FAγ51::NifD::HA that was targeted to mitochondria (lane labeled M), and the 6×His::NifD::HA that was predicted to be localized to the cytoplasm and whose size corresponds to that of MPP-processed polypeptides (lane labeled C). [Figure 13] Schematic map (not drawn to scale) of the gene construct encoding the NifD::linker(HA)::NifK fusion polypeptide. mMTP-FAγ denotes a mutant MTP with an alanine substitution that prevents cleavage by MTP. Y100Q denotes the presence of this amino acid substitution in the NifD sequence. [Figure 14] Solubility of the NifD-linker(HA)-NifK polypeptide after expression in N. benthamiana. Protein from infiltrated leaf samples was isolated as "total" protein or fractionated into insoluble and soluble fractions as described in Example 1. Protein ladder markers shown are ThermoFisher Prestained Benchmark ladder used on blots for "total" and "insoluble" samples and Invitrogen PageRuler ladder used on blots for "soluble" samples. [Figure 15] A schematic diagram of the metaxin fusion polypeptide encoded by a single gene in SN197, and how the majority of this polypeptide is located on the outer mitochondrial membrane with its N-terminus inserted into the cytoplasm. This construct uses the N. benthamiana metaxin sequence. [Figure 16]Western blot showing that purification of mitochondrially targeted MTP-FAγ51::NifU::TS from SN166 resulted in the production of the processed form of the NifU polypeptide. Upper panel: probed with anti-Strep antibody. Lower panel: gel stained with Coomassie blue. [Figure 17] Western blot showing copurification of scar9::GG::NifS::HA during purification of mitochondrial-targeted scar9::GG::NifU::TS. Samples from steps (i) to (v) of the purification process in the first purification experiment were subjected to SDS-PAGE and Western blotting using an anti-Strep antibody to detect the NifU polypeptide or an anti-HA antibody to detect the NifS polypeptide. The two bands for NifS correspond to the unprocessed and processed forms. The presence of the processed NifS form in the eluate indicated that copurification had occurred. [Figure 18] Western blot of NifU purified from N. benthamiana in the third purification experiment, showing that NifS copurifies with NifU. Panel A) Schematic of the construct infiltrated into N. benthamiana (not drawn to scale). B) Western blot analysis of the purified product. P = pellet, S = supernatant, FT = flow-through, E = eluate. All samples were loaded in duplicate and immunodetected with strep antibody (α-strep) or HA antibody (α-HA). C) Coomassie staining of the eluate showing a major band for NifU and a weaker band for NifS. [Figure 19]Western blot showing that purification of mitochondrially targeted MTP-FAγ51::NifS::TS resulted in co-purification of scar9::GG::NifU::HA. Samples from steps (ii)–(v) were subjected to SDS-PAGE and Western blotting using either an anti-Strep antibody to detect the NifS polypeptide or an anti-HA antibody to detect the NifU polypeptide. The two bands for NifS correspond to the unprocessed and processed forms. The presence of the processed NifU form in the eluate indicated that co-purification had occurred. [Figure 20]ClustalW alignment of the first 300 amino acid residues of NifV / HCS-like amino acid sequences selected in this study along with translations of N. benthamiana P72026 (SEQ ID NO: 221) and P20586 (SEQ ID NO: 222), K. oxytoca NifV (SEQ ID NO: 13), Lotus japonicus FEN1 (SEQ ID NO: 215), and Mycobacterium tuberculosis α-isopropylmaleate synthase (MtLeuA, SEQ ID NO: 223). Other HCS sequences are from Thermoanaerobacter brockii (TbHCS; SEQ ID NO:206), Thermincola potens (TpHCS; SEQ ID NO:207), Saccharomyces cerevisiae (ScHCS; SEQ ID NO:208), Nodularia spumigena (NsHCS; SEQ ID NO:209), Methanosarcina acetivorans (MaHCS; SEQ ID NO:210), Chlorobaculum tepidum (CtHCS; SEQ ID NO:211), and Methanocaldococcus infernus (MiHCS1, SEQ ID NO:212; MiHCS2, SEQ ID NO:213; MiHCS, SEQ ID NO:214). Conserved residues in the active site of LeuA are identified by *. Four amino acid residues at positions R81, D82, H291, and H293 harbor Zn2+, and two amino acid residues, E224 and T260, form the substrate-binding pocket of MtLeuA together with the native Zn2+ (Koon et al., 2004). [Figure 21] Western blot analysis of the total, insoluble, and soluble fractions of the NifV / HCS-like fusion polypeptide (MTP-FAγ51::HA::NifV / HCS) after expression in N. benthamiana leaves using an anti-HA antibody. T, total protein; I, insoluble (pellet) fraction of total protein; S, soluble (supernatant) fraction of total protein. m, mitochondrial-targeted polypeptide; c, cytoplasmic-targeted polypeptide. [Figure 22]Western blot analysis using anti-HA antibody of total, insoluble, and soluble fractions of a cytoplasmically localized NifV / HCS-like fusion polypeptide (HA::NifV / HCS) (used as a comparison standard with the corresponding mitochondrially localized fusion polypeptide) after expression in N. benthamiana leaves. T, total protein; I, insoluble (pellet) fraction of total protein; S, soluble (supernatant) fraction of total protein. c, cytoplasmically targeted polypeptide; m, mitochondrially targeted polypeptide. [Figure 23] Homocitrate target ion peak area after baseline subtraction (log10 scale) [Figure 24] Western blot analysis of soluble NifH fusion polypeptides in a transient leaf expression system in N. benthamiana leaves using an anti-Strep antibody to detect polypeptides bearing the TwinStrep epitope. All NifH gene constructs were co-infiltrated with SN44, encoding a NifM fusion polypeptide from K. oxytoca. Protein samples were prepared under aerobic conditions. [Figure 25] Western blot showing the results of purifying the NifH fusion polypeptide encoded by SL6 in stably transformed tobacco. The NifH gene encoded the MTP-CoxIV::TwinStrep::KoNifH::HA fusion polypeptide. Five-microliter samples from various stages of the purification process were analyzed by Western blot and probed with antibodies recognizing the Strep or HA epitope. Samples from the total, insoluble, and soluble fractions are shown above the lanes. The white arrowhead indicates the unprocessed NifH polypeptide, and the black arrowhead indicates the processed form. [Figure 26]Western blot analysis of the expression and processing of Anf fusion polypeptides after transient introduction of gene constructs into N. benthamiana leaves. The blot had three sets of adjacent lanes for (from left to right) the AnfD, AnfK, AnfH, and AnfG fusion polypeptides. Each set contained the fusion polypeptide under investigation, MTP-FAγ51::HA::Anf, and two control polypeptides, HA::Anf and mFAγ51::HA::Anf, as molecular weight markers. L, molecular weight marker ladder (kDa). [Figure 27] Western blots showing the expression and processing of all four AnfD, AnfK, AnfH, and AnfG fusion polypeptides when expressed from multigene constructs in N. benthamiana leaves. A. Western blot analysis of the mitochondrially targeted AnfD, AnfK, AnfH, and AnfG fusion polypeptides expressed from SL26 and the unprocessed polypeptides expressed from SL31 (detected in total protein extracts from transient leaf assays). B. Western blot analysis of proteins obtained from expression of the mitochondrially targeted AnfD, AnfK, AnfH, and AnfG fusion polypeptides from SL26 and the unprocessed polypeptides from SL31. C. Western blot showing the expression and processing of fusion polypeptides from the multigene constructs SL26, SL27, and SL28, the single-gene construct SL29, and a mixture (Mix) of the four single-gene constructs SN161, SN129, SN130, and SN131. AnfK, when present, showed an upper unprocessed band and a lower processed band. [Figure 28]Western blots showing the solubility of individual Anf polypeptides expressed from single-gene vectors in N. benthamiana leaf cells when localized to the cytoplasm or mitochondria. Upper panel, soluble fractions of AnfD, AnfK, AnfH, and AnfG fusion polypeptides; lower panel, insoluble fractions of AnfD, AnfK, AnfH, and AnfG fusion polypeptides. C, cytoplasmic localization; M, mitochondrial localization; A, alanine-substituted mMTP-FAγ51. Black arrowheads indicate the position of the MPP-cleaved protein, and white arrowheads indicate the unprocessed polypeptide. See Table 20 for the predicted molecular weights of each Anf in the unprocessed and MPP-processed polypeptides. [Figure 29] Homology model of the AnfDKHG complex with Fe-nitrogenase based on the A. vinelandii Anf amino acid sequence with a linker joining the AnfD and AnfK polypeptides. Initial coordination before 20 nanoseconds of simulation. Predicted structure of the AnfD::linker::AnfK polypeptide with a 16 amino acid linker complexed with an AnfH dimer and AnfG. The AnfH dimer is designated AnfHH. [Figure 30] Western blot analysis of total protein extracts from N. benthamiana leaves infiltrated with genetic constructs for expression of fused or separate AnfD and AnfK polypeptides. Blots were probed with an HA antibody. Expression of the AnfD-linker-AnfK fusion polypeptide from vectors SN272–SN275 was compared with expression from separate genes on vectors SL26 and SL28. SN161 and SN129 served as controls for the individual expression of AnfD and AnfK, respectively. [Figure 31]Western blot analysis of (A) soluble and (B) insoluble protein fractions from N. benthamiana leaves infiltrated with gene constructs for expression of the AnfD and AnfK genes. SN272–SN275 each encoded an AnfD-linker-AnfK fusion polypeptide, whereas SL26 and SL28 expressed separate polypeptides. [Figure 32] Western blot analysis of polypeptides produced by SL42 in N. benthamiana leaves, including total (T), insoluble (I), and soluble (S) fractions, using anti-HA (Panel A) or anti-Strep (Panel B) antibodies for detection. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. Panel B, probed with anti-Strep, shows the processed NifB polypeptide. [Figure 33] Western blot analysis of polypeptides produced by SL43 in N. benthamiana leaves, including total (T), insoluble (I), and soluble (S) fractions, using anti-HA (Panel A) or anti-Strep (Panel B) antibodies for detection. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. Panel B, probed with anti-Strep, shows the processed AnfK polypeptide. [Figure 34]Western blot analysis of polypeptides (including total (T), insoluble (I), and soluble (S) fractions) produced by SL42 and SL43 transfected collectively into N. benthamiana leaves using anti-HA antibody (Panel A) or anti-Strep antibody (Panel B). Numbers on the sides of panels A and B indicate the molecular weight (kDa) of the marker in the first lane. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in the mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. [Figure 35] Western blot analysis of polypeptides produced by SL48 in N. benthamiana leaves, including total (T), insoluble (I), and soluble (S) fractions, using anti-HA (Panel A) or anti-Strep (Panel B) antibodies for detection. Numbers at the sides of panels A and B indicate the molecular weight (kDa) of the marker in the first lane. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. Panel B, probed with anti-Strep, shows the processed NifB polypeptide. [Figure 36] Western blot analysis of polypeptides produced by SL49 in N. benthamiana leaves, including total (T), insoluble (I), and soluble (S) fractions, using anti-HA (Panel A) or anti-Strep (Panel B) antibodies for detection. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. Panel B, probed with anti-Strep, shows the processed AnfK polypeptide. [Figure 37]Western blot analysis of polypeptides (including total (T), insoluble (I), and soluble (S) fractions) produced by SL48 and SL49 transfected together into N. benthamiana leaves using anti-HA antibody (Panel A) or anti-Strep antibody (Panel B). The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in the mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. [Figure 38] Western blot analysis of polypeptides (including total, panel A), insoluble, panel B), and soluble, panel C) fractions produced by SN292, SN291, SN299, and SN300 in N. benthamiana leaves using anti-HA for detection. The numbers on the sides indicate the molecular weight (kDa) of the marker in the first lane. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in the mitochondria, the white arrowhead indicates the unprocessed polypeptide band, and the * indicates a potential dimer of FdxN protein. [Figure 39] Western blot analysis of polypeptides (including total (Panel A), insoluble (Panel B), and soluble (Panel C) fractions) produced by SN192, SL50, and SL54 individually transfected and SL50 and SL54 transfected together into N. benthamiana leaves using anti-HA for detection. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in the mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. [Figure 40] Western blot analysis of polypeptides produced by SL50 in N. benthamiana leaves (including total, panel A), insoluble, panel B), and soluble, panel C) fractions, using anti-HA for detection. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. [Figure 41] Western blot analysis of polypeptides (including total, panel A), insoluble, panel B), and soluble, panel C) fractions produced by SL50 and SL49 in N. benthamiana leaves using anti-HA for detection. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in the mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. [Figure 42] Western blot analysis of polypeptides produced by SL47 and SL55, separately or in combination, in N. benthamiana leaves using anti-HA for detection. The first lane shows molecular weight (kDa) markers. The black arrowhead indicates the position of the processed polypeptide band after cleavage by MPP in the mitochondria, and the white arrowhead indicates the unprocessed polypeptide band. [Figure 43] Western blot of proteins extracted from leaf samples of transgenic Arabidopsis plants transformed with SL49 and probed with anti-HA antibody. The positions of NifJ, NifB, NifU, and NifF fusion polypeptides are indicated by arrows, based on the positions of the same polypeptides (Benth control) after transient expression of SL49 in N. benthamiana leaves.

[0365] Array List Key

[0366] SEQ ID NO: 1: Amino acid sequence of the NifH polypeptide of K. oxyyoca, 293 amino acids.

[0367] SEQ ID NO: 2: Amino acid sequence of the wild-type NifD polypeptide from K. oxyyoca according to accession number X13303.1, 483 amino acids (Temme sequence is SEQ ID NO: 18).

[0368] SEQ ID NO: 3: Amino acid sequence of the NifK polypeptide from K. oxyyoca according to Temme et al. (2012); 520 amino acids.

[0369] SEQ ID NO: 4: Amino acid sequence of the NifB polypeptide from K. oxyyoca; 468 amino acids.

[0370] SEQ ID NO: 5: Amino acid sequence of the NifE polypeptide from K. oxyyoca; 457 amino acids.

[0371] SEQ ID NO: 6: Amino acid sequence of the NifF polypeptide from K. oxyyoca; 176 amino acids; NCBI accession number X03214.

[0372] SEQ ID NO: 7: Amino acid sequence of the NifJ polypeptide from K. oxyyoca; 1171 amino acids; NCBI accession number 43862; Cannon et al., 1988 Nucleic Acids Res. 16:11379).

[0373] SEQ ID NO: 8: Amino acid sequence of the NifM polypeptide from K. oxyyoca; 266 amino acids; NCBI accession number X05887; Paul and Merrick (1987).

[0374] SEQ ID NO: 9: Amino acid sequence of the NifN polypeptide from K. oxyyoca, NCBI accession number P08738; 461 amino acids; (Arnold et al., 1988). This sequence is identical to K. michiganensis sequence accession number WP_064371582 and is 85% identical to the sequence annotated as K. oxyyoca NifN, accession number WP_061153953.

[0375] SEQ ID NO: 10: Amino acid sequence of the NifQ polypeptide from Klebsiella. NCBI accession number WP_004138772. This sequence is 95% identical to another K. oxytoca sequence annotated as NifQ, accession number AAA25108.1.

[0376] SEQ ID NO: 11: Amino acid sequence of the NifS polypeptide from K. oxyyoca; 400 amino acids.

[0377] SEQ ID NO: 12: Amino acid sequence of the NifU polypeptide from K. oxyyoca; 274 amino acids. NCBI accession number P05343.2 (Arnold et al., 1988). This sequence is identical to accession number WP_004138782 and is also identical to another K. oxyyoca sequence, accession number AAA25155, at positions 272 / 273.

[0378] SEQ ID NO: 13: Amino acid sequence of the NifV polypeptide from K. oxyyoca; 381 amino acids. NCBI accession number CAA31119.1 (Arnold et al., 1988).

[0379] SEQ ID NO: 14: Amino acid sequence of the NifX polypeptide from K. oxyyoca; 156 amino acids (accession number P09136).

[0380] SEQ ID NO: 15: Amino acid sequence of the NifY polypeptide from K. oxyyoca; 220 amino acids. NCBI accession number CAA31670 (Arnold et al., 1988)

[0381] SEQ ID NO: 16: Amino acid sequence of the NifZ polypeptide from K. oxyyoca; 148 amino acids. NCBI accession number P0A3U2 (Arnold et al., 1988)

[0382] SEQ ID NO: 17: Amino acid sequence of the NifW polypeptide from K. oxyyoca.

[0383] SEQ ID NO: 18: Amino acid sequence of wild-type K. oxyyoca NifD according to Temme et al. (2012).

[0384] SEQ ID NO: 19: Amino acid sequence of wild-type K. oxyyoca NifS according to Temme et al. (2012).

[0385] SEQ ID NO: 20: Amino acid sequence of the N-terminal extension comprising MTP-FAγ77 (amino acids 1-77) and the amino acid triplet GAP (78-80). Cleavage by MPP occurs between amino acid residues 42 and 43.

[0386] SEQ ID NO: 21: Amino acid sequence of the MTP-FAγ51 polypeptide with an additional N-terminal Met and C-terminal GG. Cleavage by MPP occurs between amino acid residues 43 and 44.

[0387] SEQ ID NO: 22: Amino acid sequence of the FAγ-scar9 polypeptide.

[0388] SEQ ID NO: 23: Amino acid sequence of the MTP-FAγ77::NifH::HA fusion polypeptide encoded by pRA10. Amino acids 1-77 correspond to MTP-FAγ77, amino acids 78-80 are GAP, amino acids 81-372 correspond to K. oxyyoca NifH amino acids (SEQ ID NO: 1 without the initiating Met), and amino acids 373-389 contain the HA epitope.

[0389] SEQ ID NO: 24: Amino acid sequence of the MTP-FAγ51::NifH::HA fusion polypeptide encoded by pRA34. Amino acids 1-51 correspond to MTP-FAγ51, amino acids 52-54 are GAP, amino acids 55-346 correspond to K. oxyyoca NifH amino acids (SEQ ID NO: 1 without the initiating Met), and amino acids 347-363 contain the HA epitope.

[0390] SEQ ID NO: 25: Amino acid sequence of the MTP-FAγ51::NifH::HA fusion polypeptide encoded by SN18. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-347 correspond to K. oxyyoca NifH amino acids (SEQ ID NO: 1), and amino acids 348-358 contain the HA epitope.

[0391] SEQ ID NO: 26: Amino acid sequence of the MTP-FAγ51::NifH::HA fusion polypeptide encoded by SN29. Amino acids 1-53 correspond to MTP-FAγ51 with GG, amino acids 54-64 contain the HA epitope, amino acids 65-357 correspond to K. oxyyoca NifH amino acids (SEQ ID NO: 1), and amino acids 358-371 are the C-terminal extension.

[0392] SEQ ID NO: 27: 6xHis used in place of the MTP sequence, with N-terminal Met and C-terminal GG.

[0393] SEQ ID NO: 28: Amino acid sequence of CPN60 MTP.

[0394] SEQ ID NO: 29: Amino acid sequence of CPN60 / GG linker-less MTP.

[0395] SEQ ID NO: 30: Amino acid sequence of superoxide dismutase (SOD) MTP.

[0396] SEQ ID NO: 31: Amino acid sequence of dual superoxide dismutase (2SOD) MTP.

[0397] SEQ ID NO: 32: Amino acid sequence of modified superoxide dismutase (SODmod) MTP.

[0398] SEQ ID NO: 33: Amino acid sequence of modified dual superoxide dismutase (2SODmod) MTP.

[0399] SEQ ID NO: 34: Amino acid sequence of L29 MTP (At1G07830).

[0400] SEQ ID NO: 35: Amino acid sequence of Neurospora crassa F0 ATPase subunit 9 (SU9) MTP.

[0401] SEQ ID NO: 36: Amino acid sequence of gATPase gamma subunit (FAγ51) MTP without the additional N-terminal Met (SEQ ID NO: 21 has the additional N-terminal Met). Cleavage by MPP occurs between amino acid residues 42 and 43.

[0402] SEQ ID NO: 37: Amino acid sequence of CoxIV twin strep (ABM97483) MTP.

[0403] SEQ ID NO: 38: Amino acid sequence of CoxIV 10xHis(ABM97483) MTP.

[0404] SEQ ID NO: 39: Amino acid sequences of the scar predicted for superoxide dismutase (SOD) MTP with GG and dual superoxide dismutase (2SOD) MTP with GG.

[0405] SEQ ID NO: 40: Amino acid sequence of the scar predicted for L29 MTP with GG.

[0406] SEQ ID NO: 41: Amino acid sequence of the predicted scar of Neurospora crassa F0 ATPase subunit 9 (SU9) MTP with GG.

[0407] SEQ ID NO: 42: Amino acid sequence of the predicted scar of gATPase gamma subunit (FAγ51) MTP with GG.

[0408] SEQ ID NO: 43: Amino acid sequence of the scar predicted by CoxIV twin strep MTP with GG.

[0409] SEQ ID NO: 44: Amino acid sequence of the scar predicted by CoxIV 10xHis MTP with GG.

[0410] SEQ ID NO: 45: Oligonucleotide primer MIT_V2.1_SbfInifH_FW2.

[0411] SEQ ID NO: 46: Oligonucleotide primer MIT_V2.1_SbfInifJ_RV2.

[0412] SEQ ID NO: 47: Oligonucleotide primer MIT_V2.1_SbfInifB_FW.

[0413] SEQ ID NO: 48: Oligonucleotide primer MIT_V2.1_SbfIori_RV.

[0414] SEQ ID NO: 49: Amino acid sequence of mscar9 from MTP-FAγ51 in which the N-terminal Ile residue is replaced with Met for translation initiation.

[0415] SEQ ID NO: 50: Tryptic peptide.

[0416] SEQ ID NO: 51: Amino acid sequence of the MTP-FAγ9 scar without the N-terminal Met and with the C-terminal Met.

[0417] SEQ ID NOs: 52 to 54: oligonucleotide primers.

[0418] SEQ ID NO: 55: Tryptic peptide.

[0419] SEQ ID NO: 56: Tryptic peptide.

[0420] SEQ ID NO: 57: Amino acid sequence of the MTP-FAγ77::NifK fusion polypeptide (pRA25) lacking the C-terminal extension. Amino acids 1-77 correspond to MTP-FAγ77, amino acids 78-80 are GAP, and amino acids 81-599 correspond to K. oxyyoca NifK lacking the initiating Met.

[0421] SEQ ID NO: 58: Amino acid sequence of the last four amino acid residues at the C-terminus of the NifK polypeptide from K. oxyyoca.

[0422] SEQ ID NO: 59: Amino acid sequence of a mutant MTP-FA γ51 polypeptide that is not cleaved by MPP.

[0423] SEQ ID NOs: 60-107: peptide sequences.

[0424] SEQ ID NOs: 108 to 113: oligonucleotide primers.

[0425] SEQ ID NO: 114: Amino acid sequence of an 11-residue section from the linker region from Hypocrea jecorina cellobiohydrolase II (Accession number AAG39980.1).

[0426] SEQ ID NO: 115: Amino acid sequence of the 9 residue HA epitope.

[0427] SEQ ID NO: 116: Amino acid sequence of the linker for the NifD::linker::NifK fusion polypeptide. This linker is 30 residues in length and has SEQ ID NO: 114 with the final arginine substituted with alanine followed by a 9-residue HA epitope (SEQ ID NO: 115) followed by another copy of SEQ ID NO: 114 with the arginine substituted with alanine.

[0428] SEQ ID NO: 117: Oligonucleotide primer.

[0429] SEQ ID NO: 118: Oligonucleotide primer.

[0430] SEQ ID NO: 119: Scar peptide sequence.

[0431] SEQ ID NO: 120: Scar peptide sequence.

[0432] SEQ ID NO: 121: Amino acid sequence of the metaxin fusion polypeptide encoded by construct SN197. The TwinStrep epitope corresponds to amino acids 1-31, mTurquoise to amino acids 32-273, the TEV cleavage site to amino acids 274-282, and the metaxin sequence to amino acids 283-603.

[0433] SEQ ID NO: 122: Amino acid sequence of the MTP-FAγ51::NifD::HA fusion polypeptide encoded by SN10. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-536 correspond to K. oxyyoca NifD (SEQ ID NO: 18) with an initiating Met, and amino acids 537-547 contain the HA epitope.

[0434] SEQ ID NO: 123: Amino acid sequence of the MTP-FAγ51::NifM::HA fusion polypeptide encoded by SN30. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-320 correspond to K. oxyyoca NifM (SEQ ID NO: 8) with an initiating Met, and amino acids 321-331 contain the HA epitope.

[0435] SEQ ID NO: 124: Amino acid sequence of the MTP-FAγ51::NifS::HA fusion polypeptide encoded by SN31. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-454 correspond to K. oxyyoca NifS (SEQ ID NO: 19) with an initiating Met according to Temme et al. (2012), and amino acids 455-465 contain the HA epitope.

[0436] SEQ ID NO: 125: Amino acid sequence of the MTP-FAγ51::NifU::HA fusion polypeptide encoded by SN32. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-328 correspond to K. oxyyoca NifU (SEQ ID NO: 12) with an initiating Met, and amino acids 329-339 contain the HA epitope.

[0437] SEQ ID NO: 126: Amino acid sequence of the MTP-FAγ51::NifE::HA fusion polypeptide encoded by SN38. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-511 correspond to K. oxyyoca NifE with an initiating Met according to Temme et al. (2012), and amino acids 512-522 contain the HA epitope.

[0438] SEQ ID NO: 127: Amino acid sequence of the MTP-FAγ51::NifN::HA fusion polypeptide encoded by SN39. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-515 correspond to K. oxyyoca NifN (SEQ ID NO: 9) with an initiating Met, and amino acids 516-526 contain the HA epitope.

[0439] SEQ ID NO: 128: Amino acid sequence of the MTP-CoxIV-Twin-Strep::NifH::HA fusion polypeptide encoded by SN42. Amino acids 1-61 correspond to MTP-CoxIV-Twin-Strep with a C-terminal GG, amino acids 62-354 correspond to K. oxyyoca NifH (SEQ ID NO: 1) with an initiator Met, and amino acids 355-365 contain the HA epitope.

[0440] SEQ ID NO: 129: Amino acid sequence of the MTP-Su9::NifK fusion polypeptide encoded by SN46. Amino acids 1-70 correspond to MTP-Su9 with a C-terminal GG, and amino acids 71-590 correspond to K. oxyyoca NifK (SEQ ID NO: 3) with an initiator Met.

[0441] SEQ ID NO: 130: Amino acid sequence of the MTP-L29::NifV::HA fusion polypeptide encoded by SN51. Amino acids 1-34 correspond to MTP-L29 with a C-terminal GG, amino acids 35-415 correspond to K. oxyyoca NifV (SEQ ID NO: 13) with an initiating Met, and amino acids 416-426 contain the HA epitope.

[0442] SEQ ID NO: 131: Amino acid sequence of the MTP-FAγ51::NifD::linker(HA)::NifK fusion polypeptide encoded by SN68. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminus of GG, amino acids 55-536 correspond to wild-type K. oxyyoca NifD amino acids (SEQ ID NO: 18 without the N-terminal Met), amino acids 537-566 correspond to the linker containing the HA epitope, and amino acids 567-1085 correspond to NifK with a wild-type C-terminus without the N-terminal Met (SEQ ID NO: 3).

[0443] SEQ ID NO: 132: Amino acid sequence of the MTP-FAγ51::HA::NifD::HA fusion polypeptide encoded by SN75. Amino acids 1-53 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 54-64 correspond to the first HA epitope, amino acids 65-546 correspond to wild-type K. oxyyoca NifD amino acids (SEQ ID NO: 18), and amino acids 547-557 contain the HA epitope.

[0444] SEQ ID NO: 133: Amino acid sequence of the MTP-FAγ51::NifD::HA fusion polypeptide encoded by SN99. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-536 correspond to K. oxyyoca NifD containing an alanine substitution mutation at amino acids 148-152, and amino acids 537-547 contain the HA epitope.

[0445] SEQ ID NO: 134: Amino acid sequence of the MTP-FAγ51::NifD::HA fusion polypeptide encoded by SN100. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-536 correspond to K. oxyyoca NifD amino acids containing an alanine substitution mutation at amino acids 153-157, and amino acids 537-547 contain the HA epitope.

[0446] SEQ ID NO: 135: Amino acid sequence of the MTP-Su9::NifW fusion polypeptide encoded by SN104. Amino acids 1-70 correspond to MTP-Su9 with a C-terminal GG, amino acids 71-158 correspond to K. oxyyoca NifW (SEQ ID NO: 17) with an initiator Met, and amino acids 159-167 contain the HA epitope.

[0447] SEQ ID NO: 136: Amino acid sequence of the MTP-FAγ51::NifD::HA fusion polypeptide encoded by SN114. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-536 correspond to K. oxyyoca NifD containing a Y100Q substitution mutation at amino acid 154, and amino acids 537-547 contain the HA epitope.

[0448] SEQ ID NO: 137: Amino acid sequence of the MTP-FAγ51::NifF::HA fusion polypeptide encoded by SN138. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-230 correspond to K. oxyyoca NifF (SEQ ID NO: 6), and amino acids 231-241 contain the HA epitope.

[0449] SEQ ID NO: 138: Amino acid sequence of the MTP-FAγ51::NifJ::HA fusion polypeptide encoded by SN139. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-1225 correspond to K. oxyyoca NifJ (SEQ ID NO: 7), and amino acids 1226-1236 contain the HA epitope.

[0450] SEQ ID NO: 139: Amino acid sequence of the MTP-FAγ51::HA::NifK fusion polypeptide encoded by SN140. Amino acids 1-53 correspond to MTP-FAγ51 with GG, amino acids 54-64 contain the HA epitope, and amino acids 65-584 correspond to K. oxyyoca NifK with the wild-type C-terminus (SEQ ID NO: 3).

[0451] SEQ ID NO: 140: Amino acid sequence of the MTP-FAγ51::NifQ::HA fusion polypeptide encoded by SN141. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-221 correspond to K. oxyyoca NifQ (SEQ ID NO: 10), and amino acids 222-232 contain the HA epitope.

[0452] SEQ ID NO: 141: Amino acid sequence of the MTP-FAγ51::NifV::HA fusion polypeptide encoded by SN142. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-435 correspond to K. oxyyoca NifV (SEQ ID NO: 13), and amino acids 436-446 contain the HA epitope.

[0453] SEQ ID NO: 142: Amino acid sequence of the MTP-FAγ51::NifW::HA fusion polypeptide encoded by SN143. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-140 correspond to K. oxyyoca NifW (SEQ ID NO: 17), and amino acids 141-151 contain the HA epitope.

[0454] SEQ ID NO: 143: Amino acid sequence of the MTP-FAγ51::NifX::HA fusion polypeptide encoded by SN144. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-210 correspond to K. oxyyoca NifX (SEQ ID NO: 14), and amino acids 211-221 contain the HA epitope.

[0455] SEQ ID NO: 144: Amino acid sequence of the MTP-FAγ51::NifY::HA fusion polypeptide encoded by SN145. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-274 correspond to K. oxyyoca NifY according to Temme et al. (2012), and amino acids 275-285 contain the HA epitope.

[0456] SEQ ID NO: 145: Amino acid sequence of the MTP-FAγ51::NifZ::HA fusion polypeptide encoded by SN146. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-202 correspond to K. oxyyoca NifZ (SEQ ID NO: 16), and amino acids 203-213 contain the HA epitope.

[0457] SEQ ID NO:146: Amino acid sequence of the MTP-FAγ51::NifD(Y100Q)::linker(HA)::NifK fusion polypeptide encoded by SN159. Amino acids 1-54 correspond to MTP-FAγ51 with a C-terminal GG, amino acids 55-536 correspond to K. oxyyoca NifD with a Y100Q substitution, amino acids 537-566 correspond to the linker containing the HA epitope, and amino acids 567-1085 correspond to NifK (SEQ ID NO:3) lacking the N-terminal Met and with a wild-type C-terminus.

[0458] SEQ ID NO: 147: Amino acid sequence of the MTP-FAγ51::NifB::HA fusion polypeptide encoded by SN192. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-522 correspond to K. oxyyoca NifB according to Temme et al. (2012), and amino acids 523-533 contain the HA epitope.

[0459] SEQ ID NO: 148: Amino acid sequence of wild-type Azospirillum brasilense NifD polypeptide, UniProt A0A060DN91; 479 amino acids.

[0460] SEQ ID NO: 149: Amino acid sequence of wild-type Azotobacter vinelandii NifD polypeptide, UniProt C1DGZ7; 492 amino acids.

[0461] SEQ ID NO: 150: Amino acid sequence of wild-type Sinorhizobium fredii NifD polypeptide, 504 amino acids.

[0462] SEQ ID NO: 151: Amino acid sequence of wild-type Chlorobium tepidum NifD polypeptide, Uniprot Q8KC89; 543 amino acids.

[0463] SEQ ID NO: 152: Amino acid sequence of wild-type Desulfovibrio vulgaris NifD polypeptide, Uniprot B8DR77; 544 amino acids.

[0464] SEQ ID NO: 153: Amino acid sequence of wild-type Desulfotomaculum ferrireducens NifD polypeptide, 539 amino acids.

[0465] SEQ ID NO: 154: Peptide sequence (wherein X is any amino acid except Tyr).

[0466] SEQ ID NO: 155: Tryptic peptide from NifM.

[0467] SEQ ID NO: 156: Tryptic peptide from NifM.

[0468] SEQ ID NO: 157: Tryptic peptide from CAT.

[0469] SEQ ID NO: 158: Tryptic peptide from CAT.

[0470] SEQ ID NO: 159: Tryptic peptide from CAT.

[0471] SEQ ID NO: 160: Amino acid sequence of the MTP-FAγ51::NifU::TwinStrep fusion polypeptide encoded by SN166. Amino acids 1-54 are the MTP-FAγ51 sequence with an additional translation initiation methionine and C-terminal GG, amino acids 55-328 are the NifU sequence, and amino acids 329-358 are the sequence containing the Twinstrep motif.

[0472] SEQ ID NO: 161: Amino acid sequence of the MTP-FAγ51::NifS::TwinStrep fusion polypeptide encoded by SN231. Amino acids 1-54 are the MTP-FAγ51 sequence with an additional translation initiation methionine and the C-terminal GG, amino acids 55-454 are the NifS sequence, and amino acids 455-484 are the sequence containing the Twinstrep motif.

[0473] SEQ ID NO: 162: Tryptic peptide sequence from scar9

[0474] SEQ ID NO: 163: Amino acid sequence of the NifV polypeptide from A. vinelandii (AvNifV; accession number WP_012698855).

[0475] SEQ ID NO: 164: Amino acid sequence of KoNifV variant sequence (accession number WP_004138778).

[0476] SEQ ID NO: 165: N-terminal ScHCS extension (scar sequence).

[0477] SEQ ID NO: 166: N-terminal AvNifV extension (scar sequence).

[0478] SEQ ID NO: 167: Amino acid sequence of the MTP-FAγ51::HA::KoNifM polypeptide encoded by SN43. Amino acids 1-53 correspond to the MTP-FAγ51 sequence with a C-terminal GG, amino acids 54-64 correspond to the HA epitope with a C-terminal GG, and amino acids 65-330 correspond to the NifM sequence from K. oxyyoca.

[0479] SEQ ID NO: 168: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN178. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-354 correspond to the NifH sequence from Azospirillum brasilense (Accession No. WP_014239786).

[0480] SEQ ID NO: 169: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN179. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-356 correspond to the NifH sequence from Mastigocladus laminosus (accession number WP_016865872).

[0481] SEQ ID NO: 170: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN180. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-348 correspond to the NifH sequence from Frankia casurinae (accession number WP_0011438842).

[0482] SEQ ID NO: 171: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN181. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-354 correspond to the NifH sequence from Marichromatium gracile biovar thermosufidiphilum (accession number WP_062275270).

[0483] SEQ ID NO: 172: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN182. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-345 correspond to the NifH sequence from Methanocaldococcus infernus (accession number WP_013099459).

[0484] SEQ ID NO: 173: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN183. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-345 correspond to the NifH sequence from Heliobacterium modesticaldum (accession number WP_012282218).

[0485] SEQ ID NO: 174: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN184. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-335 correspond to the NifH sequence from Chlorobium tepidum (accession number WP_010933198).

[0486] SEQ ID NO: 175: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN185. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-350 correspond to the NifH sequence from Geobacter sp. M21 (accession number WP_015837436).

[0487] SEQ ID NO: 176: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN186. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-355 correspond to the NifH sequence from Bradyrhizobium diazoefficans (accession number AHY57040).

[0488] SEQ ID NO: 177: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN187. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-336 correspond to the NifH sequence from Methanobacterium thermoautotrophicum (accession number AAB86034).

[0489] SEQ ID NO: 178: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN188. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-334 correspond to the NifH sequence from Methanosarcina (accession number WP_048121466).

[0490] SEQ ID NO: 179: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN189. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-336 correspond to the NifH sequence from Desulfotomaculum acetoxidans (accession number WP_015756624).

[0491] SEQ ID NO: 180: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN190. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-336 correspond to the NifH sequence from Carboxydothermus pertinax (Accession No. WP_075859892).

[0492] SEQ ID NO: 181: Amino acid sequence of the MTP-CoxIV::TwinStrep::NifH polypeptide encoded by SN191. Amino acids 1-31 correspond to the MTP-CoxIV sequence, amino acids 32-61 correspond to the TwinStrep sequence containing a C-terminal GG, and amino acids 62-335 correspond to the NifH sequence from Nostoc calcicole (accession number WP_073644321).

[0493] SEQ ID NO: 182: Amino acid sequence of the MTP-FAγ51::AnfD::HA polypeptide encoded by SN81. Amino acids 1-54 correspond to the MTP-FAγ51 sequence including the C-terminal GG linker, amino acids 55-572 correspond to the AnfD sequence from A. vinelandii, and amino acids 573-583 correspond to the HA epitope.

[0494] SEQ ID NO: 183: Amino acid sequence of the HA::AnfD polypeptide encoded by SN82. Amino acids 1-12 correspond to the HA epitope with a C-terminal GG linker, and amino acids 13-530 correspond to the AnfD sequence from A. vinelandii.

[0495] SEQ ID NO: 184: Amino acid sequence of the MTP-FAγ51::HA::AnfK polypeptide encoded by SN129. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including the C-terminal GG linker, amino acids 54-64 correspond to the HA epitope, and amino acids 65-526 correspond to the AnfK sequence from A. vinelandii.

[0496] SEQ ID NO: 185: Amino acid sequence of the MTP-FAγ51::HA::AnfH polypeptide encoded by SN130. Amino acids 1-53 correspond to the MTP-FAγ51 sequence with a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-339 correspond to the AnfH sequence from A. vinelandii.

[0497] SEQ ID NO: 186: Amino acid sequence of the MTP-FAγ51::HA::AnfG polypeptide encoded by SN131. Amino acids 1-53 correspond to the MTP-FAγ51 sequence with a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-196 correspond to the AnfG sequence from A. vinelandii.

[0498] SEQ ID NO: 187: Amino acid sequence of the HA::AnfK polypeptide encoded by SN152. Amino acids 1-12 correspond to the HA epitope with a C-terminal GG linker, and amino acids 13-474 correspond to the AnfK sequence from A. vinelandii.

[0499] SEQ ID NO: 188: Amino acid sequence of the HA::AnfH polypeptide encoded by SN153. Amino acids 1-12 correspond to the HA epitope with a C-terminal GG linker, and amino acids 13-287 correspond to the AnfH sequence from A. vinelandii.

[0500] SEQ ID NO: 189: Amino acid sequence of the HA::AnfG polypeptide encoded by SN154. Amino acids 1-12 correspond to the HA epitope with a C-terminal GG linker, and amino acids 13-144 correspond to the AnfG sequence from A. vinelandii.

[0501] SEQ ID NO: 190: Amino acid sequence of the mFAγ51::HA::AnfK polypeptide encoded by SN155. Amino acids 1-53 correspond to the mutant mFAγ51 sequence containing a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-526 correspond to the AnfK sequence from A. vinelandii.

[0502] SEQ ID NO: 191: Amino acid sequence of the mFAγ51::HA::AnfH polypeptide encoded by SN156. Amino acids 1-53 correspond to the mutant mFAγ51 sequence containing a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-339 correspond to the AnfH sequence from A. vinelandii.

[0503] SEQ ID NO: 192: Amino acid sequence of the mFAγ51::HA::AnfG polypeptide encoded by SN157. Amino acids 1-53 correspond to the mutant mFAγ51 sequence containing a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-196 correspond to the AnfG sequence from A. vinelandii.

[0504] SEQ ID NO: 193: Amino acid sequence of the mFAγ51::HA::AnfD polypeptide encoded by SN158. Amino acids 1-53 correspond to the mutant mFAγ51 sequence containing a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-582 correspond to the AnfD sequence from A. vinelandii.

[0505] SEQ ID NO: 194: Amino acid sequence of the MTP-FAγ51::HA::AnfD polypeptide encoded by SN161. Amino acids 1-53 correspond to the MTP-FAγ51 sequence with a C-terminal GG linker, amino acids 54-64 correspond to the HA epitope with a C-terminal GG linker, and amino acids 65-582 correspond to the AnfD sequence from A. vinelandii.

[0506] SEQ ID NO: 195: Amino acid sequence of the MTP-FAγ51::AnfD::Twin Strep polypeptide encoded by SN177. Amino acids 1-54 correspond to the cMTP-FAγ51 sequence including the C-terminal GG linker, amino acids 55-572 correspond to the AnfD sequence from A. vinelandii, and amino acids 573-604 correspond to the Twin Strep epitope.

[0507] SEQ ID NO: 196: Amino acid sequence of the MTP-CoxIV::Twin Strep::AnfK polypeptide encoded by SN195. Amino acids 1-41 correspond to the MTP-CoxIV sequence with a C-terminal GG linker, amino acids 42-61 correspond to the TwinStrep epitope with a C-terminal GG linker, and amino acids 62-523 correspond to the AnfK sequence from A. vinelandii.

[0508] SEQ ID NO: 197: Peptide sequence.

[0509] SEQ ID NO: 198: Linker sequence.

[0510] SEQ ID NO: 199: Amino acid sequence of the AnfD::linker16::AnfK polypeptide used for structural modeling (Example 20). Amino acids 1-509 correspond to the AnfD sequence with the N-terminal methionine omitted (A. vinelandii), amino acids 510-525 correspond to the 16 amino acid linker, and amino acids 526-984 correspond to AnfK (A. vinelandii).

[0511] SEQ ID NO: 200: Linker sequence.

[0512] SEQ ID NO: 201: Amino acid sequence of the AnfD::linker26(HA)::AnfK polypeptide. Amino acids 1 to 517 correspond to the AnfD sequence, amino acids 518 to 543 correspond to the 26 amino acid linker, and amino acids 544 to 1004 correspond to AnfK.

[0513] SEQ ID NO:202: Amino acid sequence of the MTP-FAγ51::AnfD::linker26(HA)::AnfK polypeptide encoded by SN272. Amino acids 1-64 correspond to the MTP-FAγ51-HA sequence including the C-terminal GG linker, amino acids 65-581 correspond to the AnfD sequence (A. vinelandii), amino acids 582-607 correspond to the 26 amino acid linker (linker26(HA)), and amino acids 608-1068 correspond to AnfK (A. vinelandii).

[0514] SEQ ID NO:203: Amino acid sequence of the MTP-CoxIV::AnfD::linker26(HA)::AnfK polypeptide encoded by SN273. Amino acids 1-61 correspond to the MTP-CoxIV sequence including the C-terminal GG linker, amino acids 62-578 correspond to the AnfD sequence (A. vinelandii), amino acids 579-604 correspond to the 26 amino acid linker (linker26(HA)), and amino acids 605-1065 correspond to AnfK (A. vinelandii).

[0515] SEQ ID NO:204: Amino acid sequence of the mFAγ51::AnfD::linker26(HA)::AnfK polypeptide encoded by SN274. Amino acids 1-64 correspond to the mFAγ51 sequence including the alanine substitutions and C-terminal GG that prevent cleavage by MPP, amino acids 65-581 correspond to the AnfD sequence (A. vinelandii), amino acids 582-607 correspond to the 26 amino acid linker (linker26(HA)), and amino acids 608-1068 correspond to AnfK (A. vinelandii).

[0516] SEQ ID NO:205: Amino acid sequence of the HISx6::AnfD::linker26(HA)::AnfK polypeptide (which lacks the MTP sequence and is believed to be located in the cytoplasm) encoded by SN275. Amino acids 1-9 correspond to the HISx6 sequence with a C-terminal GG, amino acids 10-526 correspond to the AnfD sequence (A. vinelandii), amino acids 527-552 correspond to the 26 amino acid linker (linker26(HA)), and amino acids 553-1013 correspond to AnfK (A. vinelandii).

[0517] SEQ ID NO: 206: Amino sequence of TbHCS polypeptide (Accession No. CP002466).

[0518] SEQ ID NO: 207: Amino sequence of TpHCS polypeptide (Accession No. CP002028).

[0519] SEQ ID NO: 208: Amino sequence of the SchHCS polypeptide (Accession No. CP036483).

[0520] SEQ ID NO: 209: Amino sequence of NsHCS polypeptide (Accession No. CP007203).

[0521] SEQ ID NO: 210: Amino sequence of MaHCS polypeptide (Accession No. AE010299).

[0522] SEQ ID NO: 211: Amino sequence of CtHCS polypeptide (Accession No. AE006470).

[0523] SEQ ID NO: 212: Amino sequence of MiHCS1 polypeptide (Accession No. ADG13125).

[0524] SEQ ID NO: 213: Amino sequence of MiHCS2 polypeptide (Accession No. ADG13175).

[0525] SEQ ID NO: 214: Amino sequence of MiHCS3 polypeptide (Accession No. ADG14004).

[0526] SEQ ID NO: 215: Amino sequence of LjFEN1 polypeptide (Accession number BAI49592).

[0527] SEQ ID NO: 216: Amino acid sequence of AnfD from A. vinelandii (accession number WP_012703361); 518 amino acids.

[0528] SEQ ID NO: 217: Amino acid sequence of AnfK from A. vinelandii (accession number WP_012703359); 462 amino acids.

[0529] SEQ ID NO: 218: Amino acid sequence of AnfH from A. vinelandii (Accession No. WP_012703362); 275 amino acids

[0530] SEQ ID NO: 219: Amino acid sequence of AnfG from A. vinelandii (accession number WP_012703360); 132 amino acids.

[0531] SEQ ID NO: 220: Peptide sequence.

[0532] SEQ ID NO: 221: Amino acid sequence of N. benthamiana P72026; 606 amino acids.

[0533] SEQ ID NO: 222: Amino acid sequence of N. benthamiana P20586; 470 amino acids.

[0534] SEQ ID NO: 223: Amino acid sequence of Mycobacterium tuberculosis α-isopropylmaleate synthase (MtLeuA); 644 amino acids.

[0535] SEQ ID NO: 224: Amino acid sequence of NifH polypeptide from A. vinelandii (AvNifH; accession number WP_012698831); 290 amino acids.

[0536] SEQ ID NO: 225: Peptide sequence, AnfH motif I (wherein X represents any amino acid).

[0537] SEQ ID NO: 226: Peptide sequence, AnfH motif II.

[0538] SEQ ID NO: 227: Peptide sequence, AnfH motif III.

[0539] SEQ ID NO: 228: Peptide sequence, AnfH motif IV.

[0540] SEQ ID NO: 229: Peptide sequence, AnfH motif V (wherein X represents any amino acid).

[0541] SEQ ID NO: 230: Peptide sequence, AnfH motif VI.

[0542] SEQ ID NO: 231: Peptide sequence, AnfH motif VII (wherein X represents any amino acid).

[0543] SEQ ID NO: 232: Amino acid sequence of the FdxN protein of A. vinelandii; Accession number WP_012703542; 92 amino acids.

[0544] SEQ ID NO: 233: Amino acid sequence of the MTP-FAγ51-FdxN-HA fusion polypeptide of SN291; 157 amino acids. Amino acids 1 to 54 correspond to the MTP-FAγ51 sequence with the GG linker, amino acids 55 to 145 correspond to the FdxN sequence without the N-terminal methionine, and amino acids 146 to 157 correspond to the HA epitope.

[0545] SEQ ID NO: 234: Amino acid sequence of the MTP-FAγ51-HA-FdxN fusion polypeptide of SN292; 156 amino acids. Amino acids 1 to 53 correspond to the MTP-FAγ51 sequence with a GG linker, amino acids 54 to 64 correspond to the HA epitope with a GG linker, and amino acids 65 to 156 correspond to the FdxN sequence without the N-terminal methionine.

[0546] SEQ ID NO: 235: Amino acid sequence of the mFAγ51-HA-FdxN fusion polypeptide of SN299; 156 amino acids. Amino acids 1 to 53 correspond to the MTP-FAγ51 sequence with a GG linker, amino acids 54 to 64 correspond to the HA epitope with a GG linker, and amino acids 65 to 156 correspond to the FdxN sequence without the N-terminal methionine.

[0547] SEQ ID NO: 236: Amino acid sequence of the HA-FdxN fusion polypeptide of SN300; 104 amino acids. Amino acids 1 to 12 correspond to the HA epitope with a GG linker, and amino acids 13 to 104 correspond to the FdxN sequence without the N-terminal methionine.

[0548] SEQ ID NO: 237: Amino acid sequence of the MTP-FAγ51-HA-NifV fusion polypeptide of SN254; 448 amino acids. Amino acids 1 to 53 correspond to the MTP-FAγ51 sequence with a GG linker, amino acids 54 to 64 correspond to the HA epitope with a GG linker, and amino acids 65 to 448 correspond to the NifV sequence from A. vinelandii.

[0549] SEQ ID NO: 238: Amino acid sequence of the NafV polypeptide from A. vinelandii (AvNafY; accession number AGK13761).

[0550] SEQ ID NO: 239: C-terminal amino acid sequence of the NifK polypeptide.

[0551] SEQ ID NO: 240: C-terminal amino acid sequence of the NifK polypeptide.

[0552] SEQ ID NO: 241: C-terminal amino acid sequence of the NifK polypeptide.

[0553] SEQ ID NO: 242: C-terminal amino acid sequence of the NifK polypeptide.

[0554] SEQ ID NO: 243: C-terminal amino acid sequence of the NifK polypeptide.

[0555] SEQ ID NO: 244: C-terminal amino acid sequence of the AnfK polypeptide.

[0556] SEQ ID NO: 245: C-terminal amino acid sequence of the AnfK polypeptide.

[0557] SEQ ID NO: 246: C-terminal amino acid sequence of the AnfK polypeptide.

[0558] SEQ ID NO: 247: C-terminal amino acid sequence of the AnfK polypeptide.

[0559] SEQ ID NO: 248: C-terminal amino acid sequence of the AnfK polypeptide. DETAILED DESCRIPTION OF THE INVENTION

[0560] General Techniques and Definitions

[0561] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art (e.g., in cell culture, molecular genetics, plant molecular biology, protein chemistry, and biochemistry).

[0562] Unless otherwise indicated, the recombinant protein, cell culture, and immunological techniques utilized in the present invention are standard procedures, well known to those skilled in the art. Such techniques are described and explained throughout the literature in a variety of sources, such as J. Perbal, A Practical Guide to Molecular Cloning, John Wiley and Sons (1984); J. Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbour Laboratory Press (1989); T. A. Brown (editor), Essential Molecular Biology: A Practical Approach, Vols. 1 and 2, IRL Press (1991); D. M. Glover and B. D. Hames (editors), DNA Cloning: A Practical Approach, Vols. 1-4, IRL Press (1995 and 1996); and F. M. Ausubel et al. (editors), Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience (1988, including all revisions to date); Harlow and David Lane (editor), Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory (1988), and JE Coligan et al. (editors), Current Protocols in Immunology, John Wiley & Sons (including all current editions).

[0563] The term "and / or," e.g., "X and / or Y," shall be understood to mean "X and Y" or "X or Y" and shall be construed as giving explicit support for both meanings or for either meaning.

[0564] As used herein, the term about means ±10%, more preferably ±5% of the specified value, unless stated to the contrary.

[0565] Throughout this specification, the word "comprise" or its variations "comprises" or "comprising" will be understood to include a stated element, integer, or step or group of elements, integers, or steps, but not to exclude any other element, integer, or step or group of elements, integers, or steps.

[0566] nitrogenase

[0567] Nitrogenase is an enzyme found in eubacteria and archaea that catalyzes the reduction of the strong triple bond of nitrogen (N2) to produce ammonia (NH3). Nitrogenase is found naturally only in bacteria. Nitrogenase is a complex of two enzymes, dinitrogenase and dinitrogenase reductase, which can be purified separately. Dinitrogenase, also known as component I or molybdenum-iron (MoCo) protein, is a tetramer (α2β2) of two NifD polypeptides and two NifK polypeptides, and also contains two "P clusters" and two "FeMo cofactors" (FeMo-co). Each NifD-NifK subunit pair contains one P cluster and one FeMo-co. FeMo-co is a metallocluster consisting of a MoFe3-S3 cluster complexed with a homocitrate molecule, which is coordinated to the molybdenum atom and bridged to the Fe4-S3 cluster by three sulfur ligands. FeMo-co is assembled separately within the cell and then incorporated into the apo-MoFe protein. The P cluster, also a metallocluster, contains eight Fe atoms and seven sulfur atoms and has a similar but distinct structure to FeMo-co. The P cluster is located at the αβ subunit interface of dinitrogenase and is coordinated by cysteinyl residues from both subunits. Dinitrogenase reductase, also known as component II or the "iron protein," is a dimer of NifH polypeptides and contains a single Fe4-S4 cluster at the subunit interface and two Mg-ATP binding sites (one per subunit). This enzyme is the essential electron donor for dinitrogenase, where electrons are transferred from the Fe4-S4 cluster to the P cluster and then to FeMo-co, the site for reduction.

[0568] Although Mo-containing nitrogenase is the most common nitrogenase in bacteria, two homologous nitrogenases exist that share similar cofactor and subunit composition but differ in their genes: vanadium-containing nitrogenase and Fe-only nitrogenase, encoded by the Vnf (vanadium nitrogen fixation) and Anf (alternative nitrogen fixation) genes, respectively. Some bacteria in nature possess all three types of nitrogenase, while others (e.g., Klebsiella pneumoniae) contain only Mo- and V-containing enzymes or only Mo-containing enzymes.

[0569] Various nitrogen fixation (Nif) genes are required for the biosynthesis of FeMo-co and the maturation of nitrogenase components into their catalytically active forms. The roles of NifB, NifE, NifH, NifN, NifQ, NifV, and NifX polypeptides in FeMo-co synthesis have been described previously ( Rubio and Ludden, 2008 ).

[0570] Biological N2 fixation, catalyzed by the prokaryotic enzyme nitrogenase, is an alternative to the use of synthetic N2 fertilizers. The sensitivity of nitrogenase to oxygen is a major obstacle to engineering biological nitrogen fixation by direct introduction of Nif genes into plants (e.g., cereal crops).

[0571] The present inventors hypothesized that targeting Nif polypeptides to the mitochondrial matrix (MM) of plant cells might overcome the oxygen sensitivity problem. MM harbors oxygen-consuming enzymes, which allow other enzymes containing oxygen-sensitive Fe-S clusters to function. The mitochondrial Fe-S cluster assembly machinery resembles its counterpart in nitrogen-fixing bacteria (Balk and Pilon, 2011; Lill and Muhlenhoff, 2008). Therefore, some of the components required for nitrogenase biosynthesis are already appropriately located within the MM, potentially reducing the number of Nif genes required for reconstitution. A large ATP reduction capacity and concentration are also present (Geigenberger and Fernie, 2014; Mackenzie and McIntosh, 1999), both of which are prerequisites for the catalytic activity of the nitrogenase enzyme. Additionally, the presence of glutamate synthase in mitochondria provides an entry point for any ammonia fixed by nitrogenase to enter plant metabolism. These characteristics, and the fact that mitochondria themselves are of α-proteobacterial origin, led the inventors of the present invention to believe that this organelle was well suited as a site for attempting functional reconstitution of nitrogenase.

[0572] As a first step toward reconstituting nitrogenase in plant cell mitochondria, evidence was needed that individual Nif proteins could be correctly targeted to the MM. To this end, we selected the model plant Nicotiana benthamiana as an expression platform (Wood et al., 2009) to express transgenes singly or, more importantly, in combination. Because most proteins located in the MM are nuclear-encoded, we relied on recent advances in our understanding of intracellular signaling and transport processes (Huang et al., 2009; Murcha et al., 2014) and utilized previously characterized N-terminal peptide targeting signals (Lee et al., 2012).

[0573] The model nitrogen-fixing bacterium Klebsiella pneumoniae utilizes 16 unique proteins for nitrogenase biosynthesis and catalytic function. We reengineered all 16 Nif proteins from K. pneumoniae to target them to plant MM and evaluated their expression and processing in N. benthamiana leaves. We transiently expressed all 16 Nif polypeptides and examined sequence-specific MM processing. We demonstrated that all 16 Nif polypeptides can be expressed in plant leaf cells as MTP:Nif fusion polypeptides. Furthermore, we present evidence that these proteins can be targeted to the mitochondrial matrix (MM), the intracellular location where they potentially exert nitrogenase function, and cleaved by mitochondrial processing protease (MPP). This represents a major advance toward the goal of engineering endogenous nitrogen fixation in plants.

[0574] Mitochondrial protein import in plants

[0575] Nearly all mitochondrial proteins are nuclear-encoded and translated in the cytosol, requiring their import into mitochondria. Signal sequences within these polypeptides direct their import into four distinct locations within mitochondria: the outer membrane (OM), the intermembrane space (IS), the inner membrane (IM), or the matrix (MM). These signal sequences are distinguished by their biochemical properties and guide the transport of polypeptides through at least four distinct import pathways that target polypeptides to one or more of these locations (Chacinska et al., 2009). These four pathways are: (1) the general import pathway (also known as the "classical" presequence pathway, which targets polypeptides to the MM, IS, or IM); (2) the carrier-mediated transport pathway (used for import into the IM); (3) the mitochondrial intermembrane space (MIA) assembly pathway; and (4) the sorting and assembly machinery (SAM) pathway, which is used to transport polypeptides into the OM. The general import pathway imports polypeptides that contain a cleavable presequence (also known as a signal sequence). These polypeptides may also contain a hydrophobic sorting signal (HSS). The carrier uptake pathway imports polypeptides with an internal presequence and hydrophobic region, including the signal. The MIA pathway imports polypeptides with twin cysteine ​​residues. The SAM pathway imports polypeptides containing the β signal and a putative TOM20 signal. All of these pathways utilize the translocase of the outer membrane (TOM), and the first and second pathways also utilize the TIM23 translocase of the intermembrane complex. Only the first pathway utilizes a matrix processing peptidase (matrix processing protease, MPP).

[0576] One common feature of all mitochondria-targeted polypeptides is the presence of at least one domain within the polypeptide that guides transport to a precise location. The most well-studied of these is the "classical" N-terminal presequence domain, which is cleaved by MPP within the matrix (Murcha et al., 2004). It is estimated that approximately 70% of plant and animal mitochondrial proteins have a cleavable presequence, although both internal and C-terminal signal sequences have also been found (reviewed in Pfanner and Geissler, 2001; Schleiff and Soll, 2000). In Arabidopsis, these presequences range in length from 11 to 109 amino acid residues, with an average length of 50 amino acid residues. While there is no consensus sequence that completely defines the presequence for the first pathway, presequences tend to contain a large proportion of hydrophobic and positively charged amino acids. A further characteristic is its ability to form amphipathic α-helices, which usually begin within the first 10 amino acid residues (Roise et al., 1986). These domains are rich in hydrophobic amino acid residues (Ala, Leu, Phe, Val), hydroxylated amino acid residues (Ser, Thr), and positively charged amino acid residues (Arg, Lys), and lack acidic amino acids. Across numerous mitochondrial proteins, serine (16–17%) and alanine (12–13%) are overrepresented in mitochondrial signal peptides, while arginine is abundant (12%). The MPP cleavage point is defined by the presence of a conserved arginine residue for most presequences, usually at the P2 (−2 amino acids from the scissile bond) or P3 position in most other cases (Huang et al., 2009).

[0577] Mitochondrial presequences interact with the Tom20 receptor through hydrophobic residues. Studies have shown that the hydrophobic surface of the α-helix facilitates peptide recognition by the TOM20 component of the TOM import complex, whereas the positive charge is recognized by the TOM22 subunit (Abe et al., 2000). Finally, most presequences associate with Hsp70 to guide polypeptide transport, and nearly all plant presequences contain at least one binding motif for the Hsp70 molecular chaperone (Zhang and Glaser, 2002). The chaperone Hsp70 is involved in protein folding, prevents protein aggregation, and functions as a molecular motor, pulling precursors across the mitochondrial membrane. The membrane potential (ΔΨ) across the inner membrane (approximately 100 mV, negative on the inside) also drives the translocation of positively charged presequences through an electrophoretic effect.

[0578] The majority of proteins with cleavable presequences are destined for the mitochondrial matrix via a common transport pathway that utilizes the transporter of the outer membrane (TOM) complex and the transporter of the inner membrane 23 complex (TIM23). However, some proteins with cleavable presequences can assemble in the inner membrane (Murcha et al., 2004) or, if they also contain a hydrophobic sorting signal (HSS), in the inner membrane lumen (Glick et al., 1992). There are very few examples of matrix-localized proteins that do not have a cleavable presequence. In Arabidopsis, only glutamate dehydrogenase has been found in the matrix with a full-length, unprocessed presequence (Huang et al., 2009).

[0579] For proteins that are not targeted to the matrix, a variety of internal non-cleavable localization signals are used. These are typically associated with specific transport pathways, further tailoring them to that particular class of protein. In plants, no previous studies have clearly identified what constitutes an internal signal sequence for inner membrane lumen proteins. However, a motif containing twin cysteine ​​residues appears to be associated with transport through the mitochondrial inner membrane assembly pathway (MIA) (Carrie et al., 2010; Darshi et al., 2012). Finally, non-cleavable internal sequences are also used by proteins destined for the inner membrane via a carrier pathway that utilizes the TOM and TIM22 apparatus for the insertion of proteins with multiple transmembrane domains (Kerscher et al., 1997; Sirrenberg et al., 1996). These sequences typically contain a hydrophobic region followed by a presequence-like internal sequence, resembling an N-terminal presequence, but are distinguished by the internal location of these sequences within their cognate proteins.

[0580] In photosynthetic organisms, nuclear-encoded mitochondrial proteins must differentiate between import into chloroplasts and into mitochondria, yet there are many similarities between these two organelles and their proteomes. Alpha helices, which are often present in mitochondrial presequences, are typically absent from chloroplast presequences (Zhang and Glaser, 2002). Chloroplast presequences tend to be less structured and exhibit a high degree of beta-sheet domain organization (Bruce, 2001).

[0581] In plants, MPP is anchored to the inner membrane-associated Cytbc1 complex, but the functions of these two proteins are independent, as the active MPP site is located in front of the matrix ( Glaser and Dessi, 1999 ).

[0582] Mitochondrial Targeting Peptides

[0583] As used herein, the term "mitochondrial targeting peptide" or "MTP" refers to an amino acid sequence at least 10 amino acids in length, preferably 10 to about 80 amino acid residues, that targets a target protein to mitochondria and can be used heterologously in an MTP-target protein translational fusion to target a selected target protein (e.g., Nif polypeptide, Gus, GFP) to mitochondria.

[0584] MTPs typically contain at their N-terminus the translation initiation methionine of the polypeptide from which they are derived. MTPs are translationally fused to Nif polypeptides or "target proteins" by a peptide bond to the Met residue corresponding to the initiation Met of the target protein. Alternatively, the Met residue can be omitted, and the peptide bond is fused directly to an amino acid residue that is the second amino acid of the target protein in the wild type. MTPs are typically rich in basic and hydroxylated amino acids and usually lack acidic amino acids or extended hydrophobic stretches. MTPs can form amphipathic helices.

[0585] Without wishing to be bound by theory, MTPs typically contain an import targeting sequence that binds to a receptor on the outer membrane of mitochondria. Upon binding to the outer membrane, the fusion polypeptide preferably undergoes membrane translocation to a transport channel protein and passes through the mitochondrial double membrane into the mitochondrial matrix (MM). The import targeting sequence is then typically cleaved, allowing the mature fusion protein to fold.

[0586] The MTP can contain additional signals that subsequently target the protein to various regions of the mitochondrion, such as the mitochondrial matrix (MM). In one embodiment, the uptake targeting sequence is a matrix targeting sequence.

[0587] MTP may be cleavable or non-cleavable when translationally fused to a Nif polypeptide. Thus, in one embodiment, the MTP-Nif fusion polypeptide is at least partially cleaved. In this regard, the term "at least partially cleaved" refers to a detectable amount of cleavage of the MTP-Nif fusion polypeptide when expressed in a plant cell. In one embodiment, at least 50% of the MTP-Nif fusion polypeptide produced in the cell is cleaved within the MTP sequence, preferably at least 75% is cleaved, and more preferably at least 90% is cleaved. In an alternative embodiment, less than 50% of the MTP-Nif fusion polypeptide is cleaved in the cell, e.g., MTP is not cleaved. In one embodiment, MTP does not contain a cleavage site for MPP. MTP may contain a cleavage site. The N-terminal part of the resulting processed product (i.e., mature NP) may include one or more C-terminal amino acids of MTP (also referred to herein as a "scar sequence" or "scar peptide"), or may not include any C-terminal amino acids of MTP. The scar sequence, when present, is preferably 1 to 45 amino acids in length, more preferably 1 to 20 amino acids, and even more preferably 1 to 12 amino acids in length. Alternatively, a cleavage site may be located within the fusion polypeptide such that the entire MTP sequence is cleaved and removed, e.g., the linker may contain the cleavage sequence.

[0588] Natural mitochondrial targeting peptides are located at the N-terminus of precursor proteins, and the N-terminal portion is typically cleaved and removed during or after import into mitochondria. Cleavage is typically catalyzed by a general matrix processing protease (MPP). In plants, MPP is incorporated into the bc1 complex of the respiratory chain. This protease recognizes cleavage sites in nearly 1,000 precursor proteins with widely varying amino acid sequences that are largely unconserved. In one embodiment, the MTP contains a protease cleavage site for the MPP. In a further embodiment, the processed product is generated by cleavage of the fusion protein by the MPP within or immediately after the MPP. In this context, the term "immediately after" refers to after cleavage by the MPP, and no residual amino acids from the MTP fused to the Nif polypeptide are present. Thus, if the fusion polypeptide is cleaved "immediately after" the MTP, the MPP cleavage site is immediately after the C-terminal amino acid of the MTP.

[0589] The term "cleavage product" or "cleavage product," as used herein in the context of an MTP fusion polypeptide, refers to a polypeptide resulting from protease cleavage within or immediately after the MTP amino acid sequence. In this regard, a cleavage product of an MTP fusion polypeptide can be obtained by cleavage with MPP. The cleavage product may retain one or more amino acids from the cleaved MTP (i.e., the scar peptide), or may have no residual amino acids from the cleaved MTP. In one embodiment, the cleavage product of a Nif fusion polypeptide of the invention comprises at least 95%, or all, of the amino acids present in the Nif fusion polypeptide sequence.

[0590] In one embodiment, the MTP is not cleaved. The present inventors have demonstrated that incorporation of the MTP does not necessarily lead to complete processing of the Nif protein. In some cases (NifX-FLAG, NifD-HA opt1Both processed and unprocessed Nif proteins were observed in the NifDK-HA (NifDK-HA). Given that there is no general consensus sequence for MTPs and that internal protein sequences can influence mitochondrial targeting ( Becker et al., 2012 ), it is perhaps not surprising that we found differences in processing efficiency among Nif proteins.

[0591] Non-limiting examples of suitable MTPs that can be used in the context of the present invention include peptides having the general structure as defined by von Heijne (1986) or Roise and Schatz (1988). Non-limiting examples of MTPs are the mitochondrial targeting peptides defined in Table I of von Heijne (1986) or disclosed herein.

[0592] In one embodiment, the MTP is an F1-ATPase gamma-subunit (MTP-FAγ). One example of a suitable FAγ MTP is from A. thaliana (Lee et al., 2012). In one embodiment, the MTP-FAγ is 77 amino acids in length, which, when cleaved by an MMP, leaves 35 MTP residues at the N-terminus of the fusion polypeptide. In a preferred embodiment, the MTP-FAγ is less than 77 amino acids in length. For example, the MTP-FAγ can be approximately 51 amino acids in length, which, when cleaved by an MMP, leaves 9 MTP residues at the N-terminus of the fusion polypeptide.

[0593] Those skilled in the art will recognize that software exists for predicting mitochondrial proteins and their target sequences (eg, MitoProtII, PSORT, TargetP, and NNPSL).

[0594] MitoProtII is a program that predicts the mitochondrial localization of a sequence based on several physiochemical parameters (e.g., amino acid composition of the N-terminus or highest overall hydrophobicity over a 17-residue window). TargetP and NPSORT are programs that predict subcellular locations based on various sequence-derived features (e.g., the presence of sequence motifs and amino acid composition). They predict the subcellular location of eukaryotic proteins based on the predicted presence of either a terminal presequence (i.e., a chloroplast transit peptide, a mitochondrial targeting peptide, or a secretory pathway signal peptide). TargetP utilizes initial binary predictors, SignalP and ChloroP, and requires the N-terminal sequence as input to two layers of an artificial neural network (ANN). For sequences predicted to contain an N-terminal presequence, potential cleavage sites can also be predicted. NNPSL is another ANN-based method that uses amino acid composition to assign one of four subcellular localizations (cytosolic, extracellular, nuclear, and mitochondrial) to a query sequence.

[0595] Those skilled in the art will be able to easily determine whether the selected MTP targeted the fusion polypeptide to the mitochondrial matrix based on routine methods and the methods disclosed herein. To aid in the detection of processed proteins, the present inventors selected relatively long targeting peptides, which have previously been demonstrated to be capable of transporting GFP into Arabidopsis chloroplasts (Lee et al., 2012). As shown in the Examples herein, the selected MTP targeted all of the selected nitrogenase proteins to the MM. This conclusion is based on several lines of evidence. First, the observed size of the Nif polypeptides expressed by N. benthamiana was consistent with the expected size resulting from MM peptidase processing. This was also reflected in the size difference observed between bacteria (full-length, unprocessed) and the smaller Nifs (NifF and NifZ) expressed by plant mitochondria. Additionally, a mutation in MTP that prevented it from being processed by the mitochondrial import machinery resulted in broader bands for both NifD and GFP fusions. This is consistent with the difference in size between the processed and unprocessed proteins. Finally, mass spectrometry of a representative fusion polypeptide revealed that MTP-NifH was cleaved between residues 42 and 43 of MTP, as expected for differential processing within the matrix.

[0596] In some embodiments of the present invention, it may be useful to use multiple tandem copies of a selected MTP. The coding sequence for a duplexed or multiplexed targeting peptide can be obtained from an existing MTP by genetic engineering. The amount of MTP can be measured, for example, by quantitative immunoblot analysis after cell fractionation. Thus, in the present invention, the term "mitochondrial targeting peptide" or "MTP" encompasses one or more copies of a single amino acid peptide that targets a target Nif protein to mitochondria. In a preferred embodiment, the MTP contains two copies of the selected MTP. In another embodiment, the MTP contains three copies of the selected MTP. In another embodiment, the MTP contains four or more copies of the selected MTP.

[0597] Those skilled in the art will appreciate that the MTP sequence is not limited to the native MTP sequence, but may include amino acid substitutions, deletions, and / or insertions relative to native MTP, provided that the sequence variant maintains mitochondrial targeting function.

[0598] Those skilled in the art will understand that an MTP may be flanked at its N- or C-terminus by amino acids as a result of the cloning strategy and can function as linkers. These additional amino acids can be considered to form part of the MTP.

[0599] Those skilled in the art will also understand that the N- or C-terminus of MTP may be fused to an oligopeptide linker and / or tag (such as an epitope tag). In a preferred embodiment, one or more or all of the Nif fusion polypeptides of the invention produced in plant cells lack an added epitope tag compared to the corresponding wild-type Nif polypeptide.

[0600] Mitochondrial targeting peptide (MTP)-Nif fusion polypeptide

[0601] The present invention relates to mitochondrial targeting peptide (MTP)-Nif fusion polypeptides and their cleaved polypeptide products. When the MTP-Nif fusion polypeptides of the invention are expressed in plant cells, the MTP-Nif fusion polypeptides and / or cleaved polypeptide products are targeted to the mitochondrial matrix (MM). Preferably, the fusion polypeptides confer nitrogenase reductase and / or nitrogenase activity to plant cells, or the same activity as that conferred by the corresponding wild-type Nif polypeptide in bacteria.

[0602] As used herein, the term "fusion polypeptide" refers to a polypeptide comprising two or more polypeptide domains covalently linked by a peptide bond. Typically, a fusion polypeptide is encoded as a single polypeptide chain by a polynucleotide of the invention. In one embodiment, a fusion polypeptide of the invention comprises a mitochondrial targeting peptide (MTP) and a Nif polypeptide (NP). In this embodiment, the C-terminus of MTP is translationally fused to the N-terminus of NP. In an alternative embodiment, a fusion polypeptide of the invention comprises the C-terminal portions of MTP and NP, which C-terminal portions result from cleavage of MTP by MPP. Such a C-terminal portion of MTP is referred to herein as a "scar" sequence. In this embodiment, the C-terminal amino acid of the C-terminal part of MTP is translationally fused to the N-terminal amino acid of NP. In these embodiments, the fusion polypeptide may include one or more additional amino acids (such as a Gly-Gly sequence) between MTP and NP, and / or an additional methionine as the translation initiation amino acid. In one embodiment, the fusion polypeptide comprises two Nif polypeptides, preferably a NifD polypeptide translationally fused to a NifK polypeptide via a linker sequence or a NifE polypeptide translationally fused to a NifN polypeptide via a linker sequence. Both of these fused polypeptides can be present. In these embodiments, the second Nif polypeptide in the fusion polypeptide has its wild-type C-terminus, i.e., lacks a C-terminal extension.

[0603] As used herein, the phrase "translationally fused to the N-terminus" means that the C-terminus of an MTP polypeptide or a linker polypeptide is covalently linked to the N-terminus of NP by a peptide bond, resulting in a fusion polypeptide. In one embodiment, NP does not contain its native translation initiation methionine (Met) residue or its two N-terminal Met residues compared to the corresponding wild-type NP. In an alternative embodiment, NP contains the translation initiation Met of the wild-type NP polypeptide, e.g., for NifD, or contains one or both of the two N-terminal Met residues.

[0604] Such polypeptides are typically produced by expressing a chimeric protein coding region in which the translational reading frame of nucleotides encoding MTP is joined in frame to the reading frame of nucleotides encoding NP. Those skilled in the art will recognize that the C-terminal amino acid of MTP can be translationally fused to the N-terminal amino acid of NP without a linker or via a linker consisting of one or more amino acid residues (e.g., 1 to 5 amino acid residues). Such linkers can also be considered part of MTP. After expression of the protein coding region, MTP can be cleaved in the MM of the plant cell, and such cleavage (if it occurs) is included in the concept of producing the fusion polypeptide of the present invention.

[0605] Preferably, the fusion polypeptide or processed Nif polypeptide has functional Nif activity. In a preferred embodiment, the activity resembles that of the corresponding wild-type Nif polypeptide. The functional activity of the fusion polypeptide or processed Nif polypeptide can be determined in bacterial and biochemical complementation assays. In a preferred embodiment, the fusion polypeptide or processed Nif polypeptide has about 70-100% of the wild-type Nif activity. Nif polypeptides without Nif function are still useful as research tools, for example, to examine expression levels from genetic constructs or to examine association with other Nif polypeptides.

[0606] A fusion polypeptide can contain two or more MTPs and / or two or more NPs; for example, a fusion polypeptide can contain an MTP, a NifD polypeptide, and a NifK polypeptide. The fusion polypeptide can also contain, for example, an oligopeptide linker connecting two NPs. The linker is preferably long enough to allow two or more functional domains, such as two NPs (e.g., NifD and NifK, or NifE and NifN), to associate into a functional configuration in a plant cell. In a preferred embodiment, the NifD polypeptide is an AnfD polypeptide, and the NifK polypeptide is an AnfK polypeptide. Such a linker can be 8 to 50 amino acid residues in length, preferably about 25 to 35 amino acids in length, and more preferably about 30 amino acid residues in length, or about 26 amino acid residues in length, for an AnfD-linker-AnfK fusion polypeptide. Fusion polypeptides can be obtained by conventional means, such as by genetically expressing a polynucleotide sequence encoding the fusion polypeptide in a suitable cell.

[0607] As used herein, a "substantially purified polypeptide" refers to a polypeptide that is substantially free from other components (e.g., lipids, nucleic acids, carbohydrates) that normally accompany the polypeptide, e.g., in a cell. A substantially purified polypeptide is at least 90% free from such components.

[0608] Plant cells, transgenic plants, and parts thereof of the present invention contain a polynucleotide encoding a polypeptide of the present invention. Because a polypeptide of the present invention does not naturally occur in a plant cell, particularly in the mitochondria of a plant cell, a polynucleotide encoding the polypeptide can be referred to herein as an exogenous polynucleotide because it does not naturally occur in the plant cell but has been introduced into the plant cell or a progenitor cell. Thus, cells, plants, and plant parts of the present invention that produce a polypeptide of the present invention can be said to produce a recombinant polypeptide. The term "recombinant" in the context of a polypeptide refers to a polypeptide that, when produced by a cell, is encoded by an exogenous polynucleotide, which polynucleotide has been introduced into the cell or a progenitor cell by recombinant DNA or RNA techniques (e.g., transformation). Typically, a plant cell, plant, or plant part contains a non-endogenous gene that causes the production of a certain amount of the polypeptide at least at some point during the life cycle of the plant cell or plant. The exogenous polynucleotide is integrated into the respective genome of the plant cell and / or is transcribed in the nucleus of the cell.

[0609] In one embodiment, the polypeptide of the invention is not a naturally occurring polypeptide, hi an alternative embodiment, the polypeptide of the invention is naturally occurring but present in a plant cell where it does not naturally occur, preferably in the mitochondria of the plant cell.

[0610] In one embodiment, a polypeptide of the invention (e.g., an MTP fusion polypeptide or a cleavage product thereof) is at least partially soluble in the mitochondria of a plant cell. In this context, the term "at least partially soluble" means that the polypeptide is detectable in the soluble fraction of a homogenized sample containing the mitochondria of a plant cell. Suitable methods for detecting the solubility of a polypeptide are known in the art and include the method described in Example 1. In one embodiment, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the polypeptide present in the cell is soluble.

[0611] Nif polypeptides

[0612] As used herein, the terms "Nif polypeptide" and "Nif protein" are used interchangeably and refer to peptides related in amino acid sequence to naturally occurring polypeptides involved in nitrogenase activity. Nif polypeptides of the invention are selected from the group consisting of NifD polypeptides, NifH polypeptides, NifK polypeptides, NifB polypeptides, NifE polypeptides, NifN polypeptides, NifF polypeptides, NifJ polypeptides, NifM polypeptides, NifQ polypeptides, NifS polypeptides, NifU polypeptides, NifV polypeptides, NifW polypeptides, NifX polypeptides, NifY polypeptides, and NifZ polypeptides, each of which is defined herein. Nif polypeptides of the invention include "Nif fusion polypeptides," which, as used herein, refer to polypeptide homologs of naturally occurring Nif polypeptides that have additional amino acid residues grafted to either or both the N-terminus and C-terminus relative to the corresponding naturally occurring Nif polypeptide. As noted above, Nif fusion polypeptides may lack a translation initiation Met or two N-terminal Met residues relative to the corresponding wild-type Nif polypeptide. Amino acid residues of a Nif fusion polypeptide that correspond to those of a native Nif polypeptide (i.e., a Nif polypeptide without additional amino acid residues attached to either or both of the N-terminus and C-terminus) are also referred to herein as Nif polypeptides (abbreviated as "NP") or NifD polypeptides ("ND"), etc. In a preferred embodiment, the "additional amino acid residues attached to either or both of the N-terminus and C-terminus" comprise a mitochondrial targeting peptide (MTP) or processed MTP attached to the N-terminus of NP, or an epitope sequence ("tag") that is N-terminal and / or C-terminal to NP, or both the MTP or processed MTP and the epitope sequence.

[0613] Naturally occurring Nif polypeptides occur only in certain bacteria, including nitrogen fixers, which include free-living, associative, and symbiotic nitrogen fixers. Free-living nitrogen fixers can fix significant levels of nitrogen without direct interaction with other organisms. Non-limiting examples of free-living nitrogen fixers include members of the genera Azotobacter, Beijerinckia, Klebsiella, and Cyanobacteria (classified as aerobic organisms), members of the genera Clostridium and Desulfovibrio, and those designated purple sulfur bacteria, purple non-sulfur bacteria, and green sulfur bacteria. Cooperative nitrogen fixers are prokaryotes that can form close associations with some members of the grass family (herbaceous plants). These bacteria fix significant amounts of nitrogen within the rhizosphere of host plants. Members of the genus Azospirillum are representative examples of cooperative nitrogen fixers. Symbiotic nitrogen fixers are bacteria that fix nitrogen symbiotically by partnering with a host plant. The plant provides sugars from photosynthesis, which are used by the nitrogen fixers as energy for nitrogen fixation. Members of the genus Rhizobia are examples of cooperative nitrogen fixers.

[0614] The Nif polypeptide or Nif fusion polypeptide of the invention is selected from the group consisting of NifH, NifD, NifK, NifB, NifE, NifN, NifF, NifJ, NifM, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ polypeptides, the functions of which have recently been reviewed by Buren et al. (2020).

[0615] Other polypeptides of the present invention are contemplated to be VnfG and AnfG, which are involved in V-nitrogenase and Fe-nitrogenase, respectively, nitrogenase-associated factors (Naf polypeptides), such as NafY, and ferredoxin polypeptides (such as FdxN polypeptides). These polypeptides are preferably encoded and expressed as MTP fusion polypeptides for targeting to mitochondria.

[0616] A polypeptide, or class of polypeptides, can be determined by the degree of identity (% identity) of its amino acid sequence with a reference amino acid sequence and / or by the presence of certain amino acid motifs or protein family domains, or by one reference amino acid sequence having a greater % identity than another reference amino acid sequence. In addition to the degree of sequence identity, a polypeptide, or class of polypeptides, can also be determined by having the same biological activity as a naturally occurring Nif polypeptide.

[0617] Percent polypeptide identity is determined by GAP (Needleman and Wunsch, 1970) analysis (GCG program) with a gap creation penalty of 5 and a gap extension penalty of 0.3, or by Blastp version 2.5 or an updated version thereof (Altschul et al., 1997), in each case aligning two sequences, including a reference sequence, over the entire length of the reference sequence. As used herein, a reference sequence includes the sequence given for the native Nif polypeptide from K. pneumoniae (renamed K. oxytoca) (SEQ ID NOS: 1-17).

[0618] In the following definitions, the degree of amino acid sequence identity to a reference sequence, shown as a SEQ ID NO:, is determined by Blastp version 2.5 or its updated versions (Altschul et al., 1997) using default parameters, except that the maximum number of target sequences is set to 10,000, and is determined along the entire length of the reference amino acid sequence.

[0619] The NifH polypeptide in natural bacteria is one structural component of the nitrogenase complex and is often referred to as the iron (Fe) protein. NifH forms a homodimer with an Fe4S4 cluster bound between its subunit and two ATP-binding domains. NifH is an essential electron donor to the nitrogenase protein (NifD / NifK heterotetramer) and therefore functions as a nitrogenase reductase (EC 1.18.6.1). The molybdenum form of NifH is also involved in the biosynthesis of FeMo-co and the maturation of the apo-MoFe protein (Jasniewski et al., 2018). As outlined therein, NifH has three recognized major functions: (i) its role in inserting Mo and homocitrate in the synthesis of FeMo-co (involving the NifE-NifN complex), (ii) its reductase function in forming the P cluster on NifD-NifK from what has been termed the P* cluster (which may also involve the small chaperone-like polypeptide NifZ), and (iii) its role as an electron donor to the nitrogenase protein.

[0620] As used herein, "NifH polypeptide" refers to a polypeptide whose sequence contains at least 41% amino acid identity to the amino acid sequence set forth as SEQ ID NO:1 and contains one or more of the domains TIGR01287, PRK13236, PRK13233, and cd02040. The TIGR01287 domain is present in each of molybdenum-iron nitrogenase reductase (NifH), vanadium-iron nitrogenase reductase (VnfH), and iron-iron nitrogenase reductase (AnfH), but excludes homologous proteins from light-independent protochlorophyllide reductases. Thus, as used herein, NifH polypeptides include subclasses of iron-binding polypeptides whose sequence contains at least 41% amino acid identity to SEQ ID NO:1, VnfH iron-binding polypeptides, and AnfH iron-binding polypeptides. Naturally occurring NifH polypeptides are typically 260-300 amino acids in length, and the naturally occurring monomer has a molecular weight of approximately 30 kDa. Numerous NifH polypeptides have been identified, and many sequences are available in public databases. For example, NifH polypeptides are found in Klebsiella michiganensis (accession number WP_049123239.1, 99% identical to SEQ ID NO: 1), Brenneria goodwinii (WP_048638817.1, 93% identical), Sideroxydans lithotrophicus (WP_013029017.1, 84% identical), Denitrovibrio acetiphilus (WP_013010353.1, 80% identical), Desulfovibrio africanus (WP_014258951.1, 72% identical), Chlorobium phaeobacteroides (WP_011744626.1, 69% identical), and Methanosaeta concilii. (WP_013718497.1, 64% identity), Rhodobacter (WP_009565928.1, 61% identity), Methanocaldococcus infernus (WP_013099472.1, 42% identity), and Desulfosporosinus youngiae (WP_007781874.1, 41% identity).NifH polypeptides are described and reviewed in Thiel et al. (1997), Pratte et al. (2006), Boison et al. (2006), and Staples et al. (2007).

[0621] As used herein, a functional NifH polypeptide is a NifH polypeptide that can assemble with other required subunits (e.g., NifD and NifK, and FeMo-, FeV-, or FeFe-cofactors) to form a functional nitrogenase complex.

[0622] As used herein, an "AnfH polypeptide" refers to a NifH polypeptide (TIGR01287), a member of the nitrogenase conserved superfamily cl25403, that contains the conserved domain PRK13233 and is at least 69% identical to the Azotobacter vinelandii AnfH polypeptide (SEQ ID NO:218; Accession No. WP_012703362) when measured along the entire length of SEQ ID NO:218. This amino acid sequence is used herein as the reference sequence for AnfH. TIGR01287: AnfH is an all-iron variant of nitrogenase component II, also known as nitrogenase reductase. As used herein, AnfH polypeptides are a subset of NifH polypeptides. AnfH polypeptides do not include molybdenum-type NifH polypeptides or vanadium-type NifH polypeptides (VnfH). The amino acid sequences of AnfH polypeptides in sequence databases have typically been annotated as AnfH polypeptides. As of January 2020, there were 314 specific amino acid sequences in the AnfH set in the NCBI protein database, all of which contained amino acid residues specific to AnfH and were therefore distinct from the molybdenum forms of NifH and VnfH. A subset of the molybdenum forms of NifH and VnfH were more similar but still distinct. Examples of naturally occurring AnfH polypeptides include AnfH polypeptides from Rhodocyclus tenuis (accession number WP_153472986; 92.36% identity), Dickeya paradisiaca (accession number WP_015854293; 88.36% identity), Thermodesulfitimonas autotrophica (accession number WP_123927773; 78.91% identity), Clostridium kluyveri (accession number WP_073538802; 76.36% identity), and Methanophagales archaeon (accession number RCV64832; 69.37% identity), each of which references SEQ ID NO: 218.

[0623] As described in Example 23 herein, 16 amino acids were identified at positions in the AnfH sequence that were conserved compared to the molybdenum-type NifH sequence of AvNifH and characteristic of AnfH polypeptides. These can be used to distinguish AnfH polypeptides from other NifH sequences that do not share all 16 amino acids. Because AvNifH, KoNifH (SEQ ID NO: 1), and other molybdenum-type NifH sequences share motifs III and IV but not motifs I, II, and V-VII, these motifs (SEQ ID NOs: 225-231) could also be used to distinguish the AnfH subset from other NifH polypeptides.

[0624] The functional AnfH polypeptide, like other functional NifH polypeptides, can function as nitrogenase reductase, an essential electron donor for the FeFe complex. AnfH, like the molybdenum-type NifH, may be involved in the biosynthesis of FeFe and the maturation of the apo-FeFe complex (AnfD-AnfK-AnfG).

[0625] As used herein, "NifD polypeptide" refers to a polypeptide whose sequence is set forth in SEQ ID NO:2 and contains amino acids that are at least 33% identical to an amino acid sequence containing (i) one or both of the domains TIGR01282 and COG2710 (both of which are found in iron-molybdenum-binding polypeptides, including polypeptides having the amino acid sequence set forth in SEQ ID NO:2), or (ii) the iron-vanadium-binding domain TIGR01860 (in which case the NifD polypeptide is within the subclass of VnfD polypeptides), or (iii) the iron-iron-binding domain TIGR1861 (in which case the NifD polypeptide is within the subclass of AnfD polypeptides). The NifD polypeptide can be part of a fusion polypeptide, e.g., fused to MTP and / or NifK. Alternatively, the NifD polypeptide may not contain any N- or C-terminal extensions. In a preferred embodiment, the NifD polypeptide binds to the FeMo cofactor when associated with a NifK polypeptide.

[0626] As used herein, NifD polypeptides include a subclass of iron-molybdenum (FeMo-co) binding polypeptides whose sequences contain at least 33% amino acid identity to SEQ ID NO:2, VnfD iron-vanadium polypeptides, and AnfD polypeptides. Naturally occurring NifD polypeptides are typically 470-540 amino acids in length. Many NifD polypeptides have been identified, and many sequences are available in public databases. For example, the NifD polypeptide is also found in Raoultella ornithinolytica (accession number WP_044347161.1, 96% identical to SEQ ID NO: 2), Kluyvera intermedia (WP_047370273.1, 93% identical), Dickeya dadantii (WP_038902190.1, 89% identical), Tolumonas sp. BRL6-1 (WP_024872642.1, 81% identical), Magnetospirillum gryphiswaldense (WP_024078601.1, 68% identical), Thermoanaerobacterium thermosaccharolyticum (WP_013298320.1, 42% identical), Methanothermobacter thermautotrophicus (WP_010877172.1, 38% identical), Desulfovibrio africanus (WP_014258953.1, 37% identity), Desulfotomaculum sp. LMa1 (WP_066665786.1, 37% identity), Desulfomicrobium baculatum (WP_015773055.1, 36% identity), a VnfD polypeptide from Fischerella muscicola (WP_016867598.1, 34% identity), and an AnfD polypeptide from the Opitutaceae bacterium TAV5 (WP_009512873.1, 33% identity).The NifD polypeptide has been reported and reviewed in Lawson and Smith (2002), Kim and Rees (1994), Eady (1996), Robson et al. (1989), Dilworth et al. (1988), Dilworth et al. (1993), Miller and Eady (1988), Chiu et al. (2001), Mayer et al. (1999), and Tezcan et al. (2005).

[0627] The NifD polypeptide of the iron-molybdenum subclass is a key subunit of the nitrogenase complex (the α subunit of the αβMoFe protein complex at the center of nitrogenase) and is the site of substrate reduction by the FeMo cofactor. As used herein, a functional NifD polypeptide is one that can form a functional nitrogenase protein with other required subunits (e.g., NifH and NifK) and FeMo or other cofactors.

[0628] As used herein, a "NifD polypeptide (ND) that is resistant to protease cleavage" is one that is resistant to cleavage at a predetermined site or within a predetermined region (e.g., within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO: 18) when the ND is introduced into plant mitochondria using MTP. As used herein, "resistant to protease cleavage" means that less than 10% of the NifD polypeptide is cleaved at the site or within the region when the NifD polypeptide is introduced into plant mitochondria using MTP. In preferred embodiments, less than 5% of the NifD polypeptide is cleaved at the site or within the region, and more preferably, substantially no cleavage or no detectable cleavage occurs. This NifD polypeptide can be "relatively resistant to cleavage" compared to a NifD polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 18, and is cleaved at least 1 / 5, preferably at least 1 / 10, of the frequency of a NifD polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 18.

[0629] As used herein, "an amino acid sequence other than RRNY (SEQ ID NO: 101) at positions corresponding to amino acids 97 to 100 of SEQ ID NO: 18" means a sequence containing four residues that are not RRNY at positions corresponding to amino acids 97 to 100 of SEQ ID NO: 18.

[0630] As used herein, an "AnfD polypeptide" refers to a NifD polypeptide, particularly a member of the oxidoreductase nitrogenase conserved superfamily cl30843, that contains the conserved domain TIGR01861 and shares at least 71% amino acid sequence identity with the Azotobacter vinelandii AnfD polypeptide (SEQ ID NO:216; Accession No. WP_012703361) when measured along the entire length of SEQ ID NO:216. This amino acid sequence is used herein as the reference sequence for AnfD. TIGR01861:AnfD represents the all-iron variant of the nitrogenase component Iα chain. Thus, as used herein, AnfD polypeptides are a subset of NifD polypeptides. AnfD polypeptides do not include molybdenum-type NifD polypeptides or vanadium-type NifD polypeptides (VnfD), nor do they include protochlorophyllide or chlorophyllide reductase polypeptides (Boyd and Peters, 2013). The amino acid sequences of AnfD polypeptides in protein sequence databases are usually annotated as AnfD polypeptides. As of January 2020, there were 156 specific amino acid sequences in the AnfD set of the NCBI protein database. Examples of naturally occurring AnfD polypeptides include those from Desulfovibrio sp. DV (accession number WP_075356167; 87.47% identity), Paenibacillus sp. FSL H7-0357 (accession number WP_038590013; 85.52% identity), Rhodobacter capsulatus (accession number WP_023922817; 80.31% identity), Methanosarcina acetivorans C2A (accession number WP_011021232; 77.13% identity), and Bacteroidales bacterium Barb7 (accession number OAV73823; 71.25% identity), each of which references SEQ ID NO: 216. Further examples are reported in McRose et al. (2017).

[0631] The functional AnfD polypeptide, like other functional NifD polypeptides, can function as the α protein structural element of an α2β2δ2 heterohexameric nitrogenase with a β protein (AnfK) and a δ protein (AnfG), providing the catalytic complex-bound FeFe-co for dinitrogen reduction.

[0632] As used herein, "NifK polypeptide" refers to a polypeptide whose sequence contains at least 31% amino acid identity to the amino acid sequence set forth as SEQ ID NO:3 and includes one or more of the conserved domains cd01974, TIGR01286, or cd01973 (in which case the NifK polypeptide is within the subclass of VnfK polypeptides) or cl02775, which contains the conserved domain TIGR02931 (in which case the NifK polypeptide is within the subclass of AnfK polypeptides). As used herein, NifK polypeptides include VnfK polypeptides from iron-vanadium nitrogenases and AnfK iron-binding polypeptides. Naturally occurring NifK polypeptides are typically 430-530 amino acids in length. Numerous NifK polypeptides have been identified, and many sequences are available in public databases. For example, the NifK polypeptide is isolated from Klebsiella michiganensis (accession number WP_049080161.1, 99% identical to SEQ ID NO: 3), Raoultella ornithinolytica (WP_044347163.1, 96% identical), Klebsiella variicola (SBM87811.1, 94% identical), Kluyvera intermedia (WP_047370272.1, 89% identical), Rahnella aquatilis (WP_014333919.1, 82% identical), Tolumonas auensis (WP_012728880.1, 75% identical), Pseudomonas stutzeri (WP_011912506.1, 68% identical), Vibrio natriegens (WP_065303473.1, 65% identity), Azoarcus toluclasticus (WP_018989051.1, 54% identity), Frankia sp. (prf||2106319A, 50% identity), and Methanosarcina acetivorans (WP_011021239.1, 31% identity).Although there are several examples of polypeptides annotated as "NifK" in the databases, these polypeptides share less than 31% identity with SEQ ID NO: 3, but do not contain any of the domains listed above and are therefore not included in the NifK polypeptides herein. NifK polypeptides have been reported and reviewed in Kim and Rees (1994), Eady (1996), Robson et al. (1989), Dilworth et al. (1988), Dilworth et al. (1993), Miller and Eady (1988), Igarashi and Seefeldt (2003), Fani et al. (2000), and Rubio and Ludden (2008).

[0633] NifK polypeptides of the iron-molybdenum subclass are key subunits of the nitrogenase complex, i.e., the β subunit of the αβMoFe protein complex at the center of nitrogenase. As used herein, a functional NifK polypeptide is one that can form a functional nitrogenase protein complex with other necessary subunits (e.g., NifD and NifH) and FeMo or other cofactors. In a preferred embodiment, when aligned with the amino acid sequence of SEQ ID NO:3, the amino acid sequence of a NifK polypeptide of the invention has the amino acid DLVR (SEQ ID NO:58) at its C-terminus, with an arginine as the C-terminal amino acid. Thus, NifK fusion polypeptides of the invention preferably have the same C-terminus as native NifK polypeptides. That is, NifK fusion polypeptides of the invention do not have artificial additions at the C-terminus. Such preferred NifK polypeptides are better able to form functional nitrogenase complexes with NifD and NifH polypeptides.

[0634] NifK polypeptides of the iron-molybdenum subclass are key subunits of the nitrogenase complex, the β subunit of the αβMoFe protein complex at the center of nitrogenase. As used herein, a functional NifK polypeptide is one that can form a functional nitrogenase protein complex with other required subunits (e.g., NifD and NifH) and FeMo or other cofactors. In a preferred embodiment, the amino acid sequences of NifK polypeptides and truncated NifK polypeptides of the invention have the amino acid DLVR (SEQ ID NO:58) at their C-terminus, with arginine as the C-terminal amino acid, when aligned with the amino acid sequence of SEQ ID NO:3. In other preferred embodiments, the amino acid sequences of NifK polypeptides and truncated NifK polypeptides of the invention have the amino acid sequences DLIR (SEQ ID NO:239), DVVR (SEQ ID NO:240), DIIR (SEQ ID NO:241), DLTR (SEQ ID NO:242), or INVW (SEQ ID NO:243) at their C-terminus, which are typically not present in native NifK sequences. The NifK polypeptides and NifK fusion polypeptides of the invention, and the truncated NifK polypeptides derived therefrom, preferably have the same C-terminus as the native NifK polypeptide, i.e., they have no artificial additions at the C-terminus and do not lack any amino acids from the C-terminus when aligned with the native NifK polypeptide. Such preferred NifK polypeptides are better able to form functional nitrogenase complexes with NifD and NifH polypeptides.

[0635] As used herein, an "AnfK polypeptide" refers to a polypeptide that contains the conserved domain TIGR02931 and is a member of the oxidoreductase nitrogenase conserved superfamily cl02775, and that shares at least 54% amino acid sequence identity with the Azotobacter vinelandii AnfK polypeptide (SEQ ID NO:217; Accession No. WP_012703359) when measured along the entire length of SEQ ID NO:217. This amino acid sequence is used herein as the reference sequence for AnfK. TIGR02931:AnfK represents the all-iron variant of the nitrogenase component Iβ chain. As used herein, an AnfK polypeptide can also be a NifK polypeptide that shares at least 31% amino acid identity with SEQ ID NO:3. Other AnfK polypeptides may share less homology, such as only 25-31% identity with SEQ ID NO:3, but are still included within the AnfK polypeptides of the present invention. AnfK polypeptides do not include molybdenum-type NifK polypeptides or vanadium-type NifK polypeptides (VnfK). The AnfK fusion polypeptides and truncated AnfK polypeptides of the present invention preferably have the same C-terminus as native AnfK polypeptides. That is, they have no artificial additions at the C-terminus and no amino acids deleted from the C-terminus when aligned with native NifK polypeptides (e.g., SEQ ID NO: 217). In preferred embodiments, the AnfK fusion polypeptides and truncated AnfK polypeptides of the present invention have the amino acids LNVW (SEQ ID NO: 244), LNTW (SEQ ID NO: 245), LNMW (SEQ ID NO: 246), LAMW (SEQ ID NO: 247), or LSVW (SEQ ID NO: 248) at their C-terminus. The amino acid sequences of AnfK polypeptides in protein sequence databases are typically annotated as AnfK polypeptides. As of January 2020, there were 155 specific amino acid sequences in the AnfK set of the NCBI protein database that were distinct from the sequences of the molybdenum-type NifK and VnfK polypeptides.Examples of naturally occurring AnfK polypeptides include AnfK polypeptides from Azomonas agilis (accession number WP_144571040; 91.34% identity), Clostridium sp. BL-8 (accession number WP_077859050; 78.35% identity), Lucifera butyrica (accession number WP_122630336; 62.34% identity), and Rhodoblastus acidophilus (accession number WP_088520366; 54% identity), each of which references SEQ ID NO: 217.

[0636] The functional AnfK polypeptide, like other functional NifK polypeptides, can function as the β protein structural element of an α2β2δ2 heterohexameric nitrogenase that contains an α protein structural element (AnfD) and a δ protein structural element (AnfG) and can form a complex with an active site for dinitrogen reduction on FeFe-co.

[0637] In natural bacteria, the NifB polypeptide converts the [4Fe-4S] cluster into NifB-co (a more highly nucleated Fe-S cluster with a central C atom, which serves as a precursor for the synthesis of FeMo-co, FeV-co, and FeFe-co) (Guo et al., 2016). NifB is therefore essential for nitrogenase function, catalyzing the first committed step in the synthetic pathways of FeMo-co, FeV-co, and FeFe-co. The NifB-co product of NifB can bind to the NifE-NifN complex and can be shuttled from NifB to NifE-NifN by the metallocluster carrier protein NifX.

[0638] As used herein, "NifB polypeptide" refers to a polypeptide whose amino acid sequence contains at least 27% amino acids identical to the amino acid sequence set forth as SEQ ID NO:4. Most NifB polypeptides contain one or more of the conserved domains TIGR01290, the NifB conserved domain cd00852, the NifX-NifB superfamily conserved domain cl00252, and the Radical_SAM conserved domain cd01335. As used herein, NifB polypeptides include naturally occurring polypeptides annotated as having NifB function but lacking one of these domains. NifB polypeptides from Klebsiella, Azotobacter, Rhizobium, Bradyrhizobium, and other bacteria contain a C-terminal NifX-like extension, whereas most archaeal NifB polypeptides lack the NifX-like domain and are therefore referred to as "truncated NifB polypeptides." Naturally occurring NifB polypeptides are typically 440-500 amino acids in length, and the natural monomer has a molecular weight of approximately 50 kDa. Many NifB polypeptides have been identified, and many sequences are available in public databases.For example, the NifB polypeptide has been identified from Raoultella ornithinolytica (accession number WP_041145602.1, 91% identical to SEQ ID NO: 4), Kosakonia radicincitans (WP_043953592.1, 80% identical), Dickeya chrysanthemi (WP_040003311.1, 76% identical), Pectobacterium atrosepticum (WP_011094468.1, 70% identical), Brenneria goodwinii (WP_048638849.1, 63% identical), Halorhodospira halophila (WP_011813098.1, 59% identical, lacks the NifX domain), Methanosarcina barkeri (WP_048108879.1, 50% identical, lacks the NifX domain), Clostridium NifB polypeptides have been reported from Desulfovibrio purinilyticum (WP_050355163.1, 40% identity, lacking the NifX domain), and Desulfovibrio salexigens (WP_015850328.1, 27% identity). As used herein, a "functional NifB polypeptide" is a NifB polypeptide that is capable of forming NifB-co from a [4Fe-4S] cluster. Functional NifB requires S-adenosyl-methionine (SAM) for its function. NifB polypeptides are described and reviewed in Curatti et al. (2006) and Allen et al. (1995).

[0639] (2011) examined the relationship between Anf / Vnf / NifDKEN and NifB from 40 taxa and concluded that (1) horizontal gene transfer of the Nif cluster encoding NifB lacking the C-terminal NifX domain occurred from the ancestor of methanogens in the Methanosarcinales to the ancestor of anaerobic Firmicutes, allowing these two organisms to coexist in anaerobic environments and utilize molybdenum, and (2) after this horizontal gene transfer event, a fusion of NifB and NifX occurred within the Firmicutes, from which the nitrogen-fixing bacterial lineage evolved. The following evidence supports this theory: (1) none of the methanogenic archaea (Methanococcales, Methanosarcinales, and Methanobacteriales) possess a NifB with a C-terminal NifX domain; (2) NifB sequences from Methanobacteriales and Methanococcales show early divergence from Methanosarcinales and bacterial NifB sequences; and (3) some anaerobic Firmicutes, Chloroflexi, and Proteobacteria with NifB without a C-terminal NifX domain diverged from the Firmicutes lineage early, likely some time after a horizontal gene transfer event for Nif.

[0640] To determine the presence or absence of a C-terminal NifX domain in a NifB polypeptide, the NifB amino acid sequence can be aligned with representative NifB sequences (e.g., sequences from Klebsiella michiganensis NifB (accession number P10930), Klebsiella michiganensis NifX (KZT46636.1), NifY (KZT46633.1), A. vinelandii NifX (AGK13791.1), NifY (AGK13792.1), NafY (AGK13761.1), and the NifX / NifY / NafY / VnfX family of proteins (AGK14217.1)) using the Constraint-based Multiple Alignment Tool (COBALT, NCBI, www.ncbi.nlm.nih.gov / tools / cobalt / re_cobalt.cgi). The “dinitrogenase FeMo-cofactor binding site” (Pfam family PF02579) within each sequence can be identified by PfamScan (EMBL-EBI, www.ebi.ac.uk / Tools / pfa / pfamscan / ) using the Pfam-A database with an expectation value set to 10.

[0641] The NifEN complex is a scaffolding complex required for the correct assembly of dinitrogenase (it functions as a scaffold for the maturation of NifB-co to FeMo-co, a process that also requires NifH function) and is structurally similar to dinitrogenase (Fay et al., 2016). The NifEN complex consists of two subunits, NifE and NifN, forming a heterotetramer (referred to herein as ENα2β2). In natural bacteria, the NifE polypeptide is the α subunit of the ENα2β2 tetramer with the NifN polypeptide. This ENα2β2 tetramer is required for FeMo-co synthesis and has been proposed to function as a scaffold for FeMo-co synthesis on the surface.

[0642] As used herein, "NifE polypeptide" refers to a polypeptide whose sequence contains at least 32% amino acid identity to the amino acid sequence set forth as SEQ ID NO:5 and contains one or both of the domains TIGR01283 and PRK14478. Members of the TIGR01283 domain protein family are also members of the superfamily cl02775. Naturally occurring NifE polypeptides are typically 440-490 amino acids in length, and the naturally occurring monomer has a molecular weight of approximately 50 kDa. Many NifE polypeptides have been identified, and many sequences are available in public databases. For example, the NifE polypeptide is isolated from Klebsiella michiganensis (accession number WP_049114606.1, 99% identical to SEQ ID NO: 5), Klebsiella variicola (SBM87755.1, 92% identical), Dickeya paradisiaca (WP_012764127.1, 89% identical), Tolumonas auensis (WP_012728883.1, 75% identical), Pseudomonas stutzeri (WP_003297989.1, 69% identical), Azotobacter vinelandii (WP_012698965.1, 62% identical), Trichormus azollae (WP_013190624.1, 55% identical), Paenibacillus NifE polypeptides have been reported from S. durus (WP_025698318.1, 50% identity), Sulfuricurvum kujiense (WP_013460149.1, 44% identity), Methanobacterium formicicum (AIS31022.1, 39% identity), Anaeromusa acidaminophila (WP_018701501.1, 35% identity), and Megasphaera cerevisiae (WP_048514099.1, 32% identity). As used herein, a "functional NifE polypeptide" is a NifE polypeptide that can assemble with NifN to form a functional tetramer, and this complex is capable of synthesizing FeMo-co. Synthesis of FeMo-co involves other polypeptides, including NifH and NifB, and potentially NifX.The NifE polypeptide is described and reviewed in Fay et al. (2016), Hu et al. (2005), Hu et al. (2006), and Hu et al. (2008).

[0643] The NifF polypeptide in natural nitrogen-fixing bacteria is a flavodoxin, an electron donor to NifH. As used herein, "NifF polypeptide" refers to a polypeptide whose sequence contains at least 34% amino acid identity to the amino acid sequence shown as SEQ ID NO:6 and contains one or both of the flavodoxin long domain TIGR01752 and the flavodoxin FLDA domain found in Nif proteins from Azobacter and other bacterial genera PRK09267. NifF polypeptides include flavodoxins associated with pyruvate formate lyase activation and cobalamin-dependent methionine synthase activity in non-nitrogen-fixing bacteria, but exclude other flavodoxins involved in broader functions. Naturally occurring NifF polypeptides are typically 160-200 amino acids long, with the natural monomer having a molecular weight of approximately 19 kDa. Numerous NifF polypeptides have been identified, and numerous sequences are available in public databases. For example, the NifF polypeptide can be detected in Klebsiella michiganensis (accession number WP_004122417.1, 99% identical to SEQ ID NO: 6), Klebsiella variicola (WP_040968713.1, 85% identical), Kosakonia radicincitans (WP_035885760.1, 76% identical), Dickeya chrysanthemi (WP_039999438.1, 72% identical), Brenneria goodwinii (WP_048638838.1, 62% identical), Methylomonas methanica (WP_064006977.1, 56% identical), Azotobacter vinelandii (WP_012698862.1, 50% identical), Chlorobaculum tepidum (WP_010933399.1, 39% identity), Campylobacter showae (WP_002949173.1, 37% identity), and Azotobacter chromococcum (WP_039801725.1, 34% identity). As used herein, a "functional NifF polypeptide" is a NifF polypeptide that is capable of being an electron donor to a NifH polypeptide.The NifF polypeptide is described and reviewed in Drummond (1985).

[0644] As used herein, an "AnfG polypeptide" is a member of the nitrogenase conserved superfamily cl03910 (pfam03139-AnfG) that contains the conserved domain TIGR02929 and shares at least 42% amino acid sequence identity with the Azotobacter vinelandii AnfG polypeptide (SEQ ID NO:219; Accession No. WP_012703360) when measured along the entire length of SEQ ID NO:219. This amino acid sequence is used herein as the reference sequence for AnfG. TIGR02929 represents the all-iron variant of the nitrogenase component I δ chain. AnfG polypeptides do not include vanadium-type NifG polypeptides (VnfG). The amino acid sequences of AnfG polypeptides in protein sequence databases are typically annotated as AnfG polypeptides. As of January 2020, there were 150 specific amino acid sequences in the AnfG set in the NCBI protein database. Examples of naturally occurring AnfG polypeptides include AnfG polypeptides from Azomonas agilis (accession number WP_144571041; 84.73% identity), Firmicutes bacterium (accession number HBE76208; 70.37% identity), Sporomusa termitida (accession number WP_144349445; 68.75% identity), Rhodovulum viride (accession number WP_112317428; 57.14% identity), and Megasphaera cerevisiae (accession number WP_048515315; 42.86% identity), each of which references SEQ ID NO: 219.

[0645] A functional AnfG polypeptide can function as the δ protein structural element of α2β2δ2 heterohexameric nitrogenase.

[0646] The NifJ polypeptide in natural bacteria is a pyruvate:flavodoxin (ferredoxin) oxidoreductase, an electron donor to NifH. As used herein, "NifJ polypeptide" refers to a polypeptide whose sequence contains at least 40% amino acid identity to the amino acid sequence set forth as SEQ ID NO:7 and contains the conserved domain TIGR02176. Naturally occurring NifJ polypeptides are typically 1100-1200 amino acids in length, and the naturally occurring monomer has a molecular weight of approximately 128 kDa. Numerous NifJ polypeptides have been identified, and many sequences are available in public databases. For example, the NifJ polypeptide is isolated from Klebsiella michiganensis (accession number WP_024360006.1, 99% identical to SEQ ID NO: 7), Raoultella ornithinolytica (WP_044347157.1, 95% identical), Klebsiella quasipneumoniae (WP_050533844.1, 92% identical), Kosakonia oryzae (WP_064566543.1, 82% identical), Dickeya solani (WP_057084649.1, 78% identical), Rahnella aquatilis (WP_014683040.1, 72% identical), Thermoanaerobacter mathranii (WP_013149847.1, 64% identical), Clostridium NifJ polypeptides have been reported from Bacillus botulinum (WP_053341220.1, 60% identity), Spirochaeta africana (WP_014454638.1, 52% identity), and Vibrio cholerae (CSA83023.1, 40% identity). As used herein, a "functional NifJ polypeptide" is a NifJ polypeptide that is capable of being an electron donor to a NifH polypeptide. NifJ polypeptides are described and reviewed in Schmitz et al. (2001).

[0647] The NifM polypeptide in native bacteria is a polypeptide required for the maturation of some, but not all, NifH polypeptides. Without NifM, K. oxytoca NifH, present at low levels in E. coli and yeast when heterologously expressed, could not donate electrons to NifD-NifK. As used herein, "NifM polypeptide" refers to a polypeptide whose sequence contains at least 26% amino acid identity to the amino acid sequence set forth as SEQ ID NO:8 and contains the conserved domain TIGR02933. The NifM polypeptide is homologous to peptidyl-propyl cis-trans isomerases (PPIases), a group of enzymes that promote protein folding by catalyzing the cis-trans isomerization of proline-imide peptide bonds. Because it contains a PpiC-type domain, it appears to be an accessory protein for some NifH polypeptides, including at least some VnfH and AnfH polypeptides. Native NifM polypeptides are typically 240-300 amino acids in length, and the native monomer has a molecular weight of approximately 30 kDa. Many NifM polypeptides have been identified, and many sequences are available in public databases.For example, the NifM polypeptide is capable of binding to Klebsiella oxytoca (accession number WP_064342940.1, 99% identical to SEQ ID NO: 8), Klebsiella michiganensis (WP_004122413.1, 97% identical), Raoultella ornithinolytica (WP_044347181.1, 85% identical), Klebsiella variicola (WP_063105800.1, 75% identical), Kosakonia radicincitans (WP_035885759.1, 59% identical), Pectobacterium atrosepticum (WP_011094472.1, 42% identical), Brenneria goodwinii (WP_048638837.1, 33% identical), Pseudomonas aeruginosa PAO1 (CAA75544.1, 28% identity), Marinobacterium sp. AK27 (WP_051692859.1, 27% identity), and Teredinibacter turnerae (WP_018415157.1, 26% identity). As used herein, a "functional NifM polypeptide" is a NifM polypeptide that can complex with a NifH polypeptide for maturation of the NifH polypeptide. NifM polypeptides are described and reviewed in Petrova et al. (2000).

[0648] The NifN polypeptide in natural bacteria is the β subunit of an ENα2β2 tetramer with the NifE polypeptide. This ENα2β2 tetramer is required for the synthesis of FeMo-co and has been proposed to function as a scaffold for surface synthesis of FeMo-co. As used herein, "NifN polypeptide" refers to (i) a polypeptide whose sequence contains at least 76% amino acid identity to the amino acid sequence set forth in SEQ ID NO:9 and / or (ii) a polypeptide whose sequence contains at least 34% amino acid identity to the amino acid sequence set forth in SEQ ID NO:9 and contains one or more of the conserved domains TIGR01285, cd01966, and PRK14476. NifN is structurally related to the molybdenum-iron protein β chain NifK. While polypeptides containing the conserved TIGR01285 domain encompass most examples of NifN polypeptides, some NifN polypeptides, such as the putative NifN of Chlorobium tepidum, are excluded, and therefore the definition of NifN is not limited to polypeptides containing the conserved TIGR01285 domain. Members of the PRK14476 domain protein family are also members of the superfamily cl02775. Native NifN polypeptides are typically 410-470 amino acids long, but when naturally fused to NifE, they are approximately 900 amino acid residues long, and the native monomer has a molecular weight of approximately 50 kDa. Numerous NifM polypeptides have been identified, and many sequences are available in public databases.For example, the NifN polypeptide is isolated from Klebsiella oxytoca (accession number WP_064391778.1, 97% identical to SEQ ID NO: 9), Kluyvera intermedia (WP_047370268.1, 80% identical), Rahnella aquatilis (WP_014683026.1, 70% identical), Brenneria goodwinii (WP_048638830.1, 65% identical), Methylobacter tundripaludum (WP_027147663.1, 46% identical), Calothrix parietina (WP_015195966.1, 41% identical), Zymomonas mobilis (WP_023593609.1, 37% identical), Paenibacillus NifN polypeptides have been reported from Bacillus massiliensis (WP_025677480.1, 35% identity), and Desulfitobacterium hafniense (WP_018306265.1, 34% identity). As used herein, a "functional NifN polypeptide" is a NifN polypeptide that can assemble with NifE to form a functional tetramer, and this complex is capable of synthesizing FeMo-co. NifN polypeptides are described and reviewed in Fay et al. (2016), Brigle et al. (1987), Fani et al. (2000), and Hu et al. (2005).

[0649] The NifQ polypeptide in natural bacteria is likely the initial MoO4 2-This polypeptide is involved in the synthesis of FeMo-co in processing. A conserved C-terminal cysteine ​​residue may be involved in metal binding. As used herein, "NifQ polypeptide" refers to a polypeptide whose sequence contains at least 34% amino acid identity to the amino acid sequence set forth as SEQ ID NO: 10 and is a member of the CL04826 domain protein family and the pfam04891 domain protein family. Naturally occurring NifQ polypeptides are typically 160-250 amino acids in length, although lengths as long as 350 amino acid residues are possible. The naturally occurring monomer has a molecular weight of approximately 20 kDa. Numerous NifQ polypeptides have been identified, and many sequences are available in public databases. For example, NifQ polypeptides are identified from Klebsiella oxytoca (accession number WP_064391765.1, 95% identical to SEQ ID NO: 10), Klebsiella variicola (CTQ06350.1, 75% identical), Kluyvera intermedia (WP_047370257.1, 63% identical), Pectobacterium atrosepticum (WP_043878077.1, 59% identical), Mesorhizobium metallidurans (WP_008878174.1, 46% identical), Rhodopseudomonas palustris (WP_011501504.1, 42% identical), Paraburkholderia sprentiae (WP_027196569.1, 41% identical), Burkholderia stabilis (GAU06296.1, 39% identity), and Cupriavidus oxalaticus (WP_063239464.1, 34% identity). As used herein, a "functional NifQ polypeptide" refers to a polypeptide derived from MoO4 2- NifQ polypeptides are described and reviewed in Allen et al. (1995) and Siddavattam et al. (1993).

[0650] In naturally occurring bacteria, NifS polypeptides are cysteine ​​desulfurase enzymes involved in the biosynthesis of iron-sulfur (FeS) clusters, e.g., by mobilizing sulfur for the synthesis and repair of Fe-S clusters. As used herein, "NifS polypeptide" refers to (i) a polypeptide whose sequence contains at least 90% amino acid identity to the amino acid sequence set forth in SEQ ID NO: 19 and / or (ii) a polypeptide whose sequence contains at least 36% amino acid identity to the amino acid sequence set forth in SEQ ID NO: 19 and contains one or both of the conserved domains TIGR03402 and COG1104. The TIGR03402 domain protein family includes a clade that is almost always found in extended nitrogen fixation systems, as well as a second clade that is more closely related to IscS than this first clade and is also part of the NifS-like / NifU-like system. The TIGR03402 domain protein family does not extend to more distant clades found in epsilonproteobacteria (e.g., Helicobacter pylori, also named NifS in the literature), which are instead organized into TIGR03403. The COG1104 domain protein family includes cysteine ​​sulfinate desulfinases / cysteine ​​desulfurases or related enzymes. Some NifS polypeptides contain the aspartate aminotransferase domain cl18945. Native NifS polypeptides are typically 370–440 amino acids long, and the native monomer has a molecular weight of approximately 43 kDa. Numerous NifS polypeptides have been identified, and numerous sequences are available in public databases.For example, the NifS polypeptide can be detected in Klebsiella michiganensis (accession number WP_004138780.1, 99% identical to SEQ ID NO: 19), Raoultella terrigena (WP_045858151.1, 89% identical), Kluyvera intermedia (WP_047370265.1, 80% identical), Rahnella aquatilis (WP_014333911.1, 73% identical), Agarivorans gilvus (WP_055731597.1, 64% identical), Azospirillum brasilense (WP_014239770.1, 60% identical), Desulfosarcina cetonica (WP_054691765.1, 55% identical), Clostridium NifS polypeptides have been reported from Clostridium intestinale (WP_021802294.1, 47% identity), Clostridium paucivorans (WP_026894054.1, 36% identity), and Bacillus coagulans (WP_061575621.1, 42% identity and located in COG1104). As used herein, a "functional NifS polypeptide" is a NifS polypeptide that can function in the biosynthesis and / or repair of iron-sulfur (Fe-S) clusters. NifS polypeptides are described and reviewed in Clausen et al. (2000), Johnson et al. (2005), Olson et al. (2000), and Yuvaniyama et al. (2000).

[0651] The NifU polypeptide in natural bacteria is a molecular scaffold polypeptide involved in the biosynthesis of the iron-sulfur (FeS) cluster of nitrogenase components. As used herein, "NifU polypeptide" refers to a polypeptide whose sequence contains at least 31% amino acid identity to the amino acid sequence set forth as SEQ ID NO:12 and contains the TIGR02000 domain. Members of the TIGR02000 domain protein family are particularly involved in nitrogenase maturation. NifU contains an N-terminal domain (pfam01592) and a C-terminal domain (pfam01106). Three distinct but homologous Fe-S cluster assembly systems have been described: Isc, Suf, and Nif. The Nif system (of which NifU is a part) is involved in donating the Fe-S cluster to nitrogenase in many nitrogen-fixing species. Isc and Suf homologs with domain architectures equivalent to those in Helicobacter and Campylobacter are excluded from this definition of NifU, and are therefore specific to the NifU polypeptide involved in nitrogenase maturation. Members of the related TIGR01999 domain protein family, including IscU proteins (e.g., from E. coli, Saccharomyces cerevisiae, and Homo sapiens) that contain homologs of the N-terminal region of NifU, are also excluded from this definition of NifU. Naturally occurring NifU polypeptides are typically 260–310 amino acids long, and the natural monomer has a molecular weight of approximately 29 kDa. Numerous NifU polypeptides have been identified, and numerous sequences are available in public databases.For example, the NifU polypeptide is isolated from Klebsiella michiganensis (accession number WP_049136164.1, 97% identical to SEQ ID NO: 12), Klebsiella variicola (WP_050887862.1, 90% identical), Dickeya solani (WP_057084657.1, 80% identical), Brenneria goodwinii (WP_048638833.1, 73% identical), Tolumonas auensis (WP_012728889.1, 66% identical), Agarivorans gilvus (WP_055731596.1, 58% identical), Desulfocurvus vexinensis (WP_028587630.1, 54% identical), Rhodopseudomonas NifU polypeptides have been reported from Bacillus palustris (WP_044417303.1, 49% identity), Helicobacter pylori (WP_001051984.1, 31% identity), and Sulfurovum sp. PC08-66 (KIM05011.1, 31% identity). As used herein, a "functional NifU polypeptide" is a NifU polypeptide that can function as a molecular scaffold polypeptide involved in iron-sulfur (Fe-S) cluster biosynthesis. NifU polypeptides have been described and reviewed in Hwang et al. (1996), Mühlenhoff et al. (2003), and Ouzounis et al. (1994).

[0652] NifS is a pyridoxal phosphate (PLP, vitamin B6)-dependent cysteine ​​desulfurase that generates the inorganic sulfide required for the synthesis of an Fe-S cluster from cysteine. This reaction generates alanine as a byproduct. The reaction proceeds through a protein-bound cysteine ​​persulfide intermediate formed by nucleophilic attack of a highly conserved cysteine ​​residue (Cys325 in Azotobacter vinelandii) on the cysteine-PLP adduct (Zheng et al., 1994). This sulfide is donated to NifU to sequentially form an [Fe2S2] cluster and an [Fe4S4] cluster. The NifS enzyme functions as a homodimer in bacteria.

[0653] NifU functions as a homodimer and provides a scaffold for the formation of [Fe4S4] clusters. The NifU polypeptide contains three domains: an N-terminal scaffolding domain, a central domain, and a C-terminal scaffolding domain (Smith et al., 2005). The N-terminal domain shares significant sequence homology with the IscU protein from bacteria and the Isu protein from eukaryotes, whereas the C-terminal domain is homologous to the Nfu protein found in mitochondria and chloroplasts. The central domain contains the permanent redox-active [Fe2S2] that is thought to be not transferred to other Nif proteins due to its stability. 2+Each NifU subunit contains one [Fe2S2] cluster. This cluster is thought to be coordinated by four conserved cysteine ​​residues (Cys137, 139, 172, and 175 in A. vinelandii NifU) (Fu et al., 1994). NifU forms homodimers, and its N-terminal domain can bind to one [Fe2S2] cluster per monomer. The [Fe2S2] clusters in the monomers can undergo reductive fusion to form one [Fe4S4] cluster per NifU dimer. The pair of [Fe4S4] clusters is then supplied from NifU to NifB and processed into the 8Fe core on NifB, which is then used to synthesize FeMoco. In a branched pathway for Fe-S clusters, one [Fe4S4] cluster attached to either the N- or C-terminal scaffolding domain of NifU is transferred to apo-NifH to mature the NifH protein, nitrogenase reductase (Smith et al., 2005). It has also been proposed that NifU donates two [Fe4S4] clusters to the NifD-NifK protein complex (herein termed stage 0 DK), which then condenses the pair of clusters into the mature P cluster [Fe8-S7] (Dos Santos et al., 2004). These N-terminal clusters are thought to be extremely unstable and are not retained during purification (Smith et al., 2005). The C-terminal domain can retain one [Fe4S4] cluster per monomer. Unlike the N-terminal cluster, assembly of the C-terminal [Fe4S4] cluster is rapid, and no intermediate [Fe2S2] cluster has been detected (Smith et al., 2005). The C-terminal cluster is more stable than the N-terminal cluster and can therefore be retained during purification; however, upon reduction with dithionite, the C-terminal cluster is rapidly disassembled (Smith et al., 2005). Dos Santos and coworkers showed that both the N- and C-terminal clusters can be transferred to apo-NifH using a cysteine-to-alanine mutation in NifU.

[0654] Lopez-Torrejon et al. (2016) reported that the expression of both NifH and NifM could produce NifH protein in yeast mitochondria capable of donating electrons to holoNifD-NifK. They found that NifS and NifU were not required for the production of this functional NifH protein in yeast cells. They concluded that an endogenous iron-sulfur cluster assembly pathway in yeast cells (probably the mitochondrially located Nfs1 and Nfu1 proteins, which are related to each other in yeast) could donate the [Fe4S4] cluster to NifH. Therefore, NifS and NifU may not be required for the reconstitution of NifH, Fe protein, or dinitrogenase reductase in yeast, but they may be required for the maturation and function of NifB and / or NifD-NifK. It is unknown whether plant mitochondria have a similar intrinsic capacity to form [Fe4S4] clusters sufficient for nitrogenase activity.

[0655] The naturally occurring NifV polypeptide in bacteria is a homocitrate synthase (EC 2.3.3.14) that generates homocitrate by transferring an acetyl group from acetyl-coenzyme A (acetyl-CoA) to 2-oxoglutarate. Homocitrate is then used to synthesize FeMo-co, FeV-co, and FeFe-co. As used herein, "NifV polypeptide" refers to a polypeptide whose sequence contains at least 39% amino acid identity to the amino acid sequence set forth as SEQ ID NO: 13 and contains both the TIGR02660 and DRE_TIM domains. Members of the TIGR02660 domain protein family are homologous to enzymes involved in processes other than nitrogen fixation, including 2-isopropylmaleate synthase, (R)-citramaleate synthase, and homocitrate synthase. The cd07939 domain protein family also includes the NifV proteins of Heliobacterium chlorum and Gluconacetobacter diazotrophicus, which appear to be orthologs of FrbC. This family belongs to the DRE-TIM metallolyase superfamily, which includes 2-isopropylmaleate synthase (IPMS), alpha-isopropylmaleate synthase (LeuA), 3-hydroxy-3-methylglutaryl-CoA lyase, homocitrate synthase, citramaleate synthase, 4-hydroxy-2-oxovalerate aldolase, re-citrate synthase, transcarboxylase 5S, pyruvate carboxylase, AksA, and FrbC. All of these members share a conserved triosephosphate isomerase (TIM) barrel domain consisting of a core beta(8)-alpha(8) motif with eight parallel beta strands forming a closed barrel surrounded by eight alpha helices. This domain has a catalytic center containing a divalent cation binding site formed by a cluster of invariant residues that cap the core of the barrel.Additionally, this catalytic site is the basis for a domain designated "DRE-TIM" and contains three invariant residues: aspartic acid (D), arginine (R), and glutamic acid (E). Native NifV polypeptides are typically 360-390 amino acids in length, although some members are approximately 490 amino acid residues in length, and the native monomer has a molecular weight of approximately 41 kDa. Many NifV polypeptides have been identified, and many sequences are available in public databases. For example, NifV polypeptides have been identified in Klebsiella michiganensis (accession number WP_049083341.1, 95% identical to SEQ ID NO: 13), Raoultella ornithinolytica (WP_045858154.1, 86% identical), Kluyvera intermedia (WP_047370264.1, 81% identical), Dickeya dadantii (WP_038912041.1, 70% identical), Brenneria goodwinii (WP_048638835.1, 59% identical), Magnetococcus marinus (WP_011712856.1, 46% identical), Sphingomonas wittichii (WP_037528703.1, 43% identical), Frankia sp. EI5c (OAA29062.1, 41% identity), and Clostridium sp. Maddingley MBC34-26 (EKQ56006.1, 39% identity). As used herein, a "functional NifV polypeptide" is a NifV polypeptide that can function as a homocitrate synthase. NifV polypeptides are described and reviewed in Hu et al. (2008), Lee et al. (2000), Masukawa et al. (2007), and Zheng et al. (1997).

[0656] The NifX polypeptide in Azotobacter vinelandii binds to NifB-co (Fe6-S9-C), which is then transferred to NifE-NifN for the assembly of FeMo-co (Hernandez et al., 2007). It has also been shown that NifE and NifN exchange VK-clusters (Fe8-S9-C or Mo-Fe7-S9-C, Jimenez-Vincente et al., 2015). This suggests a role for NifE-NifN as a transient reservoir for the FeMo-co precursor. Hernandez et al. (2007) suggested that NifX may act as a chaperone to stabilize the NifE-NifN or NifD-NifK complex during the transfer of FeMo-co to apo-NifD-NifK and / or to regulate FeMo-co synthesis by repositioning these proteins in a favorable orientation for FeMo-co. Activation of apo-NifD-NifK by exogenous FeMo-co with dinitrogenase extracted from an A. vinelandii mutant lacking a different accessory protein combination, NifY / NafY / NifX, indicated that NifX can also support FeMo-co insertion of apo-NifD-NifK (Rubio et al., 2002). This additional function of NifX may be important for the retention of acetylene reduction activity in the Klebsiella ΔnifY mutant shown by Homer et al. (1993).

[0657] In natural bacteria, NifX polypeptides are polypeptides involved in the synthesis of FeMo-co, assisting in the transfer of at least the FeMo-co precursor from NifB to NifE-NifN or FeMo-co to NifD-NifK. As used herein, "NifV polypeptide" refers to a polypeptide whose sequence contains at least 29% amino acid identity to the amino acid sequence set forth as SEQ ID NO:14 and contains one or both of the conserved domains TIGR02663 and cd00853. Because NifX is part of a larger family of iron-molybdenum cluster-binding proteins that includes several NifB sequences and NifY, the C-terminal regions of NifX, NafY, and several NifB polypeptides all contain the pfam02579 domain, each of which is involved in the synthesis of one or more or all of FeMo-co, FeV-co, or FeFe-co. Other NifB polypeptides, particularly those from methanogenic archaea and some anaerobic Firmicutes species (including the NifBs from H. halophila, M. barkeri, and C. purinilyticum mentioned above), lack the NifX-like domain (Boyd et al., 2011). Some NifX polypeptides are annotated as NifY in databases, and vice versa. Natural NifX polypeptides are produced as part of NifB polypeptides rather than as natural fusions, and are typically 110–160 amino acids long, with the natural monomer having a molecular weight of approximately 15 kDa. Numerous NifX polypeptides have been identified, and numerous sequences are available in public databases.For example, the NifX polypeptide is effective against Klebsiella michiganensis (accession number WP_049070199.1, 97% identical to SEQ ID NO: 14), Klebsiella oxytoca (WP_064342937.1, 97% identical), Raoultella ornithinolytica (WP_044347173.1, 91% identical), Klebsiella variicola (WP_044612922.1, 83% identical), Kosakonia radicincitans (WP_043953583.1, 75% identical), Dickeya chrysanthemi (WP_039999416.1, 68% identical), Rahnella aquatilis (WP_047608097.1, 58% identical), Azotobacter NifX polypeptides have been reported from Pseudomonas chroococcum (WP_039800848.1, 34% identity), Beggiatoa leptomitiformis (WP_062149047.1, 33% identity), and Methyloversatilis discipulorum (WP_020165972.1, 29% identity). As used herein, a "functional NifX polypeptide" is a NifX polypeptide that can transfer the FeMo-co precursor from NifB to NifD-NifK. NifX polypeptides are described and reviewed in Allen et al. (1994) and Shah et al. (1999).

[0658] In natural bacteria, NifY polypeptides are polypeptides involved in the synthesis of FeMo-co, at least in transferring the FeMo-co precursor from NifB to NifE-NifN. As used herein, a "NifY polypeptide" is a polypeptide whose sequence contains at least 34% amino acid identity to the amino acid sequence set forth as SEQ ID NO:15 and contains one or both of the conserved domains TIGR02663 and cd00853. Because NifY is part of a larger family of iron-molybdenum cluster-binding proteins that includes several NifB sequences and NifY, the C-terminal regions of NifX, NafY, and several NifB polypeptides all contain the pfam02579 domain, each of which is involved in the synthesis of FeMo-co. Numerous NifY polypeptides have been identified, and many sequences are available in public databases. For example, the NifY polypeptide may be isolated from Klebsiella michiganensis (accession number WP_049089500.1, 99% identical to SEQ ID NO: 15), Klebsiella oxytoca (WP_064342935.1, 98% identical), Klebsiella quasipneumoniae (WP_044524054.1, 90% identical), Klebsiella variicola (WP_049010739.1, 81% identical), Kluyvera intermedia (WP_047370270.1, 69% identical), Dickeya chrysanthemi (WP_039999411.1, 62% identical), Serratia sp. ATCC 39006 (WP_037382461.1, 57% identical), Rahnella These NifY polypeptides have been reported from Azotobacter vinelandii (WP_012698835.1, 34% identity), ...

[0659] When apo-NifD-NifK is isolated from NifB or NifN-NifE mutant strains of K. oxytoca or A. vinelandii, it is associated with an additional polypeptide, termed the γ protein (Paustian et al., 1990; Homer et al., 1993), which forms a heterohexamer (α2β2γ2) with the NifD and NifK polypeptides. In K. oxytoca, a third polypeptide is encoded by the NifY gene (Homer et al., 1993), and addition of purified FeMo-co to the purified hexameric α2β2γ2 complex was sufficient to generate catalytically active nitrogenase. Addition of FeMo-co resulted in dissociation of NifY from the complex, resulting in the formation of the holoenzyme (α2β2). In A. vinelandii, the third polypeptide was encoded by the NafY gene (nitrogenase-associated factor Y; accession number AGK13761; Rubio et al., 2002), which is related to but distinct from the product of the NifY gene (accession number AGK13792) in A. vinelandii. The third polypeptide in each case was thought to be involved in helping insert FeMo-co to form the active enzyme. This was supported by the ability of NafY and NifY to bind FeMo-co (Homer et al., 1995).

[0660] A. vinelandii NifY and NafY bind to the α-Cys of NifD in apo-NifD-NifK at different stages of NifD-NifK holoenzyme maturation. 275 or α-His 442(Both amino acid residues of NifD covalently anchor FeMo-co) (Jimenez-Vincente et al., 2018). Thus, NifY and NafY do not bind simultaneously to apo-NifD-NifK. The order in which NifY and NafY bind to apo-NifD-NifK is currently unknown. Dissociation of NifY from NifD-NifK upon insertion of FeMo-co has been demonstrated in K. oxytoca nitrogenase (Homer et al., 1993), and dissociation of NafY from NifD-NifK upon insertion of FeMo-co has been demonstrated in A. vinelandii (Homer et al., 1995). NafY binds to His 121 It is also thought to bind FeMo-co via NifD and possibly NifB-co, suggesting a role for NafY as an FeMo-co or FeMo-co precursor insertase (Rubio et al., 2004). Because A. vinelandii NifY appears to be functionally redundant based on the lack of a single phenotype in a ΔnifY mutant (Rubio et al., 2002), it has been proposed that NafY is the primary accessory protein for apo-NifD-NifK that supports the insertion of FeMo-co. On the other hand, Klebsiella species lack the NafY gene and only possess NifY, which supports the insertion of FeMo-co into apo-NifD-NifK, yet a Klebsiella ΔnifY mutant still retained 60% of its acetylene reduction activity (Homer et al., 1993). This retention of function indicated that another accessory protein (such as NifX, described above) exists in Klebsiella that can partially cover the function of NifY in its absence.

[0661] As used herein, "NafY polypeptide" refers to a polypeptide whose sequence contains at least 50% amino acid identity over its entire length to the sequence set forth in SEQ ID NO: 238 (A. vinelandii NafY, Accession No. AGK13761, 243 amino acids) and contains the conserved domain pfam16844. This domain, approximately 91 amino acids in length, is found in several members and in the amino-terminal half of longer NafY proteins. This region is negatively charged and appears to function to recognize and interact with apo-NifD-NifK. Native NafY polypeptides are typically 230-250 amino acids in length, and the native monomer has a molecular weight of approximately 25-28 kDa. Numerous NafY polypeptides have been identified, and numerous sequences are available in public databases. Some are annotated as NifX polypeptides due to the sequence relatedness of NafY and NifX. For example, NafY polypeptides have been reported from Azotobacter beijerinckii (WP_090728988, 93% identity to SEQ ID NO: 238), Pseudomonas stutzeri (WP_011912501, 69% identity), Halomonas endophytica (WP_102654474, 68% identity), Pseudomonas linyingensis (WP_090313081, 67% identity), Acidihalobacter prosperus (WP_038093031, 56% identity), and Oscillatoriales cyanobacterium (WP_009769409, 50% identity). As used herein, a "functional NafY polypeptide" is a NafY polypeptide that is capable of binding apo-NifD-NifK and FeMo-co. The three-dimensional structure of the NafY polypeptide from A. vinelandii and a comparison and difference between the NafY polypeptide and the NifY, NifX, VnfX, and NifB polypeptides were reported in Dyer et al. (2003).

[0662] The NifZ polypeptide in natural bacteria is a polypeptide involved in the synthesis of Fe-S clusters, specifically in coupling the second Fe4S4 pair in the formation of the second P cluster of the MoFe protein. NifZ is thought to function as a chaperone, inducing a conformational change in at least the second half of the apo-MoFe protein, allowing it to combine with NifH to form the second P cluster. Deletion of NifZ in A. vinelandii reduced the activity of the MoFe protein by 66% but had no effect on NifH activity. As used herein, "NifZ polypeptide" refers to a polypeptide whose sequence contains at least 28% amino acid identity to the sequence set forth as SEQ ID NO: 16 and contains the conserved domain pfam04319. This domain, consisting of approximately 75 amino acid residues, is found isolated in some membranes and in the amino-terminal half of longer NifZ proteins. Naturally occurring NafZ polypeptides are typically 70-150 amino acids in length, and the naturally occurring monomers have molecular weights of about 9 to about 16 kDa. Many NafZ polypeptides have been identified, and many sequences are available in public databases. For example, the NifZ polypeptide is isolated from Klebsiella michiganensis (accession number WP_057173223.1, 93% identical to SEQ ID NO: 16), Klebsiella oxytoca (WP_064342939.1, 95% identical), Klebsiella variicola (WP_043875005.1, 77% identical), Kosakonia radicincitans (WP_043953588.1, 67% identical), Kosakonia sacchari (WP_065368553.1, 58% identical), Ferriphaselus amnicola (WP_062627625.1, 47% identical), Paraburkholderia xenovorans (WP_011491838.1, 41% identical), Acidithiobacillus ferrivorans (WP_014029050.1, 35% identity), and Bradyrhizobium oligotrophicum (WP_015665422.1, 28% identity).As used herein, a "functional NifZ polypeptide" is a NifZ polypeptide that is capable of coupling an Fe4S4 cluster in the synthesis of an Fe-S cluster. NifZ polypeptides are described and reviewed in Cotton (2009) and Hu et al. (2004).

[0663] In natural bacteria, NifW polypeptides associate with NifZ polypeptides to form higher-order complexes (Lee et al., 1998) and are involved in the synthesis or activity of MoFe proteins (NifD-NifK). NifW and NifZ appear to be involved in the formation or accumulation of MoFe proteins (Paul and Merrick, 1987). As used herein, "NifW polypeptide" refers to a polypeptide whose amino acid sequence contains at least 28% amino acid identity to the amino acid sequence set forth as SEQ ID NO: 17 and contains the conserved NifW superfamily protein domain (architecture ID number 10505077, in Pfamily PF03206). Many NifW polypeptides have been identified, and many sequences are available in public databases. For example, the NifW polypeptide is isolated from Klebsiella oxytoca (accession number WP_064342938.1, 98% identical to SEQ ID NO: 17), Klebsiella michiganensis (WP_049080155.1, 94% identical), Enterobacter sp. 10-1 (WP_095103586.1, 90% identical), Klebsiella quasipneumoniae (WP_065877373.1, 81% identical), Pectobacterium polaris (WP_095699971.1, 69% identical), Dickeya paradisiaca (WP_012764136.1, 58% identical), Brenneria goodwinii (WP_053085547.1, 36% identical), Aquaspirillum sp. LM1 (WP_077299824.1, 44% identity), Candidatus Muproteobacteria bacterium RBG_16_64_10 (OGI40729, 34% identity), Azotobacter vinelandii (ACO76430.1, 32% identity), and Methylocaldum marinum (BBA37427.1, 28% identity). As used herein, a "functional NifW polypeptide" is a NifW polypeptide that promotes or enhances one or more of the formation, accumulation, or activity of a MoFe protein.Functional NifW may interact with NifZ and / or play a role in oxygen protection of the MoFe protein ( Gavini et al., 1998 ).

[0664] Most organisms, including both bacteria and eukaryotes (e.g., plants), possess numerous ferredoxins. For example, the DJ and CA genomes of A. vinelandii contain 15 and 16 proteins annotated as ferredoxins or ferredoxin-like proteins, respectively. As used herein, a "ferredoxin polypeptide" is an electron-transporting protein containing one or two iron-sulfur clusters of the [2Fe-2S], [3Fe-4S], and / or [4Fe-4S] types that form the reaction center. See the review by Matsubara and Saeki (1992). Ferredoxin polypeptides are involved in a variety of metabolic processes, including those involved in nitrogen fixation and generally lower in molecular weight than those involved in nitrogenase. Based on the diversity of ferredoxins within most cells and the variation observed in several studies regarding the suitability or specificity of different ferredoxins in complementing the function of FdxN for the synthesis of NifB-co (Yates, 1972; Jimenez-Vincente et al., 2014), ferredoxins, including ferredoxins such as FdxN, are best defined based on the presence and function of an iron-sulfur cluster rather than on amino acid identity to a canonical sequence such as A. vinelandii FdxN (SEQ ID NO: 232; Accession No. WP_012703542). As used herein, an "FdxN polypeptide" is a ferredoxin or ferredoxin-like polypeptide that functions to donate electrons to the mature dinitrogenase reductase NifH and / or nitrogenase for the synthesis of NifB-co and / or acts as an intermediate carrier of the [4Fe-4S] cluster.FdxN can function by donating electrons to the mature dinitrogenase reductase NifH, which then transfers electrons to the NifD-NifK heterohexamer (Yang et al., 2017; Rhizobium japonicum FdxN, Carter et al., 1980; R. meliloti FdxN, Riedel et al., 1995; Rhodobacter capsulatus FdxN, Jouanneau et al., 1995), by donating electrons to the NifB polypeptide for NifB-co synthesis (A. vinelandii: Jimenez-Vincente et al., 2014), by acting as an intermediate carrier of the [4Fe-4S] cluster (A. vinelandii: Buren et al., 2019), or by any combination of these functions.

[0665] Representative examples of FdxN polypeptides include those identified by searching non-redundant protein databases using SEQ ID NO:232 as a query in BLASTP, with the percent identity to the sequence shown: Pseudomonas syringae (WP_065835964.1, 85.87%), Candidatus Thiodiazotropha endolucinida (WP_069124666.1, 70.65%), Uliginosibacterium sp. TH139 (WP_101942980, 64.47%), Klebsiella michiganensis (WP_049076934.1, 44.26%), Escherichia coli (WP_072048756.1, 44.26%), Rhizobium leguminosarum (WP_130674512.1, 43.86%), and Flavobacterium alvei(WP_103805005.1, 28.57%).

[0666] Sequence matching and substitution

[0667] It will be appreciated that with respect to one particular polypeptide, % identity figures greater than those shown above encompass preferred embodiments. Thus, where applicable, in light of minimum % identity figures, it is preferred that the polypeptide comprises an amino acid sequence that is at least 30%, more preferably at least 35%, more preferably at least 40%, more preferably at least 45%, more preferably at least 50%, more preferably at least 55%, more preferably at least 60%, more preferably at least 65%, more preferably at least 70%, more preferably at least 75%, more preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, and even more preferably at least 99.9% identical to the relevant SEQ ID NO.

[0668] Amino acid sequence variants of the polypeptides defined herein can be prepared by introducing appropriate nucleotide changes into the nucleic acids defined herein or by in vitro synthesis of the desired polypeptide. Such variants include, for example, deletion, insertion, or substitution of one or more amino acids. A combination of deletion, insertion, and substitution mutations can be made to arrive at the final construct, provided that the final polypeptide product possesses the desired characteristics. Preferred amino acid sequence variants have only one, two, three, four, or fewer than ten amino acid changes compared to the reference wild-type polypeptide.

[0669] Mutant (altered) polypeptides can be prepared using any technique known in the art, for example, using directed evolution or rational design strategies (see below). Products derived from mutated / altered DNA can be readily screened using the techniques described herein to determine whether their expression in a plant alters its phenotype compared to a corresponding wild-type plant (e.g., whether their expression results in increased yield, biomass, growth rate, vigor, nitrogen gain from biological nitrogen fixation, nitrogen use efficiency, abiotic stress tolerance, and / or tolerance to nutrient deficiency compared to a corresponding wild-type plant).

[0670] In designing amino acid sequence variants, the location of the mutation sites and the nature of the mutation will depend on the feature to be altered. The sites for mutation can be altered individually or sequentially, for example, by (1) first substituting selected conserved amino acids followed by more radical selection depending on the results achieved, (2) deleting the target residue, or (3) inserting other residues adjacent to the target site.

[0671] Amino acid sequence deletions generally range from about 1 to 15 contiguous residues, more preferably about 1 to 10 contiguous residues, and typically about 1 to 5 contiguous residues.

[0672] Substitutional variants have at least one amino acid residue removed from the polypeptide molecule and a different residue inserted in its place. Where it is desired to maintain a given activity, it is preferred that there be no or only conservative substitutions at amino acid positions that are highly conserved within a family of related proteins. Examples of conservative substitutions are shown in Table 1 under the heading "Representative Substitutions."

[0673] In a preferred embodiment, the mutant / variant polypeptide has one, or two, or three, or four conservative amino acid changes compared to the native polypeptide. Details of the conservative amino acid changes are provided in Table 1. In a preferred embodiment, the changes do not occur within one or more motifs or domains that are highly conserved among different polypeptides of the invention. One skilled in the art will recognize that such minor changes can be reasonably expected not to alter the activity of the polypeptide when expressed in a recombinant cell.

[0674] [Table 1]

[0675] The primary amino acid sequence of a polypeptide of the pre...

Claims

1. A plant cell comprising mitochondria and foreign polynucleotides encoding Anf fusion polypeptides, wherein each foreign polynucleotide is functionally linked to a nucleotide sequence encoding one of the Anf fusion polypeptides and comprises a promoter that expresses the nucleotide sequence in the plant cell, each Anf fusion polypeptide independently comprises a mitochondrial-targeted peptide (MTP), and the Anf fusion polypeptide comprises (i) an AnfG fusion polypeptide, or an AnfG and AnfH fusion polypeptide, (ii) an AnfD fusion polypeptide and an AnfK fusion polypeptide, or (iii) an AnfD-linker-AnfK fusion polypeptide comprising an AnfD sequence having a C terminus, an oligopeptide linker, and an AnfK sequence having an N terminus, wherein the oligopeptide linker is translated-fused to the C terminus of the AnfD sequence and the N terminus of the AnfK sequence, and at least the Anf The mitochondrial processing protease (MPP) cleavage products of the G and AnfH fusion polypeptides are, when present in the plant cell, at least partially soluble in the mitochondria of the plant cell, and the MPP cleavage product of the AnfD and AnfK fusion polypeptide in (ii) is at least partially soluble in the mitochondria of the plant cell when present in the plant cell, or the MPP cleavage product of the AnfD-linker-AnfK fusion polypeptide in (iii) is at least partially soluble in the mitochondria of the plant cell when present in the plant cell, and the MPP cleavage products of the AnfD fusion polypeptide and the AnfK fusion polypeptide in (ii), or the MPP cleavage product of the AnfD-linker-AnfK fusion polypeptide in (iii) form a protein complex in the plant cell together with the MPP cleavage product of the AnfG fusion polypeptide when present in the plant cell. plant cells.

2. A plant cell according to claim 1, wherein one or more of the following apply: (a) The plant cell comprises an exogenous polynucleotide encoding a NifM polypeptide (NM), the exogenous polynucleotide encoding the NM comprising a promoter functionally linked to a nucleotide sequence encoding the NM and expressing the nucleotide sequence within the plant cell; (b) The plant cell comprises foreign polynucleotides encoding NifS and NifU fusion polypeptides, each of which comprises a promoter functionally linked to a nucleotide sequence encoding one of the Nif fusion polypeptides and expressing the nucleotide sequence in the plant cell, and each of the NifS and NifU fusion polypeptides comprises a mitochondrial-targeted peptide (MTP); and (c) Each Anf polypeptide is produced in the plant cell as an Anf fusion polypeptide containing a mitochondrial-targeted peptide (MTP), wherein each MTP is independently identical or distinct, and the MTP is located at the N-terminus of at least one, more, or all of the Anf fusion polypeptides.

3. A plant cell according to claim 1, wherein one or more of the following apply: (a) Each Anf fusion polypeptide produced in the plant cell is independently (i) cleaved within the MTP sequence to produce an MPP-cleaved Anf polypeptide, the MPP-cleaved Anf polypeptide having a C-terminal peptide (scar peptide) derived from the MTP at its N-terminus, or (ii) cleaved immediately after the MTP, the MPP-cleaved Anf polypeptide not having a C-terminal peptide derived from the MTP; (b) Each Anf fusion polypeptide is at least partially cleaved within its MTP sequence in the plant cell to produce an MPP-cleaved Anf polypeptide, each MPP-cleaved Anf polypeptide independently comprising a peptide (scar peptide) of 1 to 45 amino acids in length derived from the MTP sequence, the peptide being translated and fused to the N-terminus of the MPP-cleaved Anf polypeptide; (c) The plant cell further comprises an exogenous polynucleotide encoding a ferredoxin fusion polypeptide, the exogenous polynucleotide encoding the ferredoxin fusion polypeptide comprising a promoter functionally linked to a nucleotide sequence encoding the ferredoxin fusion polypeptide and expressing the nucleotide sequence in the plant cell, and the ferredoxin fusion polypeptide comprising a mitochondrial-targeted peptide (MTP), the MPP cleavage product of the ferredoxin fusion polypeptide being at least partially soluble in the mitochondria of the plant cell; (d) The plant cell comprises, in order, an AnfD-linker-AnfK fusion polypeptide comprising an AnfD polypeptide (AD) amino acid sequence, an oligopeptide linker, and an AnfK polypeptide (AK) amino acid sequence, wherein the oligopeptide linker is 8 to 50 residues in length and is translated and fused to AD and AK; (e) Each Anf fusion polypeptide is cleaved in the plant cell to produce an Anf polypeptide, which is a functional Anf polypeptide; (f) The AnfK fusion polypeptide or the AnfD-linker-AnfK fusion polypeptide has an amino acid sequence such that the last four amino acids of its sequence are the same as the last four amino acids of the wild-type AnfK polypeptide; (g) The plant cell comprises an exogenous polynucleotide encoding an AnfD-linker-AnfK fusion polypeptide, the AnfD-linker-AnfK fusion polypeptide comprising an AnfD sequence having a C terminus, an oligopeptide linker, and an AnfK sequence having an N terminus, wherein the oligopeptide linker is translated-fused to the C terminus of the AnfD sequence and the N terminus of the AnfK sequence, and the oligopeptide linker has a length of at least 20 amino acids; (h) At least one or more of the foreign polynucleotides are incorporated into the nuclear genome of the plant cell or expressed in the nucleus of the plant cell; (i) At least one of the Anf fusion polypeptides comprises an MTP of about 51 amino acids in length derived from an F1-ATPase γ subunit polypeptide; and (j) The plant cell further comprises an exogenous polynucleotide encoding a NifM polypeptide (NM), the exogenous polynucleotide encoding the NM comprising a promoter functionally linked to a nucleotide sequence encoding the NM and causing the nucleotide sequence to be expressed in the plant cell, and the NM comprises a mitochondrial-targeted peptide (MTP).

4. A plant cell according to claim 1, wherein one or more of the following apply: (a) Each MPP-cleaved Anf polypeptide independently comprises a peptide (scar peptide) of 1 to 20 amino acids in length derived from the MTP sequence, the peptide being translated and fused to the N-terminus of the MPP-cleaved Anf polypeptide; (b) All of the foreign polynucleotides are incorporated into the nuclear genome of the plant cell or expressed in the nucleus of the plant cell; and (c) The plant cell comprises, in order, an AnfD-linker-AnfK fusion polypeptide comprising an AnfD polypeptide (AD) amino acid sequence, an oligopeptide linker, and an AnfK polypeptide (AK) amino acid sequence, wherein the oligopeptide linker is 16 to 50 residues in length and is translated and fused to AD and AK.

5. A plant comprising plant cells according to any one of claims 1 to 4, or a transgenic plant having an exogenous polynucleotide encoding an Anf fusion polypeptide in plant cells according to any one of claims 1 to 4.

6. A plant portion comprising plant cells according to any one of claims 1 to 4, or a transgenic plant portion having an exogenous polynucleotide encoding an Anf fusion polypeptide in plant cells according to any one of claims 1 to 4.

7. AnfD-linker-AnfK fusion polypeptide comprising a translated-fused AnfD polypeptide (AD), an oligopeptide linker, and an AnfK polypeptide (AK), wherein AD comprises an N-terminus and a C-terminus, AK comprises an N-terminus, and the oligopeptide linker is translated-fused to the C-terminus of AD and the N-terminus of AK.

8. A cleavage product of an AnfD-linker-AnfK fusion polypeptide according to claim 7, wherein the cleavage product comprises the AnfD polypeptide, the oligopeptide linker, and the AnfK polypeptide.

9. A fusion polypeptide according to claim 7, wherein the fusion polypeptide comprises a mitochondrial-targeting peptide (MTP) translated and fused to the N-terminus of AD.

10. A cleavage product according to claim 8, wherein the cleavage product comprises an MTP scar peptide translated and fused to the N-terminus of the AnfD polypeptide.

11. A polynucleotide encoding the polypeptide according to claim 7 or claim 9.

12. A polynucleotide according to claim 11, wherein one or more of the following apply: (a) The polypeptide coding region of the polynucleotide is codonally modified for expression in plant cells compared to the corresponding polypeptide coding region naturally present in bacteria; (b) The polypeptide further comprises a promoter functionally linked to the polynucleotide encoding the polypeptide; (c) The polynucleotides are present in plant cells, yeast cells or bacterial cells; and (d) The polynucleotide is either incorporated into the nuclear genome of a plant cell or expressed within the nucleus of a plant cell.

13. A method for producing transgenic plants, (i) the step of introducing the polynucleotide described in claim 11 into plant cells, and (ii) Process to regenerate transgenic plants from the cells of process (i) Methods that include...

14. A method for producing transgenic seeds, comprising harvesting seeds from the plant or its progeny described in claim 5.

15. A method for producing wheat flour, whole grain flour, starch, oil, seed meal, or other products obtained from seeds, comprising extracting the wheat flour, whole grain flour, starch, oil, or other products from the seeds of the plant described in claim 5, or producing the seed meal from the seeds.

16. A method for preparing food, comprising mixing the seeds of the plant described in claim 5, or wheat flour, whole grain flour, starch, oil, or other products derived from said seeds, with another food ingredient, or processing said seeds, wheat flour, or whole grain flour.