Nif variants

CN122535693APending Publication Date: 2026-08-07COMMONWEALTH SCI & IND RES ORG
View PDF 44 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
COMMONWEALTH SCI & IND RES ORG
Filing Date
2024-08-30
Publication Date
2026-08-07

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates, in part, to modified NifH polypeptides and NifH fusion polypeptides having improved solubility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates in part to modified NifH peptides and NifH fusion peptides with improved solubility. Background Technology

[0002] Nitrogen-fixing bacteria (diazotrophic bacteria) produce ammonia from N2 gas through biological nitrogen fixation (BNF) catalyzed by enzyme complexes and nitrogenases. However, modern agriculture's demand for nitrogen far exceeds this nitrogen fixation source, leading to the widespread use of industrially produced nitrogen fertilizers in agriculture (Smil, 2002). However, both fertilizer production and application are causes of pollution (Good and Beatty, 2011) and are considered unsustainable (Rockstrom et al., 2009). A large portion of fertilizers applied worldwide is not absorbed by crops (Cui et al., 2013; de Bruijn, 2015), resulting in fertilizer runoff, weed growth, and eutrophication of waterways (Good and Beatty, 2011). The resulting algal blooms reduce oxygen levels, causing environmental damage to local and nearshore coral reefs (De'ath et al., 2012; Glibert et al., 2014; Sutton et al., 2008). Furthermore, while over-fertilization is a problem in many developed countries, nitrogen availability limits crop yields in some regions (Mueller et al., 2012). Fertilizer production itself requires significant energy inputs and is estimated to cost $100 billion annually.

[0003] Clearly, strategies are needed to reduce nitrogen dependence in industrial production. To this end, the concept of engineered plants capable of biological nitrogen fixation has long attracted considerable interest (Merrick and Dixon, 1984) and has been a focus of recent commentary (de Bruijn, 2015; Oldroyd and Dixon, 2014). Possible approaches include i) extending the symbiotic relationship of diazotrophs from legumes to cereals (Santi et al., 2013), ii) re-engineering endosymbiotic microorganisms to enable them to perform nitrogen fixation (Geddes et al., 2015), and iii) genetically engineering nitrogenases into plant cells (Curatti and Rubio, 2014). Due to technical difficulties, all these approaches remain ambitious and speculative.

[0004] Nitrogenases are enzyme complexes that enable biological nitrogen fixation in nitrogen-fixing bacteria. Their biosynthesis and function require a multi-genome assembly pathway, which has been extensively reviewed (Hu and Ribbe, 2013; Rubo and Ludden, 2008; Seefeldt et al., 2009). Typical iron-molybdenum nitrogenases consist of catalytic proteins called NifD and NifK, and an electron donor, NifH. Approximately 12 other proteins are involved in nitrogenase assembly in nitrogen-fixing bacteria, including complex maturation, support, and cofactor insertion; specifically, NifM, NifS, NifU, NifE, NifN, NifX, NifV, NifJ, NifY, NifF, NifZ, and NifQ. Genetic mutations, complementation assays between nitrogen-fixing and non-nitrogen-fixing prokaryotes, and phylogenetic analyses (Dos Santos et al., 2012; Temme et al., 2012; Wang et al., 2013) have led to the identification of a subset of Nif proteins (NifD, NifK, NifB, NifE, and NifN) as core components, while others are considered necessary for optimizing activity and are considered auxiliary. The assembly and function of nitrogenases also require specific biochemical conditions. Most importantly, nitrogenases are extremely sensitive to oxygen (Robson and Postgate, 1980). Furthermore, the biosynthesis and function of the metalloprotein catalytic center require large amounts of ATP, reducing agents, readily available Fe, Mo, S-adenosylmethionine, and high citrate levels (Hu and Ribbe, 2013; Rubo and Ludden, 2008). All these factors contribute to the technical challenges of producing functional nitrogenase complexes in plant cells. Summary of the Invention

[0005] The inventors have identified NifH variants (such as AnfH variants) that exhibit improved solubility in the mitochondria of plant cells.

[0006] Therefore, in a first aspect, the present invention provides a plant cell comprising a modified NifH polypeptide, wherein the modified NifH polypeptide contains at least one amino acid substitution compared to a corresponding wild-type NifH polypeptide, and wherein the modified NifH polypeptide is more soluble in the mitochondria of the cell than the corresponding wild-type NifH polypeptide.

[0007] In one embodiment, the at least one amino acid substitute is located at an amino acid position selected from the group consisting of the following amino acid positions: 2, 5, 7, 19, 23, 24, 26 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 74, 76 to 78, 80 to 84, 102, 105, 107, 111 to 114, 116 to 118, 121 to 124, 139, 145, 147, 149, 158, 165, 166, 168, 169, 171, 179, 182, 183, 1 88, 191, 193 to 197, 200 to 203, 205 to 211, 214, 216, 219, 223 to 226, 228 to 235, 237, 238, 241, 242, 244 to 246, 248, 249, 251 to 253, 257, 259 to 264, 266 to 271, and 273 to 275, or located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide, preferably the modified NifH polypeptide, is compared with SEQ ID NO:37. Alternatively or additionally, in one embodiment, the at least one amino acid substitutes an amino acid position located at a position selected from the group consisting of the following amino acid positions: refer to SEQ ID NO:37. NO:39 3, 6, 8, 20, 24, 25, 27 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 75, 77 to 79, 81 to 85, 103, 106, 108, 112 to 115, 117 to 119, 122 to 125, 140, 146, 148, 150, 159, 166, 167, 169, 170, 172, 180, 183 184, 189, 192, 194 to 198, 201 to 204, 206 to 212, 215, 217, 220, 224 to 227, 229 to 236, 238, 239, 242, 243, 245 to 247, 249, 250, 252 to 254, 258, 260 to 265, 267 to 272 and 274 to 276, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:39.

[0008] In another aspect, the present invention provides a plant cell comprising a modified NifH polypeptide, wherein, compared with a corresponding wild-type NifH polypeptide, the modified NifH polypeptide, preferably the modified NifH polypeptide, comprises at least one amino acid substitution, and wherein the at least one amino acid substitution is located at an amino acid position selected from the group consisting of: (i) the group consisting of the following amino acid positions: 2, 5, 7, 19, 23, 24, 26 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 74, 76 to 78, 80 to 84, 102, 105, 107, 111 to 114, 116 to 118, 121 to 124, 139, 145, 147, 149, 158, 165, 166, 168, 169, 171, 179, 18 of SEQ ID NO:37. 2, 183, 188, 191, 193 to 197, 200 to 203, 205 to 211, 214, 216, 219, 223 to 226, 228 to 235, 237, 238, 241, 242, 244 to 246, 248, 249, 251 to 253, 257, 259 to 264, 266 to 271 and 273 to 275, or located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37, and / or (ii) the group consisting of the following amino acid positions: refer to SEQ ID NO:39 3, 6, 8, 20, 24, 25, 27 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 75, 77 to 79, 81 to 85, 103, 106, 108, 112 to 115, 117 to 119, 122 to 125, 140, 146, 148, 150, 159, 166, 167, 169, 170, 172, 180, 18 3, 184, 189, 192, 194 to 198, 201 to 204, 206 to 212, 215, 217, 220, 224 to 227, 229 to 236, 238, 239, 242, 243, 245 to 247, 249, 250, 252 to 254, 258, 260 to 265, 267 to 272, and 274 to 276, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:39. In one embodiment of this aspect, the modified NifH polypeptide is more soluble in the mitochondria of the cell than the corresponding wild-type NifH polypeptide.

[0009] In one embodiment of the foregoing, the modified NifH polypeptide comprises at least 15%, at least 20%, at least 35%, at least 40%, at least 45%, at least 50%, or 15% to 90%, 15% to 80%, 15% to 70%, or 15% to 60% of the mitochondria of the cell, preferably the modified NifH polypeptide is soluble.

[0010] In one embodiment of the above aspects, the solubility of the NifH polypeptide in the mitochondria of the cell is at least twice that of the corresponding wild-type NifH polypeptide in the mitochondria of the cell, preferably at least three times, at least four times, or at least five times, or two to ten times.

[0011] In one embodiment of the above aspects, the free energy of the modified NifH polypeptide is lower than that of the wild-type NifH polypeptide, and / or the amino acid substitution reduces the free energy of the modified NifH polypeptide relative to the wild-type NifH polypeptide.

[0012] In one embodiment of the foregoing aspects, the free energy of the modified NifH polypeptide having the substituted amino acids is reduced by at least 0.5, at least 1.0, at least 1.5, at least 2, at least 3, at least 4, at least 5, or 2 to 6 units relative to the corresponding NifH polypeptide, wherein the amino acid sequence of the corresponding NifH polypeptide is identical to that of the modified NifH polypeptide, except for the substituted amino acids. In one embodiment, the change in free energy caused by the amino acid substitution is determined by the Rosetta energy function.

[0013] In one embodiment of the foregoing aspects, when compared to the corresponding wild-type NifH polypeptide, the modified NifH polypeptide comprises at least one, preferably at least two or at least three, amino acid substitutions, wherein each substituted amino acid reduces the free energy of the modified NifH polypeptide by at least 0.5, at least 1.0, at least 1.5, at least 2, at least 3, at least 4, at least 5 units, or 2 to 6 units, and / or wherein the amino acid substitutions collectively reduce the free energy of the modified NifH polypeptide by at least 4.0, at least 5.0, at least 6.0, at least 7.0, at least 8.0, at least 9.0, at least 10.0, at least 12.0 units, or 4.0 to 15.0, 4.0 to 13.0, or 4.0 to 12.0 units.

[0014] In one embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least one amino acid substitution, or two or three amino acid substitutions, or four or more amino acid substitutions selected from the group consisting of: amino acid substitutions listed in one or more of Tables 4, 5, 9, 10 or 11, or the corresponding amino acid substitutions when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0015] In one embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, wherein the amino acid substitution is located at an amino acid position selected from the group consisting of amino acid positions 69, 168, 200, 201, 224, 228, 234, 241, 252 and 263 of reference SEQ ID NO:37, or at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0016] In one embodiment of the foregoing aspects, the at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, is selected from the group consisting of: 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H or 234C, 241R or 241A, 252M and 263E, wherein the amino acid position corresponds to the amino acid sequence provided as in SEQ ID NO:37, or is located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide, preferably the modified AnfH polypeptide, is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0017] In one embodiment of the foregoing aspects, the at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, is selected from the group consisting of positions corresponding to amino acids 69, 168, 200, 201, 224, 228, 234, 252 and 263 of SEQ ID NO:37, or is located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide, preferably the modified AnfH polypeptide, is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0018] In one embodiment of the foregoing aspects, the at least one amino acid substitution, or two or three amino acid substitutions, or four or more amino acid substitutions are selected from the group consisting of: 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H or 234C, 252M, and 263E, wherein the amino acid position corresponds to the amino acid sequence provided as in SEQ ID NO:37, or is located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide, preferably the modified AnfH polypeptide, is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0019] In one embodiment of the foregoing, when compared with the corresponding wild-type NifH polypeptide, the modified NifH polypeptide, preferably the modified AnfH polypeptide, has one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, eighteen, nineteen, twentieth, eleventh ... One, 1 to 3, 2 to 20, 2 to 15, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 2 or 3, 3 to 20, 3 to 15, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 or 4, preferably 1 to 3, 1 to 4 or 1 to 5, more preferably 2 to 4 or 2 to 5, most preferably 3 to 5 amino acid substitutions.

[0020] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least one amino acid substitution of reference SEQ ID NO:37 228V, or the same amino acid substitution at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0021] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least one amino acid substitution of reference SEQ ID NO:37 228I, or the same amino acid substitution at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0022] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least one amino acid substitution with reference SEQ ID NO:37 being 200A, or the same amino acid substitution located at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0023] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least one amino acid substitution of reference SEQ ID NO:37 234H, or the same amino acid substitution at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0024] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least two amino acid substitutions that are 200A and 228V as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0025] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least two amino acid substitutions that are 200A and 228I as referenced in SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0026] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least two amino acid substitutions that are 228V and 234H as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0027] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least two amino acid substitutions that are 228I and 234H as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0028] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least three amino acid substitutions that are 200A, 228V and 234H as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0029] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least three amino acid substitutions that are 200A, 228I and 234H as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0030] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least four amino acid substitutions that are 200A, 228V or 228I, 234H and 241R as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0031] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least four amino acid substitutions that are 168I, 200A, 228I or 228V and 234H as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0032] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least seven amino acid substitutions that are 69N, 168I, 200A, 228V or 228I, 234H, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0033] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least eight amino acid substitutions that are 69N, 168I, 200A, 201K, 228V or 228I, 234H, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0034] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least nine amino acid substitutions that are 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 252M and 263E as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0035] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least five amino acid substitutions that are 168I, 200A, 228I or 228V, 234H and 241R as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0036] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least eight amino acid substitutions of 69N, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0037] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least nine amino acid substitutions that are 69N, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0038] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least ten amino acid substitutions of 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0039] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least five amino acid substitutions that are 112L, 200A, 228V or 228I, 234H and 241R as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0040] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains at least six amino acid substitutions that are 112L, 168I, 200A, 228I or 228V, 234H and 241R as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0041] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least nine amino acid substitutions that are 69N, 112L, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0042] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least ten amino acid substitutions of 69N, 112L, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0043] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least eleven amino acid substitutions corresponding to SEQ ID NO:37 as 69N, 112L, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0044] As those skilled in the art will understand, any and all combinations of the amino acid substitutions are conceivable. It should be understood that the desired amino acid may already be present at any one or more indicated positions of the wild-type NifH polypeptide, preferably the AnfH polypeptide, and therefore does not need to be substituted, provided that the modified NifH polypeptide contains at least one amino acid substitution relative to the corresponding wild-type NifH polypeptide.

[0045] Therefore, in the embodiments, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises at least one of the following amino acid sets located at the indicated position, wherein the amino acid is either present due to substitution or is already present in the corresponding wild-type NifH polypeptide, provided that at least one of the listed amino acids is present due to substitution: i) 228V, ii) 228I, iii) 200A, or iv) 234H, v) 200A and 228V, vi) 200A and 228I, vii) 228V and 234H, viii) 228I and 234H, ix) 200A, 228V and 234H, x)200A, 228I and 234H, xi) 200A, 228V or 228I, 234H and 241R, xii) 168I, 200A, 228I or 228V and 234H, xiii) 69N, 168I, 200A, 228V or 228I, 234H, 252M and 263E, xiv) 69N, 168I, 200A, 201K, 228V or 228I, 234H, 252M and 263E, xv)69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 252M and 263E, (xvi) 168I, 200A, 228I or 228V, 234H and 241R, (xvii) 69N, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E, (xviii) 69N, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E, (xix) 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E, xx) 112L, 200A, 228V or 228I, 234H and 241R, (xxi) 112L, 168I, 200A, 228I or 228V, 234H and 241R, xxii) 69N, 112L, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E, xxiii) 69N, 112L, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E, or xxiv) 69N, 112L, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E.

[0046] In a preferred embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises amino acids 200A, 228V, and 234H as referenced in SEQ ID NO:37, or the same amino acids located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0047] In one embodiment of the foregoing aspects, preferably in combination with one embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises one or more or all of the following motifs: YGKGGIGKSTTXQN (SEQ ID NO:61), IXGCDPKAD (SEQ ID NO:62), CXESGGPEPGVGCAGRG (SEQ ID NO:63), DVLGDVVCGGFAMP (SEQ ID NO:43), VXSGEMMAXYAANNI (SEQ ID NO:64), and CNSRXXD (motif VII, SEQ ID NO:65), preferably at least DVLGDVVCGGFAMP (SEQ ID NO:43), wherein each X independently represents any amino acid.

[0048] In one embodiment of the foregoing aspects, preferably in combination with one embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, comprises one or more or all of the following motifs: YGKGGIGKSTTXQNT (SEQ ID NO:40), IHGCDPKAD (SEQ ID NO:41), CVESGGPEPGVGCAGRG (SEQ ID NO:42), DVLGDVVCGGFAMP (SEQ ID NO:43), VASGEMMAXYAANNI (SEQ ID NO:44), QSGVR (SEQ ID NO:45), and CNSRXVD (SEQ ID NO:46), preferably at least DVLGDVVCGGFAMP (SEQ ID NO:43), wherein each X independently represents any amino acid.

[0049] In one embodiment of the foregoing aspects, preferably in combination with an embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, has 12, 13, 14, 15 or all of the following amino acids: 4K, 22T, 37H, 52G, 60D, 63R, 108L, 109M, 142G, 151A, 174Q, 189V, 198E, 199F, 222F and 247I as per SEQ ID NO:37, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0050] In one embodiment of the foregoing aspects, preferably in combination with one embodiment of the foregoing aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, has 130, 131, 132, 133, 134, 135, 136 or all 137 of the following amino acids: 3R, 4K, 6A, 8Y, 9G, 10K, 11G, 12G, 13I, 14G, 15K, 16S, 17T, 18T, 20Q, 21N, 22T, 25A, 36I, 37H, 38G, 39C, 40D, 41P, 42K, 43A, 44D, 46T, 47R, 50L, 52G, 55Q, 60D, 63R, 75V, 79G, 85C, 86V, 87E, 88S, 89 G, 90G, 91P, 92E, 93P, 94G, 95V, 96G, 97C, 98A, 99G, 100R, 101G, 103I, 104T, 106I, 108L, 109M, 110E, 115Y, 119L, 120D, 125D, 126V, 127L, 128G, 129D, 130V, 131V, 132C, 133G, 134G, 135F, 136A, 137M, 13 8P, 140R, 142G, 143K, 144A, 146E, 148Y, 150V, 151A, 152S, 153G, 154E, 155M, 156M, 157A, 159Y, 160 A, 161A, 162N, 163N, 164I, 167G, 170K, 172A, 174Q, 175S, 176G, 177V, 178R, 180G, 181G, 184C, 185N, 186S, 187R, 189V, 190D, 192E, 198E, 199F, 204G, 212P, 213R, 215N, 217V, 218Q, 220A, 221E, 222F, 227V, 236Q, 239E, 240Y, 243L, 247I, 250N, 254V, 255I, 256P, 258P, 265E, and 272G, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0051] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, has one or two of the following amino acids as referenced in SEQ ID NO:37: 141D and 173K, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0052] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, has an amino acid sequence that is at least 60% identical to the amino acid sequence provided in SEQ ID NO:37 and / or SEQ ID NO:39, preferably at least 70% or at least 80% identical, more preferably at least 90% identical, and most preferably at least 95% identical.

[0053] In one embodiment of the above aspects, the modified NifH polypeptide is a modified AnfH polypeptide.

[0054] In one embodiment, the amino acid sequence of the modified AnfH polypeptide is at least 70% identical to the amino acid sequence provided in SEQ ID NO:37, preferably at least 80% identical, more preferably at least 90% identical, and most preferably at least 95% identical.

[0055] In one embodiment of the above aspects, the modified AnfH polypeptide comprises at least amino acids 2-275 of the amino acid sequence provided in SEQ ID NO:78, or comprises SEQ ID NO:78.

[0056] In one embodiment of the above aspects, the modified NifH polypeptide, preferably the modified AnfH polypeptide, contains an Fe-S cluster.

[0057] In one embodiment of the foregoing aspects, the modified NifH polypeptide is a NifH fusion polypeptide, preferably a cleavage product of an AnfH fusion polypeptide, the fusion polypeptide comprising a mitochondrial targeting peptide (MTP) translatably fused to the modified NifH polypeptide, wherein the MTP is preferably translatably fused at the N-terminus of the modified NifH polypeptide, and wherein the modified NifH polypeptide is generated in the plant cell by protease cleavage of the NifH fusion polypeptide within or adjacent to the MTP. In this embodiment, the plant cell contains an exogenous polynucleotide encoding the NifH fusion polypeptide, which is also referred to herein as an encoded polypeptide.

[0058] Therefore, in one embodiment, the present invention provides a plant cell comprising a foreign polynucleotide encoding a NifH fusion polypeptide, preferably an AnfH fusion polypeptide, wherein the foreign polynucleotide is operatively linked to a promoter guiding the expression of the foreign polynucleotide to produce a NifH fusion polypeptide (encoded polypeptide) in the plant cell, wherein the encoded polypeptide comprises a translationally fused MTP at the N-terminus of a modified NifH polypeptide, optionally having an intercalated oligopeptide linker, wherein the modified NifH polypeptide, preferably the modified AnfH polypeptide, is produced in the plant cell by protease cleavage of the NifH fusion polypeptide within or adjacent to the MTP, and wherein the solubility of the modified NifH polypeptide, preferably the modified AnfH polypeptide, is increased in the plant cell. Preferably, the modified NifH polypeptide comprises at least one amino acid set from the amino acid sets listed in the above embodiments. Preferably, the foreign polynucleotide is integrated into the genome of the plant cell.

[0059] In one embodiment of the foregoing aspects, the NifH fusion polypeptide is cleaved within the MTP by a mitochondrial processing protease (MPP) to produce the modified NifH polypeptide, wherein the modified NifH polypeptide comprises (i) a C-terminal peptide (scar peptide) from the MTP located at its N-terminus, or (ii) does not contain a C-terminal peptide from the MTP.

[0060] In one embodiment of the foregoing aspects, the modified NifH polypeptide comprises two NifH polypeptides covalently linked by an oligopeptide linker in the order NifH::linker::NifH from the N-terminus to the C-terminus, optionally further comprising a C-terminal peptide (scar peptide) from a mitochondrial targeting peptide (MTP) located at its N-terminus, optionally having an intercalating oligopeptide linker located between the scar peptide and the first NifH polypeptide.

[0061] In one embodiment of the foregoing aspects, the modified NifH polypeptide is a modified AnfH polypeptide comprising two modified AnfH polypeptides covalently linked by an oligopeptide linker in the order AnfH::linker::AnfH from the N-terminus to the C-terminus, optionally further comprising a C-terminal peptide from a mitochondrial targeting peptide (MTP) located at its N-terminus, optionally having an intercalated oligopeptide linker located between the MTP and the first AnfH polypeptide.

[0062] In one embodiment of the above aspects, the linker between the two NifH or AnfH peptides is 10-50 residues long, preferably 16-50 residues or 20-35 residues long, more preferably about 25 or about 30 residues long, or most preferably 25 or 30 residues long. In one embodiment, the linker between the MTP or scar peptide and the first NifH / AnfH peptide is less than 20 amino acids long, for example, only two amino acids long, preferably a GlyGly dipeptide.

[0063] In one embodiment of the foregoing aspects, the modified NifH polypeptide, in combination with NifD and NifK polypeptides or with AnfD, AnfK, and AnfG polypeptides, is capable of (i) reducing N2 gas to produce ammonia, and / or (ii) reducing acetylene to ethylene, preferably both (i) and (ii). In this context, the ability of the polypeptide to isolate the polypeptide from the plant cells can be demonstrated when it is not treated in vitro with NifU polypeptides having Fe-S clusters, or with such treatment, or both.

[0064] In a preferred embodiment of the foregoing aspects and embodiments, the modified NifH polypeptide is a modified AnfH polypeptide.

[0065] In one embodiment of the foregoing aspects, the plant cell further comprises exogenous polynucleotides encoding: (a) NifM polypeptide, NifS polypeptide, NifU polypeptide, and (i) NifD and NifK polypeptides or (ii) NifD-NifK fusion polypeptide, or (b) NifS polypeptide, NifU polypeptide, and AnfG polypeptide, and (iii) AnfD and AnfK polypeptide or (iv) AnfD-AnfK fusion polypeptide. In a preferred embodiment, the plant cell further comprises NifB polypeptide, NifF polypeptide, NifJ polypeptide, and NifV polypeptide, and optionally FdxN polypeptide. In a most preferred embodiment, at least some of each of the polypeptides are soluble in the mitochondria of the plant cell.

[0066] In one embodiment, the plant cells further comprise an exogenous polynucleotide encoding a NifM polypeptide.

[0067] In another embodiment, the present invention provides a plant cell comprising a NifH fusion polypeptide comprising two NifH polypeptides covalently linked by an oligopeptide linker in the order NifH::linker::NifH from N-terminus to C-terminus, preferably in the order AnfH::linker::AnfH, preferably two AnfH polypeptides, optionally further comprising a C-terminal peptide from a mitochondrial targeting peptide (MTP) located at its N-terminus.

[0068] In one embodiment, the two NifH polypeptides, preferably the two AnfH polypeptides, are the same, or differ only in that one of the NifH polypeptides has methionine at amino acid position 1.

[0069] In one embodiment, the NifH fusion polypeptide, preferably the AnfH fusion polypeptide, is more soluble in the mitochondria of the cell than the corresponding NifH or AnfH polypeptide having only a single NifH or AnfH polypeptide.

[0070] In one embodiment, the two NifH polypeptides, preferably the two AnfH polypeptides, are each at least 70% identical to the amino acid sequence provided in SEQ ID NO:37 and / or SEQ ID NO:39, preferably at least 80% identical, more preferably at least 90% identical, and most preferably at least 95% identical.

[0071] In one embodiment, the NifH fusion polypeptide is a cleavage product of an encoded polypeptide comprising a mitochondrial targeting peptide (MTP) translatably fused to the NifH::linker::NifH polypeptide, wherein the MTP is preferably translatably fused at the N-terminus of the NifH polypeptide, wherein the NifH fusion polypeptide is generated in the plant cell by protease cleavage of the encoded polypeptide within or adjacent to the MTP, wherein the NifH fusion polypeptide comprises (i) a C-terminal peptide from the MTP located at its N-terminus, or (ii) does not contain a C-terminal peptide from the MTP.

[0072] In one embodiment of the above aspects, the modified NifH polypeptide or the NifH fusion polypeptide is a cleavage product of a polypeptide encoded by an exogenous polynucleotide in the plant cell, i.e., a precursor polypeptide containing the modified NifH polypeptide or the NifH fusion polypeptide, wherein the encoded polypeptide contains a mitochondrial targeting peptide (MTP) translatorily fused to the modified NifH polypeptide or the NifH::linker::NifH polypeptide, wherein the MTP is preferably translatorily fused at the N-terminus of the modified NifH polypeptide or the NifH::linker::NifH polypeptide, wherein the modified NifH polypeptide or the NifH fusion polypeptide is generated in the plant cell by protease cleavage of the encoded polypeptide within or adjacent to the MTP.

[0073] Therefore, in one embodiment, the present invention provides a plant cell comprising an exogenous polynucleotide encoding a NifH fusion polypeptide, preferably an AnfH fusion polypeptide, wherein the exogenous polynucleotide is operatively linked to a promoter guiding the expression of the exogenous polynucleotide in the plant cell, wherein the encoded polypeptide comprises a NifH::linker::NifH polypeptide, preferably a modified NifH::linker::modified NifH polypeptide N-terminus fused translationally to the N-terminus of a modified NifH::linker::modified NifH polypeptide, preferably as described in the embodiments above, wherein each NifH sequence contains at least one amino acid-substituted N-terminus, optionally having an intercalated oligopeptide linker located between the N-terminus and a first NifH sequence, wherein the NifH fusion polypeptide, preferably an AnfH fusion polypeptide, is generated in the plant cell by protease cleavage of the encoded polypeptide within or adjacent to the N-terminus, and wherein the cleavage product NifH fusion polypeptide has increased solubility in the plant cell. When the protease cleaves the protein within the MTP sequence, the NifH fusion polypeptide has the following structure: scar::NifH::linker::NifH, or scar::AnfH::linker::AnfH, optionally having an intercalated oligopeptide linker located between the scar sequence and the first NifH / AnfH sequence. Preferably, the two NifH sequences or the two AnfH sequences comprise at least one amino acid set from the amino acid sets listed in the above embodiments, i.e., both are modified NifH / AnfH sequences.

[0074] In one embodiment, the exogenous polypeptide is integrated into the genome of the plant cell, preferably into the nuclear genome, and / or the modified NifH polypeptide is a modified AnfH polypeptide, and / or the NifH fusion polypeptide is an AnfH fusion polypeptide.

[0075] In one embodiment, the modified NifH polypeptide, the NifH fusion polypeptide, or the encoded polypeptide is capable of transferring electrons to NifDK to reduce acetylene and / or N2 gas. In another embodiment, the modified NifH polypeptide, the NifH fusion polypeptide, or the encoded polypeptide is capable of transferring electrons to NifDK to reduce acetylene and / or N2 gas in plant cells.

[0076] In addition, a modified NifH polypeptide as defined herein is provided.

[0077] A NifH fusion peptide as defined herein is also provided.

[0078] Additionally, an encoded polypeptide as defined herein is provided.

[0079] In one embodiment, the MTP of the encoded polypeptide is translatorily fused at the N-terminus of the modified NifH polypeptide or the NifH::linker::NifH polypeptide.

[0080] In one embodiment, the modified NifH polypeptide, the NifH fusion polypeptide, or the encoded polypeptide is a modified AnfH polypeptide, an AnfH fusion polypeptide, or an encoded polypeptide that respectively contains one or two AnfH polypeptides.

[0081] In one embodiment, the encoded polypeptide comprises an AnfH polypeptide and a mitochondrial targeting peptide (MTP) translatorily fused to the AnfH polypeptide. In one embodiment, the MTP is translatorily fused to the N-terminus of the AnfH polypeptide.

[0082] On the other hand, the present invention provides an exogenous polynucleotide comprising a promoter operatively linked to a nucleotide sequence, the exogenous polynucleotide encoding the modified NifH polypeptide of the present invention, the NifH fusion polypeptide of the present invention, or the encoded polypeptide of the present invention, wherein the promoter directs the expression of the nucleotide sequence in cells, preferably plant cells.

[0083] In one embodiment, the modified NifH polypeptide is a modified AnfH polypeptide.

[0084] In one embodiment, the NifH fusion polypeptide is an AnfH fusion polypeptide.

[0085] In one embodiment, the encoded polypeptide is an AnfH polypeptide.

[0086] In one embodiment, the exogenous polynucleotide is integrated into the genome of the cell, preferably into the nuclear genome of the plant cell.

[0087] In one embodiment, the protein-coding region of the polynucleotide has been codon-modified for expression in plant cells.

[0088] In another aspect, the present invention provides a carrier comprising the exogenous polynucleotide of the present invention.

[0089] On the other hand, the present invention provides a transgenic plant or a portion thereof, comprising one or more of the following: plant cells of the present invention, polypeptides of the present invention, or polynucleotides of the present invention, or preferably the transgenic plant or a portion thereof is a transgenic plant containing exogenous polynucleotides of the present invention.

[0090] In one embodiment, the modified NifH polypeptide, the NifH fusion polypeptide, or the encoded polypeptide is a modified AnfH polypeptide, an AnfH fusion polypeptide, or an encoded polypeptide that respectively contains one or two AnfH polypeptides.

[0091] In one embodiment, the portion is a genetically modified seed.

[0092] In one embodiment, the plant is a cereal plant or a portion thereof. Examples include, but are not limited to, wheat, rice, corn, rye, oats, or barley plants.

[0093] In another aspect, the present invention provides a method for selecting the modified NifH polypeptide of the present invention, preferably the modified AnfH polypeptide or the NifH fusion polypeptide, preferably the AnfH fusion polypeptide, the method comprising... i) Expressing the exogenous polynucleotides of the present invention in plant cells, ii) Extract a protein containing the modified NifH polypeptide or the NifH fusion polypeptide from the mitochondria of the plant cells. iii) Determine the solubility level of the modified NifH peptide or the NifH fusion peptide in the extracted protein, and iv) Select the modified NifH polypeptide or the NifH fusion polypeptide, wherein at least 15%, preferably at least 50%, of the modified NifH polypeptide or the NifH fusion polypeptide in the cells is soluble.

[0094] In one embodiment of the foregoing aspects, the modified NifH has one or more of the characteristics of a modified NifH polypeptide as defined herein, or the NifH fusion polypeptide has one or more of the characteristics of a NifH fusion polypeptide as defined herein. In one embodiment, the NifH fusion polypeptide comprises two NifH polypeptides, preferably two AnfH polypeptides, covalently linked by an oligopeptide linker in the order NifH::linker::NifH from N-terminus to C-terminus, preferably in the order AnfH::linker::AnfH, optionally further comprising a C-terminal peptide from a mitochondrial targeting peptide (MTP) located at its N-terminus.

[0095] In one embodiment, the modified NifH polypeptide has at least one amino acid substitution that imparts a lower free energy to the modified NifH polypeptide compared to a corresponding NifH polypeptide lacking the at least one amino acid substitution.

[0096] In another aspect, the present invention provides a method for producing the modified NifH polypeptide of the present invention, preferably the modified AnfH polypeptide or the NifH fusion polypeptide, preferably the AnfH fusion polypeptide, said method comprising expressing the exogenous polynucleotide of the present invention in plant cells or transgenic plants.

[0097] In another aspect, the present invention provides a method for selecting plant cells that produce the modified NifH polypeptide of the present invention, preferably the modified AnfH polypeptide or the NifH fusion polypeptide, preferably the AnfH fusion polypeptide, the method comprising... i) Expressing the exogenous polynucleotides of the present invention in plant cells, ii) Determine whether the modified NifH peptide or the NifH fusion peptide is produced at the desired level and / or has the desired activity. iii) Select the plant cells based on the results of step ii). iv) Optionally, transgenic plants are produced from selected plant cells, and v) Optionally, the transgenic plant produces transgenic progeny plants and / or transgenic seeds.

[0098] In another aspect, the present invention provides a method for selecting plants that produce the modified NifH polypeptide of the present invention, preferably a modified AnfH polypeptide or a NifH fusion polypeptide, preferably an AnfH fusion polypeptide, the method comprising... i) Expressing the exogenous polynucleotide of the present invention in one or more plants. ii) Determine whether the modified NifH polypeptide or the NifH fusion polypeptide is produced at the desired level and / or has the desired activity in the one or more plants. iii) Select a plant from step ii) that produces the modified NifH polypeptide or the NifH fusion polypeptide at a desired level and / or has the desired activity in the plant or a portion thereof, and iv) Optionally, transgenic progeny plants and / or transgenic seeds may be produced from transgenic plants.

[0099] In one embodiment of the above aspects, at least 15%, preferably at least 50%, of the modified NifH polypeptide or the NifH fusion polypeptide in the cell or the plant or a portion thereof is soluble.

[0100] In one embodiment of the above aspects, the modified NifH polypeptide or the NifH fusion polypeptide contains an MTP preferably located at the N-terminus of the polypeptide for localization to the plant cell or the mitochondria of the plant.

[0101] In one embodiment of the above aspects, the NifH fusion polypeptide comprises two NifH polypeptides, preferably two AnfH polypeptides, covalently linked by an oligopeptide linker in the order NifH::linker::NifH from N-terminus to C-terminus, preferably in the order AnfH::linker::AnfH, and optionally further comprising a C-terminal peptide from a mitochondrial targeting peptide (MTP) located at its N-terminus. In one embodiment, the linker between the two NifH or AnfH polypeptides is 10-50 residues long, preferably 16-50 residues long or 20-35 residues long, more preferably about 25 or about 30 residues long, or most preferably 25 or 30 residues long.

[0102] In one embodiment of the above aspects, the exogenous polynucleotide is integrated into the plant cell or the genome of the plant, preferably into the nuclear genome of the plant cell or the plant.

[0103] The invention also provides the use of the exogenous polynucleotide of the present invention and / or the vector of the present invention for the production of transgenic plant cells.

[0104] Additionally, a method for producing transgenic plants is provided, the method comprising the following steps: i) Introducing one or more exogenous polynucleotides of the present invention and / or the vectors of the present invention into plant cells. ii) Regenerating the transgenic plant of the present invention from the cells described in step i), and iii) Optionally, transgenic seeds and / or progeny plants are produced from the transgenic plant regenerated in step ii).

[0105] In another aspect, the present invention provides a method for producing transgenic seeds, the method comprising: i) Collecting seeds from the transgenic plants of the present invention, and / or ii) Collect seeds from one or more transgenic progeny plants produced by the method of the present invention.

[0106] In another aspect, the present invention provides a method for producing flour, whole wheat flour, starch, oil, seed flour or other products obtained from seeds, the method comprising extracting flour, whole wheat flour, starch, oil or other products from the seeds of the present invention, or producing the seed flour from the seeds.

[0107] In another aspect, the present invention provides a product produced from the genetically modified plant of the present invention or a portion thereof, such as the seed of the present invention, wherein said product comprises one or more of the following: the modified NifH polypeptide of the present invention, the NifH fusion polypeptide of the present invention, the encoded polypeptide of the present invention, and the exogenous polynucleotide of the present invention.

[0108] In another aspect, the present invention provides a method for preparing food, the method comprising mixing the seeds of the present invention or flour, whole wheat flour, starch, oil or other products obtained from said seeds with another food ingredient.

[0109] In another aspect, the present invention provides a method of feeding an animal, the method comprising providing the animal with a plant or a portion thereof, such as the seeds of the present invention, or a product of the present invention.

[0110] Unless otherwise specified, any embodiments described herein should be considered applicable to any other embodiments after necessary modifications.

[0111] The scope of this invention is not limited to the specific embodiments described herein, which are intended for illustrative purposes only. Functionally equivalent products, compositions, and methods as described herein are clearly within the scope of this invention.

[0112] Throughout this specification, unless expressly stated otherwise or the context otherwise requires, references to a single step, a composition of substance, a group of steps, or a group of compositions of substance shall be deemed to cover one or more (i.e., one or more) of such steps, compositions of substance, groups of steps, or groups of compositions of substance.

[0113] The invention is described below by way of the following non-limiting examples and with reference to the accompanying drawings. Attached Figure Description

[0114] Figure 1 Wild-type Vignelandian nitrogen-fixing bacteria ( A. vinelandii AlphaFold predictions of the tertiary structure of the NifU protein show the N-terminal, central, and C-terminal domains.

[0115] Figure 2 : Structural model of FeFe nitrogenase. The polypeptide subunits that form the AnfDKGH complex are labeled with their Anf letters. For example, members of the polypeptide pair are shown as K and K'. The portion of AnfD that appears to span AnfK' to AnfD is a segment without a corresponding structure for constructing the model and may not be an accurate description of these residues, but it is included for completeness.

[0116] Figure 3Western blot analysis of HA-labeled AnfH fusion peptides produced in plant cells and targeting the mitochondrial matrix. Lanes labeled T were loaded with total protein extracts, while lanes labeled S were loaded with soluble protein extracts from the same infiltration. WT refers to wild-type AnfH fusion peptides, and... The labels correspond to variants V1-V5. The top image shows the signal after detecting HA-labeled peptides, and the bottom image shows total protein staining of the same gel with Coomassie blue.

[0117] Figure 4 Coomassie staining gel and Western blot analysis of different steps in the purification process of the scar::TS::AnfH variant V1 fusion peptide after targeting plant mitochondria. Samples are (from left to right): protein size markers, total protein extract, supernatant packed onto the column, precipitate after centrifugation of the total sample, flow-through from the column, and eluted protein. Western blots show the TwinStrep-labeled AnfH-V1 peptide in these fractions, with arrows indicating positions between 33 kDa and 40 kDa (AnfH).

[0118] Figure 5 Top: A scatter plot of the log-likelihood of amino acid substitutions in the AnfH polypeptide, focusing on non-conserved residues that have not yet become the highest log-likelihood substitutions at the wild-type position. The x-axis is the change in Rosetta energy units from wild-type to substituted AnfH, and the y-axis is the log-likelihood of the substitution. All potential substitutions are shown as circles. Among the potential substitutions, those selected from all five variants in Table 4 by PROSS are shown as triangles. Bottom: A scatter plot of possible stable substitutions, i.e., the upper left quadrant of the top plot, where the substitutions from wild-type to monosubstituted AnfH... Rosetta energy units are negative, and the probability of substitution is positive. Potential substitutions are shown as circles, and the selected substitutions of variant V1 are shown as triangles.

[0119] Figure 6 The image above shows the wild-type and variant AnfH fusion peptides from two biological replication experiments in *Nicotiana sapiens* (Bernobyl). N. benthamiana Western blot and Coomassie staining gels were used to analyze the total extract (T) and soluble extract (S) from the leaves after expression. Lanes were labeled for the construct encoding the fusion peptide as follows: SN675 (WT), SN676 (…). 5) SN710 ( 3R), SN708 4) and SN709 ( 3H). Below: Box-and-whisker plot of the integrated signal intensity of the bands in the protein blot above.

[0120] Figure 7 Coomassie staining gel and Western blot analysis of samples from different steps of the purification process following targeting of plant mitochondria for the fusion peptides derived from wild-type scar::TS::AnfH (A), variant V7 from SN709 (B), and variant V8 from SN710 (C). Samples are (from left to right): total protein extract (T), precipitate after centrifugation of the total sample (P), supernatant packed onto the column (S), flow-through buffer from the column (FT), and eluted protein (E). The scar::TS::AnfH peptides in these fractions are labeled. The positions of molecular weight markers (kDa) are shown.

[0121] Figure 8 The average acetylene reduction (A and B) and N2 reduction δ- of AnfH-V7 and AnfH-V8 peptides isolated from plant mitochondria when combined in vitro with NifDK from Venetian violaceum. 15 N value (C) activity, regardless of whether it is from Escherichia coli ( E. coli The purified NifU peptides were processed in vitro. Symbols indicate data from repeated experiments, and bars show the average values ​​across multiple experiments. A: AnfH protein was measured 'as assoluted' without NifU treatment; B: Acetylene reduction activity of wild-type AnfH and AnfH variants isolated from *Nicotiana benthamiana* after NifU treatment; C: δ-N2 reduction of isolated wild-type AnfH and AnfH-V7 and AnfH-V8 peptides in vitro after NifU treatment in combination with purified AnfDKG. 15 N value. For wild-type AnfH, only one experiment was performed due to the limited availability of the isolated peptides. Labels: AvAnfH refers to the presence of purified wild-type Azotobacter venereideum AnfH, AvNifDK refers to purified wild-type Azotobacter venereideum NifD, and AnfDKG refers to purified wild-type Azotobacter venereideum AnfDKG, all of which were used for positive control reactions; EcNifU refers to TS::NifU protein purified from Escherichia coli, which was reloaded with Fe-S clusters before use and added at the indicated ratio; NbAnfH-wt, NbAnfH-V7, and NbAnfH-V8 refer to wild-type V7 and V8 AnfH peptides isolated from the mitochondria of Nicotiana benthamiana and added at the indicated ratio, respectively.

[0122] Figure 9 Comparison of the structures of methyl viologen (MV, a) and its sulfonated derivative (S2V, b).

[0123] Figure 10 This is a time-dependent, S2V-based in vitro colorimetric assay used to measure the activity of a combination of purified bacterial components, namely AvNifH, AvAnfH, AvNifD-NifK (NifDK), and AvAnfD-AnfK-AnfG (AnfDKG). The ratios of AvNifH / AvNifDK and AvAnfDKG were maintained at 20:1, 30:1, and 30:1 respectively.

[0124] Figure 11 : such as the δ- of three plant samples determined by IRMS 15 N. A: Increase the amount of unpurified 15 N2 gas is added to the plant in a sealed container at a specified percentage, and samples are taken after 24 hours in the dark. B: Use unpurified or purified gas. 15 N2 gas was added to leaf discs in a sealed container, and samples were taken after 24 hours for IRMS. C: Leaves were infiltrated with live Venetian nitrogen-fixing bacteria cells (P sample) or with non-nitrogenous Agrobacterium spp. ( Agrobacterium The sample was infiltrated (A sample) and analyzed by IRMS after 24 hours in the dark.

[0125] Figure 12 : DNA modules used to construct expression vectors, indicating the following element sequences: A, a non-replicating vector with a single transcription terminator (Tm); B, a non-replicating vector with a hybrid transcription terminator (TTm); and C, a replicating vector with a GV replication region. Select the MTP, protein-coding region (CDS), and epitope TAG sequences as needed. Restriction enzymes Bsa I is cut just inside the end of the module to provide a compatible end for attachment to the T-DNA of the binary vector, thereby providing a gene construct for introduction into plant cells.

[0126] Figure 13Western blot and Coomassie staining gels of Nif fusion peptides expressed by standard non-replicating vectors (C), GV-based replicating vectors (GVc or GVt). As indicated, the gene constructs encode MTP::NifD::HA, MTP::HA:NifH, MTP::NifU::HA (SN466, SN499, SN500), or MTP::GFP::HA (SN493), or MTP::GFP (pRA1) fusion peptides. The molecular weight (kDa) of the blotted peptides is shown. The lower left and lower right panels show the gels after Coomassie staining. The bands of the NifD and NifH fusion peptides expressed by the GVc and GVt vectors are outlined. The positions of the dominant protein bands of rubisco are indicated by arrows.

[0127] Figure 14 Gel electrophoresis and Western blotting were performed on NifH::linker::NifH fusion peptides expressed by a standard non-replicating vector with a single terminator (SN655), a non-replicating vector using a hybridization terminator (TTm, SN656), and a GV-based replicating vector (SN657). Each gene construct was co-infiltrated with a construct encoding the MTP::NifM fusion peptide (SN360). Extracts from tobacco leaves after 4 days were processed for total protein (T), soluble protein (S), and insoluble protein (I) as described herein. The inset figure below shows the gel after Coomassie staining, and the outer lanes show molecular weight marker proteins.

[0128] Figure 15 Analysis of the solubility of the scar::NifH::linker::NifH dimer fusion peptide after production and localization to mitochondria in plant cells by Western blotting (top figure). A gene construct (SN283) for producing the MTP::NifH::linker::NifH::HA dimer peptide (SEQ ID NO:103) was introduced into *Nicotiana benthamiana* leaf cells. The construct was not used to express accessory proteins (lanes 1-3, HH only), or MTP::GFP was used as a control (lanes 4-6), or MTP::NifM was used (lanes 7-9). Lanes labeled T were from total protein extracts, lanes labeled S were from soluble protein extracts, and lanes labeled I were from insoluble protein extracts. Lane 7 shows the peptide size marker (kDa). Western blotting was performed using an anti-HA antibody. The bottom figure shows Coomassie staining gels as controls for protein loading in each lane.

[0129] Figure 16Analysis of the processing of the MTP::NifH::linker::NifH dimer fusion peptide in plant mitochondria by Western blotting (above figure), using an anti-HA antibody for detection. Based on Klebsiella acidogenic bacteria ( K. oxytoca The NifH sequence was used to introduce a gene construct (SN283) encoding the MTP::KoNifH::linker::KoNifH::HA dimer polypeptide (SEQ ID NO:103) into tobacco leaf cells using MTP-KoNifM. The left lane shows the polypeptide size marker (kDa). Lane 1, 35S-p19 only; 2, MTP-FAγ51::KoNifH::Connector::KoNifH::HA (SN283); 3, Unprocessed alaMTP-FAγ51::KoNifH::Connector::KoNifH::HA (SN807); 4, 6xHis::KoNifH::Connector::KoNifH::HA (SN303); 5, MTP-FAγ51::HA::KoNifH::Connector (TS)::KoNifH (SN656); 6, Unprocessed alaMTP-FAγ51HA::KoNifH::Connector (TS)::KoNifH (SN808). Hollow triangles indicate strips corresponding to unprocessed MTP::NifH::Connector::NifH, and solid triangles indicate strips corresponding to MPP-processed scar::NifH::Connector::NifH. The figure below shows Coomassie stained gels as controls for protein loading in each lane. Lanes 1 and 2 were loaded with less protein than lanes 3-6.

[0130] Figure 17 Analysis of the solubility of the scar::AvNifH::linker::NifH dimer fusion peptide after its production and localization to mitochondria in plant cells by Western blotting (top image). Based on the NifH sequence of *Azotobacter venereum*, a gene construct (SN678) for producing the MTP::NifH::linker::NifH dimer peptide (SEQ ID NO: 109) was introduced into *Nicotiana benthamiana* leaf cells lacking NifM (lanes 2-4, -AvNifM) or containing MTP::NifM from SN605 (lanes 5-7). Lanes labeled T were from total protein extracts, lanes labeled S were from soluble protein extracts, and lanes labeled I were from insoluble protein extracts. Lane 1 shows the peptide size marker (kDa). The bottom image shows Coomassie staining gels as controls for protein loading in each lane.

[0131] Figure 18The amount of ethylene produced in an in vitro acetylene reduction assay (ARA) when NifD and NifK from *Venelandia diffusa* are combined with scar::TS::KoNifH::linker::KoNifH isolated from SN638 (strips 1 and 2) or SN650 (strips 3-5) or scar::HA::KoNifH::linker (TS)::KoNifH) isolated from SN656 (strips 6 and 7) (isolated from mitochondria in each case). As indicated in the figure below, these constructs were co-expressed with the gene constructs encoding MTP::KoNifM (KoNifM), MTP::AvNifS, and MTP::AvNifU (AvNifSU), MTP::AvFdxN::HA (AvFdxN), and MTP::KoNifD::linker (HA)::NifK (KoNifD::HA::K). Some leaves were collected after a certain period of growth in darkness (dark collection). Fe and cysteine ​​were added to some separation buffers. The ratio of the amount of NifH dimer:AvNifDK protein in the assay is indicated.

[0132] Figure 19 The dimer NifH protein scar::HA::NifH::linker (TS)::NifH), expressed and isolated from plant mitochondria, is functional in an in vitro acetylene reduction assay using purified bacterial NifD and NifK (AvNifDK). This figure shows the amount of ethylene produced by plant-expressed scar::HA::NifH::linker (TS)::NifH), produced by bacterial NifD and NifK combined with bacterial NifH (AvNifH-AnNifDK) as a positive control for full-functioning NifH, bacterial NifD and NifK without bacterial NifH (AnNifDK only) as a negative control, or in the presence of NifM (KoHH_KoM) or additionally in the presence of NifM, NifS, and NifU (KoHH_KoM_AvSU). The three right bars show the ethylene produced after in vitro treatment of the NifH peptide with NifU protein loaded with Fe-S clusters purified from E. coli, with scar::HA::NifH::linker (TS)::NifH protein or apoNifH from Azotobacter venerealis as a positive control for Fe-S recombinant experiments. Each bar represents a biological replicate.

[0133] Figure 20Analysis of the solubility of the scar::AnfH-V7::linker::AnfH-V7 dimer fusion peptide after production and localization to mitochondria in plant cells, performed by Western blotting (left panel) or Coomassie staining (right panel). Gene constructs SN775 (e35S lane), SN776 (TTm), and SN777 (GV) with three different expression systems were expressed in plant leaves to produce the scar::AnfH-V7::linker::AnfH-V7 dimer peptide. Proteins extracted from leaves were run on gels to analyze total protein extract (T), soluble protein extract (S), and insoluble protein formulation (I). Lanes 1 (top) or 4 (bottom) on the Western blot show peptide size markers (kDa). The Western blot was probed using an anti-Strep antibody. The right panel shows a Coomassie-stained gel as a control for protein loading in each lane.

[0134] Figure 21 Western blot and Coomassie staining gels of samples from different steps of the affinity purification process of the scar::AnfH-V7::linker (TS)::AnfH-V7) peptide generated by SN776 after targeting plant mitochondria. Samples from (from left to right): total protein extract (T), precipitate after centrifugation of total sample (P), supernatant packed onto the column (I), flow-through buffer from the column (FT), and eluted protein (E). The positions of molecular weight markers (kDa) are shown.

[0135] Figure 22 Top image: The scar::AnfH-V7::linker (TS)::AnfH-V7 peptide generated by SN776, after targeting plant mitochondria and combining with NifDK from *Azotobacter venerealis*, produces ethylene from acetylene in vitro. Bottom image: The same V7 peptide combined with AnfDKG from *Azotobacter venerealis*. 15 δ- in N2 reduction determination 15N. Lanes labeled with AvNifDK were purified and supplemented with bacteria AvNifDK; lanes labeled with AvAnfDKG were purified and supplemented with bacteria AvAnfDKG; lanes labeled with AvAnfH were purified and supplemented with bacteria AvAnfH at a ratio of 20:1 or 40:1 as a positive control; lanes labeled with SN776 AvNifDK / AvAnfDKG were supplemented at a ratio of 20:1. An mixture of plant-produced and isolated AnfH-V7 dimer peptide (untreated with NifU) and purified bacteria AvNifDK or AvAnfDKG was added at a ratio of 10:1; and lanes labeled with SN776AvNifDKU / AvAnfDKGU were filled with isolated AnfH-V7 dimer peptide (treated with NifU) and purified bacteria AvNifDK or AvAnfDKG at a ratio of 20:1 or 10:1.

[0136] Figure 23 Amino acid sequence variations between wild-type NifH sequences from natural sources and thermophilic bacteria. The numbers at the top of each figure are position numbers in the consensus alignment. For example, referring to the consensus sequences, significant differences can be seen at the positions shown in Figures (A) 71 and 75, (B) 125, (C) 153, and (D) 200 and 204.

[0137] Figure 24 Western blot analysis was performed on transgenic Nicotiana benthamiana plants transformed with T-DNA from SL149. The blot was probed with an anti-TS antibody to detect the approximately 38 kDa scar::TS::AvAnfH-V7 peptide. A weak upper band indicates binding of the anti-TS antibody to a nonspecific endogenous protein.

[0138] Symbol explanation of sequence lists SEQ ID NO:1 is the amino acid sequence of the NifH polypeptide from Klebsiella acidogenic bacteria, 293aa.

[0139] SEQ ID NO:2 The amino acid sequence of wild-type NifD polypeptide from Klebsiella pneumoniae according to accession number X13303.1; 483aa.

[0140] SEQ ID NO:3 Amino acid sequence of NifK polypeptide from Klebsiella acidogenic bacteria according to Temme et al., (2012); 520aa.

[0141] The amino acid sequence of SEQ ID NO:4 is 468aa, derived from the NifB polypeptide of Klebsiella acidogenic bacteria.

[0142] The amino acid sequence of the NifE polypeptide from Klebsiella acidogenic bacteria, SEQ ID NO:5, is 457aa.

[0143] SEQ ID NO:6 Amino acid sequence of NifF polypeptide from Klebsiella acidogenic bacteria, 176 aa; NCBI accession number X03214.

[0144] SEQ ID NO:7 Amino acid sequence of NifJ polypeptide from Klebsiella acidogenic bacteria, 1171 aa; NCBI accession number WP_064371580, Cannon et al., (1988).

[0145] SEQ ID NO:8 Amino acid sequence of NifM polypeptide from Klebsiella acidogenic bacteria, 266 aa; NCBI accession number X05887; Paul and Merrick (1987).

[0146] SEQ ID NO:9 is the amino acid sequence of the NifN polypeptide from Klebsiella acidogenic bacteria, NCBI accession number P08738; 461aa; (Arnold et al., 1988).

[0147] SEQ ID NO:10 is from the genus Klebsiella ( Klebsiella The amino acid sequence of the NifQ polypeptide. NCBI accession number WP_004138772.

[0148] SEQ ID NO:11 is the amino acid sequence of the NifS polypeptide from Klebsiella acidogenic bacteria, 400aa.

[0149] SEQ ID NO:12 Amino acid sequence of NifU polypeptide from Klebsiella acidogenic bacteria; 274aa. NCBI accession number P05343.2 (Arnold et al., 1988).

[0150] SEQ ID NO:13 Amino acid sequence of NifV polypeptide from Klebsiella acidogenic bacteria; 381aa. NCBI accession number CAA31119.1 (Arnold et al., 1988) SEQ ID NO:14 The amino acid sequence of the NifX polypeptide from Klebsiella acidogenic bacteria, 156aa (accession number P09136).

[0151] SEQ ID NO:15 The amino acid sequence of the NifY polypeptide from Klebsiella acidogenic bacteria, 220aa; NCBI accession number CAA31670 (Arnold et al., 1988).

[0152] SEQ ID NO:16 Amino acid sequence of NifZ polypeptide from Klebsiella acidogenic bacteria, 148aa; NCBI accession number P0A3U2 (Arnold et al., 1988).

[0153] SEQ ID NO:17. Amino acid sequence of NifW polypeptide from Klebsiella acidogenic bacteria.

[0154] SEQ ID NO:18. Based on the amino acid sequence of wild-type Klebsiella acid-producing NifD by Temme et al. (2012).

[0155] SEQ ID NO:19. Based on the amino acid sequence of wild-type Klebsiella acid-producing NifS by Temme et al. (2012).

[0156] SEQ ID NO:20. The amino acid sequence of the MTP-FAγ51::NifH::HA fusion polypeptide encoded by SN18 and SN27. Amino acids 1-54 correspond to MTP-FAγ51 having an additional initiating methionine and a C-terminal GG, amino acids 55-347 are NifH amino acids from Klebsiella acidogenic bacteria (SEQ ID NO:1), and amino acids 348-358 include the HA epitope.

[0157] SEQ ID NO:21. The amino acid sequence of the NifD::linker::NifK fusion polypeptide linker. The linker is 30 residues long and is derived from *Carya paliurus* (…). Hypocrea jecorina It consists of an 11-residue segment of cellobiase II (accession number AAG39980.1; SEQ ID NO:38), in which the last arginine is replaced by alanine, followed by a 9-residue HA epitope (SEQ ID NO:39), and then another copy of the 11-residue segment in which arginine is replaced by alanine.

[0158] SEQ ID NO:22. Peptide sequence.

[0159] SEQ ID NO:23. Peptide sequence.

[0160] SEQ ID NO:24. The amino acid sequence of the MTP-FAγ51::NifD(Y100Q)::linker(HA)::NifK fusion polypeptide encoded by SN159. Amino acids 1-54 correspond to MTP-FAγ51 with GG at its C-terminus, amino acids 55-536 correspond to Klebsiella acidogenic NifD with Y100Q substitution, amino acids 537-566 correspond to the linker including the HA epitope, and amino acids 567-1085 correspond to NifK (SEQ ID NO:3) without its N-terminus Met and with its wild-type C-terminus.

[0161] SEQ ID NO:25. Amino acid sequence of NifV polypeptide from *AvNifV* (accession number WP_012698855).

[0162] SEQ ID NO:26. The amino acid sequence of the MTP-FAγ51::AnfD::HA polypeptide encoded by SN81. Amino acids 1-54 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, amino acids 55-572 correspond to the AnfD sequence from *Venerendae*, and amino acids 573-583 correspond to the HA epitope.

[0163] SEQ ID NO:27. The amino acid sequence of the MTP-FAγ51::HA::AnfK polypeptide encoded by SN129. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, amino acids 54-64 correspond to the HA epitope, and amino acids 65-526 correspond to the AnfK sequence from *Venelandia diffusa*.

[0164] SEQ ID NO:28. The amino acid sequence of the MTP-FAγ51::HA::AnfH polypeptide encoded by SL48 and SL79. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, amino acids 54-64 correspond to the HA epitope having a GG linker at its C-terminus, and amino acids 65-339 correspond to the AnfH sequence from *Venerendae*.

[0165] SEQ ID NO:29. The amino acid sequence of the MTP-FAγ51::HA::AnfG polypeptide encoded by SL48 and SL79. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, amino acids 54-64 correspond to the HA epitope having a GG linker at its C-terminus, and amino acids 65-196 correspond to the AnfG sequence from *Azotobacter venerealis*.

[0166] SEQ ID NO:30. The amino acid sequence of the MTP-FAγ51::HA::AnfD polypeptide encoded by SN161. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, amino acids 54-64 correspond to the HA epitope having a GG linker at its C-terminus, and amino acids 65-582 correspond to the AnfD sequence from *Venerendae*.

[0167] SEQ ID NO:31. The amino acid sequence of the MTP-FAγ51::AnfD::Twin Strep polypeptide encoded by SN177. Amino acids 1-54 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, amino acids 55-572 correspond to the AnfD sequence from *Azotobacter venerealis*, and amino acids 573-604 correspond to the TwinStrep epitope.

[0168] SEQ ID NO:32. The amino acid sequence of the MTP-CoxIV::TwinStrep::AnfK polypeptide encoded by SN195. Amino acids 1-41 correspond to the MTP-CoxIV sequence including a GG linker at its C-terminus, amino acids 42-61 correspond to the TwinStrep epitope including GG at its C-terminus, and amino acids 62-523 correspond to the AnfK sequence from *Azotobacter venerealis*.

[0169] SEQ ID NO:33. The amino acid sequence of the MTP-FAγ51::AnfD::linker 26(HA)::AnfK polypeptide encoded by SN272. Amino acids 1-64 correspond to the MTP-FAγ51::HA sequence including GG at its C-terminus, amino acids 65-581 correspond to the AnfD sequence (Venerend azotobacter), amino acids 582-607 correspond to the 26-amino acid linker (linker 26(HA)), and amino acids 608-1068 correspond to AnfK (Venerend azotobacter).

[0170] SEQ ID NO:34. The amino acid sequence of the MTP-CoxIV::AnfD::linker 26(HA)::AnfK polypeptide encoded by SL48 and SL79. Amino acids 1-61 correspond to the MTP-CoxIV sequence including GG at its C-terminus, amino acids 62-578 correspond to the AnfD sequence (Venerelandia diffusa), amino acids 579-604 correspond to the 26-amino acid linker (linker 26(HA)), and amino acids 605-1065 correspond to AnfK (Venerelandia diffusa).

[0171] SEQ ID NO:35. Amino acid sequence of AnfD from *Venelandia diffusa* (accession number WP_012703361); 518aa.

[0172] SEQ ID NO:36. Amino acid sequence of AnfK from *Venelandia diffusa* (accession number WP_012703359); 462aa.

[0173] SEQ ID NO:37. Amino acid sequence of AnfH from *Venelandia diffusa* (accession number WP_012703362); 275aa.

[0174] SEQ ID NO:38. Amino acid sequence of AnfG from *Venelandia diffusa* (accession number WP_012703360); 132aa.

[0175] SEQ ID NO:39. Amino acid sequence of NifH polypeptide from *Zytobacter venerealis* (AvNifH; accession number WP_012698831); 290aa.

[0176] SEQ ID NO:40. Peptide sequence, AnfH motif I, where X represents any amino acid.

[0177] SEQ ID NO:41. Peptide sequence, AnfH motif II.

[0178] SEQ ID NO:42. Peptide sequence, AnfH motif III.

[0179] SEQ ID NO:43. Peptide sequence, AnfH motif IV.

[0180] SEQ ID NO:44. Peptide sequence, AnfH motif V, where X represents any amino acid.

[0181] SEQ ID NO:45. Peptide sequence, AnfH motif VI.

[0182] SEQ ID NO:46. Peptide sequence, AnfH motif VII, where X represents any amino acid.

[0183] SEQ ID NO:47. Amino acid sequence of FdxN protein from *Zytobacter venerealis*; accession number WP_012703542; 92aa.

[0184] SEQ ID NO:48. Amino acid sequence of NafY polypeptide from *Azotobacter venerelandus* (AvNafY; accession number AGK13761).

[0185] SEQ ID NO:49-53. C-terminal amino acid sequence of NifK polypeptide.

[0186] The C-terminal amino acid sequence of the AnfK polypeptide, SEQ ID NO:54-58.

[0187] SEQ ID NO:59. Amino acid sequence of NifU polypeptide from wild-type Azotobacter venerealis, Genbank accession number WP_012698853.1; 312aa.

[0188] SEQ ID NO:60. Amino acid sequence of NifS polypeptide from wild-type Azotobacter venerelandus (Johnson et al., 2005; Yuvaniyama et al., 2000); NCBI reference sequence: WP_012698854.1; 402aa.

[0189] SEQ ID NO:61. Peptide sequence, NifH motif I, where X represents any amino acid.

[0190] SEQ ID NO:62. Peptide sequence, NifH motif II, where X represents any amino acid.

[0191] SEQ ID NO:63. Peptide sequence, NifH motif III, where X represents any amino acid.

[0192] SEQ ID NO:64. Peptide sequence, NifH motif V, where X represents any amino acid.

[0193] SEQ ID NO:65. Peptide sequence, NifH motif VII, where X represents any amino acid.

[0194] SEQ ID NO:66. The amino acid sequence of the MTP-FAγ51::HA::MiNifB fusion polypeptide encoded by SL78. Amino acids 1-53 correspond to MTP-FAγ51 having GG at its C-terminus, amino acids 54-64 include the HA epitope having GG, and amino acids 65-366 correspond to *Methanococcus hellebore* having its initiator Met. Methanocaldococcus infernus The corresponding value for NifB is 366aa.

[0195] SEQ ID NO:67. The amino acid sequence of a TS::NifU fusion polypeptide containing a NifU sequence from *Azotobacter venerealis* produced in *Escherichia coli*. Amino acids 10-41 correspond to the TwinStrep epitope, followed by GG, and amino acids 42-353 correspond to the NifU sequence from *Azotobacter venerealis*; 353aa.

[0196] SEQ ID NO:68. The amino acid sequence of the MTP-FAγ51::NifS fusion polypeptide encoded by SL122. Amino acids 1-54 correspond to MTP-FAγ51 with GG at its C-terminus, while amino acids 55-456 correspond to Nitrosporium viniferum NifS (SEQ ID NO:60), including its initiator Met;456aa.

[0197] SEQ ID NO:69. Amino acid sequence of the mitochondrial scar::NifS fusion polypeptide after MPP processing. Amino acids 1-11 correspond to the scar sequence from MTP-FAγ51, while amino acids 12-413 correspond to NifS from Venetian vinifera (SEQ ID NO:60); 413aa.

[0198] SEQ ID NO:70. The amino acid sequence of the MTP-CoxIV::TwinStrep::NifU fusion polypeptide encoded by SL122. Amino acids 1-29 correspond to MTP-CoxIV, amino acids 30-61 correspond to the TwinStrep epitope with GG at its C-terminus, and amino acids 62-373 correspond to NifU (SEQ ID NO:59) of the Venetian violaceae bacterium, which includes its initiator Met; 373aa.

[0199] SEQ ID NO:71. Amino acid sequence of the mitochondrial scar::TwinStrep::NifU fusion polypeptide after MPP processing. Amino acids 1-4 correspond to the scar sequence from MTP-CoxIV, amino acids 5-36 correspond to the TwinStrep epitope with GG at its C-terminus, and amino acids 37-348 correspond to NifU from Venetian violaceum (SEQ ID NO:59); 348aa.

[0200] SEQ ID NO:72. Amino acid sequence of variant V1 of AnfH from *Azotobacter venerelandus*, with 5 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0201] SEQ ID NO:73. Amino acid sequence of variant V2 of AnfH from *Azotobacter venerelandus*, with 6 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0202] SEQ ID NO:74. Amino acid sequence of variant V3 of AnfH from *Azotobacter venerealis*, with 9 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0203] SEQ ID NO:75. Amino acid sequence of variant V4 of AnfH from *Azotobacter venerealis*, with 10 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0204] SEQ ID NO:76. Amino acid sequence of variant V5 of AnfH from *Azotobacter venerealis*, with 11 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0205] SEQ ID NO:77. Amino acid sequence of variant V6 of AnfH from *Azotobacter venerealis*, with 4 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0206] SEQ ID NO:78. Amino acid sequence of variant V7 of AnfH from *Azotobacter venerealis*, with 3 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0207] SEQ ID NO:79. Amino acid sequence of variant V8 of AnfH from *Azotobacter venerelandus*, with 3 amino acid substitutions compared to wild-type AnAnfH (Table 4); 275 aa.

[0208] SEQ ID NO:80. The amino acid sequence of the MTP-CoxIV::TS::AvAnfH fusion polypeptide encoded by SN675. Amino acids 1-29 are derived from *Saccharomyces cerevisiae* cleaved by MPP between amino acids 25 and 26. S. cerevisiae The MTP-CoxIV of cytochrome c oxidase subunit 4, amino acids 30-61 are TwinStrep epitope tags including the C-terminal GG, amino acids 62-336 are the AnfH sequence of wild-type Azotobacter vinifera; 336aa.

[0209] SEQ ID NO:81. The amino acid sequence of the MTP-CoxIV::TS::AvAnfH-V1 fusion polypeptide encoded by SN676. Amino acids 1-29 are MTP-CoxIV, amino acids 30-61 are a TwinStrep epitope tag including a C-terminal GG, and amino acids 62-336 are the AvAnfH-V1 sequence; 336aa.

[0210] SEQ ID NO:82. The amino acid sequence of the MTP-CoxIV::TS::AvAnfH-V6 fusion polypeptide encoded by SN708. Amino acids 1-29 are MTP-CoxIV, amino acids 30-61 are a TwinStrep epitope tag including a C-terminal GG, and amino acids 62-336 are the AvAnfH-V6 sequence; 336aa.

[0211] SEQ ID NO:83. The amino acid sequence of the MTP-CoxIV::TS::AvAnfH-V7 fusion polypeptide encoded by SN709. Amino acids 1-29 are MTP-CoxIV, amino acids 30-61 are a TwinStrep epitope tag including a C-terminal GG, and amino acids 62-336 are the AvAnfH-V7 sequence; 336aa.

[0212] SEQ ID NO:84. The amino acid sequence of the MTP-CoxIV::TS::AvAnfH-V8 fusion polypeptide encoded by SN710. Amino acids 1-29 are MTP-CoxIV, amino acids 30-61 are a TwinStrep epitope tag including a C-terminal GG, and amino acids 62-336 are the AvAnfH-V8 sequence; 336aa.

[0213] SEQ ID NO:85. The amino acid sequence of the scar::TS::AvAnfH-V1 fusion polypeptide generated by SN676. Amino acids 1-4 are the scar sequence, C-terminal amino acids derived from MTP-CoxIV after MPP processing in mitochondria; amino acids 5-34 are TwinStrep epitope tags including C-terminal GG; and amino acids 35-311 are the AvAnfH-V1 sequence; 311aa.

[0214] SEQ ID NO:86. The amino acid sequence of the scar::TS::AvAnfH-V6 fusion polypeptide generated by SN708. Amino acids 1-4 are the scar sequence, C-terminal amino acids derived from MTP-CoxIV after MPP processing in mitochondria; amino acids 5-34 are TwinStrep epitope tags including C-terminal GG; and amino acids 35-311 are the AvAnfH-V6 sequence; 311aa.

[0215] SEQ ID NO:87. Amino acid sequence of the scar::TS::AvAnfH-V7 fusion polypeptide generated by SN709. Amino acids 1-4 are the scar sequence, C-terminal amino acids derived from MTP-CoxIV after MPP processing in mitochondria; amino acids 5-34 are TwinStrep epitope tags including C-terminal GG; and amino acids 35-311 are the AvAnfH-V7 sequence; 311aa.

[0216] SEQ ID NO:88. The amino acid sequence of the scar::TS::AvAnfH-V8 fusion polypeptide generated by SN710. Amino acids 1-4 are the scar sequence, C-terminal amino acids derived from MTP-CoxIV after MPP processing in mitochondria; amino acids 5-34 are TwinStrep epitope tags including C-terminal GG; and amino acids 35-311 are the AvAnfH-V8 sequence; 311aa.

[0217] SEQ ID NO:89. The amino acid sequence of the MTP-CPN60::AvNifS fusion polypeptide encoded by SL133. Amino acids 1-33 are MTP-CPN60 cleaved by MPP between amino acids 31 and 32, having a C-terminal GG, and amino acids 34-435 are the NifS sequence of *Venerendae*; 435aa.

[0218] SEQ ID NO:90. The amino acid sequence of the MTP-SU9::AvNifU fusion polypeptide encoded by SL133. Amino acids 1-68 are MTP-SU9 cleaved by MPP between amino acids 66 and 67, followed by GG, and amino acids 71-382 are the NifU sequence of *Venerendica vegana*; 336aa.

[0219] SEQ ID NO:91. The amino acid sequence of the MTP-FAγ51::AvFdxN fusion polypeptide encoded by SN291. Amino acids 1-52 are MTP-FAγ51 cleaved by MPP between amino acids 43 and 44, followed by GG, and amino acids 55-145 are the Venetian azeotropic bacteria FdxN sequence without its initiating methionine; 145aa.

[0220] SEQ ID NO:92. Nucleotide sequence of a DNA fragment comprising the P module of the LIR from BeYDV; fragment number EN38510. Bsa The I restriction sites are located at nucleotides 9-14 and 319-324, flanking the LIR region, including the 9 invariant nucleotides of the LIR located at positions 179-187; 332 nt.

[0221] SEQ ID NO:93. Nucleotide sequence of a DNA fragment containing the CaMV e35S promoter and the U module of the TMV 5'UTR region, fragment number EN38509. BsaThe I restriction site is located at nucleotides 9-14 and 864-869, the e35S promoter is located at nucleotides 13-767, and the TMV 5'UTR is located at position 796-852; the translation initiation ATG of the S module polypeptide is located at nucleotides 860-862; 877 nt.

[0222] SEQ ID NO:94. The nucleotide sequence of the DNA fragment of the T module used in the GV vector of this article, wherein the T module contains the CaMV 35S Tm transcription terminator and the SIR / Rep / RepA / LIR region from the benzivirus BeYDV, numbered EN38511. Bsa The I restriction site is located at nucleotides 9-14 and 1766-1771, the 35S Tm site is located at nucleotides 20-223, the SIR site is located at nucleotides 224-375, the Rep / RepA coding region is located in reverse orientation relative to nucleotides 1466-376, and the LIR site is located at nucleotides 1467-1760. The thymidine at nucleotide position 1469 is replaced by guanosine. Nucleotide position 1469 in GVc is G, and the nucleotide position in GVt is A; 1779 nt.

[0223] SEQ ID NO:95. Nucleotide sequence of a DNA fragment containing the T module of the TTm hybridization transcription terminator. Bsa The I restriction site is located at nucleotides 9-14 and 1155-1160, the EU Tm is located at nucleotides 20-499, the 35S Tm is located at positions 500-710, and the simplified Rb7 MAR sequence is located at positions 711-1149; 1168 nt.

[0224] SEQ ID NO:96. Amino acid sequence of the MTP-FAγ51::NifD(Y100Q)::HA fusion polypeptide. Amino acids 1-54 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, having a C-terminal GG; amino acids 55-482 are the NifD sequence based on the Klebsiella acidogenic sequence with Y100Q substitution, followed by GG; and amino acids 539-547 are HA epitopes; 547aa.

[0225] SEQ ID NO:97. Amino acid sequence of the MTP-FAγ51::HA::NifH fusion polypeptide. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by a GG at its C-terminus. Amino acids 52-64 are HA epitope tags including the C-terminal GG. Amino acids 65-353 are from *Geotrichum* (…). Geobacter)NifH sequence; 353aa.

[0226] SEQ ID NO:98. The amino acid sequence of the MTP-FAγ51::AvNifU::HA fusion polypeptide encoded by SN466, SN499, and SN500. Amino acids 1-54 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, with a C-terminal GG; amino acids 55-368 are the NifU sequence of *Venerellynidae* followed by GG; and amino acids 369-377 are the HA epitope; 377aa.

[0227] SEQ ID NO:99. Amino acid sequence of the MTP-FAγ51::GFP::HA fusion polypeptide encoded by SN493. Amino acids 1-54 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, with a C-terminal GG, amino acids 55-295 are the GFP sequence, followed by GG, and amino acids 296-304 are the HA epitope; 304aa.

[0228] SEQ ID NO:100. Amino acid sequence of the MTP-FAγ77::GFP fusion polypeptide encoded by pRA1. Amino acids 1-77 are MTP-FAγ77 cleaved by MPP between amino acids 42 and 43, with a C-terminal GAP, and amino acids 81-319 are the GFP sequence; 319aa.

[0229] SEQ ID NO:101. The amino acid sequence of the MTP-FAγ51::HA::NifH::linker (TS)::NifH) fusion polypeptide encoded by SN655, SN656, and SN657. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by a GG at its C-terminus; amino acids 52-64 are HA epitope tags including the C-terminal GG; amino acids 65-357 and 388-680 are two Klebsiella acidogenic NifH sequences; and amino acids 358-387 are oligopeptide linkers including the TwinStrep epitope; 680aa.

[0230] SEQ ID NO:102. The amino acid sequence of the MTP-SU9::KoNifM fusion polypeptide encoded by SN360. Amino acids 1-70 contain amino acids from Neurospora crassa (… Neurospora crassa The first 68 amino acid residues of the subunit 9 F0-ATPase of MTP-SU9 (Buren et al., 2017a) are cleaved by MPP between amino acids 66 and 67, have a C-terminal GG, and amino acids 71-336 are the NifM sequence of Klebsiella acidogenic bacteria; 336aa.

[0231] SEQ ID NO:103. The amino acid sequence of the MTP-FAγ51::KoNifH::linker::NifH::HA fusion polypeptide encoded by SN283. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, amino acids 54-345 and 371-663 are two Klebsiella acid-producing NifH sequences, amino acids 346-370 are a 25-amino acid oligopeptide linker, and amino acids 666-674 are HA epitope tags; 674aa.

[0232] SEQ ID NO:104. The amino acid sequence of the MTP-CoxIV::TS::KoNifH::linker::NifH fusion polypeptide encoded by SN638, SN649, and SN650. Amino acids 1-29 are from the MTP-CoxIV of Saccharomyces cerevisiae cytochrome c oxidase subunit 4 (Jiang et al., 2021) cleaved by MPP between amino acids 25 and 26; amino acids 30-61 are a TwinStrep epitope tag including a C-terminal GG; amino acids 62-353 and 379-671 are two Klebsiella acidogenic NifH sequences; and amino acids 354-378 are an oligopeptide linker; 671 is amino acid.

[0233] SEQ ID NO:105. The amino acid sequence of the MTP-FAγ51::HA::KoNifH::TS::NifH fusion polypeptide encoded by SN656. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by a GG at its C-terminus; amino acids 52-64 are HA epitope tags including the C-terminal GG; amino acids 65-357 and 388-680 are two Klebsiella acidogenic NifH sequences; and amino acids 358-387 are 30 amino acid oligopeptide linkers including the TwinStrep epitope; 680aa.

[0234] SEQ ID NO:106. The amino acid sequence of the MTP-FAγ51::KoNifD(Y100Q)::HA linker::KoNifK fusion polypeptide encoded by SN159. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by GG; amino acids 54-535 are the Klebsiella acidogenic NifD sequence; amino acids 537-1084 are the Klebsiella acidogenic NifK sequence without its initiating Met and without C-terminal extension; and amino acids 536-565 are the oligopeptide linker including the HA epitope; 1084aa.

[0235] SEQ ID NO:107. The amino acid sequence of the MTP-FAγ51::AvFdxN::HA fusion polypeptide encoded by SN291. Amino acids 1-52 are MTP-FAγ51 cleaved by MPP between amino acids 43 and 44, followed by GG; amino acids 55-146 are the FdxN sequence of *Venerenella vinifera*; and amino acids 147-157 are HA epitopes; 157aa.

[0236] SEQ ID NO:108. The amino acid sequence of the MTP-FAγ51::KoNifH::HA fusion polypeptide encoded by SN18. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by GG; amino acids 54-346 are the NifH sequence of Klebsiella acidogenic bacteria; and amino acids 347-357 are HA epitopes; 357aa.

[0237] SEQ ID NO:109. The amino acid sequence of the MTP-FAγ51::HA::AvNifH::linker::NifH fusion polypeptide encoded by each of SN678, SN679, and SN680. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by GG; amino acids 54-62 are HA epitopes, followed by GG; amino acids 65-352 and 378-665 are the NifH sequence of *Azotobacter venerealis* (accession number AOG20751); and amino acids 353-377 are oligopeptide linkers; 665aa.

[0238] SEQ ID NO:110. The amino acid sequence of the MTP-FAγ51::KoNifH::linker (MPP10)::NifH::HA fusion polypeptide encoded by SN285. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, amino acids 54-345 and 356-648 are two Klebsiella acidogenic NifH sequences, amino acids 346-355 are a 10-amino acid oligopeptide linker, and amino acids 651-659 are HA epitope tags; 659aa.

[0239] SEQ ID NO:111. The amino acid sequence of the MTP-FAγ51::AvNifM fusion polypeptide encoded by SN605. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by GG; amino acids 54-346 are the NifM sequence of *Venerendae*; 346aa.

[0240] SEQ ID NO:112. The amino acid sequence of the MTP-FAγ51::AvFdxN fusion polypeptide encoded by SN465. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by GG; amino acids 54-145 are the Vinerlandia vinifera FdxN sequence; 145aa.

[0241] SEQ ID NO:113. The amino acid sequence of the scar::TS::KoNifH::linker::NifH fusion polypeptide generated by SN638 and SN650. Amino acids 1-4 are scar sequences from MTP-CoxIV, amino acids 5-36 are TwinStrep epitope tags including a C-terminal GG, amino acids 37-328 and 354-646 are two Klebsiella acidogenic NifH sequences, and amino acids 339-353 are oligopeptide linkers; 646aa.

[0242] SEQ ID NO:114. The amino acid sequence of the MTP-FAγ51::AvAnfH-V7::linker (TS)::AvAnfH-V7 fusion polypeptide encoded by SN775, SN776, and SN777. Amino acids 1-51 are MTP-FAγ51 cleaved by MPP between amino acids 42 and 43, followed by GG at its C-terminus; amino acids 54-328 and 359-633 are two modified Venetian Azotobacter AnfH-V7 sequences; and amino acids 329-358 are a 30-amino acid oligopeptide linker including a TwinStrep epitope; 633aa.

[0243] SEQ ID NO:115. The amino acid sequence of the scar::AvAnfH-V7::linker (TS)::AvAnfH-V7 fusion polypeptide produced in plant mitochondria by SN775, SN776, and SN777. Amino acids 1-11 are the C-terminal scar sequence remaining after MPP cleavage of the MTP sequence, followed by GG; amino acids 12-286 and 317-591 are two modified AnfH-V7 sequences from *Azotobacter venereum*; and amino acids 287-316 are a 30-amino acid oligopeptide linker including a TwinStrep epitope; 591aa.

[0244] SEQ ID NO:116. The amino acid sequence of the MTP-FAγ51::AnfG polypeptide encoded by SN737. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, and amino acids 54-185 correspond to the wild-type AnfG sequence from *Venelandia diffusa*.

[0245] SEQ ID NO:117. The amino acid sequence of the MTP-FAγ51::NifB polypeptide encoded by SN598. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, and amino acids 54-341 correspond to those from thermoautotrophic methanotherapeutic bacteria (…). Methanothermobacter thermautotrophicus The wild-type NifB sequence corresponds to this.

[0246] SEQ ID NO:118. The amino acid sequence of the MTP-FAγ51::AnfH-V8 polypeptide encoded by SN714. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, and amino acids 54-328 correspond to the AnfH-V8 sequence having substituted amino acids T253A, T281V, and G294R, corresponding to the substituted amino acids T200A, T228V, and T241R of the AnfH-V8 sequence of *Azotobacter vinifera*.

[0247] SEQ ID NO:119. The amino acid sequence of the MTP-FAγ51::AnfH-V7 polypeptide encoded by SN754. Amino acids 1-53 correspond to the MTP-FAγ51 sequence including a GG linker at its C-terminus, and amino acids 54-328 correspond to the AnfH-V7 sequence having substituted amino acids T253A, T281V, and E287H, corresponding to the substituted amino acids T200A, T228V, and E234H of the AnfH-V7 sequence of *Azotobacter vinifera*.

[0248] SEQ ID NO:120. The amino acid sequence of the MTP-FAγ51::AvAnfD::linker (HA)::AvAnfK fusion polypeptide encoded by SN312, SN734, and SN735. Amino acids 1-53 correspond to MTP-FAγ51 with GG at its C-terminus, amino acids 54-570 correspond to the sequence of wild-type Azotobacter venereidesis AnfD (AvAnfD), amino acids 571-596 correspond to the linker including the HA epitope, and amino acids 597-1057 correspond to AnfK (SEQ ID NO:36) without its N-terminus Met and with its wild-type C-terminus.

[0249] SEQ ID NO:121. The amino acid sequence of the MTP-FAγ51::NifJ fusion polypeptide encoded by SL138 and SN748. Amino acids 1-54 correspond to MTP-FAγ51 with GG, and amino acids 55-1225 correspond to Klebsiella acidogenic NifJ (SEQ ID NO:7).

[0250] The amino acid sequence of the MTP-FAγ51::NifV fusion polypeptide of SEQ ID NO:122.SN532. Amino acids 1-54 correspond to the MTP-FAγ51 sequence with a GG linker, and amino acids 55-438 correspond to the NifV sequence from *Venerendae*.

[0251] SEQ ID NO:123. The amino acid sequence of the MTP-FAγ51::AvFdxN::HA fusion polypeptide encoded by SN736. Amino acids 1-54 are MTP-FAγ51, followed by GG; amino acids 55-148 are the Venetian azeotropic bacteria FdxN sequence without its initiating methionine, followed by GG; and amino acids 149-157 are HA epitopes; 157aa.

[0252] SEQ ID NO:124. The amino acid sequence of the MTP-FAγ51::NifF fusion polypeptide encoded by SL138 and SN747. Amino acids 1-54 correspond to MTP-FAγ51 with GG, and amino acids 55-230 correspond to Klebsiella acidogenic NifF (SEQ ID NO:6).

[0253] SEQ ID NO:125. The amino acid sequence of the MTP-FAγ51::AnfD fusion polypeptide encoded by SN733. Amino acids 1-54 correspond to MTP-FAγ51 with GG, and amino acids 55-572 correspond to Azotobacter venereum AnfD (SEQ ID NO:35).

[0254] SEQ ID NO:126. The amino acid sequence of the MTP-FAγ51::AnfK fusion polypeptide encoded by SN528. Amino acids 1-54 correspond to MTP-FAγ51 with GG, and amino acids 55-572 correspond to AnfK of Azotobacter venerealis (SEQ ID NO:36).

[0255] SEQ ID NO:127-150. Peptide sequence.

[0256] SEQ ID NO:151. The amino acid sequence of the last four amino acid residues at the C-terminus of the NifK polypeptide from Klebsiella acidogenic bacteria.

[0257] The amino acid sequence of the MTP-FAγ51::HA::FdxN fusion polypeptide of SEQ ID NO:152.SL79; 156aa. Amino acids 1-53 correspond to the MTP-FAγ51 sequence with a GG linker, amino acids 54-64 correspond to the HA epitope with a GG linker, and amino acids 65-156 correspond to the FdxN sequence without an N-terminal methionine.

[0258] SEQ ID NO:153. Amino acid sequence of the MTP-FAγ51::HA::NifV fusion polypeptide of SL48 and SL79; 448aa. Amino acids 1-53 correspond to the MTP-FAγ51 sequence with a GG linker, amino acids 54-64 correspond to the HA epitope with a GG linker, and amino acids 65-448 correspond to the NifV sequence from *Venerendae*.

[0259] SEQ ID NO:154. The amino acid sequence of the MTP-FAγ51::NifS::HA fusion polypeptide encoded by SL78. Amino acids 1-54 correspond to MTP-FAγ51 having GG at its C-terminus, amino acids 55-454 correspond to the acid-producing Klebsiella NifS (SEQ ID NO:19) having its initiator Met, according to Temme et al., (2012); and amino acids 455-465 include the HA epitope.

[0260] SEQ ID NO:155. The amino acid sequence of the MTP-FAγ51::NifU::HA fusion polypeptide encoded by SL78. Amino acids 1-54 correspond to MTP-FAγ51 having GG at its C-terminus, amino acids 55-328 correspond to the acid-producing Klebsiella NifU (SEQ ID NO:12) having its initiator Met; and amino acids 329-339 include the HA epitope.

[0261] SEQ ID NO:156. The amino acid sequence of the MTP-FAγ51::NifF::HA fusion polypeptide encoded by SL78. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-230 correspond to Klebsiella acidogenic NifF (SEQ ID NO:6); and amino acids 231-241 include the HA epitope.

[0262] SEQ ID NO:157. The amino acid sequence of the MTP-FAγ51::NifJ::HA fusion polypeptide encoded by SL78. Amino acids 1-54 correspond to MTP-FAγ51 with GG, amino acids 55-1225 correspond to Klebsiella acidogenic NifJ (SEQ ID NO:7); and amino acids 1226-1236 include the HA epitope. Detailed Implementation

[0263] General techniques and definitions Unless otherwise specified, all technical and scientific terms used herein should be regarded as having the same meaning as commonly understood by a person skilled in the art (e.g., in molecular genetics, plant molecular biology, nitrogen fixation, protein chemistry, and biochemistry).

[0264] Unless otherwise indicated, the recombinant proteins, cell culture, and immunological techniques used in this invention are standard procedures well known to those skilled in the art. Such techniques are described and explained in the following sources: J. Perbal, *A Practical Guide to Molecular Cloning*, John Wiley and Sons (1984); J. Sambrook et al., *Molecular Cloning: A Laboratory Manual*, Cold Spring Harbor Laboratory Press (1989); TA Brown (ed.), *Essential Molecular Biology: A Practical Approach*, Volumes 1 & 2, IRL Press (1991); DM Glover and BD Hames (ed.), *DNA Cloning: A Practical Approach*, Volumes 1–4, IRL Press (1995 and 1996); and FMAusubel et al. (ed.), *Current Protocols in Molecular Biology*. Molecular Biology, Greene Pub. Associates and Wiley-Interscience (1988, including all updates to date); Ed Harlow and David Lane (eds.), Antibodies: A Laboratory Manual, Cold Spring Harbour Laboratory (1988); and JE Coligan et al. (eds.), Current Protocols in Immunology, John Wiley & Sons (including all updates to date).

[0265] The term “and / or”, such as “X and / or Y”, should be understood to mean “X and Y” or “X or Y”, and should be regarded as providing explicit support for both meanings or either meaning.

[0266] As used herein, unless otherwise stated, the term approximately means + / - 10% of the specified value, or more preferably + / - 5%, or even more preferably + / - 1%.

[0267] Throughout this specification, the word “comprise” or variations such as “comprises” or “comprising” will be understood to imply inclusion of the stated elements, integers or steps, or groups of elements, groups of integers or groups of steps, but does not exclude any other elements, integers or steps, or groups of elements, groups of integers or groups of steps.

[0268] Nitrogenase Nitrogenase is an enzyme found in eubacteria and archaea that catalyzes the reduction of the strong triple bond of nitrogen (N2) to produce ammonia (NH3). Nitrogenase exists only naturally in bacteria. It is a complex of two separately purified enzymes: dinitrogenase and dinitrogenase reductase. Dinitrogenase, also known as component I or the molybdenum-iron (MoFe) protein, is a tetramer of two NifD and two NifK polypeptides (α2β2) containing two "P clusters" and two "FeMo-co" cofactors (FeMo-co). Each pair of NifD-NifK subunits contains one P cluster and one FeMo-co. FeMo-co is a cluster of metal atoms composed of a MoFe3-S3 cluster complexed with a homocitrate molecule coordinated to a molybdenum atom and bridged to the Fe4-S3 cluster by three sulfur ligands. FeMo-co is assembled separately in the cell and then incorporated into the apo-MoFe protein. The P cluster is also a metal cluster containing 8 Fe atoms and 7 sulfur atoms, with a structure similar to but different from FeMo-co. The P cluster is located at the αβ subunit interface of dinitrogenase and is coordinated by cysteine ​​residues from both subunits. Dinitrogenase reductase, also known as component II or "Fe protein," is a dimer of the NifH polypeptide, which also contains a single Fe4-S4 cluster and two Mg-ATP binding sites at the subunit interface, one for each subunit. This enzyme is the forced electron donor for dinitrogenase, where electrons are transferred from the Fe4-S4 cluster to the P cluster and then to the N2 reduction site FeMo-co.

[0269] Although nitrogenases containing Mo are the most common nitrogenases in bacteria, there are two genetically different homologous nitrogenases with similar cofactor and subunit compositions, namely, those composed of Mo and Mo respectively. Vnf (vanadium nitrogen fixation) and AnfThe (alternating nitrogen fixation) gene encodes vanadium-containing nitrogenases and Fe-containing nitrogenases only. Some bacteria in nature possess all three types of nitrogenases, while others contain only Mo- and V-containing enzymes or only Mo-containing enzymes, such as Klebsiella pneumoniae. Klebsiella pneumoniae ).

[0270] The biosynthesis of FeMo-co and the maturation of nitrogenase components into catalytically active forms require multiple nitrogen fixation (Nif) genes. The roles of NifB, NifE, NifH, NifN, NifQ, NifV, and NifX peptides in FeMo-co synthesis have been described (Rubio and Ludden, 2008).

[0271] Biological N2 fixation catalyzed by prokaryotic nitrogenase is an alternative to using synthetic N2 fertilizers. The sensitivity of nitrogenase to oxygen is a key feature of engineered biological nitrogen fixation through direct... Nif The main barrier to gene transfer into plants (e.g., into cereal crops).

[0272] Thermal stability and its relationship with solubility Many proteins misfold and aggregate when removed from their native environment (Goldenzweig et al., 2016). Protein aggregation is more likely to occur when nascent peptides misfold and fail to form a structurally stable conformational state (Wang and Roberts, 2018). It has also been proposed that natively occurring proteins are only slightly stable because there is no greater energy stability selection pressure within their cellular environment (Magliery, 2015). Therefore, the compatibility of heterologous protein production and stability in the host is difficult to predict and is often influenced by the intrinsic energy stability of the protein. It has been reported that when the respiratory chain is fully functional, the internal temperature of mitochondria operates at a higher temperature than the rest of the organism (Chretien et al., 2018), at least in endothermic organisms. If this is also the case in the plant mitochondrial matrix, the relatively high local operating temperature may promote protein aggregation and thereby reduce the solubility of proteins such as wild-type AnfH, which are exogenous to plant cells.

[0273] Proteins typically exist in cells in a dynamic range of conformations. Each conformation has an associated Gibbs energy that can be theoretically calculated from several parameters, including intermolecular interactions, covalent bond angles, and protein-solvent interactions (Onuchic et al., 1997). The native fold of a protein is usually located in a local energy minimum or energy 'pore'. This native fold is thought to be stabilized by hydrophobic collapse and intermolecular bonds. Hydrophobic collapse removes the hydrophobic residues of the protein from the solvent because disrupting the hydrogen-bonded network of water is energy-disadvantageous (Sadqi et al., 2003). Intermolecular interactions involve favorable intermolecular bonds between the protein and the solvent, as well as within the protein itself (Jaenicke. 2000). For example, α-helices and β-sheets occur when the main chain amino hydrogen and carboxyl oxygen atoms form a hydrogen-bonded network. The lower the energy of the native fold of a protein, the more energy is required to change its folded state or 'unfold' the protein (Liu et al., 2000). Some proteins exhibit higher thermal stability at ambient temperatures, enabling them to maintain their structure and function at high temperatures (Modarres et al., 2016). However, even with current understanding of protein structure and folding, the effects of amino acid substitutions on protein stability and function are difficult to predict (Wilding et al., 2019). Despite this uncertainty, the inventors used protein engineering methods to reduce the total energy of AnfH in order to improve its stability and solubility.

[0274] Analysis of thermophilic proteins has shown that they contain more hydrophobic and charged amino acids compared to less thermophilic proteins (Modarres et al., 2016). These characteristics suggest that thermophilic proteins have better hydrophobic stacking and more solvent or salt bridge interactions. Shared design has also been used to increase protein stability (Porebski and Buckle, 2016). In this approach, homologous sequences are aligned, and the most common amino acid at each position is adopted into a novel protein sequence for expression. It is assumed that natural selection more frequently removes unstable amino acids, thus stabilizing the more common amino acids at the same positions across homologous sequences (Georgoulis et al., 2020).

[0275] In one embodiment, the polypeptide of the present invention (e.g., an MTP fusion polypeptide or a cleaved product thereof) is at least partially soluble in the mitochondria of plant cells. In this context, the phrase "at least partially soluble" means that the polypeptide is detectable in the soluble fraction of a homogenate sample containing the mitochondria of plant cells. Suitable methods for detecting the solubility of the polypeptide are known in the art and include those described in Example 1. In one embodiment, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the polypeptide present in the mitochondria of cells is soluble.

[0276] Nif peptides As used herein, the terms "Nif polypeptide" and "Nif protein" are used interchangeably and refer to polypeptides whose amino acid sequences relate to naturally occurring polypeptides involved in nitrogenase activity, wherein the Nif polypeptides of the present invention are selected from the group consisting of: NifD polypeptide, NifH polypeptide, NifK polypeptide, NifB polypeptide, NifE polypeptide, NifN polypeptide, NifF polypeptide, NifJ polypeptide, NifM polypeptide, NifQ polypeptide, NifS polypeptide, NifU polypeptide, NifV polypeptide, NifW polypeptide, NifX polypeptide, NifY polypeptide, and NifZ polypeptide, each as defined herein. Specifically, the present invention relates to modified NifH polypeptides and / or NifH fusion polypeptides, preferably modified AnfH polypeptides and / or AnfH fusion polypeptides. As used herein, the Nif polypeptides of the present invention include “Nif fusion polypeptides,” which are polypeptide homologs of naturally occurring Nif polypeptides that, relative to the corresponding naturally occurring Nif polypeptide, have additional amino acid residues attached to an N-terminus or a C-terminus or both, and / or have at least one amino acid substitution compared to the corresponding wild-type polypeptide. In this context, “corresponding wild-type polypeptide” means a wild-type polypeptide naturally occurring in bacteria that is, for all naturally occurring sequences known up to the date of this application, most closely related in amino acid sequence to the Nif polypeptide of the present invention in terms of sequence identity. The Nif fusion polypeptide may lack a translation initiation Met or two N-terminal Met residues relative to the corresponding wild-type Nif polypeptide. The amino acid residues of the Nif fusion polypeptide corresponding to the naturally occurring Nif polypeptide, i.e., lacking additional amino acid residues attached to an N-terminus or a C-terminus or both, are also referred to herein as Nif polypeptides, abbreviated as “NP” in this case, or as NifH polypeptides (“NH”), etc. In a preferred embodiment, the “additional amino acid residues linked to the N-terminus or C-terminus or both” comprises a mitochondrial targeting peptide (MTP) or the amino acid remaining after protease cleavage (processed MTP, or “scar sequence”) of an MTP linked to the N-terminus or C-terminus or both of the NP, or an epitope sequence (“tag”) of the N-terminus or C-terminus or both of the NP, or an MTP or both a processed MTP and an epitope sequence.

[0277] Naturally occurring Nif polypeptides are found only in certain bacteria, including nitrogen-fixing bacteria, specifically free-living, associative, and symbiotic nitrogen-fixing bacteria. Free-living nitrogen-fixing bacteria can fix significant levels of nitrogen without direct interaction with other organisms. Without limitation, the free-living nitrogen-fixing bacteria include the genus *Azobacteria* (…). Azotobacter ), spp. of Bayerlink ( Beijerinckia Klebsiella spp., Cyanobacteria spp. Cyanobacteria(Classified as aerobic organisms) members and Clostridium ( Clostridium ), Desulfuric Vibrio spp. Desulfovibrio (And members named purple sulfur bacteria, purple non-sulfur bacteria, and green sulfur bacteria.) Associating nitrogen-fixing bacteria are those prokaryotes that can form close associations with several members of the Poaceae (weeds). These bacteria fix a considerable amount of nitrogen in the rhizosphere of the host plant. *Azospirillum* genus ( Azospirillum Members of the symbiotic nitrogen-fixing bacteria (Rhizobium) are representatives of symbiotic nitrogen-fixing bacteria. Symbiotic nitrogen-fixing bacteria are those bacteria that fix nitrogen through symbiotic cooperation with a host plant. The plant provides sugars through photosynthesis, and the nitrogen-fixing bacteria use these sugars to provide the energy needed for nitrogen fixation. Rhizobia The members of ) are representatives of associative nitrogen-fixing bacteria.

[0278] The Nif peptides or Nif fusion peptides used in this invention include peptides selected from the group consisting of: NifH, NifD, NifK, NifB, NifE, NifN, NifF, NifJ, NifM, NifQ, NifS, NifU, NifV, NifW, NifX, NifY, and NifZ peptides. The functions of these peptides have recently been reviewed by Burén et al. (2020).

[0279] Other peptides used in this invention are considered to be VnfG and AnfG, respectively, involving V-nitrogenase and Fe-nitrogenase, nitrogenase-associated factors (Naf peptides), such as NafY, and ferroredoxin peptides such as FdxN peptide. These peptides preferably encode and are expressed as MTP fully fusion peptides for mitochondrial targeting.

[0280] A polypeptide or polypeptide class can be defined by the degree of identity (identity %) of its amino acid sequence with a reference amino acid sequence and / or by the presence of certain amino acid motifs or protein family domains, or by having a higher percentage of identity with one reference amino acid sequence compared to another. In addition to sequence identity, a polypeptide or polypeptide class can also be defined by having the same biological activity as naturally occurring Nif polypeptides.

[0281] The identity percentage of the polypeptide was determined by GAP (Needleman and Wunsch, 1970) analysis (GCG procedure), where the void production penalty = 5 and the void extension penalty = 0.3, or by Blastp version 2.5 or a later version thereof (Altschul et al., 1997), wherein in each case, the analysis compared two sequences of the reference sequence, including the entire length of the reference sequence. As used herein, the reference sequence includes the reference sequence provided for naturally occurring Nif polypeptides from Klebsiella pneumoniae (renamed Klebsiella acidogenic), SEQ ID NO:1-17, or the reference sequence provided for AnfH from Azotobacter venerealis, SEQ ID NO:37.

[0282] In the following definitions, the degree of identity between the amino acid sequence and the reference sequence as provided in SEQ ID NO is determined by Blastp version 2.5 or a later version (Altschul et al., 1997) using default parameters (except that the maximum number of target sequences is set to 10,000) and along the full length of the reference amino acid sequence.

[0283] As used herein, the phrase "one amino acid substitution" or "amino acid substitution" refers to replacing an amino acid in a wild-type NifH polypeptide with a single different amino acid. For example, in one embodiment, threonine at position 200 of a NifH polypeptide having the amino acid sequence provided in SEQ ID NO:37 is replaced with alanine. As used herein, amino acid substitution does not include the insertion or deletion of the amino acid. Therefore, a comparison of a polypeptide with amino acid substitutions or multiple amino acid substitutions to a corresponding unmodified polypeptide will not show any vacancies due to insertions or deletions, i.e., the lengths will be the same. As used herein, the "number of amino acid substitutions" in the NifH polypeptide of the present invention is counted along the full length of the comparison of the amino acid sequence of the polypeptide to the amino acid sequence of the corresponding wild-type NifH polypeptide.

[0284] As used herein, the phrase “reference SEQ ID NO” refers to the same amino acid position as the defined polypeptide.

[0285] As used herein, the phrase "located at the corresponding amino acid position when the sequence of the modified NifH polypeptide is aligned with SEQ ID NO:37 (or SEQ ID NO:39)" or variations thereof refers to the relative position of the amino acid to its surrounding amino acids. More specifically, not all wild-type NifH polypeptides have the same length (number of amino acid residues), although any two wild-type NifH polypeptides may have the same length, and therefore this example relates to the case where a particular wild-type NifH polypeptide has or does not have the same length as SEQ ID NO:37. However, those skilled in the art can determine the relevant corresponding amino acid position using protein standard alignment procedures, or even visually.

[0286] The naturally occurring bacterial NifH polypeptide is a structural component of the nitrogenase complex and is commonly referred to as the iron (Fe) protein. It forms a homodimer in which the Fe4S4 cluster binds between the subunit and two ATP-binding domains. NifH is a dedicated electron donor for the nitrogenase protein (NifD / NifK heterotetramer) and thus acts as the nitrogenase reductase of molybdenum-type nitrogenase (EC 1.18.6.1). AnfH polypeptides are considered a subclass of NifH polypeptides in this study, and AnfH is a dedicated electron donor for the AnfD-AnfH-AnfG nitrogenase protein (AnfDKG, iron-only nitrogenase), although it can also function in conjunction with the NifD / NifK heterotetramer. Molybdenum-type NifH is also involved in FeMo-co biosynthesis and apo-MoFe protein maturation (Jasniewski et al., 2018), and AnfH is considered to have corresponding functions in FeFe-co biosynthesis and apo-FeFe protein maturation. As reviewed by Jasniewski et al. (2018), NifH has three main recognized functions: (i) involvement in the insertion of Mo and high citrate in the synthesis of FeMo-co, as well as the NifE-NifN complex; (ii) reductase function in the formation of P clusters from so-called P* clusters on NifD-NifK, which may also involve the small chaperone-like peptide NifZ; and (iii) electron donor for nitrogenase proteins.

[0287] As used herein, “NifH polypeptide” means a polypeptide comprising amino acids whose sequence is at least 41% identical to the amino acid sequence provided in SEQ ID NO:1, and the polypeptide comprising one or more of the domains TIGR01287, PRK13236, PRK13233, and cd02040. The TIGR01287 domain is present in each of molybdenum-iron nitrogenase reductase (NifH), vanadium-iron nitrogenase reductase (VnfH), and iron-iron nitrogenase reductase (AnfH), but does not include light-independent homologs from light-independent protochlorophyll esters. Therefore, as used herein, NifH polypeptides comprise a subclass of iron-binding polypeptides comprising amino acids whose sequence is at least 41% identical to the sequence of NifH (SEQ ID NO:1), VnfH iron-binding polypeptide, and AnfH iron-binding polypeptide from Klebsiella pneumoniae. Naturally occurring NifH peptides are typically 260 to 300 amino acids in length, and the molecular weight of the natural monomers is approximately 30 kDa. Numerous NifH peptides have been identified, and many sequences are available in publicly available databases. For example, NifH peptides from *Klebsiella micrantha* have been reported. Klebsiella michiganensis (Accession number WP_049123239.1, 99% identical to SEQ ID NO:1), *Bukhnerella guillier* ( Brenneria goodwinii (WP_048638817.1, 93% identical), stony nutritional iron-oxidizing bacteria ( Sideroxydans lithotrophicus (WP_013029017.1, 84% identical), Acetophilic nitrate-reducing Vibrio ( Denitrovibrio acetiphilus (WP_013010353.1, 80% identical), African desulfurization vibrio ( Desulfovibrio africanu s)(WP_014258951.1, 72% identical), brown rod-shaped green bacteria ( Chlorobium phaeobacteroides (WP_011744626.1, 69% identical), and Methanogens fasciatus ( Methanosaeta concilii (WP_013718497.1, 64% identical), Rhodopseudomonas ( Rhodobacter (WP_009565928.1, 61% identical), *Methemoglobinococcus hellenophilus* (WP_013099472.1, 42% identical) and *Campylobacter juvenile desulfurized* ( Desulfosporosinus youngiae(WP_007781874.1, 41% identical). Of particular importance and frequently used in this paper as the corresponding wild-type NifH sequence is NifH from *Azotobacter venereum* (SEQ ID NO:39; 290 amino acids), which is 89% identical to SEQ ID NO:1 (293 amino acids). NifH peptides have been described and reviewed in the following journals: Thiel et al. (1997), Pratte et al. (2006), Boison et al. (2006), and Staples et al. (2007).

[0288] As used herein, a functional NifH peptide is a NifH peptide capable of forming a functional nitrogenase protein complex with other desired subunits (e.g., NifD and NifK), or, in the case of an AnfH peptide, with AnfDKG and FeMo-cofactor, FeV-cofactor, or FeFe-cofactor. In this context, the functional nitrogenase complex is capable of reducing N2 gas to ammonia and / or reducing acetylene to ethylene, which can be determined in vitro in reactions as described herein.

[0289] As used herein, “AnfH polypeptide” is a NifH polypeptide that is a member of the conserved nitrogenase superfamily cl25403 (TIGR01287) containing the PRK13233 conserved domain and having at least 69% amino acid sequence identity with the Venetian violaceum AnfH polypeptide (SEQ ID NO:37; accession number WP_012703362) when measured along the full length of SEQ ID NO:37. This amino acid sequence is used herein as a reference sequence for AnfH. TIGR01287:AnfH represents the all-iron variant of nitrogenase component II, also known as nitrogenase reductase. As used herein, AnfH polypeptide is a subset of NifH polypeptides. AnfH polypeptides do not include molybdenum-type NifH polypeptides and vanadium-type NifH polypeptides (VnfH). The amino acid sequence of AnfH polypeptide in sequence databases is generally annotated as AnfH polypeptide. As of January 2020, 314 specific amino acid sequences exist in the NCBI protein database within the AnfH set. All of these sequences contain amino acid residues specific to AnfH and, unlike molybdenum-type NifH and VnfH, this subset appears more similar but is still distinct. Examples of naturally occurring AnfH peptides include those from *Rhodotorula slenderis* (…). Rhodocyclus tenuis (Accession ID WP_153472986; 92.36% identical), Decardia banana ( Dickeya paradisiaca (Accession number WP_015854293; 88.36% same), autotrophic thermophilic sulfite-reducing bacteria ( Thermodesulfitimonas autotrophica(Accession number WP_123927773; 78.91% identical), Clostridium kraniliforme ( Clostridium kluyveri (Accession number WP_073538802; 76.36% identical) and Methanogenic Archaea ( Methanophagales archaeon (Registration number RCV64832; 69.37% identical), each refer to SEQ ID NO:37.

[0290] As described in Example 2 of this document, 16 amino acids were identified at the defined position in SEQ ID NO:37 or at corresponding positions in other AnfH sequences. These amino acids are conserved relative to molybdenum-type NifH sequences and are characteristic of AnfH polypeptides. These 16 amino acid positions can be used to distinguish AnfH polypeptides from other NifH sequences that do not have all 16 common amino acids. AvNifH (SEQ ID NO:39), KoNifH (SEQ ID NO:1), and other molybdenum-type NifH sequences have motif IV, but lack motifs I, II, III, and V-VII due to substitutions of one or more amino acids in each of these sequences, and therefore these motifs (SEQ ID NO:40-46) can also be used to distinguish AnfH subsets from other NifH polypeptides.

[0291] Similar to other functional NifH peptides, functional AnfH peptides can function as nitrogenase reductases and are specific electron donors for the FeFe complex (AnfDKG). Similar to molybdenum-type NifH, AnfH is believed to participate in the biosynthesis of FeFe-co and the maturation of the apo-FeFe complex (AnfDKG).

[0292] As used herein, “NifD polypeptide” means a polypeptide comprising amino acids whose sequence is at least 33% identical to the amino acid sequence provided in SEQ ID NO:2, and the polypeptide comprises (i) one or both of the domains TIGR01282 and COG2710, both of which are present in an iron-molybdenum binding polypeptide comprising a polypeptide having the amino acid sequence shown in SEQ ID NO:2, or (ii) the iron-vanadium binding domain TIGR01860, in which case the NifD polypeptide belongs to the subclass of VnfD polypeptides, or (iii) the iron-iron binding domain TIGR1861, in which case the NifD polypeptide belongs to the subclass of AnfD polypeptides. The NifD polypeptide may be part of a fusion polypeptide, for example, fused with MTP and / or NifK, or alternatively may not contain any N-terminal or C-terminal extensions. In a preferred embodiment, the NifD polypeptide binds to the FeMo cofactor upon association with a NifK polypeptide.

[0293] As used herein, NifD peptides include subclasses of iron-molybdenum (FeMo-co)-binding peptides, VnfD iron-vanadium peptides, and AnfD peptides containing amino acids whose sequences are at least 33% identical to SEQ ID NO:2. Naturally occurring NifD peptides are typically 470 to 540 amino acids in length. Numerous NifD peptides have been identified, and many sequences are available in publicly available databases. For example, NifD peptides from *Rauvolobacterium ornithine* have been reported. Raoultella ornithinolytica (Accession number WP_044347161.1, 96% identical to SEQ ID NO:2), Kluyveromyces intermedia ( Kluyvera intermedia ), (WP_047370273.1, 93% identical), Datnidadica ( Dickeya dadantii ), (WP_038902190.1, 89% identical), Toxocobalamins ( Tolumonas sp.) BRL6-1 (WP_024872642.1, 81% identical), Greifswald magnetotactic spirochetes ( Magnetospirillum gryphiswaldense ), (WP_024078601.1, 68% identical), pyrolytic sugar heat anaerobic bacteria ( Thermoanaerobacterium thermosaccharolyticum ), (WP_013298320.1, 42% identical), thermoautotrophic methanotherapeutic bacillus (WP_010877172.1, 38% identical), African desulfurized vibrio (WP_014258953.1, 37% identical), desulfurized enterobacteria ( Desulfotomaculum sp.) LMa1 (WP_066665786.1, 37% identical), rod-shaped desulfurizing microorganisms ( Desulfomicrobium baculatum (WP_015773055.1, 36% identical), from moss resident Fishery bacteria ( Fischerella muscicola (WP_016867598.1, 34% identical) VnfD polypeptide, and bacteria from the family Aestheticae ( Opitutaceae bacterium AnfD peptides of TAV5 (WP_009512873.1, 33% identical). AnfD peptides have been described and reviewed in the following: Lawson and Smith (2002), Kim and Rees (1994), Eady (1996), Robson et al. (1989), Dilworth et al. (1988), Dilworth et al. (1993), Miller and Eady (1988), Chiu et al. (2001), Mayer et al. (1999), and Tezcan et al. (2005).

[0294] The iron-molybdenum subclass of NifD peptides is a key subunit of the nitrogenase complex, the α subunit of the α2β2MoFe protein complex at the core of nitrogenase, and the site of substrate reduction by the FeMo cofactor. As used herein, a functional NifD peptide is a NifD peptide capable of forming a functional nitrogenase protein complex with other desired subunits (e.g., NifH and NifK) and FeMo or other cofactors.

[0295] As used herein, a “protease-resistant NifD polypeptide (ND)” is resistant to cleavage at a defined site or region (e.g., within the amino acid sequence corresponding to amino acids 97-100 of SEQ ID NO:18) when introduced into plant mitochondria using an MTP. As used herein, “protease-resistant” means that <10% cleavage occurs when the NifD polypeptide is introduced into plant mitochondria using an MTP. In preferred embodiments, less than 5% of the NifD polypeptide is cleaved at the site or region, more preferably substantially uncleaved or undetected. Compared to a NifD polypeptide containing the amino acid sequence provided in SEQ ID NO:18, the NifD polypeptide may be “relatively resistant to cleavage”, typically cleaved at a frequency at most 1 / 5, preferably at most 1 / 10, that of a NifD polypeptide containing the amino acid sequence provided in SEQ ID NO:18.

[0296] As used herein, "the amino acid sequence other than RRNY (SEQ ID NO:150) located at the position corresponding to amino acids 97-100 of SEQ ID NO:18" refers to a sequence containing four residues located at the position corresponding to amino acids 97-100 of SEQ ID NO:18 and not RRNY.

[0297] As used herein, “AnfD polypeptide” is a NifD polypeptide that specifically contains the conserved TIGR01861 domain and, when measured along the full length of SEQ ID NO:35, is a member of the conserved superfamily cl30843 of nitrogenase, an oxidoreductase, and has at least 71% amino acid sequence identity with the AnfD polypeptide of *Azotobacter venereum* (SEQ ID NO:35; accession number WP_012703361). This amino acid sequence is used herein as a reference sequence for AnfD. TIGR01861:AnfD represents an all-ferric variant of the nitrogenase component Iα chain. As used herein, AnfD polypeptide is therefore a subset of NifD polypeptides. AnfD polypeptides do not include molybdenum-type NifD polypeptides and vanadium-type NifD polypeptides (VnfD), nor do they include prochlorophyllin or chlorophyllin reductase polypeptides (Boyd and Peters, 2013). The amino acid sequence of AnfD polypeptides is commonly annotated as AnfD polypeptide in protein sequence databases. As of January 2020, 156 specific amino acid sequences exist in the NCBI protein database within the AnfD collection. Examples of naturally occurring AnfD peptides include those from *Desulfovibrio* (…). Desulfovibrio sp.) DV (accession number WP_075356167; 87.47% identical), Bacillus subtilis ( Paenibacillus sp.) FSL H7-0357 (accession number WP_038590013; 85.52% identical), capsular red bacteria ( Rhodobacter capsulatus (Accession number WP_023922817; 80.31% identical), Methanococcus acetate ( Methanosarcina acetivorans C2A (accession number WP_011021232; 77.13% identical) and Bacteroidetes ( Bacteroidal bacteria Barb7 (accession number OAV73823; 71.25% identical), each refer to SEQ ID NO:35. McRose et al. (2017) reported another instance.

[0298] Similar to other functional NifD peptides, functional AnfD peptides can function as an α-protein structural component of α2β2δ2 heterohexameric nitrogenase, working together with β-protein (AnfK) and δ-protein (AnfG) to provide a FeFe-co-binding catalytic complex for molecular nitrogen reduction.

[0299] As used herein, “NifK polypeptide” means a polypeptide containing amino acids whose sequence is at least 31% identical to the amino acid sequence provided as SEQ ID NO:3, and which contains one or more conserved domains of cd01974, TIGR01286, or cd01973, in which case the NifK polypeptide belongs to the subclass of VnfK polypeptides, or cl02775 containing the conserved domain of TIGR02931, in which case the NifK polypeptide belongs to the subclass of AnfK polypeptides. As used herein, NifK polypeptides include VnfK polypeptides derived from iron-vanadium nitrogenase and AnfK iron-binding polypeptides. Naturally occurring NifK polypeptides are typically 430 to 530 amino acids in length. Numerous NifK polypeptides have been identified, and many sequences are available in publicly available databases. For example, NifK peptides have been reported from the following: *Klebsiella micranthae* (accession number WP_049080161.1, 99% identical to SEQ ID NO:3), *Rauvolfia ornithine-lysinophila* (WP_044347163.1, 96% identical), and *Klebsiella variegata* (…). Klebsiella variicola (SBM87811.1, 94% identical), Kluyveromyces intermedia (WP_047370272.1, 89% identical), Laenia aquaticus ( Rahnella aquatica (WP_014333919.1, 82% identical), Toxocobalamin aurea ( Tolumonas auensis (WP_012728880.1, 75% identical), *Pseudomonas stearothermiae* ( Pseudomonas stutzeri (WP_011912506.1, 68% identical), sodium-dependent Vibrio ( Vibrio natriuretic (WP_065303473.1, 65% identical), Toluene destroys nitrogen-fixing bacteria ( Azoarcus toluclasticus (WP_018989051.1, 54% identical), Franklinella ( Frankishsp. (prf||2106319A, 50% identical) and methanogenic taenia (WP_011021239.1, 31% identical). Several instances of peptides exist in the database labeled “NifK”, which share less than 31% identity with SEQ ID NO:3 but do not contain any of the domains listed above and are therefore not included here as NifK peptides. NifK peptides have been described and reviewed in the following: Kim and Rees (1994), Eady (1996), Robson et al. (1989), Dilworth et al. (1988), Dilworth et al. (1993), Miller and Eady (1988), Igarashi and Seefeldt (2013), Fani et al. (2000), and Rubio and Ludden (2008).

[0300] The iron-molybdenum subclass of NifK peptides is a key subunit of the nitrogenase complex, being the β subunit of the α2β2MoFe protein complex at the core of nitrogenase. As used herein, a functional NifK peptide is one capable of forming a functional nitrogenase protein complex with other desired subunits (e.g., NifD and NifH) and FeMo or other cofactors. In a preferred embodiment, the amino acid sequence of the NifK peptide of the present invention, when compared with the amino acid sequence SEQ ID NO:3, has the amino acid DLVR (SEQ ID NO:151) at its C-terminus, with arginine being the C-terminal amino acid. That is, the NifK peptide and NifK fusion peptide of the present invention preferably have the same C-terminus as the natural NifK peptide, i.e., without any artificial additions to the C-terminus. Such preferred NifK peptides are better able to form functional nitrogenase complexes with NifD and NifH peptides.

[0301] The iron-molybdenum subclass of NifK peptides is a key subunit of the nitrogenase complex, being the β subunit of the α2β2MoFe protein complex at the core of nitrogenase. As used herein, a functional NifK peptide is one capable of forming a functional nitrogenase protein complex with other desired subunits (e.g., NifD and NifH) and FeMo or other cofactors. In a preferred embodiment, when compared with the amino acid sequence SEQ ID NO:3, the amino acid sequences of the NifK fusion peptide and the cleaved NifK peptide of the present invention have the amino acid DLVR (SEQ ID NO:151) at their C-terminus, where arginine is the C-terminal amino acid. In other preferred embodiments, the amino acid sequences of the NifK fusion peptide and the cleaved NifK peptide of the present invention have the amino acid sequences DLIR (SEQ ID NO:49), DVVR (SEQ ID NO:50), DIIR (SEQ ID NO:51), DLTR (SEQ ID NO:52), or INVW (SEQ ID NO:53) at their C-terminus, which are typically not present in the native AnfK sequence. The NifK peptides and NifK fusion peptides of the present invention, as well as the NifK peptides cleaved therefrom, preferably have the same C-terminus as the natural NifK peptides, i.e., they do not have any artificial additives to the C-terminus, and when compared with the natural NifK peptides, they do not have any amino acids missing from the C-terminus. Such preferred NifK peptides are better able to form functional nitrogenase complexes with NifD and NifH peptides.

[0302] As used herein, “AnfK polypeptide” is a polypeptide containing the conserved domain TIGR02931 and having at least 54% amino acid sequence identity with the Venetian Azotobacter vinifera AnfK polypeptide (SEQ ID NO:36; accession number WP_012703359) when measured along the full length of SEQ ID NO:36. This amino acid sequence is used herein as a reference sequence for AnfK. TIGR02931:AnfK represents an all-ferric variant of the nitrogenase component I β chain. As used herein, AnfK polypeptides may be NifK polypeptides having at least 31% amino acid identity with SEQ ID NO:3. Other AnfK polypeptides have lower homology and are only 25-31% identical to SEQ ID NO:3, but are still included in the AnfK polypeptides of the present invention. AnfK polypeptides do not include molybdenum-type NifK polypeptides and vanadium-type NifK polypeptides (VnfK). The AnfK fusion peptide and cleaved AnfK peptide of the present invention preferably have the same C-terminus as the natural AnfK peptide, i.e., they do not have any artificial additions to the C-terminus, and when compared with the natural AnfK peptide (SEQ ID NO:36), they do not have any amino acids missing from the C-terminus. In preferred embodiments, the amino acid sequences of the AnfK fusion peptide and cleaved AnfK peptide of the present invention have the amino acid sequences LNVW (SEQ ID NO:54), LNTW (SEQ ID NO:55), LNMW (SEQ ID NO:56), LAMW (SEQ ID NO:57), or LSVW (SEQ ID NO:58) at their C-terminus. The amino acid sequences of AnfK peptides in protein sequence databases are generally annotated as AnfK peptides. As of January 2020, there are 155 specific amino acid sequences in the AnfK collection protein database that are different from the molybdenum-type NifK and VnfK peptide sequences. Examples of naturally occurring AnfK peptides include those from *Azomonas flexneri* (…). Azomonas agilis (Accession ID WP_144571040; 91.34% identical), Clostridium ( Clostridium sp.) BL-8 (accession number WP_077859050; 78.35% identical), butyric acid-associated matchstick-shaped luminescent bacteria ( Butterflies (Accession number WP_122630336; 62.34% identical) and Rhodopsinophilus acidophilus ( Rhodoblastus acidophilus (Registration number WP_088520366; 54% identical), each refer to SEQ ID NO:36.

[0303] Similar to other functional NifK peptides, functional AnfK peptides can function as a β-protein structural component of α2β2δ2 heterohexameric nitrogenase, working together with α-protein (AnfD) and δ-protein (AnfG) to form a complex with an active site for reducing molecular nitrogen on FeFe-co.

[0304] The naturally occurring bacterial NifB polypeptide is a protein that converts the [4Fe-4S] cluster into NifB-co, where the central C atom with high nuclear polarity acts as a precursor for the synthesis of FeMo-co, FeV-co, and FeFe-co (Guo et al., 2016). Therefore, NifB catalyzes the first key step in the FeMo-co, FeV-co, and FeFe-co synthesis pathways and is thus essential for nitrogenase function. The NifB-co product can bind to the NifE-NifN complex and can shuttle from NifB to NifE-NifN via the metal cluster carrier protein NifX.

[0305] As used herein, “NifB polypeptide” means a polypeptide whose amino acid sequence contains at least 27% of the amino acids identical to the amino acid sequence provided in SEQ ID NO:4. Most NifB polypeptides contain one or more of the conserved domains TIGR01290, NifB conserved domain cd00852, NifX-NifB superfamily conserved domain cl00252, and Radical_SAM conserved domain cd01335. As used herein, NifB polypeptides include naturally occurring polypeptides that have been annotated as having NifB function but do not possess one of these domains. They originate from the genera *Klebsiella*, *Azotobacter*, and *Rhizobium*. Rhizobium ), Slow-growing rhizobia ( Bradyrhizobium NifB peptides from bacteria such as *Rauvolf. spp.* and others have a C-terminal NifX-like extension, while most archaea NifB peptides lack the NifX-like domain and are referred to as "truncated NifB peptides." Naturally occurring NifB peptides are typically 440 to 500 amino acids in length, and the molecular weight of the natural monomer is approximately 50 kDa. A large number of NifB peptides have been identified, and many sequences are available in publicly available databases. For example, NifB peptides from *Rauvolf. spp.* (accession number WP_041145602.1, 91% identical to SEQ ID NO:4), *Sacchariformis spp.* (…), and *Rauvolf. spp.* (…) have been reported. Kosakonia radicincitans (WP_043953592.1, 80% identical), Chrysodendron ( Dickeya chrysanthemum (WP_040003311.1, 76% identical), *Pectinobacterium niger* ( Pectococcus atroseptic(WP_011094468.1, 70% identical), *Buchnerella guinea* (WP_048638849.1, 63% identical), *Rhodospirillum halophilus* ( Halorhodospira halophila (WP_011813098.1, 59% identical, lacking NifX domain), *Bacillus thuringiensis* ( Methanosarcina barkeri (WP_048108879.1, 50% identical, lacking the NifX domain), Clostridium purine-degrading ( Clostridium purinilyticum (WP_050355163.1, 40% identical, lacking the NifX domain) and salt-demyogenic Vibrio ( Desulfovibrio salexigens (WP_015850328.1, 27% identical). As used herein, “functional NifB peptides” are NifB peptides capable of forming NifB-co from [4Fe-4S] clusters. Functional NifB requires S-adenosylmethionine (SAM) to function. NifB peptides have been described and reviewed in Curatti et al. (2006) and Allen et al. (1995).

[0306] Boyd et al. (2011) studied the phylogenetic relationships of Anf / Vnf / NifDKEN and NifB from 40 taxa and drew the following conclusions: (1) NifB encoding a lack of the C-terminal NifX domain... Snow Lateral gene transfer of clusters occurs in the order Methanosarcoptera ( Methanosarcinales The ancestor of methanogens in the phylum Methanogens ( ) descended to the phylum Firmicutes ( ) Firmicutes (1) Ancestors in which two organisms coexisted in an anaerobic environment where molybdenum was available; and (2) following this lateral gene transfer event, the fusion of NifB and NifX occurred in Firmicutes, from which the nitrogen-fixing bacterial lineage evolved. Evidence supporting this theory includes: (1) the absence of methanogenic archaea (Methanococciles (...) Methanococcales Methanocytococcidales and Methanobacteria ( Methanobacteria (2) NifB sequences from Methanobacteria and Methanococci order indicate a relationship with Methanocytiformes and bacteria ( ) Bacteria Early divergence of the NifB sequence, and (3) some anaerobic Firmicutes, Green Curvatures ( Chloroflexi ) and Proteobacteria ( Proteobacteria NifB, which lacks a C-terminal NifX domain, is presumably formed shortly after a Nif lateral gene transfer event by Firmicutes (Bacteria). Firmly Early differentiation of lineage.

[0307] To determine the presence of a C-terminal NifX domain in the NifB peptide, constraint-based multiple alignment tools (COBALT, NCBI, etc.) can be used. www.ncbi.nlm.nih.gov / tools / cobalt / re_cobalt.cgi The amino acid sequence of NifB was compared with representative NifB sequences, such as Klebsiella micranthae NifB (accession number P10930), Klebsiella micranthae NifX (KZT46636.1), NifY (KZT46633.1), Azotobacter vinifera NifX (AGK13791.1), NifY (AGK13792.1), NafY (AGK13761.1), and NifX / NifY / NafY / VnfX family protein (AGK14217.1). The 'dinitrogenase FeMo-cofactor binding site' (Pfam family PF02579) in each sequence can be identified using the Pfam-A database with an expected value set to 10 via PfamScan (EMBL-EBI, www.ebi.ac.uk / Tools / pfa / pfamscan / ).

[0308] The NifEN complex is a scaffold complex required for the proper assembly of dinitrogenase, acting as a scaffold for the maturation of NifB-co to FeMo-co (a process also requiring NifH function), and is structurally similar to dinitrogenase (Fay et al., 2016). The NifEN complex consists of two subunits from each of NifE and NifN, forming a heterotetramer, referred to here as ENα2β2. The naturally occurring bacterial NifE polypeptide is a polypeptide containing the α subunit of the ENα2β2 tetramer of the NifN polypeptide, and this ENα2β2 tetramer is required for FeMo-co synthesis and has been proposed to serve as a scaffold for FeMo-co synthesis.

[0309] As used herein, “NifE polypeptide” means a polypeptide containing amino acids whose sequence is at least 32% identical to the amino acid sequence provided in SEQ ID NO:5, and which contains one or both of the domains TIGR01283 and PRK14478. Members of the TIGR01283 domain protein family are also members of the superfamily cl02775. Naturally occurring NifE polypeptides are typically 440 to 490 amino acids in length, and the molecular weight of the natural monomer is approximately 50 kDa. Numerous NifE polypeptides have been identified, and many sequences are available in publicly available databases. For example, NifE peptides from the following strains have been reported: *Klebsiella micrantha* (accession number WP_049114606.1, 99% identical to SEQ ID NO:5), *Klebsiella variegata* (SBM87755.1, 92% identical), *Descarcass bananas* (WP_012764127.1, 89% identical), *Toxococcus aureus* (WP_012728883.1, 75% identical), *Pseudomonas schrenckii* (WP_003297989.1, 69% identical), *Azotobacter ventriculatus* (WP_012698965.1, 62% identical), and *Azotobacter ventriculatus* (…). Trichormus azollae (WP_013190624.1, 55% identical), Bacillus sclerotiorum ( Paenibacillus durus (WP_025698318.1, 50% identical), Campylobacter jiuci sulfur-producing bacteria ( Sulfuricurvum kujiense (WP_013460149.1, 44% identical), formic acid-associated methanobacteria ( Methanobacterium formicum (AIS31022.1, 39% identical), amino acid-loving anaerobic banana bacteria ( Anaeromus acidaminophila (WP_018701501.1, 35% identical) and Saccharomyces macrococcus ( Megasphaera cerevisiae (WP_048514099.1, 32% identical). As used herein, a “functional NifE peptide” is a NifE peptide capable of forming a functional tetramer with NifN, enabling the synthesis of FeMo-co from the complex. The synthesis of this FeMo-co involves other peptides, including NifH and NifB, and possibly NifX. NifE peptides have been described and reviewed in: Fay et al. (2016), Hu et al. (2005), Hu et al. (2006), and Hu et al. (2008).

[0310] NifF polypeptides in naturally occurring nitrogen-fixing organisms are flavin-redoxins that act as electron donors for NifH. As used herein, “NifF polypeptide” means a polypeptide containing amino acids whose sequence is at least 34% identical to the amino acid sequence provided as SEQ ID NO:6, and said polypeptide is contained in nitrogen-fixing bacteria ( Azobacter NifF peptides are found on Nif proteins of the flavin-oxidizing protein long domain TIGR01752 and the flavin-oxidizing protein FLDA domain, or both. NifF peptides encompass flavin-oxidizing proteins in non-nitrogen-fixing bacteria associated with pyruvate-formate lyase activation and cobalamin-dependent methionine synthase activity, but exclude other flavin-oxidizing proteins involved in broader functions. Naturally occurring NifF peptides are typically 160 to 200 amino acids in length, and the molecular weight of the natural monomer is approximately 19 kDa. Numerous NifF peptides have been identified, and many sequences are available in publicly available databases. For example, NifF peptides have been reported from the following: *Klebsiella micrantha* (accession number WP_004122417.1, 99% identical to SEQ ID NO:6), *Klebsiella variegata* (WP_040968713.1, 85% identical), *Sacchariformis Root-Promoting Bacteria* (WP_035885760.1, 76% identical), *Cyclocarya guillier* (WP_039999438.1, 72% identical), *Buchneria guillier* (WP_048638838.1, 62% identical), and *Methane-Associated Methylmonas* (…). Methylomonas methanica (WP_064006977.1, 56% identical), *Vibrio vulnificus* (WP_012698862.1, 50% identical), *Corynebacterium pulveratum* ( Chlorobaculum tepidum (WP_010933399.1, 39% identical), Campylobacter Showa ( Campylobacter show (WP_002949173.1, 37% identical) and Azotobacter brownii ( Azotobacter chroococcum (WP_039801725.1, 34% identical). As used herein, “functional NifF peptides” are NifF peptides that can act as electron donors for NifH peptides. NifF peptides have been described and reviewed in Drummond (1985).

[0311] As used herein, “AnfG polypeptide” is a member of the conserved nitrogenase superfamily cl03910 (pfam03139-AnfG) containing the conserved domain TIGR02929 and having at least 42% amino acid sequence identity with the AnfG polypeptide from *Azotobacter venereum* (SEQ ID NO:38; accession number WP_012703360) when measured along its full length along SEQ ID NO:38. This amino acid sequence is used herein as a reference sequence for AnfG. TIGR02929 represents an all-ferric variant of the nitrogenase component I δ chain. AnfG polypeptide does not include vanadium-type NifG polypeptide (VnfG). The amino acid sequence of AnfG polypeptide is generally annotated as AnfG polypeptide in protein sequence databases. As of January 2020, 150 specific amino acid sequences exist in protein databases containing the AnfG collection. Examples of naturally occurring AnfG peptides include those from: *Azomonas aflexis* (accession number WP_144571041; 84.73% identical), *Firmite* bacteria (… Firmicutes bacteria (Accession number HBE76208; 70.37% identical), *Termite Species* ( Sporomusa termitida (Accession number WP_144349445; 68.75% identical), Green small red oomycetes ( Green rhododendron (accession number WP_112317428; 57.14% identical) and Saccharomyces cerevisiae (accession number WP_048515315; 42.86% identical), each refer to SEQ ID NO:38.

[0312] The functional AnfG peptide can function as a structural component of the δ protein of the α2β2δ2 heterohexamer nitrogenase.

[0313] Naturally occurring bacterial NifJ peptides are pyruvate:flavin oxidoreductases (ferroreductin) oxidoreductases, which are electron donors for NifH. As used herein, “NifJ peptide” means a peptide containing amino acids whose sequence is at least 40% identical to the amino acid sequence provided as SEQ ID NO:7, and which contains the conserved domain TIGR02176. Naturally occurring NifJ peptides are typically 1100 to 1200 amino acids in length, and the molecular weight of the natural monomer is approximately 128 kDa. Numerous NifJ peptides have been identified, and many sequences are available in publicly available databases. For example, NifJ peptides from *Klebsiella micranthae* (accession number WP_024360006.1, 99% identical to SEQ ID NO:7), *Rauvolfia ornithine-lysinus* (WP_044347157.1, 95% identical), and *Klebsiella pneumoniae* (…) have been reported. Klebsiella quasipneumoniae(WP_050533844.1, 92% identical), Sacchariformis oryzae ( Kosakonia rice (WP_064566543.1, 82% identical), Solanum spp. ( Dickeya solani (WP_057084649.1, 78% identical), Aquatic Raenella (WP_014683040.1, 72% identical), Maslania heat anaerobic bacillus ( Thermoanaerobacter mathranii (WP_013149847.1, 64% identical), Clostridium botulinum ( Clostridium botulinum (WP_053341220.1, 60% identical), African spirochetes ( Spirochaete African (WP_014454638.1, 52% identical) and Vibrio cholerae ( Cholera vibrio (CSA83023.1, 40% identical). As used herein, “functional NifJ peptides” are NifJ peptides that can act as electron donors for NifH peptides. NifJ peptides have been described and reviewed in Schmitz et al. (2001).

[0314] Naturally occurring bacterial NifM peptides are peptides required for the maturation of some, but not all, NifH peptides. In the absence of NifM, Klebsiella pneumoniae NifH exists only at low levels in Escherichia coli and yeast when heterologously expressed and cannot donate electrons to NifD-NifK. As used herein, “NifM peptide” means a peptide containing amino acids whose sequence is at least 26% identical to the amino acid sequence provided as SEQ ID NO:8, and which contains the TIGR02933 domain. NifM peptides are homologous to peptidyl-prolyl cis-trans isomerases (PPIs), a group of enzymes that promote protein folding by catalyzing the cis-trans isomerization of proline imine peptide bonds, possess a PpiC-type domain, and appear to be accessory proteins of some NifH peptides, including at least some VnfH and AnfH peptides. Naturally occurring NifM peptides are typically 240 to 300 amino acids in length, and the molecular weight of the natural monomer is approximately 30 kDa. Numerous NifM peptides have been identified, and many sequences are available in publicly available databases. For example, NifM peptides from the following strains have been reported: *Klebsiella acidogenica* (accession number WP_064342940.1, 99% identical to SEQ ID NO:8), *Klebsiella micrantha* (WP_004122413.1, 97% identical), *Rauvolfia ornithine-lysinophila* (WP_044347181.1, 85% identical), *Klebsiella variegata* (WP_063105800.1, 75% identical), *Sacchariformis radicans* (WP_035885759.1, 59% identical), *Pectinobacter niger* (WP_011094472.1, 42% identical), *Buchneria guinea* (WP_048638837.1, 33% identical), and *Pseudomonas aeruginosa* (…). Pseudomonas aeruginosa PAO1 (CAA75544.1, 28% identical), Micrococcus pyogenes ( Marine bacteria sp.) AK27 (WP_051692859.1, 27% identical) *Bostomia tumefaciens* ( Teredinibacter turner (WP_018415157.1, 26% identical). As used herein, “functional NifM peptide” is a NifM peptide capable of complexing with NifH peptide to mature NifH peptide. NifM peptides have been described and reviewed in Petrova et al. (2000).

[0315] Naturally occurring bacterial NifN polypeptides are β subunits of the ENα2β2 tetramer of NifE polypeptides, and the ENα2β2 tetramer is required for FeMo-co synthesis and is suggested to serve as a scaffold for FeMo-co synthesis. As used herein, “NifN polypeptide” means (i) a polypeptide containing at least 76% of the amino acids whose sequence is identical to the sequence provided as SEQ ID NO:9 and / or (ii) a polypeptide containing at least 34% of the amino acids whose sequence is identical to the sequence provided as SEQ ID NO:9, and which contains one or more of the conserved domains TIGR01285, cd01966, and PRK14476. NifN is structurally associated with the molybdenum-ferritin β chain NifK. Polypeptides containing the conserved TIGR01285 cover most examples of NifN polypeptides, but exclude some NifN polypeptides, such as those from *Vibrio vulnificus* (…). Chlorobium tepidum The presumed NifN is defined as follows, and therefore, NifN is not limited to peptides containing the conserved TIGR01285 domain. Members of the PRK14476 domain protein family are also members of the superfamily cl02775. Naturally occurring NifN peptides are typically 410 to 470 amino acids in length, although they can have approximately 900 amino acid residues when naturally fused with NifE, and the molecular weight of the natural monomer is approximately 50 kDa. A large number of NifN peptides have been identified, and many sequences are available in publicly available databases. For example, NifN peptides have been reported from the following: Klebsiella pneumoniae (accession number WP_064391778.1, 97% identical to SEQ ID NO:9), Kluyveromyces intermedius (WP_047370268.1, 80% identical), Laenia aquaticus (WP_014683026.1, 70% identical), Buchenella guleri (WP_048638830.1, 65% identical), and Methylobacterium tundrae (…). Methylobacter tundripaludum (WP_027147663.1, 46% identical), *Cyclocarya paliurus* ( Calothrix parietina (WP_015195966.1, 41% identical), *Mammotrophic motility* ( Zymomonas mobilis (WP_023593609.1, 37% identical), Bacillus simietus ( Paenibacillus massiliensis (WP_025677480.1, 35% identical) and Copenhagen desulfurization bacteria ( Desulfitobacterium hafniense(WP_018306265.1, 34% identical). As used herein, “functional NifN peptides” are NifN peptides capable of forming functional tetramers with NifE, enabling the complex to synthesize FeMo-co. NifN peptides have been described and reviewed in the following publications: Fay et al. (2016), Brigle et al. (1987), Fani et al. (2000), and Hu et al. (2005).

[0316] The NifQ polypeptide, found in naturally occurring bacteria, is a polypeptide involved in FeMo-co synthesis and may be present in MoO4. 2- Early stages of processing. Conserved C-terminal cysteine ​​residues may be involved in metal binding. As used herein, “NifQ peptide” means a peptide containing amino acids whose sequence is at least 34% identical to the amino acid sequence provided as SEQ ID NO:10, and which is a member of the CL04826 domain protein family and the pfam04891 domain protein family. Naturally occurring NifQ peptides are typically 160 to 250 amino acids in length, although they can be as long as 350 amino acid residues, and the molecular weight of the natural monomer is approximately 20 kDa. A large number of NifQ peptides have been identified, and many sequences are available in publicly available databases. For example, NifQ peptides from the following strains have been reported: Klebsiella acidogenica (accession number WP_064391765.1, 95% identical to SEQ ID NO:10), Klebsiella variegata (CTQ06350.1, 75% identical), Kluyveromyces intermedius (WP_047370257.1, 63% identical), Bacillus niger (WP_043878077.1, 59% identical), and metal-resistant slow-growing rhizobia (… Mesorhizobium metallidurans (WP_008878174.1, 46% identical), Rhodopseudomonas palustris ( Rhodopseudomonas palustris (WP_011501504.1, 42% identical), Burkholderia stesii ( Paraburkholderia sprentiae (WP_027196569.1, 41% identical), Stable Burkholderia ( Burkholderia stable (GAU06296.1, 39% identical) and copper oxalate scavengers ( Copper-loving oxalate (WP_063239464.1, 34% identical). As used herein, the “functional NifQ peptide” is capable of processing MoO4. 2- The NifQ peptide has been described and reviewed in Allen et al. (1995) and Siddavattam et al. (1993).

[0317] Naturally occurring bacterial NifS polypeptides are cysteine ​​desulfurases involved in the biosynthesis of iron-sulfur (FeS) clusters, such as their role in sulfur mobilization for Fe-S cluster synthesis and repair. As used herein, “NifS polypeptide” means (i) a polypeptide containing at least 90% of the amino acids whose sequence is identical to the amino acid sequence provided as in SEQ ID NO:19 and / or (ii) a polypeptide containing at least 36% of the amino acids whose sequence is identical to the amino acid sequence provided as in SEQ ID NO:19, and which contains one or both of the conserved domains TIGR03402 and COG1104. The TIGR03402 domain protein family includes a branch almost always found in extended nitrogen fixation systems plus a second branch more closely related to the first than to IscS, and a portion of the NifS-like / NifU-like system. The TIGR03402 domain protein family does not extend to species such as Helicobacter pylori (… Helicobacter pylori The more distant branch found in the phylum Epsiloneus (also referred to as NifS in the literature) is constructed in TIGR03403. The COG1104 domain protein family includes cysteine ​​sulfinate desulfinase / cysteine ​​desulfonase or related enzymes. Some NifS peptides include the aspartate aminotransferase domain cl18945. Naturally occurring NifS peptides are typically 370 to 440 amino acids in length, and the molecular weight of the natural monomer is approximately 43 kDa. A large number of NifS peptides have been identified, and many sequences are available in publicly available databases. For example, NifS peptides from the following have been reported: Klebsiella micranthae (accession number WP_004138780.1, 99% identical to SEQ ID NO:19), Arnebrio micranthae (… Raoultella native to the earth (WP_045858151.1, 89% identical), Kluyveromyces intermedia (WP_047370265.1, 80% identical), Laenia aquaticus (WP_014333911.1, 73% identical), pale yellow agar-eating bacteria ( Agarivorous gilvus (WP_055731597.1, 64% identical), Azotobacter brasiliensis ( Azospirillum brasilense (WP_014239770.1, 60% identical), Oxidized Ketone Desulfurized Micrococcus ( Desulfosarcine ketone (WP_054691765.1, 55% identical), Clostridium ( Intestinal Clostridium (WP_021802294.1, 47% identical), less halophilic bacteria ( Clostridisalibacter paucivorans (WP_026894054.1, 36% identical) and Bacillus coagulans ( Bacillus coagulans(WP_061575621.1, 42% identical and located in COG1104). As used herein, “functional NifS peptides” are NifS peptides capable of playing a role in the biosynthesis and / or repair of iron-sulfur (FeS) clusters. NifS peptides have been described and reviewed in the following: Clausen et al. (2000), Johnson et al. (2005), Olson et al. (2000), and Yuvaniyama et al. (2000).

[0318] The naturally occurring bacterial NifU polypeptide is a molecular scaffold polypeptide involved in the biosynthesis of iron-sulfur (FeS) clusters, components of nitrogenase. As used herein, "NifU polypeptide" means a polypeptide containing amino acids whose sequence is at least 31% identical to the sequence provided as SEQ ID NO:12, and which contains the TIGR02000 domain. Members of the TIGR02000 domain protein family are specifically involved in nitrogenase maturation. NifU contains an N-terminal domain (pfam01592) and a C-terminal domain (pfam01106). Three distinct but partially homologous Fe-S cluster assembly systems have been described: Isc, Suf, and Nif. The Nif system (in which NifU is a component) associates with nitrogenases that provide Fe-S clusters to various nitrogen-fixing species. [The text then abruptly shifts to a completely unrelated topic:] ...from Helicobacter pylori (... Helicobacter ) and Campylobacter ( Campylobacter Homologs of Isc and Suf with equivalent domain architectures of NifU are excluded from the definition of NifU in this paper. Therefore, NifU is specific for NifU peptides involved in nitrogenase maturation. Members of the related TIGR01999 domain protein family are also excluded from the definition of NifU in this paper; these members are IscU proteins (e.g., from *E. coli*, *Saccharomyces cerevisiae*, and *Homo sapiens*) containing homologs of the N-terminal region of NifU. Naturally occurring NifU peptides are typically 260 to 310 amino acids in length, and the molecular weight of the natural monomer is approximately 29 kDa. A large number of NifU peptides have been identified, and many sequences are available in publicly available databases. For example, NifU peptides have been reported from the following: *Klebsiella micranthae* (accession number WP_049136164.1, 97% identical to SEQ ID NO:12), *Klebsiella variegata* (WP_050887862.1, 90% identical), *Diccardia solanacea* (WP_057084657.1, 80% identical), *Buchnerella guillierii* (WP_048638833.1, 73% identical), *Toxococcus aureus* (WP_012728889.1, 66% identical), and *Agaricus aureus* (…). Methylomonas methanica (WP_055731596.1, 58% identical), Veksen desulfurized aspergillus ( Desulfocurvus vexinensis (WP_028587630.1, 54% identical), Rhodopseudomonas palustris (WP_044417303.1, 49% identical), Helicobacter pylori (WP_001051984.1, 31% identical) and Thiooomie ( Sulfur sp.) PC08-66 (KIM05011.1, 31% identical). As used herein, “functional NifU peptides” are NifU peptides that can function as molecular scaffold peptides involved in the biosynthesis of iron-sulfur (FeS) clusters. NifU peptides have been described and reviewed in the following: Hwang et al. (1996), Mühlenhoff et al. (2003), and Ouzounis et al. (1994).

[0319] NifS is a pyridoxal phosphate (PLP, vitamin B6)-dependent cysteine ​​desulfurase that produces inorganic sulfides required for the synthesis of Fe-S clusters from cysteine. Alanine is produced as a byproduct of the reaction. The reaction proceeds via a protein-bound cysteine ​​persulfide intermediate formed by nucleophilic attack on a highly conserved cysteine ​​residue (Cys325 in Venetian violaceae) on the cysteine-PLP adduct (Zheng et al., 1994). The sulfide is supplied to NifU for the sequential formation of [Fe2S2] and [Fe4S4] clusters. The NifS enzyme functions as a homodimer in the bacteria.

[0320] NifU provides a scaffold for [Fe4S4] cluster formation, functioning as a homodimer. The NifU polypeptide contains three domains: an N-terminal scaffold domain, a central domain, and a C-terminal scaffold domain (Smith et al., 2005). The N-terminal domain shows high sequence homology with IscU proteins from bacteria and Isu proteins from eukaryotes, while the C-terminal domain is homologous to Nfu proteins found in mitochondria and chloroplasts. Each NifU subunit in the central domain contains a permanent redox activity [Fe2S2]. 2+Clusters, due to their stability, are considered not to be transferred to other Nif proteins. These clusters are thought to be coordinated by four conserved cysteine ​​residues (Cys137, 139, 172, and 175 in NifU from Venetian davidii) (Fu et al., 1994). In bacteria, NifU forms homodimers, and its N-terminal domain can bind one [Fe2S2] cluster to each monomer. The [Fe2S2] clusters in the monomers can be reduced and fused to form one [Fe4S4] cluster per NifU dimer. A pair of [Fe4S4] clusters is then delivered from NifU to NifB and processed on NifB into an 8Fe core, subsequently used for FeMoco synthesis. In different pathways involving the Fe-S clusters, a [Fe4S4] cluster bound to the N-terminal or C-terminal scaffold domain of NifU is transferred to apo-NifH to mature the nitrogenase reductase, i.e., the NifH protein (Smith et al., 2005). It has been proposed that NifU also provides two [Fe4S4] clusters to the NifD-NifK protein complex (designated herein as stage 0 DK), and that NifH condenses this pair of clusters into a mature P cluster [Fe8-S7] (Dos Santos et al., 2004). These N-terminal clusters are considered extremely unstable and are not retained during purification (Smith et al., 2005). C-terminal domains can retain one [Fe4S4] cluster per monomer. In contrast to the N-terminal clusters, the assembly of the C-terminal [Fe4S4] clusters is rapid and the intermediate [Fe2S2] cluster has not been detected (Smith et al., 2005). C-terminal clusters are more stable than N-terminal clusters and can be retained during purification. However, C-terminal clusters are rapidly degraded upon reduction with dithionite (Smith et al., 2005). Using a cysteine-to-alanine mutation in NifU, Dos Santos and colleagues showed that both N-terminal and C-terminal clusters can be transferred to apo-NifH.

[0321] López-Torrejón et al. (2016) reported that a NifH protein capable of donating electrons to holoNifD-NifK could be generated in yeast mitochondria by expressing both NifH and NifM. These authors found that the generation of a functional NifH protein in yeast cells did not require NifS and NifU. They concluded that the endogenous iron-sulfur cluster assembly pathway in yeast cells, presumably involving the mitochondrial localization of related proteins Nfs1 and Nfu1, is capable of donating [Fe4S4] clusters to NifH. Therefore, the reconstructing of NifH, Fe proteins, or nitrogenase reductase in yeast may not require NifS and NifU, but the maturation and function of NifB and / or NifD-NifK may require both NifS and NifU. Whether plant mitochondria possess a similar endogenous capacity to form sufficient [Fe4S4] clusters for nitrogenase activity is unknown.

[0322] The naturally occurring bacterial NifV polypeptide is a hypercitrate synthase (EC 2.3.3.14), which produces hypercitrate by transferring an acetyl group from acetyl-CoA (acetyl-CoA) to 2-oxoglutarate. The hypercitrate is then used for the synthesis of FeMo-co, FeV-co, and FeFe-co. As used herein, “NifV polypeptide” means a polypeptide containing amino acids whose sequence is at least 39% identical to the amino acid sequence provided as SEQ ID NO:13, and which contains one or both of the domains TIGR02660 and DRE_TIM. Members of the TIGR02660 domain protein family are homologous to enzymes including 2-isopropylmalate synthase, (R)-citrate synthase, and hypercitrate synthase associated with processes other than nitrogen fixation. The cd07939 domain protein family also includes *Helicobacter aeruginosa*, which appears to be orthologous to *FrbC*. Heliobacterium chlorine ) and food-eating acetic acid bacteria ( Gluconacetobacter diazotrophicusThe NifV protein belongs to the DRE-TIM metallolytic enzyme superfamily. DRE-TIM metallolytic enzymes include 2-isopropylmalate synthase (IPMS), α-isopropylmalate synthase (LeuA), 3-hydroxy-3-methylglutaryl-CoA lyase, homocitrate synthase, citrate synthase, 4-hydroxy-2-oxovalerate aldolase, re-citrate synthase, transcarboxylase 5S, pyruvate carboxylase, AksA, and FrbC. These members all share a conserved triose phosphate isomerase (TIM) barrel domain, which consists of a core β(8)-α(8) motif, in which eight parallel β chains form a closed barrel structure surrounded by eight α helices. The domain has a catalytic center containing a divalent cation-binding site formed by a cluster of invariant residues covering the barrel core. Additionally, the catalytic site includes three invariant residues—aspartic acid (D), arginine (R), and glutamic acid (E)—which form the basis of the domain name “DRE-TIM.” Naturally occurring NifV peptides are typically 360 to 390 amino acids in length, although some members are approximately 490 amino acid residues long, and the molecular weight of the natural monomer is approximately 41 kDa. Numerous NifV peptides have been identified, and many sequences are available in publicly available databases. For example, NifV peptides have been reported from the following: Klebsiella micranthae (accession number WP_049083341.1, 95% identical to SEQ ID NO:13), Ornithine-lysine-laurella (WP_045858154.1, 86% identical), Kluyveromyces intermedius (WP_047370264.1, 81% identical), Tidicarbium dadanense (WP_038912041.1, 70% identical), Buchenella guleri (WP_048638835.1, 59% identical), and Magnetococcus spp. ( Marine Magnetococcus (WP_011712856.1, 46% identical), Vitigrymistomonas sphingosine monospora ( Sphingomonas wittichii (WP_037528703.1, 43% identical), Frankincense ( Frankish sp.) EI5c (OAA29062.1, 41% identical) and Clostridium martini ( Clostridium sp. Maddingley) MBC34-26 (EKQ56006.1, 39% identical). As used herein, “functional NifV peptide” refers to a NifV peptide capable of functioning as a hypercitrate synthase. NifV peptides have been described and reviewed in the following: Hu et al. (2008), Lee et al. (2000), Masukawa et al. (2007), and Zheng et al. (1997).

[0323] NifX peptides in *Azotobacter venereum* bind to NifB-co (Fe6-S9-C), which is then transferred to NifE-NifN for FeMo-co assembly (Hernandez et al., 2007). Exchange of VK clusters (Fe8-S9-C or Mo-Fe7-S9-C) between NifE-NifN has also been shown, indicating their role as transient reservoirs of FeMo-co precursors. Hernandez et al. (2007) reported that NifX can act as a molecular chaperone stabilizing the NifE-NifN or NifD-NifK complex during FeMo-co transfer to apo-NifD-NifK, and / or reorient the protein to orientations favorable to FeMoco transfer and thus regulate FeMoco synthesis. Activation of apo-NifD-NifK by exogenous FeMo-co with a dinitrogenase complex extracted from Venelandia vinifera mutants lacking different combinations of NifY / NafY / NifX accessory proteins suggests that NifX can also assist in the FeMo-co insertion of apo-NifD-NifK (Rubio et al., 2002). This additional function of NifX may be as shown by Homer et al. (1993) in Klebsiella spp. nifY The reason why the acetylene reducing activity is maintained in the mutant.

[0324] Naturally occurring bacterial NifX peptides are peptides involved in FeMo-co synthesis, at least facilitating the transfer of FeMo-co precursors from NifB to NifE-NifN or from FeMo-co to NifD-NifK. As used herein, “NifX peptide” means a peptide containing amino acids whose sequence is at least 29% identical to the amino acid sequence provided as SEQ ID NO:14, and which contains one or both of the conserved domains TIGR02663 and cd00853. NifX is included in a larger family of iron-molybdenum cluster-binding proteins that includes some NifB sequences and NifY, because the C-terminal regions of NifX, NafY, and some NifB peptides all contain the pfam02579 domain, and each is involved in the synthesis of one or more or all of FeMo-co, FeV-co, or FeFe-co. Other NifB peptides, particularly those from methanogenic archaea and some anaerobic Firmicutes, lack the NifX-like domain (Boyd et al., 2011), including NifB peptides from the aforementioned *Rhodospirillum halospirillum*, *Bacillus methanogenus*, and *Clostridium purine-decomposing*. Some NifX peptides are annotated as NifY in databases, and vice versa. Naturally occurring NifX peptides are generated independently, rather than as part of a NifB peptide as a natural fusion, and are typically 110 to 160 amino acids in length, with the natural monomer having a molecular weight of approximately 15 kDa. A large number of NifX peptides have been identified, and many sequences are available in publicly available databases. For example, NifX peptides from the following strains have been reported: *Klebsiella micranthae* (accession number WP_049070199.1, 97% identical to SEQ ID NO:14), *Klebsiella acidogenica* (WP_064342937.1, 97% identical), *Rauvolfia ornithine-lysinophila* (WP_044347173.1, 91% identical), *Klebsiella variegata* (WP_044612922.1, 83% identical), *Sacchariformis radicans* (WP_043953583.1, 75% identical), *Cyclocarya spp.* (WP_039999416.1, 68% identical), *Laenia aquatica* (WP_047608097.1, 58% identical), *Azotobacter chrysogenum* (WP_039800848.1, 34% identical), and *Berigerella filamentosa* (…). Beggiatoa leptomitoformis (WP_062149047.1, 33% identical) and Student Methylvariella ( Methyloversatilis students(WP_020165972.1, 29% identical). As used herein, “functional NifX peptide” is a NifX peptide capable of transferring the FeMo-co precursor from NifB to NifE-NifN. NifX peptides have been described and reviewed in Allen et al. (1994) and Shah et al. (1999).

[0325] Naturally occurring bacterial NifY polypeptides are polypeptides involved in FeMo-co synthesis, at least contributing to the transfer of FeMo-co precursors from NifB to NifE-NifN. As used herein, “NifY polypeptide” means a polypeptide containing amino acids whose sequence is at least 34% identical to the amino acid sequence provided as in SEQ ID NO:15, and which contains one or both of the conserved domains TIGR02663 and cd00853. NifY is included in the larger family of iron-molybdenum cluster-binding proteins, which includes NifB and NifX, because the C-terminal regions of NifX, NafY, and NifB all contain the pfam02579 domain, and each is involved in FeMo-co synthesis. Numerous NifY polypeptides have been identified, and many sequences are available in publicly available databases. For example, NifY peptides have been reported from the following: *Klebsiella micranthae* (accession number WP_049089500.1, 99% identical to SEQ ID NO:15), *Klebsiella acidogenica* (WP_064342935.1, 98% identical), *Klebsiella pneumoniae* (WP_044524054.1, 90% identical), *Klebsiella variegata* (WP_049010739.1, 81% identical), *Kluyveromyces intermedius* (WP_047370270.1, 69% identical), *Cyclocarya pallida* (WP_039999411.1, 62% identical), and *Serratia marcescens* (…). Serratia sp.) ATCC 39006 (WP_037382461.1, 57% identical), Aquatic Raenella (WP_014683024.1, 47% identical), Pseudomonas putida ( Pseudomonas putida (AEX25784.1, 37% identical) and Venetian violaceum (WP_012698835.1, 34% identical). As used herein, “functional NifY peptide” is a NifY peptide capable of transferring the FeMo-co precursor from NifB to NifE-NifN.

[0326] When from acid-producing Klebsiella or Venetian violaceum NifB or NifN - NifEWhen isolated from mutant strains, apo-NifD-NifK associates with another polypeptide called γ protein (Paustian et al., 1990; Homer et al., 1993), forming a heterohexamer (α2β2γ2) with the NifD and NifK polypeptides. In Klebsiella acidogenic bacteria, the third polypeptide is formed by… NifY The gene encodes (Homer et al., 1993), and the addition of purified FeMo-co to a purified heterohexamer α2β2γ2 complex is sufficient to produce a catalytically active nitrogenase. The addition of FeMo-co causes NifY to dissociate from the complex and form the holoenzyme (α2β2). In *Venelandia diffusa*, the third polypeptide is generated by... NafY The gene (nitrogenase-associated factor Y; accession number AGK13761, Rubio et al., 2002) encodes a gene that is distinct from but related to that in *Azotobacter venerealis*. NifY The product of the gene (accession number AGK13792) is associated. In each case, the third polypeptide is thought to be involved in assisting FeMo-co insertion to form the active enzyme. This is supported by the ability of NafY and NifY to bind to FeMo-co (Homer et al., 1995).

[0327] At different stages of NifD-NifK holoenzyme maturation, Venetian nitrogen-fixing bacteria NifY and NafY bind to apo-NifD-NifK and α-Cys of NifD. 275 or α-His 442 The two amino acid residues of NifY are covalently anchored to FeMo-co (Jimenez-Vincente et al., 2018). That is, NifY and NafY do not bind simultaneously to apo-NifD-NifK. The binding order of NifY and NafY to apo-NifD-NifK is currently unknown. It has been demonstrated that for Klebsiella acidogenic nitrogenase, NifY dissociates from NifD-NifK upon FeMo-Co insertion (Homer et al., 1993), and for Viennelandia nitrogenase, NafY dissociates from NifD-NifK upon FeMo-Co insertion (Homer et al., 1995). NafY is also thought to bind via His... 121 It may also bind to FeMo-co via NifB-co, suggesting its role as a FeMo-co or FeMo-co precursor insertase (Rubio et al., 2004). Based on nifYThe mutant lacks the phenotype that *Venerendica* NifY appears to be functionally redundant (Rubio et al., 2002), and NafY is considered a major accessory protein supporting the FeMo-co insertion of apo-NifD-NifK. On the other hand, *Klebsiella* species do not possess... NafY Genes, and only the NifY gene to support the insertion of FeMo-co into apo-NifD-NifK, although Klebsiella spp. nifY The mutant retains 60% of the acetylene reduction activity (Homer et al., 1993). This retention of function suggests the existence of another accessory protein in Klebsiella that could partially cover the function of NifY in its absence, such as NifX as described above.

[0328] As used herein, “NafY polypeptide” refers to a polypeptide containing amino acids whose sequence along its full length is at least 50% identical to the sequence provided in SEQ ID NO:48 (Azotobacter venerealis NafY, accession number AGK13761, 243aa), and the polypeptide contains a conserved domain pfam16844. This domain, approximately 91 amino acid residues in length, has been found to be present in some members and in the N-terminal half of longer NafY proteins. This region is negatively charged and appears to have the function of recognizing and interacting with apo-NifD-NifK. Naturally occurring NafY polypeptides are typically 230 to 250 amino acids in length, and the molecular weight of the natural monomer is approximately 25-28 kDa. A large number of NafY polypeptides have been identified, and many sequences are available in publicly available databases. Due to the correlation between NafY and NifX sequences, some have been annotated as NifX polypeptides. For example, NafY polypeptides have been reported from the following: Azotobacter venerealis (… Azotobacter beijerinckii (WP_090728988, 93% identical to SEQ ID NO:48), *Pseudomonas stearothermiae* (WP_011912501, 69% identical), *Haloxylon ammodendron* ( Halomonas endophyticus (WP_102654474, 68% identical), *Pseudomonas linyingense* ( Pseudomonas linyingensis (WP_090313081, 67% identical), Proliferating acidophilus and halophilic bacteria ( Acidihalobacter prosperus (WP_038093031, 56% identical), Cyanobacteria Oscillatoria ( Oscillatorial cyanobacteria(WP_009769409, 50% identical). As used herein, “functional NafY peptides” are NafY peptides capable of binding to apo-NifD-NifK and FeMo-co. Dyer et al. (2003) reported the three-dimensional structures of NafY peptides from Venetian violaceum and compared and differentiated the sequences of NafY and NifY, NifX, VnfX and NifB peptides.

[0329] The naturally occurring bacterial NifZ polypeptide is a peptide involved in the synthesis of Fe-S clusters, particularly playing a role in the coupling of the second Fe4S4 pair in the formation of the second P cluster of the MoFe protein. NifZ is thought to act as a molecular chaperone, inducing a conformational change in at least the latter half of the apo-MoFe protein, thereby allowing the formation of the second P cluster with NifH. This is particularly relevant to *Zytobacter venereum*. NifZ The deletion reduced MoFe protein activity by 66% but had no effect on NifH activity. As used herein, “NifZ polypeptide” refers to a polypeptide containing amino acids whose sequence is at least 28% identical to the sequence provided as SEQ ID NO:16, and which contains the conserved domain pfam04319. This approximately 75-amino acid residue domain was found to be independently present in some members and in the N-terminal half of longer NifZ proteins. Naturally occurring NifZ polypeptides are typically 70 to 150 amino acids in length, and the molecular weight of the natural monomers is approximately 9 kDa to approximately 16 kDa. Numerous NifZ polypeptides have been identified, and many sequences are available in publicly available databases. For example, NifZ peptides have been reported from the following: *Klebsiella micranthae* (accession number WP_057173223.1, 93% identical to SEQ ID NO:16), *Klebsiella acidogenica* (WP_064342939.1, 95% identical), *Klebsiella variegata* (WP_043875005.1, 77% identical), *Sacchariformis* (WP_043953588.1, 67% identical), and *Sacchariformis* (…). Kosakonia sacchari (WP_065368553.1, 58% identical), River Iron Bean Fungus ( Ferriphaselus amnicola (WP_062627625.1, 47% identical), Burkholderia paraphylla digesting foreign compounds ( Paraburkholderia xenovorans (WP_011491838.1, 41% identical), iron-feeding acidophilic thiobacterium ( Acidithiobacillus ferrivorans (WP_014029050.1, 35% identical) and oligotrophic slow-growing rhizobia ( Bradyrhizobium oligotrophicum(WP_015665422.1, 28% identical). As used herein, “functional NifZ peptide” is a NifZ peptide capable of coupling with the Fe4S4 cluster in Fe-S cluster synthesis. NifZ peptides have been described and reviewed in Cotton (2009) and Hu et al. (2004).

[0330] Naturally occurring bacterial NifW peptides are peptides that associate with NifZ peptides to form higher-order complexes (Lee et al., 1998) and participate in the synthesis or activity of MoFe proteins (NifD-NifK). NifW and NifZ appear to be involved in the formation or accumulation of MoFe proteins (Paul and Merrick, 1987). As used herein, “NifW peptide” means a peptide whose amino acid sequence contains at least 28% identical amino acid sequences to those provided in SEQ ID NO:17, and that the peptide contains a conserved NifW superfamily protein domain, architecture ID 10505077, and is located in the P family PF03206. A variety of NifW peptides have been identified, and many sequences are available in publicly available databases. For example, NifW peptides have been reported from the following: Klebsiella pneumoniae (accession number WP_064342938.1, 98% identical to SEQ ID NO:17), Klebsiella micranthae (WP_049080155.1, 94% identical), Enterobacteriaceae ( Enterobacter sp. ) 10-1 (WP_095103586.1, 90% identical), Klebsiella pneumoniae (WP_065877373.1, 81% identical), Polar pectinobacterium ( Pectobacterium polaris (WP_095699971.1, 69% identical), *Diccardia bananaensis* (WP_012764136.1, 58% identical), *Bukhnerella guillier* ( WP_053085547.1, 36% identical), Water bream ( Aquaspirillum sp. ) LM1 (WP_077299824.1, 44% identical), tentatively classified as μ-Proteobacteria ( Candidatus Muproteobacteria bacterium ) RBG_16_64_10)( OGI40729 (34% identical), Vignelandia diffusa ( ACO76430.1, 32% identical) and marine thermophilic methyltrophic bacteria ( Methylocaldum marinum(BBA37427.1, 28% identical). As used herein, a “functional NifW peptide” is a NifW peptide that promotes or enhances one or more of the formation, accumulation, or activity of MoFe proteins. Functional NifWs may interact with NifZ and / or play a role in the oxygen protection of MoFe proteins (Gavini et al., 1998).

[0331] Most organisms, including both bacteria and eukaryotes (such as plants), possess numerous ferrugins. For example, in the genomes of *Azotobacter venereum* DJ and CA, 15 and 16 proteins, respectively, are annotated as ferrugins or ferrugin-like proteins. As used herein, a “ferrugin polypeptide” is an electron carrier protein with one or two iron-sulfur clusters of the [2Fe-2S], [3Fe-4S], and / or [4Fe-4S] type forming its reaction center; see Matsubara and Saeki (1992) review. They participate in various metabolic processes, including nitrogen fixation, and typically have a lower molecular weight than those not involved in nitrogenase. Given the extensive diversity of ferredoxins in most cells and the variations in compatibility or specificity of different ferredoxins in the function of supplemented FdxN in NifB-co synthesis observed in several studies (Yates, 1972; Jimenez-Vincente et al., 2014), ferredoxins (including ferredoxins such as FdxN) are best defined based on the presence and function of the iron-sulfur cluster, rather than on amino acid identity with the standard sequence (SEQ ID NO:47; accession number WP_012703542) of *Azotobacter vinifera* FdxN. As used herein, “FdxN polypeptide” is a ferredoxin or ferredoxin-like polypeptide that functions to provide electrons to mature diazotase reductase NifH and / or to NifB-co synthesis of nitrogenase and / or to act as an intermediate carrier of the [4Fe-4S] cluster. FdxN functions by donating electrons to the mature dinitrogenase reductase NifH, which then transfers electrons to the NifD-NifK heterohexamer (see Yang et al., 2017; Rhizobium japonicum ( Rhizobium japonicum FdxN, Carter et al., 1980; Alfalfa rhizobia ( R. meliloti (FdxN, Riedel et al., 1995; Rhodops capsulatum FdxN, Jouanneau et al., 1995), or donate electrons to NifB polypeptides for NifB synthesis (Venelandian nitrogen-fixing bacteria: Jimenez-Vincente et al., 2014), or act as an intermediate carrier for [4Fe-4S] clusters (Venelandian nitrogen-fixing bacteria: Burén et al., 2019), or any combination of these functions.

[0332] Representative examples of FdxN peptides include the following peptides identified by searching a non-redundant protein database using SEQ ID NO:47 as a query in BLASTP and showing the percentage of identity with the sequence: *Pseudomonas syringae* (… Pseudomonas syringae (WP_065835964.1, 85.87%), tentatively identified as full-moon clam disulfide feeding bacteria ( Candidatus Thiodiazotropha endolucinida (WP_069124666.1, 70.65%), Peat moss ( Uliginosibacterium sp.) TH139 (WP_101942980, 64.47%), Klebsiella micranthae (WP_049076934.1, 44.26%), Escherichia coli (WP_072048756.1, 44.26%), Rhizobium peatum ( Rhizobium leguminosarum (WP_130674512.1, 43.86%) and Flavobacterium brevicornu ( Flavobacterium alvei ) (WP_103805005.1, 28.57%).

[0333] Sequence identity and substitution Regarding the definition of polypeptides, it should be understood that higher identity percentages than those provided above will cover preferred embodiments. Therefore, where applicable, based on the minimum percentage of identity, it is preferred that the polypeptide contains at least 30%, more preferably at least 35%, more preferably at least 40%, more preferably at least 45%, more preferably at least 50%, more preferably at least 55%, more preferably at least 60%, more preferably at least 65%, more preferably at least 70%, more preferably at least 75%, more preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, and even more preferably at least 99.9% identical amino acid sequence to the associated specified SEQ ID NO.

[0334] As used in this article, "free energy" or "free energy" "G" refers to the Gibbs free energy in protein folding, which is directly related to the enthalpy and entropy of a polypeptide folding into its active structure. For negative... The presence of G and the thermodynamically favorable folding of proteins necessitate that enthalpy, entropy, or both be favorable. Methods for determining the free energy of a polypeptide are known in the art and include those described in Leman et al. (2020).

[0335] As used in this article, " G” is the energy change between the folded state and the unfolded state. G) and when amino acid substitutions are present A measure of the changes in G-folding.

[0336] As used herein, wild-type polypeptides are polypeptides that exist in nature, and examples of such polypeptides are provided herein. In one example, the wild-type NifH polypeptide has an amino acid sequence as provided in SEQ ID NO:1 or 39.

[0337] Amino acid sequence variants of the peptides defined herein can be prepared by introducing appropriate nucleotide changes into nucleic acids as defined herein. Such variants include, for example, one or more amino acid deletions, insertions, or substitutions. Combinations of deletion, insertion, and substitution mutations can be made to obtain the final construct, provided that the final peptide product has the desired properties. Preferred amino acid sequence mutants have only one, two, three, four, or fewer than 10 amino acid changes relative to a reference wild-type peptide. However, a more preferred NifH peptide of the present invention is a NifH peptide having at least one amino acid substitution (e.g., 2-5 or 3-5 substitutions) and no deletions or insertions relative to the corresponding wild-type NifH peptide.

[0338] Mutant (altered) or variant peptides can be prepared using any technique known in the art (e.g., using directed evolution or rational design strategies) (see below). Products derived from mutated / altered DNA can be readily screened using the techniques described herein to determine whether the expression of the product in plants alters the properties of the mutant or variant peptide, such as the solubility of the product in the mitochondria of plant cells, or the plant phenotype relative to the corresponding wild-type plant, for example, if the expression of the product results in an increase in yield, biomass, growth rate, vigor, nitrogen gain from biological nitrogen fixation, nitrogen use efficiency, tolerance to abiotic stresses, and / or tolerance to nutrient deficiencies relative to the corresponding wild-type plant.

[0339] When designing amino acid sequence variants, the location of the variant site and the nature of the variant will depend on the characteristics to be modified. The variant site can be modified individually or sequentially, for example by (1) first selecting with a conserved amino acid and then replacing it with a more radical selection based on the results obtained; (2) deleting the target residue or (3) inserting other residues adjacent to the target site.

[0340] The polypeptides of the present invention having at least one amino acid substitution have at least one removed amino acid residue in the wild-type polypeptide molecule, and different residues are inserted at the position of each removed amino acid, i.e., exchange or substitution of each amino acid in at least one amino acid. In the case of multiple amino acid substitutions, each new amino acid is selected independently of the others. Where it is desired to maintain a certain activity, it is preferred that no substitution is made at highly conserved amino acid positions in the relevant protein family, or only conserved substitutions are made. Examples of conserved substitutions are shown under the heading “Exemplary Substitutions” in Table 1. Preferred amino acid substitutions in the NifH polypeptides of the present invention, preferably in the AnfH polypeptides, are described herein in Tables 4, 5, 9, 10, and 11, which may not be conserved substitutions. In relevant cases, if there is a conflict between Table 1 and those tables, Tables 4, 5, 9, 10, and 11 shall prevail in determining suitable amino acid substitutions.

[0341] In one embodiment, the at least one amino acid substitute is located at an amino acid position selected from the group consisting of the following amino acid positions: 2, 5, 7, 19, 23, 24, 26 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 74, 76 to 78, 80 to 84, 102, 105, 107, 111 to 114, 116 to 118, 121 to 124, 139, 145, 147, 149, 158, 165, 166, 168, 169, 171, 179, 18 of SEQ ID NO:37. 2, 183, 188, 191, 193 to 197, 200 to 203, 205 to 211, 214, 216, 219, 223 to 226, 228 to 235, 237, 238, 241, 242, 244 to 246, 248, 249, 251 to 253, 257, 259 to 264, 266 to 271, and 273 to 275, or located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO: 37. Alternatively or additionally, in one embodiment, the at least one amino acid replaces an amino acid position located at a position selected from the group consisting of the following amino acid positions: refer to SEQ ID NO: 37. NO:39 3, 6, 8, 20, 24, 25, 27 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 75, 77 to 79, 81 to 85, 103, 106, 108, 112 to 115, 117 to 119, 122 to 125, 140, 146, 148, 150, 159, 166, 167, 169, 170, 172, 180, 18 3, 184, 189, 192, 194 to 198, 201 to 204, 206 to 212, 215, 217, 220, 224 to 227, 229 to 236, 238, 239, 242, 243, 245 to 247, 249, 250, 252 to 254, 258, 260 to 265, 267 to 272 and 274 to 276, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:39.

[0342] In one embodiment, the modified NifH polypeptide comprises at least one amino acid substitution, or two or three amino acid substitutions, or four or more amino acid substitutions selected from the group consisting of: amino acid substitutions listed in one or more of Tables 4, 5, 9, 10 or 11, or corresponding amino acid substitutions when the modified NifH polypeptide sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0343] In one embodiment, the modified NifH polypeptide comprises at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, wherein the amino acid substitution is located at an amino acid position selected from the group consisting of amino acid positions 69, 168, 200, 201, 224, 228, 234, 241, 252 and 263 of reference SEQ ID NO:37, or at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0344] In one embodiment, the at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, is selected from the group consisting of: 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H or 234C, 241R or 241A, 252M and 263E, wherein the amino acid position corresponds to the amino acid sequence provided as in SEQ ID NO:37, or is located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0345] In one embodiment, the at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, is selected from the group consisting of positions corresponding to amino acids 69, 168, 200, 201, 224, 228, 234, 252 and 263 of SEQ ID NO:37, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0346] In one embodiment, the at least one amino acid substitution, or two or three amino acid substitutions, or four or more amino acid substitutions are selected from the group consisting of: 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H or 234C, 252M, and 263E, wherein the amino acid position corresponds to the amino acid sequence provided as in SEQ ID NO:37, or is located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0347] In one embodiment, when compared with the corresponding wild-type NifH polypeptide, the modified NifH polypeptide has one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, one to twenty, one to fifteen, one to eleven, one to ten, one to nine, one to eight, one to seven, one to six, one to five, one to four, one to three, or two. Up to 20, 2 to 15, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 2 or 3, 3 to 20, 3 to 15, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 or 4, preferably 1 to 3, 1 to 4 or 1 to 5, more preferably 2 to 4 or 2 to 5, and most preferably 3 to 5 amino acid substitutions.

[0348] In one embodiment, the modified NifH polypeptide contains at least one amino acid substitution of reference SEQ ID NO:37 228V, or the same amino acid substitution located at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0349] In one embodiment, the modified NifH polypeptide contains at least one amino acid substitution of reference SEQ ID NO:37 228I, or the same amino acid substitution located at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0350] In one embodiment, the modified NifH polypeptide contains at least one amino acid substitution of reference SEQ ID NO:37 200A, or the same amino acid substitution at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0351] In one embodiment, the modified NifH polypeptide contains at least one amino acid substitution of reference SEQ ID NO:37 234H, or the same amino acid substitution located at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0352] In one embodiment, the modified NifH polypeptide contains at least two amino acid substitutions that are 200A and 228V as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0353] In one embodiment, the modified NifH polypeptide contains at least two amino acid substitutions, namely 200A and 228I, as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0354] In one embodiment, the modified NifH polypeptide contains at least two amino acid substitutions, namely 228V and 234H, as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0355] In one embodiment, the modified NifH polypeptide contains at least two amino acid substitutions, namely 228I and 234H, as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0356] In one embodiment, the modified NifH polypeptide contains at least two amino acid substitutions that are 200A, 228V, and 234H as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0357] In one embodiment, the modified NifH polypeptide contains at least two amino acid substitutions that are 200A, 228I, and 234H as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0358] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 200A, 228V or 228I, 234H and 241R as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0359] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 168I, 200A, 228I or 228V and 234H as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0360] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 69N, 168I, 200A, 228V or 228I, 234H, 252M and 263E as referenced in SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0361] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 69N, 168I, 200A, 201K, 228V or 228I, 234H, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0362] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 252M and 263E as referenced in SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0363] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 168I, 200A, 228I or 228V, 234H and 241R as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0364] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions of 69N, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0365] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions of 69N, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E as reference SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0366] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions of reference SEQ ID NO:37, namely 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0367] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 112L, 200A, 228V or 228I, 234H and 241R as reference SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0368] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions that are 112L, 168I, 200A, 228I or 228V, 234H and 241R as referenced in SEQ ID NO:37, or the same amino acid substitutions located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0369] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions of reference SEQ ID NO:37, namely 69N, 112L, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0370] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions of 69N, 112L, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E as referenced in SEQ ID NO:37, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0371] In one embodiment, the modified NifH polypeptide comprises at least two amino acid substitutions of reference SEQ ID NO:37, namely 69N, 112L, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E, or the same amino acid substitutions at the corresponding amino acid positions when the modified NifH sequence is aligned with SEQ ID NO:37 and / or SEQ ID NO:39.

[0372] In one embodiment, the modified NifH polypeptide comprises amino acids 200A, 228V, and 234H as referenced in SEQ ID NO:37, or the same amino acids located at the corresponding amino acid positions when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:11539.

[0373] In one embodiment, the modified NifH polypeptide comprises one or more of the following motifs: YGKGGIGKSTTXQN (SEQ ID NO:61), IXGCDPKAD (SEQ ID NO:62), CXESGGPEPGVGCAGRG (SEQ ID NO:63), DVLGDVVCGGFAMP (SEQ ID NO:43), VXSGEMMAXYAANNI (SEQ ID NO:64), and CNSRXXD (motif VII, SEQ ID NO:65), preferably at least DVLGDVVCGGFAMP (SEQ ID NO:43), wherein each X independently represents any amino acid.

[0374] In one embodiment, the modified NifH polypeptide comprises one or more of the following motifs: YGKGGIGKSTTXQNT (SEQ ID NO:40), IHGCDPKAD (SEQ ID NO:41), CVESGGPEPGVGCAGRG (SEQ ID NO:42), DVLGDVVCGGFAMP (SEQ ID NO:43), VASGEMMAXYAANNI (SEQ ID NO:44), QSGVR (SEQ ID NO:45), and CNSRXVD (SEQ ID NO:46), preferably at least DVLGDVVCGGFAMP (SEQ ID NO:43), wherein each X independently represents any amino acid.

[0375] In one embodiment, the modified NifH polypeptide has 12, 13, 14, 15, or all of the following amino acids: 4K, 22T, 37H, 52G, 60D, 63R, 108L, 109M, 142G, 151A, 174Q, 189V, 198E, 199F, 222F, and 247I as per SEQ ID NO:37, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0376] In one embodiment, the modified NifH polypeptide has 130, 131, 132, 133, 134, 135, 136, or all 137 of the following amino acids: 3R, 4K, 6A, 8Y, 9G, 10K, 11G, 12G, 13I, 14G, 15K, 16S, 17T, 18T, 20Q, 21N, 22T, 25A, 36I, 37H, 38G, 39C, 40D, 41P, 42K, 43A, 44D, 46T, 47R, 50L, 52G, 55Q, 60D, 63R, 75V, 79G, 85C, 86V, 87E, 88S, 89G ... G, 90G, 91P, 92E, 93P, 94G, 95V, 96G, 97C, 98A, 99G, 100R, 101G, 103I, 104T, 106I, 108L, 109M, 110E, 115Y, 119L, 120D, 125D, 126V, 127L, 128G, 129D, 130V, 131V, 132C, 133G, 134G, 135F, 136A, 137M, 13 8P, 140R, 142G, 143K, 144A, 146E, 148Y, 150V, 151A, 152S, 153G, 154E, 155M, 156M, 157A, 159Y, 160 A, 161A, 162N, 163N, 164I, 167G, 170K, 172A, 174Q, 175S, 176G, 177V, 178R, 180G, 181G, 184C, 185N, 186S, 187R, 189V, 190D, 192E, 198E, 199F, 204G, 212P, 213R, 215N, 217V, 218Q, 220A, 221E, 222F, 227V, 236Q, 239E, 240Y, 243L, 247I, 250N, 254V, 255I, 256P, 258P, 265E, and 272G, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0377] In one embodiment, the modified NifH polypeptide has one or two of the following amino acids as referenced in SEQ ID NO:37: 141D and 173K, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39.

[0378] In one embodiment, the amino acid sequence of the modified NifH polypeptide is at least 60% identical to the amino acid sequence provided in SEQ ID NO:37 and / or SEQ ID NO:39, preferably at least 70% or at least 80%, more preferably at least 90%, and most preferably at least 95%.

[0379] In one embodiment, the modified NifH polypeptide is a modified AnfH polypeptide.

[0380] In one embodiment, the amino acid sequence of the modified AnfH polypeptide is at least 70% identical to the amino acid sequence provided in SEQ ID NO:37, preferably at least 80% identical, more preferably at least 90% identical, and most preferably at least 95% identical.

[0381] In one embodiment, the modified NifH polypeptide comprises at least amino acids 2-275 of the amino acid sequence provided in SEQ ID NO:78, or comprises SEQ ID NO:78.

[0382] In one embodiment, the modified peptides of the present invention have one, two, three, or four conserved amino acid changes compared to naturally occurring (wild-type) peptides. Detailed information on these conserved amino acid changes is provided in Table 1. In a preferred embodiment, as described in Examples 2, 8, 9, and 18 herein, the changes are not in one or more highly conserved motifs or domains among the different peptides of the present invention. As those skilled in the art will appreciate, it is reasonable to predict that such minor changes will not alter the activity of the peptide when expressed in recombinant cells.

[0383] Table 1. Exemplary conservative substitutions. The primary amino acid sequence of the polypeptide of the present invention can be used to design its variants / mutants or modified forms based on comparisons with closely related polypeptides. Those skilled in the art will understand that highly conserved residues in closely related proteins are less likely to be altered, especially substituted with non-conserved residues, and activity is maintained compared to less conserved residues (see above). A more stringent test for identifying conserved amino acid residues is compared to more distantly related polypeptides with the same function. Highly conserved residues should be maintained to preserve function, while non-conserved residues are more readily substituted or deleted without loss of function.

[0384] Also included within the scope of this invention are the polypeptides of this invention, which are differentially modified during or after cellular synthesis, for example by glycosylation, acetylation, phosphorylation, or proteolytic cleavage.

[0385] Introduction of mitochondrial proteins in plants Almost all mitochondrial proteins are nuclear-encoded and translated in the cytosol, thus requiring translocation of these proteins into the mitochondria. A signal sequence within the polypeptide directs it to one of four distinct intramitochondrial locations: the outer membrane (OM), the intermembrane space (IS), the inner membrane (IM), or the matrix (MM). These signal sequences differ in their biochemical properties and are directed via at least four different import pathways that guide the polypeptide to one or more of these four locations (Chacinska et al., 2009). These four pathways are: (1) the general import pathway, also known as the “classical” pre-sequence pathway, which directs the polypeptide to the MM, IS, or IM; (2) the carrier import pathway for transport to the IM; (3) the mitochondrial intermembrane space (MIA) assembly pathway; and (4) the sorting and assembly machine (SAM) pathway for transporting the polypeptide to the OM. The general import pathway imports polypeptides with a cleavable pre-sequence (also known as a signal sequence). These polypeptides may also have a hydrophobic sorting signal (HSS). The carrier import pathway imports polypeptides with an internal pre-sequence-like signal and a hydrophobic region. The MIA pathway delivers peptides containing biscysteine ​​residues. The SAM pathway delivers peptides containing β-signaling and putative TOM20 signaling. All of these pathways utilize outer membrane translocases (TOMs), and both the first and second pathways also utilize the TIM23 translocase of the intermembrane complex. Only the first pathway uses matrix-processing peptidases (MPPs).

[0386] A common feature of all mitochondrial-targeting peptides is the presence of at least one domain within the peptide that guides transport to the correct location. Among these, the "classic" N-terminal pre-sequence domain, cleaved by MPP in the matrix, is best studied (Murcha et al., 2004). It is estimated that approximately 70% of plant and animal mitochondrial proteins possess a cleavable pre-sequence, but both internal and C-terminal signaling sequences have also been found (reviewed in Pfanner and Geissler (2001), Schleiff and Soll (2000)). In Arabidopsis thaliana (… ArabidopsisIn these pre-sequences, the length ranges from 11 to 109 amino acid residues, with an average length of 50 amino acid residues. Although there is no consensus sequence that fully defines the pre-sequence of the first pathway, it tends to contain a high proportion of hydrophobic and positively charged amino acids. Another characteristic is its ability to form an amphiphilic α-helix, usually starting within the first 10 amino acid residues (Roise et al., 1986). These domains are rich in hydrophobic (Ala, Leu, Phe, Val), hydroxylated (Ser, Thr), and positively charged (Arg, Lys) amino acid residues, and lack acidic amino acids. In a large number of mitochondrial proteins, serine (16-17%) and alanine (12-13%) are significantly overexpressed in mitochondrial signal peptides, and arginine is abundant (12%). For most pre-sequences, the MPP cleavage site is defined by the presence of conserved arginine residues, usually located at position P2 (approximately -2 amino acids from the cleavable bond), or in most other cases at position P3 (Huang et al., 2009).

[0387] Mitochondrial presequences interact with the Tom20 receptor via hydrophobic residues. Studies have shown that the hydrophobic surface of the α-helix facilitates recognition of the Tom20 component of the Tom20-to-M-20 receptor peptide by the Tom20-to-M-20 receptor complex, while the positively charged portion is recognized by the Tom22 subunit (Abe et al., 2000). Finally, most presequences guide the transport of peptides associated with Hsp70, and therefore almost all plant presequences contain at least one Hsp70 chaperone binding motif (Zhang and Glaser, 2002). The molecular chaperone Hsp70 is involved in protein folding, prevents protein aggregation, and acts as a molecular motor, pulling the precursor across the mitochondrial membrane. The transmembrane potential (TMP) of the inner membrane... ψ (approximately 100 mV, negative inside) also drives the translocation of positively charged presequences through electrophoresis.

[0388] Most proteins with cleavable pre-sequences reach the mitochondrial matrix via a general import pathway utilizing transporters of the outer membrane (TOM) complex and the inner membrane 23 complex (TIM23). However, some proteins with cleavable pre-sequences can assemble in the inner membrane (Murcha et al., 2004) or the intermembrane space if they also contain a hydrophobic separation signal (HSS) (Glick et al., 1992). Instances of matrix-localized proteins whose pre-sequences are uncleavable are rare. In Arabidopsis, only glutamate dehydrogenases have been found in the matrix with unprocessed, full-length pre-sequences (Huang et al., 2009).

[0389] For proteins that are not matrix-targeted, various internal, cleavable localization signals are employed. These are often associated with specific transport pathways and are additionally tailored to specific protein species. In plants, no studies have yet determined precisely what constitutes the internal signaling sequence of intermembranous space proteins. However, motifs with biscysteine ​​residues appear to be associated with transport via the mitochondrial intermembranous space assembly pathway (MIA) (Carrie et al., 2010; Darshi et al., 2012). Finally, cleavable internal sequences are also utilized by proteins reaching the inner membrane via a carrier pathway that uses TOM and TIM22 devices to insert into proteins with multiple transmembrane regions (Kerscher et al., 1997; Sirrenberg et al., 1996). These sequences typically contain hydrophobic regions followed by a proto-sequence-like internal sequence and are therefore similar to N-terminal proto-sequences, but differ in their internal location within their homologous proteins.

[0390] In photosynthetic organisms, nuclear-encoded mitochondrial proteins need to be distinguished between chloroplasts and mitochondria, despite the many similarities between these two organelles and their proteomes. α-helices, which are predominantly found in mitochondrial pre-sequences, are typically absent in chloroplast pre-sequences (Zhang and Glaser, 2002), which tend to be more unstructured and exhibit high β-sheet domain structures (Bruce, 2001).

[0391] In plants, MPP is anchored to the inner membrane-bound Cytbc1 complex, although the active MPP site is located in a matrix-facing position, and the functions of the two proteins are independent (Glaser and Dessi, 1999).

[0392] Mitochondrial targeting peptides As used herein, the term “mitochondrial targeting peptide” or “MTP” means an amino acid sequence containing at least 10 amino acids and preferably of a length of 10 to about 80 amino acid residues that guides a target protein to the mitochondria, and said amino acid sequence can be heterologously used in MTP-target protein translational fusion to guide a selected Nif peptide to the mitochondria.

[0393] MTPs typically contain a methionine residue, the translation initiator of the polypeptide from which they originate, at their N-terminus. MTPs are translationally fused to the Nif polypeptide or "target protein" via a peptide bond to a Met residue corresponding to the Met initiator of the target protein, or the Met residue may be omitted and the peptide bond may be directly fused to the second amino acid residue of the target protein in the wild type. MTPs are typically rich in basic and hydroxylated amino acids and generally lack acidic amino acids or extended hydrophobic extensions. MTPs can form amphiphilic helices.

[0394] To avoid being limited by theory, MTPs typically contain an uptake-targeting sequence that binds to a receptor on the outer mitochondrial membrane. Upon binding to the outer membrane, the fusion peptide preferably undergoes membrane translocation to transfer the channel protein and crosses the mitochondrial bilayer to reach the mitochondrial matrix (MM). The uptake-targeting sequence is then typically cleaved and the mature fusion protein folds.

[0395] MTPs may contain additional signals that subsequently target the protein to different regions of the mitochondria, such as the mitochondrial matrix (MM). In one embodiment, the uptake targeting sequence is a matrix targeting sequence.

[0396] When translatively fused with a Nif polypeptide, the MTP can be cleavable or non-cleavable. Therefore, in one embodiment, the MTP-Nif fusion polypeptide (such as the MTP-modified NifH polypeptide or MTP-NifH fusion polypeptide of the present invention, preferably the MTP-modified AnfH polypeptide or MTP-AnfH fusion polypeptide) is at least partially cleaved in the mitochondria of plant cells. In this regard, the phrase "at least partially cleaved" refers to a detectable amount of cleavage of the MTP-modified NifH polypeptide or MTP-Nif fusion polypeptide when expressed in plant cells. In one embodiment, at least 50% of the MTP-Nif fusion polypeptide produced in the cell is cleaved within the MTP sequence, preferably at least 75%, more preferably at least 90%. In an alternative, less preferred embodiment, less than 50% of the MTP-Nif fusion polypeptide is cleaved in the cell; for example, the MTP is not cleaved. In one embodiment, the MTP does not contain a cleavage site for the MPP. The MTP preferably contains a cleavage site for the MPP. During cleavage, the N-terminal portion of the resulting processed product (i.e., the mature NP or "cleavage product") may contain one or more C-terminal amino acids of the MTP, referred herein as a "scar sequence" or "scar peptide," or the N-terminal portion may not contain any C-terminal amino acids of the MTP. In the latter case, the cleavage occurs immediately adjacent to the MTP sequence, for example at the junction between the MTP and the NifH sequence. When present, the scar sequence is preferably 1 to 45 amino acids long, more preferably 1 to 20 amino acids, even more preferably 1 to 12 amino acids, or most preferably 4 to 12 amino acids. The scar sequence can be 4, 5, 6, 7, 8, 9, 10, 11, or 12 amino acids long, without counting the amino acids in the NifH sequence. Alternatively, the cleavage site may be located within the fusion peptide, such that the entire MTP sequence is cleaved away; for example, the linker may contain the cleavage sequence.

[0397] Natural mitochondrial-targeting peptides are located at the N-terminus of precursor proteins, and the N-terminal portion is typically cleaved during or after introduction into the mitochondria. Cleavage is usually catalyzed by a universal matrix processing protease (MPP), which, in plants, integrates into the bc1 complex of the respiratory chain. This protease recognizes cleavage sites for nearly 1000 precursor proteins with a broad range of amino acid sequences that are largely unconserved. In one embodiment, the MTP contains the protease cleavage site of the MPP. In another embodiment, the processed product is produced by the MPP cleaving the fusion protein within or immediately after the MTP. In this context, the phrase "immediately after" or "adjacent to" means that no residual amino acids remain in the MTP after cleavage by the MPP and fusion with the Nif polypeptide. Therefore, in the case where the fusion polypeptide is cleaved "immediately after" the MTP, the MPP cleavage site is immediately adjacent to the C-terminal amino acid of the MTP.

[0398] The term "cleavage product," or as used herein, in the context of an MTP-modified Nif or MTP-Nif fusion polypeptide, preferably a modified NifH polypeptide or NifH dimer polypeptide as described herein, and more preferably a modified AnfH polypeptide or AnfH dimer polypeptide, refers to a polypeptide generated by cleavage by a protease within or immediately following the amino acid sequence of the MTP. In this respect, cleavage products of MTP fusion polypeptides can be obtained by cleavage with MPP. The cleavage product may retain one or more amino acids from the MTP after cleavage (i.e., a scar peptide), or the cleavage product may retain no amino acids remaining in the MTP after cleavage. In one embodiment, the cleavage product generated from the MTP-modified Nif or MTP-Nif fusion polypeptide of the present invention comprises at least 95% or all of the amino acids present in the Nif polypeptide sequence.

[0399] In one embodiment, the MTP is not cut; preferably, the MTP is cut.

[0400] Suitable MTPs that can be used in the context of this invention include, but are not limited to, peptides having a general structure as defined by von Heijne (1986) or by Roise and Schatz (1988). Non-limiting examples of MTPs are mitochondrial targeting peptides as defined in Table I of von Heijne (1986) or disclosed herein.

[0401] In one embodiment, the MTP is the F1-ATPase γ subunit (MTP-FAγ). A suitable example of a FAγ MTP is from Arabidopsis thaliana (…). A. thalianaThe MTP-FAγ is derived from the FAγ MTP (Lee et al., 2012). In a preferred embodiment, the length of MTP-FAγ is less than 77 amino acids. For example, the length of MTP-FAγ can be about 51 amino acids, with MMP cleavage leaving 9 MTP residues at the N-terminus of the fusion polypeptide. The scar sequence can also contain the two amino acid sequence GG between the C-terminal amino acid of the MTP and the NifH sequence, as a result of the GoldenGate cloning method.

[0402] Technicians will understand that software exists for predicting mitochondrial proteins and their target sequences, such as MitoProtII, PSORT, TargetP, and NNPSL.

[0403] MitoProtII is a procedure that predicts mitochondrial localization of sequences based on several physicochemical parameters, such as the amino acid composition of the N-terminal portion or the highest total hydrophobicity of the 17-residue window. PSORT is a procedure that predicts subcellular localization based on various sequence-derived features, such as the presence of sequence motifs and amino acid composition. TargetP predicts subcellular localization of eukaryotic proteins based on the predicted presence of any N-terminal pre-sequence: chloroplast transport peptides, mitochondrial targeting peptides, or secretory pathway signaling peptides. Utilizing earlier binary predictors, SignalP, and ChloroP, TargetP requires the N-terminal sequence as input to a two-layer artificial neural network (ANN). Potential cleavage sites can also be predicted for sequences containing an N-terminal pre-sequence. NNPSL is another ANN-based method that assigns one of four subcellular localizations (cytoplasmic, extracellular, nuclear, and mitochondrial) to the query sequence using amino acid composition.

[0404] Based on conventional methods and the methods disclosed herein, those skilled in the art can easily determine whether the selected MTP will target the fusion peptide to the mitochondrial matrix.

[0405] In some embodiments of the invention, using multiple tandem copies of a selected MTP may be useful. The coding sequence of the replicated or multiplied targeting peptide can be obtained from existing MTPs through genetic engineering. The amount of MTP can be measured by cell grading, followed, for example, by quantitative immunoblotting analysis. Therefore, in this invention, the term "mitochondrial targeting peptide" or "MTP" encompasses one or more copies of a single amino acid peptide that guides a target Nif protein to the mitochondria. In a preferred embodiment, the MTP comprises two copies of the selected MTP. In another embodiment, the MTP comprises three copies of the selected MTP. In yet another embodiment, the MTP comprises four or more copies of the selected MTP.

[0406] Technicians will understand that the MTP sequence is not limited to the natural MTP sequence, but may contain amino acid substitutions, deletions and / or insertions relative to the naturally occurring MTP, provided that the sequence variant still serves a mitochondrial targeting function.

[0407] Technicians will understand that, as a result of the cloning strategy, MTPs can have amino acids side-attached at their N-terminus or C-terminus and can act as linkers. These additional amino acids can be considered as part of the formation of the MTP.

[0408] Those skilled in the art will also understand that MTPs can be fused to oligopeptide linkers and / or tags (such as epitope tags) at the N-terminus or C-terminus. In a preferred embodiment, one or more of the Nif fusion peptides of the present invention produced in plant cells lack the added epitope tag compared to the corresponding wild-type Nif peptide.

[0409] Mitochondrial Targeting Peptide (MTP)-Nif Fusion Peptide This invention includes mitochondrial targeting peptide (MTP)-Nif fusion polypeptides, also referred to herein as "encoded polypeptides," because the MTP-Nif fusion polypeptide is an instantaneous translation product encoded by the exogenous polynucleotide of this invention, and its cleaved polypeptide product. Specifically, this invention includes MTP-modified NifH polypeptides and MTP-NifH fusion polypeptides, preferably MTP-modified AnfH polypeptides or MTP-AnfH dimers, and their cleaved polypeptide products. When the MTP-modified Nif polypeptides or MTP-Nif fusion polypeptides (encoded polypeptides) of this invention are expressed in plant cells, the MTP-modified Nif polypeptides or MTP-Nif fusion polypeptides and / or cleaved polypeptide products target the mitochondrial matrix (MM). Preferably, the fusion polypeptide confers nitrogenase reductase and / or nitrogenase activity to plant cells, or the same activity conferred by the corresponding wild-type Nif polypeptide in bacteria.

[0410] As used herein, the term "fusion polypeptide" means a polypeptide comprising two or more polypeptide domains covalently linked by peptide bonds, one of which is a Nif sequence, preferably having at least one amino acid-substituted NifH sequence, or two NifH sequences linked by an oligopeptide linker, more preferably having at least one amino acid-substituted AnfH sequence, or two AnfH sequences linked by an oligopeptide linker. The fusion polypeptide is encoded by the chimeric polynucleotide of the present invention into a single polypeptide chain. In one embodiment, the fusion polypeptide of the present invention comprises a mitochondrial targeting peptide (MTP) and a Nif polypeptide (NP). In this embodiment, the C-terminus of the MTP is translatorily fused to the N-terminus of the NP via a peptide bond. In an alternative embodiment, the fusion polypeptide of the present invention comprises the C-terminal portions of the MTP and the NP, wherein the C-terminal portions are generated by cleavage of the MTP by an MPP. Such a C-terminal portion of the MTP is referred to herein as a "scar" sequence. In this embodiment, the C-terminal amino acid of the C-terminal portion of the MTP is translatorily fused to the N-terminal amino acid of the NP via a peptide bond. In these embodiments, the fusion polypeptide may contain one or more additional amino acids (such as a GlyGly sequence) between the MTP and NP, and / or an added methionine as a translation initiation amino acid. In one embodiment, the cell, plant, or part thereof of the present invention contains a fusion polypeptide comprising two NifH polypeptides, preferably AnfH polypeptides, fused translationally via a linker.

[0411] As used herein, the term "translationally fused at the N-terminus" means that the C-terminus of an MTP polypeptide or linker polypeptide is covalently linked to the N-terminus of an NP via a peptide bond, thereby becoming a fusion polypeptide. In one embodiment, the NP does not contain its native translation initiation methionine (Met) residue or either of its two N-terminal Met residues, relative to the corresponding wild-type NP. In an alternative embodiment, the NP contains the translation initiation Met of a wild-type NP polypeptide (e.g., NifH) or one or both of its two N-terminal Met residues.

[0412] Such polypeptides are typically generated by expressing a chimeric protein-coding region, wherein the reading frame of the nucleotide encoding the MTP is in-frame linked to the reading frame of the nucleotide encoding the NP. Those skilled in the art will understand that the C-terminal amino acid of the MTP can be translatedally fused to the N-terminal amino acid of the NP either in the absence of a linker or via a linker of one or more amino acid residues (e.g., 1-5 amino acid residues). Such a linker can also be considered part of the MTP. After expression of the protein-coding region, the MTP can be cleaved in the MM of plant cells, and such cleavage (if it occurs) is included in the concept of generating fusion polypeptides of the present invention.

[0413] The fusion peptide or processed Nif peptide preferably possesses functional Nif activity. In a preferred embodiment, the activity is similar to that of the corresponding wild-type Nif peptide. The functional activity of the fusion peptide or processed Nif peptide can be determined in bacterial and biochemical complementation assays. In a preferred embodiment, the activity of the fusion peptide or processed Nif peptide is about 70-100% of the wild-type Nif activity. Nif peptides without Nif function still have practical applications, for example, as research tools to test the expression levels of gene constructs or their association with other Nif peptides.

[0414] The cells, plants, or portions thereof of the present invention may comprise fusion polypeptides having more than one MTP and / or more than one NP, for example, the fusion polypeptide may comprise an MTP, a NifD polypeptide, and a NifK polypeptide. The fusion polypeptide may also comprise an oligopeptide linker, for example, linking two NPs. Preferably, the linker has sufficient length to allow two or more functional domains, for example, two NPs (such as NifH and NifH (homodimers), NifD and NifK, or NifE and NifN) to associate in a functional conformation of the plant cell. In a preferred embodiment, the NifD polypeptide is an AnfD polypeptide, and the NifK polypeptide is an AnfK polypeptide. For AnfD-linker-AnfK fusion polypeptides, the length of such linkers may be from 8 to 50 amino acid residues, preferably about 25-35 amino acids, more preferably about 30 amino acid residues, or about 26 amino acid residues. The fusion polypeptide can be obtained by conventional methods, for example, by gene expression encoding a polynucleotide sequence of the fusion polypeptide in suitable cells.

[0415] In one embodiment, the polypeptide of the present invention is a substantially purified polypeptide. As used herein, "substantially purified polypeptide" means a polypeptide that is substantially free of, for example, components that normally associate with polypeptides in cells (e.g., lipids, nucleic acids, carbohydrates). Preferably, the substantially purified polypeptide is at least 90% free of said components.

[0416] The plant cells, transgenic plants, and portions thereof of the present invention contain exogenous polynucleotides encoding the polypeptides of the present invention. The polypeptides of the present invention are not naturally present in plant cells, particularly not in the mitochondria of plant cells, and therefore the polynucleotides encoding the polypeptides may be referred to herein as exogenous polynucleotides because they are not naturally present in plant cells but have been introduced into plant cells or progenitor cells. Thus, it can be said that the cells, plants, and plant portions of the present invention that produce the polypeptides of the present invention generate recombinant polypeptides. In the context of polypeptides, the term "recombinant" refers to a polypeptide encoded by an exogenous polynucleotide produced by a cell that has been introduced into the cell or progenitor cell via recombinant DNA or RNA technology (e.g., transformation). Typically, plant cells, plants, or plant portions contain non-endogenous genes that produce a certain amount of polypeptide at least at some point in the life cycle of the plant cell or plant. Preferably, the exogenous polynucleotide is integrated into the nuclear genome of the plant cell and / or transcribed in the nucleus.

[0417] connector As used herein, in the context of peptides, the term "linker" or "oligopeptide linker" refers to one or more amino acids covalently linked to two or more functional domains (e.g., MTPs and NPs), two NPs, NPs, and a tag. As used herein, "linker" does not include amino acids of the Nif sequence (e.g., NifH or AnfH sequences), which are also present in the corresponding wild-type Nif sequences. "Linker" also does not include amino acids of the MTP sequence, if present, which are identical to the naturally occurring MTP sequence. Amino acids are covalently linked by peptide bonds, both within the linker and between the linker and the functional domains. Linkers allow free movement of one functional domain relative to another without significantly detrimental to the function of the two or more functional domains. Linkers can help facilitate proper folding and function of one or both functional domains, or, as described herein, contribute to increased solubility of the fusion peptide. Those skilled in the art will understand that linker size can be determined empirically or modeled based on protein folding information.

[0418] Linkers can contain cleavage sites for proteases such as MPP. Such linkers can also be considered part of MTP.

[0419] Technicians will understand that the C-terminus of the MTP can be translatorily fused with the N-terminal amino acid of the NP in the absence of a linker or through a linker of one or more amino acid residues (e.g., 1-5 amino acid residues).

[0420] In one embodiment, the linker comprises at least 1 amino acid, at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, at least 10 amino acids, at least 12 amino acids, at least 14 amino acids, at least 16 amino acids, at least 18 amino acids, at least 20 amino acids, at least 25 amino acids, at least 30 amino acids, at least 35 amino acids, at least 40 amino acids, at least 45 amino acids, at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, or about 100 amino acids. In one embodiment, the maximum size of the linker is 100 amino acids, preferably 60 amino acids, and more preferably 40 amino acids.

[0421] In some embodiments, the linker will allow one functional domain to move relative to another to increase the stability of the fusion peptide. If desired, the linker may encompass repeats of polyglycine or combinations of glycine, proline, and alanine residues.

[0422] The linker used to connect two Nif peptides (e.g., NifH-linker-NifH, preferably AnfH-linker-AnfH) is preferably selected based on several criteria to determine the number and sequence of amino acids in the linker. These are: the absence of cysteine ​​residues to avoid the formation of unwanted disulfide bonds; few or preferably no charged residues (Glu, Asp, Arg, Lys) to reduce the likelihood of unwanted surface salt bridge interactions; few or no hydrophobic residues (Phe, Trp, Tyr, Met, Val, Ile, Leu) because such residues can promote the tendency to penetrate the peptide surface; and the lack of amino acids that can be post-translationally modified. In this context, "few charged residues" means less than 10% of the amino acid residues in the linker, and "few hydrophobic residues" means less than 15% of the amino acid residues in the linker.

[0423] In one embodiment, the connector does not contain cysteine ​​residues.

[0424] In one embodiment, the linker comprises four, three, two, one, or no charged residues in an increasingly preferred order. Preferably, the linker comprises a total of four, three, two, one, or no glutamic acid, aspartic acid, arginine, and lysine residues.

[0425] In one embodiment, the linker comprises four, three, two, one, or no hydrophobic residues. Preferably, the linker comprises a total of four, three, two, one, or no phenylalanine, tryptophan, tyrosine, methionine, valine, isoleucine, and leucine residues.

[0426] In one embodiment, at least 70%, or at least 80%, or at least 90% of the linkers contain residues selected from threonine, serine, glycine, and alanine.

[0427] The use of oligopeptide linkers in modified peptides was reviewed in Chen et al. (2013) and Zhang et al. (2009).

[0428] Label In one particular embodiment, the fusion polypeptide comprises at least one tag suitable for detecting or purifying the fusion polypeptide or a processed product thereof. The tag typically binds to a C-terminal or N-terminal domain of the fusion polypeptide, or preferably serves as part of a linker in a NifH-linker-NifH fusion polypeptide. In a preferred embodiment, the tag is attached to the N-terminus of a Nif sequence. The tag is typically a peptide or amino acid sequence capable of binding with high affinity to one or more ligands, such as one or more ligands or antibodies of an affinity matrix (e.g., chromatographic support or beads). Those skilled in the art will understand that once the MTP is cleaved after introduction into the mitochondria, the tag should be located in the fusion protein at a position that does not cause the tag to be removed from the NP. Further, the tag should not interfere with the mitochondrial introduction mechanism. In a preferred embodiment, the polynucleotide of the present invention encodes a fusion polypeptide comprising, in N-terminal to C-terminal order, an N-terminal MTP, a Nif polypeptide, and a detection / purification tag, such as an N-terminal MTP, a NifH polypeptide, and a detection / purification tag, optionally followed by a second NifH polypeptide. In an alternative embodiment, the fusion peptide comprises, in N-terminal to C-terminal order, an N-terminal MTP, a detection / purification tag, and a Nif peptide, such as an N-terminal MTP, a detection / purification tag, such as an N-terminal MTP, a NifH peptide, and a detection / purification tag, and a NifH peptide, optionally followed by a linker and a second NifH peptide.

[0429] Further illustrative, non-limiting examples of tags used for the detection, separation, or purification of fusion peptides or their processed products include human influenza hemagglutinin (HA) tags, histidine tags containing, for example, 6 or 8 histidine residues, fluorescent tags such as fluorescein, resourfin and its derivatives, Arg-tags, FLAG-tags, Strep-tags, epitopes that can be recognized by antibodies, such as c-myc-tags (recognized by anti-c-myc antibodies), SBP-tags, S-tags, calmodulin-binding peptides, cellulose-binding domains, chitin-binding domains, glutathione S-transferase-tags, maltose-binding proteins, NusA, TrxA, DsbA, Avi tags, etc.

[0430] Translational fusion involving Nif peptides As reported in the scientific literature, several Nif peptides have undergone translational fusion. These are summarized in Table 2 and in the review by Burén and Rubio (2018). Most of these involve the artificial addition of epitopes or binding domains (such as histidine or Strep tags) to the protein for detection and purification purposes, and only a few have been expressed in plant cells. There are some reports of naturally occurring fusions between Nif peptides in bacteria. For assays in bacterial hosts, His tags of varying lengths (7–10 histidines) were added to both full-length and truncated versions of NifD (Christiansen et al., 1998), NifE (Goodwin et al., 1998), NifM (Gavini et al., 2006), and NifB (Fay et al., 2015). In each case, the modified Nif peptides retained Nif function, as demonstrated in bacterial or in vitro nitrogenase remodeling assays.

[0431] Table 2. Summary of gene fusions of Nif peptides as reported in the literature. Thiel et al. (1995) identified 29 naturally occurring deletions of nucleotides in the cyanobacterium Anabaena variegata, and consequently, the deletion of... NifE and NifN The nine amino acids in the intergenic spacer region between genes and NifE The stop codon was removed. The deletion resulted in a NifE-NifN polypeptide fusion that retained at least some of the nitrogenase functions of both the NifE and NifN polypeptides. The NifE-NifN fusion polypeptide also has 19 other amino acid substitutions in the fusion linker region, which may affect Nif function, but the mechanism is unknown. The fusion gene was expressed only under strictly anaerobic conditions. Whether the activity was reduced compared to the non-fusion gene has not been reported.

[0432] Suh et al. (2003) used methods including NifD The stop codon and NifK The deletion of the translation start codon (ATG) in the chromosome of *Venelandia diffusa* NifD and NifK Artificial links were established between the genes, forming a vector designated pBG1404. The deletion resulted in a net loss of three amino acids and seven amino acid substitutions in amino acid 2–10 of the NifK polypeptide. Compared to the corresponding wild-type bacterium, the growth of *Zytophthora venereum* host cells containing pBG1404 was impaired in low-nitrogen media.

[0433] Wiig et al. (2011) used [a specific ingredient] found in Clostridium pasteurellum. NifN and NifB A naturally occurring translational fusion between genes was identified, and its function in bacterial and biochemical complementation assays was determined to include NifN and NifB activities. This fusion was direct without any peptide linkers; that is, the C-terminus of NifN was directly covalently linked to the N-terminus of NifB.

[0434] In yeast and plant cells, translational fusions have been used to direct nuclear-encoded proteins to the mitochondrial matrix. In yeast expression assays, translational fusions of mitochondrial-targeting peptides (MTPs) and several Nif peptides (NifH, NifM, NifS, and NifU) have shown functionality under aerobic conditions (Lopez-Torrejon et al., 2016). Epitope fusions (FLAG and HIS) have also shown functionality with NifH, NifM, NifS, and NifU, although these fusions are designed to localize within the yeast cytoplasm and are only functional when yeast is grown under anaerobic conditions. Burén et al. (2017b) showed that a mitochondrial matrix-targeting version of a soluble variant of NifB is functional in an in vitro complementation assay when re-isolated from yeast mitochondria. This version of NifB includes an N-terminal MTP, a truncated variant of NifB (lacking a NifX-like domain), and a C-terminal 10xHis epitope tag. Numerous MTP-Nif fusions were also generated in yeast expression assays. However, the large collection of this co-expressed protein failed to show activity in yeast (Burén et al., 2017b).

[0435] The MTP from the CPN-60 gene was fused to the N-terminus of NifH, NifM, NifS, and NifU, and its functionality was demonstrated by in vitro complementation assays when the Fe protein was re-isolated from plants grown under reduced oxygen stress at 10% oxygen (US2016 / 0304842).

[0436] Polynucleotides The terms “polynucleotide” and “nucleic acid” are used interchangeably in this document. They refer to a polymeric form of nucleotides of any length, deoxyribonucleotides or ribonucleotides or analogues thereof. Polynucleotides as defined herein can be genomic, cDNA, semi-synthetic or synthetic in origin, single-stranded or preferably double-stranded, and due to their origin or manipulation: (1) associated with a polynucleotide in nature (e.g., not containing a natural promoter coding sequence). Nif(1) The polynucleotide is wholly or partially non-associated with a polynucleotide, (2) linked to a polynucleotide other than the one to which it is linked in nature (e.g., a Nif polynucleotide linked to an MTP-encoding nucleotide sequence and / or a non-natural promoter-encoding sequence), or (3) not found in nature (e.g., the polynucleotide encoding the MTP-Nif fusion polypeptide of the present invention). The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, multiple loci (loci / locus) defined by ligation analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, chimeric DNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain modified nucleotides (e.g., methylated nucleotides) and nucleotide analogs. If present, modifications to the nucleotide structure may be conferred before or after polymer assembly. The sequence of a nucleotide may be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with labeled components.

[0437] The "isolated polynucleotide" is substantially free of components typically linked to the polynucleotide (e.g., regulatory sequences) or associated components. Therefore, the isolated polynucleotide is substantially free of other cellular material or, when produced by recombinant technology, substantially free of culture medium or, when chemically synthesized, substantially free of chemical precursors or other chemicals. Preferably, the isolated polynucleotide is at least 60%, more preferably at least 75%, and even more preferably at least 90% free of said components.

[0438] As used herein, the phrase "exogenous polynucleotide" refers to a polynucleotide having a sequence derived from outside the cell or organism in which the exogenous polynucleotide is present. Therefore, exogenous polynucleotides are not naturally present in cells or organisms. They can be chimeric polynucleotides formed by linking two or more non-naturally present nucleic acid sequences together. For example, the promoter and the Nif protein-coding region are heterologous. Furthermore, the Nif protein-coding region may have been codon-modified relative to the Nif sequence encoding the protein in bacteria. When a Nif polypeptide contains at least one amino acid substitution, the polynucleotide encoding said Nif polypeptide must be an exogenous polynucleotide.

[0439] As used herein, the term "gene" is understood in its broadest context and includes a deoxyribonucleotide sequence comprising the transcribed and (if translated) protein-coding regions of a structural gene, and includes sequences located at the 5' and 3' ends, at a distance of at least about 2 kb from either end, that are involved in the expression of the gene. In this respect, a gene may include control signals naturally associated with a given gene, such as promoters, enhancers, translation and transcription termination and / or polyadenylation signals, or heterologous control signals, in which case the gene is referred to as a "chimeric gene." A sequence located at the 5' end of a protein-coding region and present on mRNA is referred to as a 5' untranslated sequence. A sequence located at the 3' end or downstream of a protein-coding region and present on mRNA is referred to as a 3' untranslated sequence. The term "gene" encompasses both cDNA and genomic forms of genes. A gene in genomic form or clone contains a coding region interrupted by a non-coding sequence that may be referred to as an "intron," "intercalation region," or "intercalation sequence." An intron is a segment of a gene transcribed into nuclear RNA (nRNA). Introns may contain regulatory elements such as enhancers. Introns are removed or "spliced ​​out" from nuclear transcripts or primary transcripts; therefore, introns are absent from mRNA transcripts. mRNA functions during translation to specify the sequence or order of amino acids in the nascent polypeptide. The term "gene" includes synthetic or fusion molecules that encode all or part of the proteins of the invention described herein, as well as nucleotide sequences complementary to any of the foregoing.

[0440] As used herein, “chimeric DNA” is also referred to herein as a “DNA construct”, meaning any DNA molecule that does not exist naturally in nature but is artificially linked together by two DNA parts into a single molecule, each part of which may exist in nature, but the whole molecule does not exist in nature. For example, a DNA construct encoding the MTP-modified NifH polypeptide or MTP-Nif fusion polypeptide of the present invention. Typically, chimeric DNA contains regulatory and transcriptional or protein-coding sequences that do not exist naturally in nature (e.g., sequences linked to non-natural promoter coding sequences). Nif (Polynucleotides). Therefore, chimeric DNA can contain regulatory and coding sequences derived from different sources or derived from the same source but arranged in a manner different from that found in nature. Open reading frames may or may not be linked to their natural upstream and downstream regulatory elements. Open reading frames can be incorporated into, for example, plant genomes, non-natural locations, or replicons or vectors that are not naturally occurring, such as bacterial plasmids or viral vectors. The term "chimeric DNA" is not limited to DNA molecules that can replicate in a host, but includes DNA that can be linked to replicons via, for example, specific adaptor sequences.

[0441] "Transgenic" refers to a gene introduced into the genome through a transformation process. The term includes genes introduced into the genome of daughter cells, plants, seeds, non-human organisms, or parts thereof from their progenitor cells. Such daughter cells can be at least the third or fourth generation offspring of a progenitor cell derived from a primary transformed cell. Offspring can be produced through sexual or asexual reproduction, for example, from tubers from potatoes or root cuttings from sugarcane. The term "genetically modified" and its variations, in a broader sense, include the introduction of genes into cells through transformation or transduction, mutation of genes in cells, and gene alteration or regulation in the daughter cells or any modified cells as described above.

[0442] As used herein, a “genomic region” refers to a location within the genome where a transgene or transgenome (also referred to herein as a cluster) has been inserted into a cell or its precursor. Such regions contain only nucleotides that have been incorporated through human intervention, such as by the methods described herein.

[0443] The "recombinant polynucleotide" of this invention refers to a nucleic acid molecule constructed or modified by artificial recombination methods. Compared to its native state, the recombinant polynucleotide can be present in cells in altered amounts or expressed at altered rates (e.g., in the case of mRNA). In one embodiment, the polynucleotide is introduced into cells that do not naturally contain polynucleotides. Typically, exogenous DNA is used as a template for mRNA transcription, and then the mRNA is translated into a continuous sequence of amino acid residues encoding the polypeptide of this invention within the transformed cells. In another embodiment, the polynucleotide is endogenous to the bacterial cell, and its expression is altered through recombination, for example, by introducing an exogenous control sequence upstream of the endogenous gene of interest, so that the transformed cells can express the polypeptide encoded by said gene.

[0444] The recombinant polynucleotides of the present invention comprise polynucleotides that have not yet been isolated from other components of a cell-based or cell-free expression system in which they are present, and polynucleotides subsequently purified from at least some of the other components and produced in the cell-based or cell-free system. The polynucleotides may be continuous nucleotide extensions present in nature (e.g., Nif A polynucleotide), or a polynucleotide consisting of two or more consecutive nucleotide extensions from different sources (naturally occurring and / or synthetic), linked to form a single polynucleotide (e.g., linked to an MTP-encoding nucleotide sequence and / or a non-natural promoter-encoding sequence). Nif (Polynucleotides). Typically, such chimeric polynucleotides contain at least an open reading frame encoding the polypeptide of the present invention, said chimeric polynucleotide being operatively linked to a promoter adapted to drive open reading frame transcription in the cells of interest. The term "promoter" herein encompasses a single promoter or multiple promoters.

[0445] Regarding the definition of polynucleotides, it should be understood that identity % figures higher than those provided above will cover preferred embodiments. Therefore, where applicable, according to the minimum identity % figure, it is preferred that the polynucleotide contains at least 60%, more preferably at least 65%, more preferably at least 70%, more preferably at least 75%, more preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, and even more preferably at least 99.9% identical polynucleotide sequences to the associated specified SEQ ID NO.

[0446] The polynucleotides of the present invention or useful to the present invention can selectively hybridize with polynucleotides as defined herein under stringent conditions. As used herein, stringent conditions are: (1) the use of a denaturing agent such as formamide during hybridization, for example, 50% (v / v) formamide and 0.1% (w / v) bovine serum albumin, 0.1% Ficoll, 0.1% polyvinylpyrrolidone, 50 mM sodium phosphate buffer containing 750 mM NaCl and 75 mM sodium citrate at pH 6.5 at 42°C; or (2) the use of 50% formamide, 5 x SSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5 x Denhardt's solution, sonicated salmon sperm DNA (50 g / ml), 0.2 x SSC and 0.1% dextran sulfate at 42°C. SDS%, and / or (3) washing with low ionic strength and high temperature, for example, 0.015 M NaCl / 0.0015 M sodium citrate / 0.1% SDS at 50°C.

[0447] When compared with naturally occurring molecules, the polynucleotides of the present invention may have one or more mutations, which are deletions, insertions, or substitutions of nucleotide residues. The polynucleotides with mutations relative to the reference sequence may be naturally occurring (i.e., isolated from natural sources) or synthetic (e.g., through site-directed mutagenesis of nucleic acids or DNA shuffling as described above).

[0448] The polynucleotides of the present invention can be codon-modified for expression in plant cells. Those skilled in the art will understand that protein-coding regions can be codon-optimized relative to the coding regions of polynucleotides naturally present in, for example, nitrogen-fixing bacteria.

[0449] Nucleic acid constructs This invention includes nucleic acid constructs comprising one or more polynucleotides of the invention, vectors and host cells containing such polynucleotides, methods of their production and use, and their applications. This invention relates to operably connected or linked elements. "Operably connected" or "operably linked" refers to the connection of polynucleotide elements in a functional relationship. Typically, operably connected nucleic acid sequences are contiguous, and in cases where two protein-coding regions need to be linked, they are contiguous and within the reading frame. When an RNA polymerase transcribes two coding sequences into a single RNA, the coding sequence is "operably linked" to another coding sequence, and if said single RNA is translated, it is translated into a single polypeptide having amino acids derived from both coding sequences. The coding sequences do not need to be contiguous with each other, as long as the expressed sequence is ultimately processed to produce the desired protein.

[0450] As used herein, the terms “cis-acting sequence,” “cis-acting element,” “cis-regulatory region,” or “regulatory region,” or similar terms, should be understood to mean any nucleotide sequence that, when properly positioned and connected relative to an expressible gene sequence, is capable of at least partially regulating the expression of that gene sequence. Those skilled in the art will recognize that cis-regulatory regions can be activated, silenced, enhanced, repressed, or otherwise altered at the transcriptional or post-transcriptional levels to determine the expression level of a gene sequence and / or cell type specificity and / or developmental specificity. In a preferred embodiment of the invention, the cis-acting sequence is an activator sequence that enhances or stimulates the expression of an expressible gene sequence.

[0451] "Operationally linking" a promoter or enhancer element to a transcribed polynucleotide means placing the transcribed polynucleotide (e.g., a protein-coding polynucleotide or other transcript) under the regulatory control of a promoter, which then controls the transcription of the polynucleotide. In the construction of heterologous promoter / structural gene combinations, it is generally preferred to position the promoter or a variant thereof at a distance from the transcription start site of the transcribed polynucleotide, said distance being approximately the same as the distance between the promoter and the protein-coding region it controls in its natural environment; i.e., the gene deriving the promoter. As is known in the art, some variations of this distance can be adapted without loss of function. Similarly, the preferred positioning of a regulatory sequence element (e.g., an operator, enhancer, etc.) relative to the transcribed polynucleotide to which it is to be placed is defined by the positioning of the element in its natural environment; i.e., the gene deriving said element.

[0452] As used herein, a “promoter” or “promoter sequence” refers to a gene region, typically located upstream (5') of the RNA coding region, that controls the initiation and level of transcription in the cell of interest. “Promoters” include transcriptional regulatory sequences of classic genomic genes, such as TATA-box and CCAAT-box sequences, as well as additional regulatory elements (i.e., upstream activating sequences, enhancers, and silencers) that respond to developmental and / or environmental stimuli or alter gene expression in a tissue-specific or cell-type-specific manner. Promoters are typically, but not necessarily (e.g., some PolIII promoters) located upstream of structural genes that regulate the expression of said structural genes. Furthermore, regulatory elements containing promoters are typically located within 2 kb of the transcription start site. Promoters may contain additional, more distal, specific regulatory elements located at the start site to further enhance expression in the cell and / or alter the timing or inducibility of expression of structural genes operatively linked to them.

[0453] A “constitutive promoter” is a promoter that directs the expression of an operatively linked transcriptional sequence in many or all tissues of an organism such as a plant. As used herein, preferred constitutive promoters are the CaMV 35S promoter or the CaMV e35S promoter. As used herein, the term constitutive does not necessarily indicate that a gene is expressed at the same level in all cell types, but rather that the gene is expressed in a wide range of cell types, although some variations in level are generally detectable. As used herein, “selective expression” refers to expression almost entirely in a specific organ of a plant, such as the endosperm, embryo, leaf, fruit, tuber, or root. In a preferred embodiment, the promoter is expressed selectively or preferentially in the roots, leaves, and / or stems of a plant, preferably cereals. Thus, selective expression can be contrasted with constitutive expression, which refers to expression in many or all tissues of a plant under most or all of the conditions experienced by the plant.

[0454] Selective expression can also lead to compartmentalization of gene expression products in specific plant tissues, organs, or developmental stages. Compartmentation in specific subcellular locations, such as plastids, cytosols, vacuoles, or apoplasts, can be achieved by including appropriate signals (e.g., signal peptides) in the structure of the gene product for transport to the desired cellular compartment, or, in the case of semi-autonomous organelles (plastids and mitochondria), by directly integrating transgenes with appropriate regulatory sequences into the organelle genome.

[0455] A "tissue-specific promoter" or "organ-specific promoter" is a promoter that is preferentially expressed in one tissue or organ, preferably in most (if not all) other tissues or organs, such as in plants. Typically, the expression level of the promoter in a particular tissue or organ is 10 times that in other tissues or organs.

[0456] In one embodiment, the promoter is a stem-specific promoter, a leaf-specific promoter, or a promoter that guides gene expression in the aboveground parts of a plant (at least the stem and leaves) (green tissue-specific promoter), such as the ribulose-1,5-bisphosphate carboxylase oxygenase (RUBISCO) promoter.

[0457] Examples of stem-specific promoters include, but are not limited to, the stem-specific promoters described in US 5,625,136.

[0458] In one embodiment, the promoter is a root-specific promoter. Examples of root-specific promoters include, but are not limited to, the promoter of the acid chitinase gene and the specific subdomain of the CaMV 35S promoter.

[0459] The promoters envisioned in this invention can be natural for the host plant to be transformed, or can be derived from alternative sources, wherein the region is functional in the host plant. Other sources include Agrobacterium t-DNA genes, such as promoters of genes for the biosynthesis of carmine, octopine, mannitine, or other crown gall alkaloid promoters; tissue-specific promoters (see, for example, US 5,459,252 and WO 91 / 13992); promoters from viruses (including host-specific viruses), or partially or fully synthesized promoters. Many promoters functional in monocotyledonous and dicotyledonous plants are well known in the art (see, for example, Salomon et al., 1984; Garfinkel et al., 1983; Barker et al., 1983); including various promoters isolated from plants and viruses, such as the cauliflower mosaic virus promoter (CaMV 35S, 19S). Non-limiting methods for evaluating promoter activity are disclosed by Medberry et al. (1992 and 1993); Sambrook et al. (1989, ibid.); and US 5,164,316.

[0460] Alternatively or additionally, the promoter can be an inducible promoter or a developmentally regulatory promoter, capable of driving the expression of an introduced polynucleotide at an appropriate developmental stage, for example, in a plant. Other cis-acting sequences that can be employed include transcriptional and / or translational enhancers. Enhancer regions are well known to those skilled in the art and may include an ATG translation start codon and adjacent sequences. When included, the start codon should be in phase with the reading frame of the coding sequence associated with the foreign or exogenous polynucleotide to ensure translation of the entire sequence (if to be translated). The translation initiation region may be provided by a source of transcription start region or by a foreign or exogenous polynucleotide. The sequence may also be derived from a source of promoter selected to drive transcription and may be specifically modified to enhance mRNA translation.

[0461] The nucleic acid construct of the present invention may contain a 3' untranslated sequence of about 50 to 1,000 nucleotide base pairs, which may include a transcription termination sequence. The 3' untranslated sequence may contain a transcription termination signal, which may or may not include a polyadenylation signal and any other regulatory signals capable of influencing mRNA processing. The role of the polyadenylation signal is to add a polyadenylate bundle to the 3' end of the mRNA precursor. Although variations are not uncommon, the polyadenylation signal is generally identified by the presence of homology with the canonical form 5'AATAAA-3'. Transcription termination sequences that do not include a polyadenylation signal include terminators of PolI or PolIII RNA polymerases containing a series of four or more thymidines. Examples of suitable 3' untranslated sequences are those containing a polyadenylation signal from Agrobacterium tumefaciens (…). Agrobacterium tumefaciensThe 3' untranslated transcriptional region of the polyadenylation signal of the octopine synthase (ocs) gene or the nopaline synthase (nos) gene is preferred (Bevan et al., 1983). Suitable 3' untranslated sequences can also be derived from plant genes, such as the ribulose-1,5-bisphosphate carboxylase (ssRUBISCO) gene, although other 3' elements known to those skilled in the art may also be used. The preferred transcription terminator sequence is the TTm terminator described herein, or a functionally equivalent terminator having multiple transcription terminator regions and optionally a MAR region.

[0462] Because the DNA sequence inserted between the transcription start site and the coding sequence start point—the untranslated 5' leader sequence (5'UTR)—can affect gene expression if translated and transcribed, specific leader sequences can also be used. Suitable leader sequences include those containing sequences selected to guide the optimal expression of foreign or endogenous DNA sequences. For example, such leader sequences include preferred consensus sequences that can increase or maintain mRNA stability and prevent inappropriate translation initiation, as described, for example, by Joshi (1987).

[0463] carrier This invention includes the use of vectors for manipulating or transferring gene constructs. A vector is a nucleic acid molecule, preferably a DNA molecule, that can be used to artificially carry foreign genetic material into another cell, where the vector replicates or is expressed. A vector containing foreign DNA is called a "recombinant vector." Examples of vectors include, but are not limited to, plasmids and viral vectors, such as the geminivirus-based replication vectors described herein. Vectors may contain transposable elements.

[0464] The vector is preferably double-stranded DNA and contains one or more unique restriction sites, and is capable of autonomous replication in a defined host cell, including a target cell or tissue or its progenitor cell or tissue, or of integrating into the genome of a defined host, preferably the nuclear genome, such that the cloned sequence is reproducible. Thus, the vector can be a self-replicating vector, i.e., a vector existing as an extrachromosomal entity whose replication is independent of chromosomal replication, such as linear or closed circular plasmids, extrachromosomal elements, microchromosomes, or artificial chromosomes. The vector can contain any means to ensure self-replication. Alternatively, the vector can be a vector that, when introduced into a cell, integrates into the genome of the recipient cell, preferably the nuclear genome, and replicates along with the chromosome already integrated therein. The vector system can comprise a single vector or plasmid, or two or more vectors or plasmids containing the total DNA to be introduced into a host cell or transposon. The choice of vector will generally depend on the compatibility of the vector with the cells to which it will be introduced. The vector may also include selection markers, such as antibiotic resistance genes, herbicide resistance genes, or other genes that can be used to select suitable transformants. Examples of such genes are well known to those skilled in the art.

[0465] The nucleic acid constructs of this invention can be introduced into vectors such as plasmids. Plasmid vectors typically include additional nucleic acid sequences that provide an expression cassette for easy selection, amplification, and transformation in prokaryotic and eukaryotic cells; for example, pUC-derived vectors, pSK-derived vectors, pGEM-derived vectors, pSP-derived vectors, pBS-derived vectors, or binary vectors containing one or more T-DNA regions. The additional nucleic acid sequences include a replication origin for providing autonomous replication of the vector, preferably encoding a selectable marker gene for antibiotic or herbicide resistance, a unique multiple cloning site providing multiple sites for insertion into the nucleic acid construct encoding a nucleic acid sequence or gene, and sequences that enhance transformation of prokaryotic and eukaryotic (especially plant) cells.

[0466] A "marker gene" is a gene that confers a different phenotype to cells expressing a marker gene and thus allows such transformed cells to be distinguished from cells without the marker. A selectable marker gene confers a trait that can be "selected" based on resistance to a selector (e.g., herbicides, antibiotics, radiation, heat, or other treatments that damage untransformed cells). A screenable marker gene (or reporter gene) confers a trait that can be identified by observation or testing, i.e., by "screening" (e.g., β-glucuronidase, luciferase, GFP, or other enzyme activity not present in untransformed cells). The marker gene and the nucleotide sequence of interest do not necessarily need to be linked.

[0467] To facilitate the identification of transformants, nucleic acid constructs preferably contain selectable or screenable marker genes as foreign or exogenous polynucleotides, or in addition to these. The actual choice of marker is not critical, as long as the marker is functional (i.e., selective) when combined with the host cell (preferably a plant host cell). The marker gene and the foreign or exogenous polynucleotide of interest need not be linked, as co-transformation of unlinked genes, as described, for example, in US 4,399,216, is also an efficient process in plant transformation.

[0468] Examples of selectable markers for bacteria are markers that confer antibiotic resistance (preferably kanamycin resistance) such as ampicillin, erythromycin, chloramphenicol, or tetracycline resistance. Exemplary selectable markers for selecting plant transformants include, but are not limited to, those encoding hygromycin B resistance. hyg Gene; neomycin phosphotransferase conferring resistance to kanamycin, paromomycin, and G418 ( nptII ) genes; glutathione-S-transferase genes from rat liver conferring resistance to glutathione-derived herbicides, as described, for example, in EP 256223; glutamine synthase genes that, when overexpressed, confer resistance to glutamine synthase inhibitors (such as glufosinate), as described, for example, in WO 87 / 05327; from *Streptomyces viridans* ( Streptomyces viridochromogenes Acetyltransferase genes conferring resistance to the selective agent glufosinate, such as those described in EP 275957; genes encoding 5-enolpyruvylshikimate-3-phosphate synthase (EPSPS) conferring resistance to N-phosphonomethylglycine, such as those described in Hinchee et al. (1988); genes conferring resistance to diammonium phosphate. bar Genes, such as those described in WO91 / 02071; nitrile hydrolase genes conferring resistance to bromonazine, such as those from Klebsiella odorata ( Klebsiella ozaenae )of bxn (Stalker et al., 1988); conferring resistance to methotrexate dihydrofolate reductase (DHFR) genes (Thillet et al., 1988); conferring resistance to imidazolinone, sulfonylurea, or other ALS-inhibiting chemicals to acetolactate synthase (ALS) genes (EP 154,204); conferring resistance to 5-methyltryptophan to mutated anthranilate synthase genes; or conferring resistance to herbicides to dalapon dehalogenase genes.

[0469] Preferred screenable markers include, but are not limited to, those encoding known β-glucuronidase (GUS) for various chromogenic substrates. uidAGenes; β-galactosidase genes encoding enzymes known to develop chromogenic substrates; jellyfish luminescent protein genes (Prasher et al., 1985), which can be used for calcium-sensitive bioluminescence detection; green fluorescent protein (Niedz et al., 1995) or derivatives thereof; luciferase (… luc The gene (Ow et al., 1986) allows for bioluminescent detection; and other genes known in the art. As used herein, a “reporter molecule” means a molecule that provides an analytically identifiable signal by its chemical properties, which helps in determining promoter activity by referring to the protein product.

[0470] Preferably, the nucleic acid construct is stably incorporated into, for example, the genome of a plant, and more preferably into the nuclear genome of a plant. The nucleic acid construct may also be incorporated into the plastid genome or mitochondrial genome of a plant. Thus, the nucleic acid contains appropriate elements that allow the molecule to be incorporated into the genome, or the construct is placed in an appropriate vector that can be incorporated into the chromosome of a plant cell.

[0471] One embodiment of the invention includes a recombinant vector containing at least one polynucleotide as defined herein and capable of delivering the polynucleotide into a host cell. Such vectors contain a heterologous nucleic acid sequence, i.e., a nucleic acid sequence that is not naturally occurring in a neighboring nucleic acid molecule of the invention, and preferably derived from a species different from the species from which the derived nucleic acid molecule is derived. The vector can be RNA or DNA, prokaryotic or eukaryotic, and is typically a virus or plasmid.

[0472] The recombinant vector of the present invention contains a fusion sequence that causes nucleic acid molecules to be expressed as fusion proteins.

[0473] Recombinant vectors may also include intercalated and / or untranslated sequences around and / or within the nucleic acid sequences of polynucleotides as defined herein.

[0474] Preferably, the recombinant vector is stably incorporated into the genome of a host cell (such as a plant cell). Therefore, the recombinant vector may contain appropriate elements that allow the vector to be incorporated into the genome or the chromosome of the cell.

[0475] Recombinant cells Another embodiment of the invention includes recombinant cells (e.g., recombinant plant cells) which are host cells transformed with one or more polynucleotides, constructs or vectors of the invention or their progeny cells. The term "recombinant cell" is used interchangeably with the term "transgenic cell" herein.

[0476] The conversion of nucleic acid molecules into cells can be accomplished by any method of inserting nucleic acid molecules into cells. Conversion techniques include, but are not limited to, transfection, electroporation, microinjection, lipid transfection, adsorption, and protoplast fusion. Recombinant cells can remain single cells or grow into tissues, organs, or multicellular organisms. The converted nucleic acid molecules of this invention can remain extrachromosomally or can be integrated into one or more sites within the chromosome of the converted cells in a manner that preserves their expressive ability.

[0477] The preferred host cells are plant cells, more preferably cereal cells, more preferably barley or wheat cells, and even more preferably rice cells.

[0478] The recombinant cells can be cells in a culture, cells in vitro, or cells in an organism such as a plant, or cells in an organ (e.g., root, leaf, or stem). Preferably, the cells are in a plant, more preferably in the roots, leaves, and / or stems of a plant.

[0479] In one embodiment, the expression of active NifH in plant cells requires the expression of NifM, NifS, NifU, NifD, and NifK. In one embodiment, NifD and NifK are present as NifD-NifK fusion peptides.

[0480] In another embodiment, expression of active NifH in plant cells requires the expression of NifH and NifM, and optionally NifU and / or NifN.

[0481] plant As used herein, the term "plant" as a noun refers to the whole plant and any member of the plant kingdom, but as an adjective refers to any substance present in, obtained from, derived from, or associated with a plant, such as plant organs (e.g., leaves, stems, roots, flowers), single cells (e.g., pollen), seeds, plant cells, etc. Small plantlets that have taken root and sprouted, and germinating seeds, are also included in the meaning of "plant," as used herein. The term "plant part" refers to one or more plant tissues or organs obtained from a plant and containing the plant's genomic DNA. Plant parts include vegetative structures (e.g., leaves, stems), roots, floral organs / structures, seeds (including embryos, cotyledons, and seed coats), plant tissues (e.g., vascular tissues, terrestrial tissues, etc.), cells, and their progeny. In a preferred embodiment, the plant part is a seed. As used herein, the term "plant cell" refers to a cell obtained from or within a plant and includes protoplasts or other cells derived from a plant, gamete-producing cells, and cells that regenerate the whole plant. Plant cells can be cells in culture. "Plant tissue" means differentiated tissue in or obtained from a plant ("explant"), or undifferentiated tissue derived from immature or mature embryos, seeds, roots, buds, fruits, tubers, pollen, tumor tissues such as crown galls, and various forms of plant cell aggregates in culture such as callus. Exemplary plant tissues in or derived from seeds are cotyledons, embryos, and embryonic axes. Therefore, the present invention includes plants and plant parts, as well as products comprising these.

[0482] As used herein, the term “seed” refers to a “mature seed” of a plant, which is ready for harvesting or has already been harvested from the plant, such as in commercial harvesting commonly done in the field, or a “developing seed” that is present in the plant after fertilization and before seed dormancy is established and before harvesting.

[0483] As used herein, a “transgenic plant” means a plant containing nucleic acid constructs not found in wild-type plants of the same species, variety, or cultivar. In other words, a transgenic plant (a transformed plant) contains genetic material (transgenic material) that it did not possess prior to transformation. Transgenic material can include genetic sequences or synthetic sequences obtained or derived from plant cells, other plant cells, or non-plant sources. Typically, transgenic material is introduced into plants through human manipulation, such as transformation, but those skilled in the art recognize that any method can be used. The genetic material is preferably stably integrated into the plant's genome, preferably the nuclear genome. The introduced genetic material may contain sequences naturally present in the same species but arranged in rearranged order or with different elements, such as antisense sequences. Plants containing such sequences are included herein in the term “transgenic plant.”

[0484] In a preferred embodiment, the transgenic plant is homozygous for each and every introduced gene (transgene), such that its offspring do not segregate for the desired phenotype. The transgenic plant may also be heterozygous for the introduced transgene, for example, in the F1 offspring already grown from hybrid seeds. Such plants can provide advantages well-known in the art, such as hybrid viability.

[0485] As defined in the context of this invention, a transgenic plant includes the progeny of a plant that has been genetically modified using recombinant technology, wherein the progeny contains the transgene of interest. Such progeny can be obtained by self-fertilization of a primary transgenic plant or by hybridization of such plant with another plant of the same species. This is typically for the purpose of regulating the production of at least one protein as defined herein in a desired plant or plant organ. A transgenic plant portion includes all portions and cells of the plant containing the transgene, such as cultured tissues, callus tissue, and protoplasts.

[0486] Transgenic plants can be produced using techniques known in the art, such as those commonly described in A. Slater et al., Plant Biotechnology - The Genetic Manipulation of Plants, Oxford University Press (2003), and P. Christou and H. Klee, Handbook of Plant Biotechnology, John Willie & Son Publishing (2004).

[0487] "Non-GMO plant" refers to a plant that has not been genetically modified by introducing genetic material through recombinant DNA technology. As used herein, the term "compared to isogenetic plants" or similar phrases refer to plants that are isogenetic relative to transgenic plants but do not contain the transgene of interest. Preferably, the corresponding non-GMO plant is a cultivar or variety with the same ancestor as the transgenic plant of interest, or a sibling plant line lacking a construct commonly referred to as an "isosome," or a plant of the same cultivar or variety transformed with an "empty vector" construct, and may be non-GMO. As used herein, "wild-type" refers to cells, tissues, or plants that have not been modified according to the present invention. Wild-type cells, tissues, or plants can be used as controls to compare the expression levels of exogenous nucleic acids or the degree and nature of trait modifications with modified cells, tissues, or plants as described herein.

[0488] As defined in the context of this invention, a transgenic plant includes the offspring of a plant that has been genetically modified using recombinant technology, wherein the offspring contains the transgene of interest. Such offspring can be obtained by self-fertilization of a primary transgenic plant or by hybridization of such plant with another plant of the same species. A transgenic plant portion includes all portions and cells of the plant containing the transgene, such as cultured tissues, callus tissues, and protoplasts.

[0489] The plants envisioned for use in this invention include both monocotyledonous and dicotyledonous plants. Target plants include, but are not limited to, the following: cereals (e.g., wheat, barley, rye, oats, rice, corn, sorghum, and related crops); grapes; sugar beets (sugar beets and forage beets); pome fruits, stone fruits, and seedless small fruits (apples, pears, plums, peaches, almonds, cherries, strawberries, raspberries, and blackberries); legumes (legumes, lentils, peas, soybeans); oilseed plants (rapeseed or other brassicasas, mustard, olives, sunflowers, saffron, flax, coconuts, castor oil plants, cocoa beans, peanuts); Cucumber plants (gourd, cucumber, melon); fiber plants (cotton, flax, jute); citrus fruits (orange, lemon, grapefruit, tangerine); vegetables (spinach, lettuce, asparagus, cabbage, carrot, onion, tomato, potato, pepper); Lauraceae (avocado, cinnamon, camphor); or plants such as corn, tobacco, nuts, coffee, sugarcane, tea, vines, hops, turf, bananas, and natural rubber plants, as well as ornamental plants (flowering plants, shrubs, broad-leaved trees, and evergreen plants such as conifers). Preferably, the plants are cereals, more preferably wheat, rice, corn, rye, oats, or barley, and even more preferably wheat.

[0490] As used in this article, the term "wheat" refers to wheat ( Triticum Wheat includes any species of the genus *Wheat*, including its ancestors and its offspring produced by hybridization with other species. Wheat includes the genome organization AABBDD, containing 42 chromosomes ("hexaploid wheat"), and the genome organization AABB, containing 28 chromosomes ("tetraploid wheat"). Hexaploid wheat includes common wheat (…). T. aestivum ), Spelt wheat ( T. spelta ), moga wheat ( T. macha ), dense ear wheat ( T. compactum ), round-grain wheat ( T. sphaerococcum ), Vavilov wheat ( T. vavilovii ) and its interspecific hybridization. The preferred species of hexaploid wheat is the common wheat subspecies (also known as "bread wheat"). Tetraploid wheat includes durum wheat ( T. durum (Also referred to in this article as durum wheat or durum wheat subspecies) Triticum turgidum ssp. durum ), wild emmer wheat ( T. dicoccoides ), emmer wheat ( T. dicoccum Polish wheat T. polonicum ) and its interspecific hybridization. Additionally, the term "wheat" includes potential ancestors of hexaploid or tetraploid wheat, such as *Wheat urartu* of genome A (…). T. uartu A grain of wheat ( T. monococcum ) or wild wheat ( T. boeoticum ); Goatgrass of the B genome ( Aegilops speltoides ), and jointed wheat (D genome) T. tauschii (also known as rough goat grass) Aegilops squarrosa ) or jointed barley ( Aegilops tauschii The particularly preferred ancestor is the ancestor of genome A, and even more preferably, the ancestor of genome A is a grain of wheat. The wheat cultivars used in this invention may belong to, but are not limited to, any of the species listed above. It also covers the use of wheat as a parent and non-wheat species (such as rye [naked wheat]) through conventional techniques. Secale cereale Plants produced by progressive hybridization, including but not limited to black wheat ( Triticale ).

[0491] As used in this article, the term "barley" refers to barley ( Hordeum Any species of the genus *Barley*, including its ancestors and its offspring produced through hybridization with other species. Preferably, the plant is commercially cultivated barley, such as barley (…). Hordeum vulgare ( ) strains or cultivars or varieties, or those suitable for commercial grain production.

[0492] Methods for producing transgenic plants Four common methods for delivering genes directly into cells have been described: (1) chemical methods (Graham et al., 1973); (2) physical methods, such as microinjection (Capecchi, 1980); electroporation (see, for example, WO 87 / 06614, US 5,472,869, 5,384,253, WO 92 / 09696 and WO 93 / 21335); and gene guns (see, for example, US4,945,050 and US 5,141,131); (3) viral vectors (Clapp, 1993; Lu et al., 1993; Eglitis et al., 1988); and (4) receptor-mediated mechanisms (Curiel et al., 1992; Wagner et al., 1992).

[0493] Acceleration methods that can be used include, for example, particle bombardment. One example of a method for delivering transformed nucleic acid molecules into plant cells is particle bombardment. This method has been reviewed in Yang et al., Particle Bombardment Technology for Gene Transfer, Oxford Press, Oxford, England (1994). Non-biological particles (microparticles) can be coated with nucleic acids and delivered into cells by propulsion. Exemplary particles include those containing tungsten, gold, platinum, etc. In addition to being an efficient method for reproducibly transforming monocotyledonous plants, a particular advantage of particle bombardment is that it eliminates the need for protoplast isolation and susceptibility to Agrobacterium infection. A suitable particle delivery system for this invention is the helium-accelerated PDS-1000 / He gun, available from Bio-Rad Laboratories. For bombardment, immature embryos or derived target cells, such as scutes or calluses from immature embryos, can be arranged on a solid culture medium.

[0494] In another alternative embodiment, plastids can be stably transformed. Disclosed methods for plastid transformation in higher plants include particle gun delivery of DNA containing selective markers and targeting the DNA to the plastid genome via homologous recombination (US 5,451,513, US 5,545,818, US 5,877,402, US 5,932479, and WO 99 / 05265).

[0495] Agrobacterium-mediated transfer is a widely used system for introducing genes into plant cells because DNA can be introduced throughout the plant tissue, thereby bypassing the need for regeneration of a complete plant from protoplasts. The use of Agrobacterium-mediated plant integration vectors to introduce DNA into plant cells is well-known in the art (see, for example, US 5,177,010, US 5,104,310, US 5,004,863, US 5,159,135). Furthermore, T-DNA integration is a relatively precise process resulting in minimal rearrangements. The DNA region to be transferred is defined by a boundary sequence, and intercalated DNA is typically inserted into the plant genome.

[0496] Agrobacterium-mediated transformation vectors can replicate in both *Escherichia coli* and *Agrobacterium*, allowing for convenient manipulation as described (Klee et al., *Plant DNA Infectious Agents*, Hohn and Schell, (eds.), Springer-Verlag, New York, (1985): 179-203). Furthermore, technological advancements in vectors for *Agrobacterium*-mediated gene transfer have improved the arrangement of genes and restriction sites within the vectors, facilitating the construction of vectors capable of expressing a wide range of polypeptide-encoding genes. The described vectors possess convenient multi-connector regions flanked by promoters and polyadenylation sites for the direct expression of inserted polypeptide-encoding genes and are suitable for the purposes of this invention. Additionally, *Agrobacterium* containing both armed and unarmed Ti genes can be used for transformation. This method is preferred in plant varieties where *Agrobacterium*-mediated transformation is highly efficient due to the ease and defined nature of gene transfer.

[0497] Transgenic plants created using Agrobacterium-mediated transformation typically contain a single gene locus on a chromosome. Such transgenic plants targeting the added gene can be termed hemizygous. More preferably, transgenic plants are homozygous for the added structural gene; that is, transgenic plants containing two added genes, one gene located at the same locus on each chromosome of a chromosome pair. Homozygous transgenic plants can be obtained by sexually mating (self-pollinating) independent isolates of transgenic plants containing a single added gene, germinating some of the resulting seeds, and analyzing the gene of interest in the resulting plants.

[0498] It should also be understood that two different transgenic plants can be interbred to produce offspring containing two independently segregated foreign genes. Self-pollination of appropriate offspring can produce plants in which both foreign genes are homozygous. Similar to vegetative propagation, backcrossing with the parent plant and outcrossing with non-transgenic plants are also conceivable. Descriptions of other breeding methods commonly used for different traits and crops can be found in Fehr, "Breeding Methods for Cultivar Development," J. Wilcox (ed.), American Society of Agronomy, Madison Wisconsin (1987).

[0499] Plant protoplast transformation can be achieved using methods based on calcium phosphate precipitation, polyethylene glycol treatment, electroporation, and combinations of these treatments. The application of these systems to different plant varieties depends on the ability to regenerate the specific plant line from protoplasts. Illustrative methods for regenerating cereals from protoplasts are described (Fujimura et al., 1985; Toriyama et al., 1986; Abdullah et al., 1986).

[0500] Other cell transformation methods can also be used, including but not limited to introducing DNA into plants by: directly transferring DNA into pollen, injecting DNA into the reproductive organs of plants, or directly injecting DNA into the cells of immature embryos, followed by rehydration of the dried embryos.

[0501] The regeneration, development, and culture of plants from single plant protoplast transformants or from various transformed explants is well-known in the field (Weissbach et al., Methods for Plant Molecular Biology, Academic Press, San Diego, (1988)). This regeneration and growth process typically involves selecting transformed cells and culturing these individualized cells through the typical stages of embryonic development to the rooted plantlet stage. Transgenic embryos and seeds are regenerated similarly. The resulting transgenic rooted shoots are then planted in a suitable plant growth medium, such as soil.

[0502] The development or regeneration of plants containing foreign or exogenous genes is well known in the art. Preferably, the regenerated plant is self-pollinated to provide a homozygous transgenic plant. Otherwise, pollen obtained from the regenerated plant is hybridized into seed plants of agronomically important lines. Conversely, pollen from these important lines is used to pollinate the regenerated plant. The transgenic plants containing the desired exogenous nucleic acids of the present invention are cultured using methods well known to those skilled in the art.

[0503] Methods for obtaining transgenic plants by transforming dicotyledonous plants with Agrobacterium tumefaciens have been disclosed for the following: cotton (US 5,004,863, US 5,159,135, US 5,518,908); soybean (US 5,569,834, US 5,416,011); brassica (US 5,463,174); peanut (Cheng et al., 1996); and pea (Grant et al., 1995).

[0504] Transformation methods for introducing genetic variants into plants by introducing exogenous nucleic acids, and for regenerating plants from protoplasts or immature plant embryos, such as cereals like wheat and barley, are well known in the art, see, for example, CA 2,092,588, AU 61781 / 94, AU 667939, US 6,100,447, WO 97 / 048814, US 5,589,617, US 6,541,257, and other methods listed in WO 99 / 14314. Preferably, transgenic wheat or barley plants are produced via an Agrobacterium-mediated transformation procedure. A vector carrying the desired nucleic acid construct can be introduced into regenerable wheat cells from tissue-cultured plants or explants, or into suitable plant systems such as protoplasts. Regenerable wheat cells are preferably derived from the scutellum of immature embryos, mature embryos, callus, or meristems derived from these embryos.

[0505] To confirm the presence of transgenic cells and plants, polymerase chain reaction (PCR) amplification or Southern blotting analysis can be performed using methods known to those skilled in the art. The expression products of the transgenic organism can be detected in any of a variety of ways, depending on the natu...

Claims

1. A plant cell comprising a modified NifH polypeptide, wherein the modified NifH polypeptide contains at least one amino acid substitution compared to a corresponding wild-type NifH polypeptide, and wherein the modified NifH polypeptide is more soluble in the mitochondria of the cell than the corresponding wild-type NifH polypeptide.

2. The plant cell according to claim 1, wherein the at least one amino acid substitution is located at a position selected from the group consisting of: (i) the group consisting of the following amino acid positions: 2, 5, 7, 19, 23, 24, 26 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 74, 76 to 78, 80 to 84, 102, 105, 107, 111 to 114, 116 to 118, 121 to 124, 139, 145, 147, 149, 158, 165, 166, 168, 169, 171, 179, 18 of SEQ ID NO:

37. 2, 183, 188, 191, 193 to 197, 200 to 203, 205 to 211, 214, 216, 219, 223 to 226, 228 to 235, 237, 238, 241, 242, 244 to 246, 248, 249, 251 to 253, 257, 259 to 264, 266 to 271 and 273 to 275, or located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37, and / or (ii) the group consisting of the following amino acid positions: refer to SEQ ID NO:39 3, 6, 8, 20, 24, 25, 27 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 75, 77 to 79, 81 to 85, 103, 106, 108, 112 to 115, 117 to 119, 122 to 125, 140, 146, 148, 150, 159, 166, 167, 169, 170, 172, 180, 18 3, 184, 189, 192, 194 to 198, 201 to 204, 206 to 212, 215, 217, 220, 224 to 227, 229 to 236, 238, 239, 242, 243, 245 to 247, 249, 250, 252 to 254, 258, 260 to 265, 267 to 272 and 274 to 276, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:

39.

3. A plant cell comprising a modified NifH polypeptide, wherein, when compared with a corresponding wild-type NifH polypeptide, the modified NifH polypeptide comprises at least one amino acid substitution, and wherein the at least one amino acid substitution is located at an amino acid position selected from: (i) the group consisting of the following amino acid positions: 2, 5, 7, 19, 23, 24, 26 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 74, 76 to 78, 80 to 84, 102, 105, 107, 111 to 114, 116 to 118, 121 to 124, 139, 145, 147, 149, 158, 165, 166, 168, 169, 171, 179, 18 of SEQ ID NO:

37. 2, 183, 188, 191, 193 to 197, 200 to 203, 205 to 211, 214, 216, 219, 223 to 226, 228 to 235, 237, 238, 241, 242, 244 to 246, 248, 249, 251 to 253, 257, 259 to 264, 266 to 271 and 273 to 275, or located at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37, and / or (ii) the group consisting of the following amino acid positions: refer to SEQ ID NO:39 3, 6, 8, 20, 24, 25, 27 to 35, 45, 48, 49, 51, 53, 54, 56 to 59, 61, 62, 64 to 75, 77 to 79, 81 to 85, 103, 106, 108, 112 to 115, 117 to 119, 122 to 125, 140, 146, 148, 150, 159, 166, 167, 169, 170, 172, 180, 18 3, 184, 189, 192, 194 to 198, 201 to 204, 206 to 212, 215, 217, 220, 224 to 227, 229 to 236, 238, 239, 242, 243, 245 to 247, 249, 250, 252 to 254, 258, 260 to 265, 267 to 272 and 274 to 276, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:

39.

4. The plant cell according to claim 3, wherein the modified NifH polypeptide is more soluble in the mitochondria of the cell than the corresponding wild-type NifH polypeptide, preferably wherein: a) At least 15%, at least 20%, at least 35%, at least 40%, at least 45%, at least 50%, or 15% to 90%, 15% to 80%, 15% to 70%, or 15% to 60% of the modified NifH polypeptide in the mitochondria of the cells are soluble; and / or b) The solubility of the NifH polypeptide in the mitochondria of the cell is at least twice that of the corresponding wild-type NifH polypeptide in the mitochondria of the cell, preferably at least three times, at least four times, or at least five times, or two to ten times.

5. The plant cell according to any one of claims 1 to 4, wherein: a) The free energy of the modified NifH polypeptide is lower than that of the wild-type NifH polypeptide, and / or wherein the amino acid substitution reduces the free energy of the modified NifH polypeptide relative to the wild-type NifH polypeptide, preferably wherein the free energy of the modified NifH polypeptide having the substituted amino acid is reduced by at least 0.5, at least 1.0, at least 1.5, at least 2, at least 3, at least 4, at least 5, or 2 to 6 units relative to the corresponding NifH polypeptide, and the amino acid sequence of the corresponding NifH polypeptide is the same as that of the modified NifH polypeptide except for the substituted amino acid; and / or b) When compared with the corresponding wild-type NifH polypeptide, the modified NifH polypeptide contains at least one, preferably at least two or at least three amino acid substitutions, wherein each substituted amino acid reduces the free energy of the modified NifH polypeptide by at least 0.5, at least 1.0, at least 1.5, at least 2, at least 3, at least 4, at least 5 units or 2 to 6 units, and / or wherein the amino acid substitutions collectively reduce the free energy of the modified NifH polypeptide by at least 4.0, at least 5.0, at least 6.0, at least 7.0, at least 8.0, at least 9.0, at least 10.0, at least 12.0 units, or 4.0 to 15.0, 4.0 to 13.0 or 4.0 to 12.0 units.

6. The plant cell according to any one of claims 1 to 5, wherein: a) The modified NifH polypeptide comprises at least one amino acid substitution, or two or three amino acid substitutions, or four or more amino acid substitutions, wherein the amino acid substitutions are selected from the group consisting of: amino acid substitutions listed in one or more of Tables 4, 5, 9, 10 or 11, or the corresponding amino acid substitutions when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39; and / or b) The modified NifH polypeptide comprises at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, wherein the amino acid substitution is located at an amino acid position selected from the group consisting of amino acid positions 69, 168, 200, 201, 224, 228, 234, 241, 252 and 263 as per SEQ ID NO:37, or at the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39, optionally wherein the at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, is selected from the group consisting of: 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H or 234C, 241R or 241A, 252M and 263E, wherein the amino acid position is as per SEQ ID NO:

37. The amino acid sequence provided in NO:37 corresponds to, or is located at, the corresponding amino acid position in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:39; and / or c) The at least one amino acid substitution, or preferably two or three amino acid substitutions or four or more amino acid substitutions, are selected from the group consisting of positions corresponding to amino acids 69, 168, 200, 201, 224, 228, 234, 252, and 263 of SEQ ID NO:37, or located at the corresponding amino acid positions in the modified NifH polypeptide when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:

39. Optionally, the at least one amino acid substitution, or two or three amino acid substitutions or four or more amino acid substitutions are selected from the group consisting of: 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H or 234C, 252M, and 263E, wherein the amino acid positions correspond to the amino acid sequences provided as in SEQ ID NO:37, or are located when the sequence of the modified NifH polypeptide is compared with SEQ ID NO:37 and / or SEQ ID NO:

39. NO:39 The corresponding amino acid position in the modified NifH polypeptide during the comparison; and / or d) When compared with the corresponding wild-type NifH polypeptide, the modified NifH polypeptide has one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 1 to 20, 1 to 15, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2 to 20 One, 2 to 15, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 2 or 3, 3 to 20, 3 to 15, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 or 4, preferably 1 to 3, 1 to 4 or 1 to 5, more preferably 2 to 4 or 2 to 5, most preferably 3 to 5 amino acid substitutions.

7. The plant cell according to any one of claims 1 to 6, wherein the modified NifH polypeptide comprises at least one or at least two amino acid substitutions according to SEQ ID NO:37, wherein the at least one amino acid substitution is i) 228V, ii) 228I, iii) 200A, or iv) 234H, The at least two amino acid substitutions are v) 200A and 228V, vi) 200A and 228I, vii) 228V and 234H, viii) 228I and 234H, ix) 200A, 228V and 234H, x)200A, 228I and 234H, xi) 200A, 228V or 228I, 234H and 241R, xii) 168I, 200A, 228I or 228V and 234H, xiii) 69N, 168I, 200A, 228V or 228I, 234H, 252M and 263E, xiv) 69N, 168I, 200A, 201K, 228V or 228I, 234H, 252M and 263E, xv)69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 252M and 263E, (xvi) 168I, 200A, 228I or 228V, 234H and 241R, (xvii) 69N, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E, (xviii) 69N, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E, (xix) 69N, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E, xx) 112L, 200A, 228V or 228I, 234H and 241R, (xxi) 112L, 168I, 200A, 228I or 228V, 234H and 241R, xxii) 69N, 112L, 168I, 200A, 228V or 228I, 234H, 241R, 252M and 263E, xxiii) 69N, 112L, 168I, 200A, 201K, 228V or 228I, 234H, 241R, 252M and 263E, or xxiv) 69N, 112L, 168I, 200A, 201K, 224R, 228I or 228V, 234H, 241R, 252M and 263E, Alternatively, when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39, the same amino acid substitution is located at the corresponding amino acid position.

8. The plant cell according to claim 7, wherein the modified NifH polypeptide comprises amino acids 200A, 228V and 234H of reference SEQ ID NO:37, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:

39.

9. The plant cell according to any one of claims 1 to 8, wherein: a) The modified NifH polypeptide comprises one or more of the following motifs: YGKGGIGKSTTXQN (SEQ ID NO:61), IXGCDPKAD (SEQ ID NO:62), CXESGGPEPGVGCAGRG (SEQ ID NO:63), DVLGDVVCGGFAMP (SEQ ID NO:43), VXSGEMMAXYAANNI (SEQ ID NO:64), and CNSRXXD (motif VII, SEQ ID NO:65), preferably at least DVLGDVVCGGFAMP (SEQ ID NO:43), wherein each X independently represents any amino acid; or b) The modified NifH polypeptide comprises one or more of the following motifs: YGKGGIGKSTTXQNT (SEQ ID NO:40), IHGCDPKAD (SEQ ID NO:41), CVESGGPEPGVGCAGRG (SEQ ID NO:42), DVLGDVVCGGFAMP (SEQ ID NO:43), VASGEMMAXYAANNI (SEQ ID NO:44), QSGVR (SEQ ID NO:45), and CNSRXVD (SEQ ID NO:46), preferably at least DVLGDVVCGGFAMP (SEQ ID NO:43), wherein each X independently represents any amino acid.

10. The plant cell according to any one of claims 1 to 9, wherein: a) The modified NifH polypeptide has 12, 13, 14, 15, or all of the following amino acids: 4K, 22T, 37H, 52G, 60D, 63R, 108L, 109M, 142G, 151A, 174Q, 189V, 198E, 199F, 222F, and 247I as per SEQ ID NO:37, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39; and / or b) The modified NifH polypeptide has 130, 131, 132, 133, 134, 135, 136, or all 137 of the following amino acids: 3R, 4K, 6A, 8Y, 9G, 10K, 11G, 12G, 13I, 14G, 15K, 16S, 17T, 18T, 20Q, 21N, 22T, 25A, 36I, 37H, 38G, 39C, 40D, 41P, 42K, 43A, 44D, 46T, 47R, 50L, 52G, 55Q, 60D, 63R, 75V, 79G, 85C, 86V, 87E, 88S, 89G ... G, 90G, 91P, 92E, 93P, 94G, 95V, 96G, 97C, 98A, 99G, 100R, 101G, 103I, 104T, 106I, 108L, 109M, 110E, 115Y, 119L, 120D, 125D, 126V, 127L, 128G, 129D, 130V, 131V, 132C, 133G, 134G, 135F, 136A, 137M, 13 8P, 140R, 142G, 143K, 144A, 146E, 148Y, 150V, 151A, 152S, 153G, 154E, 155M, 156M, 157A, 159Y, 160 A, 161A, 162N, 163N, 164I, 167G, 170K, 172A, 174Q, 175S, 176G, 177V, 178R, 180G, 181G, 184C, 185N, 186S, 187R, 189V, 190D, 192E, 198E, 199F, 204G, 212P, 213R, 215N, 217V, 218Q, 220A, 221E, 222F, 227V, 236Q, 239E, 240Y, 243L, 247I, 250N, 254V, 255I, 256P, 258P, 265E, and 272G, or the same amino acid at the corresponding amino acid position when the modified NifH sequence is compared with SEQ ID NO:37 and / or SEQ ID NO:39; and / or c) The amino acid sequence of the modified NifH polypeptide is at least 60% identical to the amino acid sequence provided in SEQ ID NO:37 and / or SEQ ID NO:39, preferably at least 70% or at least 80%, more preferably at least 90%, and most preferably at least 95%.

11. The plant cell according to any one of claims 1 to 10, wherein the modified NifH polypeptide is a modified AnfH polypeptide. Optionally, the amino acid sequence of the modified AnfH polypeptide is at least 70% identical to the amino acid sequence provided in SEQ ID NO:37, preferably at least 80% identical, more preferably at least 90% identical, and most preferably at least 95% identical. Preferably, the modified AnfH polypeptide comprises at least amino acids 2-275 of the amino acid sequence provided in SEQ ID NO:78, or comprises SEQ ID NO:

78.

12. The plant cell according to any one of claims 1 to 11, wherein: a) The modified NifH polypeptide contains an Fe-S cluster; and / or b) The modified NifH polypeptide is a cleavage product of a NifH fusion polypeptide comprising a mitochondrial targeting peptide (MTP) translatably fused to the modified NifH polypeptide, wherein the MTP is preferably translatably fused at the N-terminus of the modified NifH polypeptide, wherein the modified NifH polypeptide is produced in plant cells by protease cleavage of the NifH fusion polypeptide within or adjacent to the MTP, optionally wherein the NifH fusion polypeptide is cleaved within the MTP by a mitochondrial processing protease (MPP) to produce the modified NifH polypeptide, wherein the modified NifH polypeptide comprises (i) a C-terminal peptide (scar peptide) from the MTP located at its N-terminus, or (ii) does not contain a C-terminal peptide from the MTP; and / or c) The modified NifH polypeptide, in combination with NifD and NifK polypeptides or with AnfD, AnfK, and AnfG polypeptides, can (i) reduce N2 gas to produce ammonia, and / or (ii) reduce acetylene to ethylene, preferably both (i) and (ii), preferably wherein the modified NifH polypeptide is a modified AnfH polypeptide; and / or d) The plant cells further comprise exogenous polynucleotides encoding the following: (a) NifM polypeptide, NifS polypeptide, NifU polypeptide, and (i) NifD and NifK polypeptide or (ii) NifD-NifK fusion polypeptide, or (b) NifS polypeptide, NifU polypeptide and AnfG polypeptide, and (iii) AnfD and AnfK polypeptide or (iv) AnfD-AnfK fusion polypeptide.

13. The plant cell according to any one of claims 1 to 12, wherein the modified NifH polypeptide comprises two NifH polypeptides covalently linked by an oligopeptide linker in the N-terminal to C-terminal order NifH::linker::NifH, preferably two modified AnfH polypeptides covalently linked by an oligopeptide linker in the N-terminal to C-terminal order AnfH::linker::AnfH, optionally further comprising a C-terminal peptide from a mitochondrial targeting peptide (MTP) located at its N-terminus. Optionally, the length of the linker is 10-50 residues, preferably 16-50 residues or 20-35 residues, more preferably about 25 or about 30 residues, or most preferably 25 or 30 residues.

14. A plant cell comprising a NifH fusion polypeptide comprising two NifH polypeptides, preferably two AnfH polypeptides, covalently linked by an oligopeptide linker in the order NifH::linker::NifH from N-terminus to C-terminus, preferably in the order AnfH::linker::AnfH, optionally further comprising a C-terminal peptide from a mitochondrial targeting peptide (MTP) located at its N-terminus.

15. The plant cell according to claim 14, wherein... a) The two NifH polypeptides, preferably the two NifH polypeptides are identical, or differ only in that one of the NifH polypeptides contains methionine at amino acid position 1; and / or b) The NifH fusion polypeptide, preferably the AnfH fusion polypeptide, is more soluble in the mitochondria of the cells than the corresponding NifH or AnfH polypeptide having only a single NifH or AnfH polypeptide; and / or c) The two NifH polypeptides, preferably the two AnfH polypeptides, are each at least 70% identical to the amino acid sequence provided in SEQ ID NO:37 and / or SEQ ID NO:39, preferably at least 80% identical, more preferably at least 90% identical, and most preferably at least 95% identical; and / or d) The NifH fusion polypeptide is a cleavage product of an encoded polypeptide comprising a mitochondrial targeting peptide (MTP) translatably fused to the NifH::linker::NifH polypeptide, wherein the MTP is preferably translatably fused at the N-terminus of the NifH polypeptide, wherein the NifH fusion polypeptide is generated in the plant cell by protease cleavage of the encoded polypeptide within or adjacent to the MTP, wherein the NifH fusion polypeptide comprises (i) a C-terminal peptide from the MTP located at its N-terminus, or (ii) does not contain a C-terminal peptide from the MTP.

16. The plant cell according to any one of claims 1 to 15, wherein the modified NifH polypeptide or the NifH fusion polypeptide is a cleavage product of a polypeptide encoded by an exogenous polynucleotide in the plant cell, wherein the encoded polypeptide comprises a mitochondrial targeting peptide (MTP) translatorily fused to the modified NifH polypeptide or the NifH::linker::NifH polypeptide, wherein the MTP is preferably translatorily fused at the N-terminus of the modified NifH polypeptide or the NifH::linker::NifH polypeptide, wherein the modified NifH polypeptide or the NifH fusion polypeptide is produced in the plant cell by protease cleavage of the encoded polypeptide within or adjacent to the MTP. Optionally, the exogenous polypeptide is integrated into the genome of the plant cell, preferably into the nuclear genome, and / or the modified NifH polypeptide is a modified AnfH polypeptide, and / or the NifH fusion polypeptide is an AnfH fusion polypeptide.

17. A modified NifH polypeptide as defined in any one of claims 1 to 13, a NifH fusion polypeptide as defined in claim 14 or claim 15, or an encoded polypeptide as defined in claim 15 or claim 16.

18. The encoded polypeptide of claim 17, wherein the MTP is translatorily fused at the N-terminus of the modified NifH polypeptide or the NifH::linker::NifH polypeptide.

19. The modified NifH polypeptide, NifH fusion polypeptide, or encoded polypeptide according to claim 17 or claim 18, wherein the modified AnfH polypeptide, AnfH fusion polypeptide, or encoded polypeptide comprises one or two AnfH polypeptides respectively, and optionally is further characterized by one or more of the features defined in any one of claims 1 to 16.

20. The encoded polypeptide of claim 19, comprising an AnfH polypeptide and a mitochondrial targeting peptide (MTP), wherein the MTP is preferably translatorily fused to the AnfH polypeptide at the N-terminus.

21. An exogenous polynucleotide comprising a promoter operatively linked to a nucleotide sequence, said exogenous polynucleotide encoding a modified NifH polypeptide as defined in any one of claims 1 to 13, a NifH fusion polypeptide as defined in claim 14 or 15, or an encoded polypeptide as defined in claim 15 or 16, wherein said promoter directs the expression of said nucleotide sequence in cells, preferably plant cells.

22. The exogenous polynucleotide according to claim 21, wherein... a) The modified NifH polypeptide is a modified AnfH polypeptide, the NifH fusion polypeptide is an AnfH fusion polypeptide, or the encoded polypeptide is an AnfH polypeptide; and / or b) The exogenous polynucleotide is integrated into the genome of the cell, preferably into the nuclear genome of a plant cell; and / or c) The protein-coding region of the polynucleotide has been codon-modified for expression in plant cells.

23. A vector comprising the exogenous polynucleotide according to claim 21 or claim 22.

24. A transgenic plant or a portion thereof, comprising one or more of the following: plant cells according to any one of claims 1 to 16, a modified NifH polypeptide as defined in any one of claims 1 to 13, a NifH fusion polypeptide as defined in claim 14 or claim 15, an encoded polypeptide as defined in claim 15 or claim 16, or an exogenous polynucleotide as defined in claim 21 or claim 22.

25. The transgenic plant or a portion thereof according to claim 24, wherein the modified NifH polypeptide, the NifH fusion polypeptide, or the encoded polypeptide is a modified AnfH polypeptide, AnfH fusion polypeptide, or encoded polypeptide as defined in any one of claims 1 to 16, each comprising one or two AnfH polypeptides.

26. The transgenic plant or a portion thereof according to claim 24 or claim 25, wherein the portion is a transgenic seed.

27. The genetically modified plant or a portion thereof according to any one of claims 24 to 26, wherein it is a cereal plant, preferably wheat, rice, corn, rye, oats or barley or a portion thereof.

28. A method for selecting a modified NifH polypeptide, preferably a modified AnfH polypeptide or a NifH fusion polypeptide, preferably an AnfH fusion polypeptide, the method comprising... i) Expressing the exogenous polynucleotide according to claim 21 or claim 22 in plant cells. ii) Extract a protein containing the modified NifH polypeptide or the NifH fusion polypeptide from the mitochondria of the plant cells. iii) Determine the solubility level of the modified NifH peptide or the NifH fusion peptide in the extracted protein, and iv) Select the modified NifH polypeptide or the NifH fusion polypeptide, wherein at least 15%, preferably at least 50%, of the modified NifH polypeptide or the NifH fusion polypeptide in the cells is soluble.

29. A method for producing a modified NifH polypeptide, preferably a modified AnfH polypeptide or a NifH fusion polypeptide, preferably an AnfH fusion polypeptide, said method comprising expressing the exogenous polynucleotide according to claim 21 or claim 22 in plant cells or a transgenic plant.

30. A method for selecting plant cells, said plant cells producing a modified NifH polypeptide, preferably a modified AnfH polypeptide or a NifH fusion polypeptide, preferably an AnfH fusion polypeptide, said method comprising... i) Expressing the exogenous polynucleotide according to claim 21 or claim 22 in plant cells. ii) Determine whether the modified NifH peptide or the NifH fusion peptide is produced at the desired level and / or has the desired activity. iii) Select the plant cells based on the results of step ii). iv) Optionally, transgenic plants are produced from selected plant cells, and v) Optionally, the transgenic plant produces transgenic progeny plants and / or transgenic seeds.

31. A method for selecting plants, said plants producing modified NifH polypeptides, preferably modified AnfH polypeptides or NifH fusion polypeptides, preferably AnfH fusion polypeptides, said method comprising... i) Expressing the exogenous polynucleotide according to claim 21 or claim 22 in one or more plants. ii) Determine whether the modified NifH polypeptide or the NifH fusion polypeptide is produced at the desired level and / or has the desired activity in the one or more plants. iii) Select a plant from step ii) that produces the modified NifH polypeptide or the NifH fusion polypeptide at a desired level and / or has the desired activity in the plant or a portion thereof, and iv) Optionally, transgenic progeny plants and / or transgenic seeds may be produced from transgenic plants.

32. The method according to any one of claims 28 to 31, wherein a) The modified NifH has one or more of the features of a modified NifH polypeptide as defined in any one of claims 1 to 19, or the NifH fusion polypeptide has one or more of the features of a NifH fusion polypeptide as defined in any one of claims 14 to 19; and / or (b) The modified NifH polypeptide has at least one amino acid substitution, which imparts a lower free energy to the modified NifH polypeptide compared to the corresponding NifH polypeptide lacking the at least one amino acid substitution.

33. Use of the exogenous polynucleotide according to claim 21 or claim 22 and / or the vector according to claim 23 for the production of transgenic plant cells.

34. A method for producing transgenic plants, the method comprising the following steps i) Introducing one or more exogenous polynucleotides according to claim 21 or claim 22 and / or the vector according to claim 23 into plant cells. ii) Regeneration of the transgenic plant according to any one of claims 24 to 27 from the cells of step i), and iii) Optionally, transgenic seeds and / or progeny plants are produced from the transgenic plant regenerated in step ii).

35. A method for producing transgenic seeds, the method comprising: i) Collecting seeds from the transgenic plant according to any one of claims 24 to 27, and / or ii) Collect seeds from one or more transgenic progeny plants produced by the method according to claim 34.

36. A method for producing flour, whole wheat flour, starch, oil, seed flour or other products obtained from seeds, the method comprising extracting flour, whole wheat flour, starch, oil or other products from the seeds of claim 26, or producing the seed flour from the seeds.

37. A product produced from a transgenic plant or a portion thereof according to any one of claims 24 to 27, wherein the product comprises one or more of the following: a modified NifH polypeptide as defined in any one of claims 1 to 13, a NifH fusion polypeptide as defined in claim 14 or claim 15, an encoded polypeptide as defined in claim 15 or claim 16, and an exogenous polynucleotide as defined in claim 21 or claim 22.

38. A method for preparing a food, the method comprising mixing the seed of claim 26 or flour, whole wheat flour, starch, oil or other product obtained from said seed with another food ingredient.

39. A method of feeding an animal, the method comprising providing the animal with a plant or a portion thereof according to any one of claims 24 to 27, or a product according to claim 37.

Citation Information

Patent Citations

  • Enhanced regeneration system for cereals

    AU1994061781A1

  • Nif variants

    AU2023902807

  • Method of transforming monocotyledon

    AU667939B

  • Enhanced regeneration system for cereals

    CA2092588A1

  • Moisture remover of dampproof box

    CN1132106A