Artificial enzymes for biocatalytic reduction of nitrogen
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for nitrogen fixation, such as the Haber-Bosch process and bacterial nitrogenases, face challenges in efficiency, cost, and environmental impact, while heterologous expression of nitrogenases is hindered by complex maturation processes and high ATP costs.
Development of Artificial Nitrogenase (ArtNzase) enzymes with modified peptides that enhance iron-sulfur cluster cofactor binding, allowing for nitrogen reduction to ammonia without ATP consumption, using optimized vectors and expression in simpler organisms.
The ArtNzase enzymes provide a decentralized, environmentally friendly method for ammonia production with improved catalytic efficiency and reduced energy consumption compared to traditional processes.
Smart Images

Figure US2025034566_26032026_PF_FP_ABST
Abstract
Description
ARTIFICIAL ENZYMES FOR BIOCATALYTIC REDUCTION OF NITROGENCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 662,706, filed June 21 , 2024, and U.S. Provisional Application No. 63 / 665,722, filed June 28, 2024, which are each incorporated by reference herein in their entireties.REFERENCE TO SEQUENCE LISTING
[0002] The sequence listing submitted on June 20, 2025, as an .XML file entitled “10046- 623W01_ST26.xml” created on June 18, 2025, and having a file size of 185,428 bytes is hereby incorporated by reference pursuant to 37 C.F.R. § 1.52(e)(5).BACKGROUND
[0003] Nitrogen is present in most biomolecules, natural products, and synthetic compounds. The nitrogen source for these molecules can often be traced to atmospheric N2. Despite its abundance, activating the 945 kJ / mol triple bond of N2 efficiently under mild conditions has been a major challenge facing chemists for over a century. While the Haber- Bosch process has been highly effective in fixing N2 throughout the 20th century, it requires large quantities of pure H2 generated by steam reforming of natural gas methane. As a result, the process not only consumes 1-2% of global energy but also generates CO2 at an equally colossal scale (-300 Mtons annually, -1.4% of global CO2 emissions).
[0004] In contrast to the Haber-Bosch process, bacteria and archaea use nitrogenases (N2ases) to fix N2 at ambient conditions, independent of H2. However, despite efforts spanning >50 years, translating N2ases into practical catalysts has been unsuccessful due to several major challenges. First, the N2ases include multiple subunits and different metallocofactors, e.g., [Fe4S4] and FeMoco, which makes improving the catalytic subunit difficult. Enzyme engineering tools like site-directed mutagenesis and directed evolution are powerful but have proven difficult to apply to N2ases, as they are hindered by limited expression vectors and phenotypic screening methods for N2ase mutants. As a result, most N2ase mutants are made through homologous recombination. However, this approach is cumbersome and prone to reverting mutants back to wild type due to bases' essential role in organism survival. Heterologous expression of enzymes in simpler organisms (e.g., E. coli) has been a common approach to engineering enzymes from higher organisms, because of higher yield and purity'.However, heterologous expression of N2ases has proven challenging due to their complex maturation process. Finally, N2ases rely on an ATP cofactor; the high cost of ATP makes it difficult to develop practical catalysts.
[0005] Thus, there is a need for improved enzy mes for reducing nitrogen. This need and others are at least partially satisfied by the present disclosure.SUMMARY
[0006] Disclosed herein are Artificial Nitrogenase (ArtNzase) enzymes which can reduce atmospheric N2 to ammonia. These enzymes can provide a new avenue to nitrogen fixation which breaks free of the Haber Bosch process as a decentralized method of ammonia production, and which is also more environmentally friendly.
[0007] In an aspect, provided is a non-naturally occurring peptide for reducing nitrogen, including: at least one modification to a naturally occurring peptide; and a binding pocket for binding an iron-sulfur cluster cofactor; wherein the naturally occurring peptide is not a nitrogenase or a subunit or domain of a nitrogenase; and wherein the at least one modification improves binding of the iron-sulfur cluster cofactor in the binding pocket.
[0008] In another aspect, provided is a vector encoding any' of the disclosed peptides.
[0009] In yet another aspect, provided is a cell including any of the disclosed vectors.
[0010] In yet still another aspect, provided is a method of producing a non-naturally occurring peptide for reducing nitrogen, the method including: a) identifying a first group of peptides each having a binding pocket similar to a naturally occurring iron-sulfur cluster cofactor binding pocket of a naturally occurring nitrogenase, wherein each of the first group of peptides is not a nitrogenase or a subunit or domain of a nitrogenase; b) selecting, from the first group of peptides, a second group of peptides each having at least one target amino acid (e.g., at least two target amino acids, at least three target amino acids, at least four target amino acids) in the binding pocket; c) selecting, from the second group of peptides, a third group of peptides each having at least one target property; and d) modifying the sequence and / or structure of each of the third group of peptides to improve binding of an iron-sulfur cluster cofactor in the binding pocket.
[0011] In yet still another aspect, provided is a method of making any of the disclosed peptides, the method including transfecting any of the disclosed vectors into a cell.
[0012] In yet still another aspect, provided is a method of making ammonia, the method including exposing any of the disclosed peptides to nitrogen and an iron-sulfur cluster cofactor
[0013] Other systems, methods, features and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features and / or advantages be included within this description and be protected by the accompanying claims.BRIEF DESCRIPTION OF DRAWINGS
[0014] FIGURE 1 depicts an overview of the protein design pipeline. Structural features of the native MoFe FeMoco site were extracted, and the entire PDB was searched for experimentally practical proteins that are compatible with the FeMoco pocket. The selected scaffolds were then further optimized and filtered to finalize the designs. Protein figure generated from the ChimeraX software suite.
[0015] FIGURE 2 depicts SASA distribution within the Rosetta matches. A higher S ASA score means more solvent exposure. Reference values for MoFe were based on NifD subunit of 1M1N PDB (8.48 A2for cluster core (ICS) and 70.1 A2for homocitrate (HCA)). The highlighted red region of 50-100 A2for HCA and 0.25-10 A2for ICS was picked for optimizations. Figure generated by MATLAB.
[0016] FIGURES 3A-3E depict selected Rosetta Matches with different SASA profiles alongside MoFe NifD. FIG. 3A is a structure with high HCA and ICS solvent exposure. FIG. 3B is a structure with high ICS solvent exposure and HCA burial. FIG. 3C is a structure with burial of both HCA and ICS. FIG. 3D is a structure with HCA and ICS SASA score close to native MoFe FIG. 3E is the MoFe nitrogenase NifD subunit (PDB ID: 3U7Q).
[0017] FIGURE 4 depicts an acety lene reduction assay for different engineered designs compared to the free cluster. Activity' was compared to the measurement at the 5 -minute mark of the assay start. TrplO-KO stands for the knockout mutant of the Trp-10 design in which the intended His442 and Cys275 interactions were replaced with alanine.
[0018] FIGURE 5 depicts X-Band EPR of Trp-10 and its comparison to the free FeMoco and resting state MoFe.
[0019] FIGURE 6A depicts cyanide reduction product distribution for Trp-10 vs. free FeMoco. Trp-10 does not produce any three-carbon hydrocarbons. FIGURE 6B depicts frequency-selective pulse 'H-NMR spectroscopy of13N2-incubated trp-10 and evidence of N2 reduction.
[0020] FIGURE 7 depicts utilized interactions for the constraint file. The structure was taken from PDB ID 3U7Q.
[0021] FIGURE 8 depicts EPR of native MoFe (with FeMoco) in pink and trp-dFABp- 5gkb bound FeMoco in green.
[0022] FIGURE 9 depicts example iron-sulfur cluster cofactors.
[0023] FIGURE 10 depicts a comparison between Mo nitrogenase from Azotobacter yinelandii and the design of artificial nitrogenase. The native IShase FeMoco site is parameterized and the entire PDB is searched for experimentally viable proteins that are compatible with the FeMoco pocket. The selected scaffolds are further optimized and filtered to finalize the designs. Solvent accessible pockets in are displayed by a mesh surface. N2ase structure taken from PDB id 3U7Q and ArtN2ase-10 computational structure used as the ArtNzase representative. Protein figure generated via ChimeraX.
[0024] FIGURE 11 depicts a visualization of the FeMoco pocket of MoFe NifD subunit, parameterized using Fpocket. PDB ID:3U7Q.
[0025] FIGURE 12 depicts an acetylene reduction assay of ArtN2ase-10 compared to Native Nzase, FeMoco and the AA variant.
[0026] FIGURE 13 depicts X-Band EPR of ArtN2ase-10 and its AA mutant compared to the spectra of FeMoco and Native N2ase in different forms. Simulated spectra for the ArtN2ase- 10 variants are overlay ed as black lines. Spectra are normalized according to their geff ~4.5 peak, except ArtN2ase-10-AA spectra which is plotted with its original intensity relative to ArtN2ase-10 at the same conditions.
[0027] FIGURE 14 depicts frequency-selective pulse 'H- MR spectroscopy of15N2 reduction by ArtN2ase-10. The marked peaks correspond to the triplet splitting from14N ammonium which corresponds to trace amounts present within the solvent and not from the15N2 reduction.DETAILED DESCRIPTION
[0028] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate aspects, can also be provided in combination with a single aspect. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single aspect, can also be provided separately or in any suitable subcombination. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure.DEFINITIONS
[0029] In this specification and in the claims that follow, reference will be made to a number of terms, which shall be defined to have the following meanings:
[0030] As used herein, “comprising’’ is to be interpreted as specifying the presence of the stated features, integers, steps, or components as referred to, but does not preclude the presence or addition of one or more features, integers, steps, or components, or groups thereof. Moreover, each of the terms “by”, “comprising,” “comprises”, “comprised of,” “including,” “includes,” “included,” “involving,” “involves,” “involved,” and “such as” are used in their open, non-limiting sense and may be used interchangeably. Further, the term “comprising” is intended to include examples and aspects encompassed by the terms “consisting essentially of’ and “consisting of.” Similarly, the term “consisting essentially of’ is intended to include examples encompassed by the term “consisting of.
[0031] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a compound”, “a composition”, or “a cancer”, includes, but is not limited to, two or more such compounds, compositions, or cancers, and the like.
[0032] It should be noted that ratios, concentrations, amounts, and other numerical data can be expressed herein in a range format. It can be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it can be understood that the particular value forms a further aspect. For example, if the value “about 10” is disclosed, then “10” is also disclosed.
[0033] When a range is expressed, a further aspect includes from the one particular value and / or to the other particular value. For example, where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, e.g. the phrase “x to y” includes the range from ‘x’ to ‘y’ as well as the range greater than ‘x’ and less than ‘y’. The range can also be expressed as an upper limit, e.g. ‘about x, y, z. or less’ and should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges of Tess than x’. less than y’. and Tess than z’. Likewise, the phrase 'about x, y, z, or greater’ should be interpreted to include the specific ranges of ‘aboutx’, ‘about y’, and ‘about z’ as well as the ranges of ‘greater than x', greater than y', and ‘greater than z’. In addition, the phrase “about ‘x’ to ‘y’”. where ‘x’ and ‘y’ are numerical values, includes “about ‘x’ to about ‘y’”.
[0034] It is to be understood that such a range format is used for convenience and brevity, and thus, should be interpreted in a flexible manner to include not only the numerical values explicitly recited as the limits of the range, but also to include all the individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly recited. To illustrate, a numerical range of “about 0.1% to 5%” should be interpreted to include not only the explicitly recited values of about 0.1% to about 5%, but also include individual values (e.g., about 1%, about 2%, about 3%, and about 4%) and the subranges (e.g., about 0.5% to about 1.1%: about 5% to about 2.4%; about 0.5% to about 3.2%, and about 0.5% to about 4.4%, and other possible sub-ranges) within the indicated range.
[0035] As used herein, the terms “about,” “approximate,” “at or about,” and “substantially” mean that the amount or value in question can be the exact value or a value that provides equivalent results or effects as recited in the claims or taught herein. That is, it is understood that amounts, sizes, formulations, parameters, and other quantities and characteristics are not and need not be exact, but may be approximate and / or larger or smaller, as desired, reflecting tolerances, conversion factors, rounding off, measurement error and the like, and other factors known to those of skill in the art such that equivalent results or effects are obtained. In some circumstances, the value that provides equivalent results or effects cannot be reasonably determined. In such cases, it is generally understood, as used herein, that “about” and “at or about” mean the nominal value indicated ±10% variation unless otherwise indicated or inferred. In general, an amount, size, formulation, parameter or other quantity or characteristic is “about,” “approximate,” or “at or about” whether or not expressly stated to be such. It is understood that where “about,” “approximate,” or “at or about” is used before a quantitative value, the parameter also includes the specific quantitative value itself, unless specifically stated otherwise.
[0036] As used herein, the terms “optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0037] As used herein, the term “nucleic acid” or “nucleic acid sequence” refers to the order or sequence of nucleotides along a strand of nucleic acids. In some cases, the order of these nucleotides may determine the order of the amino acids along a corresponding polypeptide chain. The nucleic acid sequence thus codes for the amino acid sequence. Thenucleic acid sequence may be single-stranded or double-stranded, as specified, or contain portions of both double-stranded and single-stranded sequences. The nucleic acid sequence may be composed of DNA, both genomic and cDNA, RNA, or a hybrid, where the sequence comprises any combination of deoxyribo- and ribo-nucleotides, and any combination of bases, including uracil (U), adenine (A), thymine (T), cytosine (C), guanine (G), inosine, xanthine hypoxanthine, isocytosine, isoguanine, etc. It may include modified bases, including locked nucleic acids, peptide nucleic acids and others known to those skilled in the art.
[0038] As used herein, “amino acid” refers to a compound containing both amino ( — NH2) and carboxyl ( — COOH) groups generally separated by one carbon atom. The central carbon atom may contain a substituent which can be either charged, ionizable, hydrophilic or hydrophobic. Any of 22 basic building blocks of proteins having the formula NH2 — CHR — COOH, where R is different for each specific amino acid, and the stereochemistry is in the ‘L’ configuration. Additionally, the term “amino acid” can optionally include those with an unnatural ‘D’ stereochemistry and modified forms of the ‘D’ and ‘L’ amino acids.
[0039] As used herein, "protein," "peptide," and "polypeptide" are used interchangeably to denote an amino acid polymer or a set of two or more interacting or bound ammo acid polymers.
[0040] As used herein, the term “enzy me” refers to a protein which can catalyze or facilitate a chemical reaction or biological process. As used herein, the term “cofactor” refers to a non-protein compound that operates in combination with an enzyme that catalyzes a reaction of interest.PEPTIDES
[0041] In one aspect, provided is a non-naturally occurring peptide for reducing nitrogen, including: at least one modification to a naturally occurring peptide; and a binding pocket for binding an iron-sulfur cluster cofactor; wherein the naturally occurring peptide is not a nitrogenase or a subunit or domain of a nitrogenase; and wherein the at least one modification improves binding of the iron-sulfur cluster cofactor in the binding pocket.
[0042] In some aspects, the iron-sulfur cluster cofactor can be 7Fe-9S-C-Mo-R- homocitrate (FeMoco), 7Fe-9S-C-V-R-homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof. For example, in some aspects, the iron-sulfur cluster cofactor can have a structure as depicted in FIG. 9.
[0043] In some aspects, the at least one modification can include a substitution of at least one amino acid (e.g., at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, or at least six amino acids, or more) in the binding pocket.
[0044] In some aspects, the substitution can create at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; the at least one new interaction can replicate or mimic a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and the at least one naturally occurring amino acid can be, for example, His442, Cys275, Hisl95. Glnl91, Arg96, Arg359. Val70. Ser278, Tyr229. and / or Phe381 of SEQ ID NO: 67. In some such aspects, the at least one naturally occurring amino acid can be, for example, His442, Cys275, Hisl95, and / or Glnl91 of SEQ ID NO: 67.SEQ ID NO: 67 (Alpha chain of PDB ID 1M1N, EC: 1.18.6.1, UNIPROT ID P07329) MTGMSREEVESLIQEVLEVYPEKARKDRNKHLAVNDPAVTQSKKCIISNKKSQPGL MTIRGCAYAGSKGVVWGPIKDMIHISHGPVGCGQYSRAGRRNYYIGTTGVNAFVTM NFTSDFQEKDIVFGGDKKLAKLIDEVETLFPLNKGISVQSECPIGLIGDDIESVSKVKG AELSKTIVPVRCEGFRGVSQSLGHHIANDAVRDWVLGKRDEDTTFASTPYDVAIIGD YNIGGDAWSSRILLEEMGLRCVAQWSGDGSISEIELTPKVKLNLVHCYRSMNYISRH MEEKYGIPWMEYNFFGPTKTIESLRAIAAKFDESIQKKCEEVIAKYKPEWEAVVAKY RPRLEGKRVMLYIGGLRPRHVIGAYEDLGMEVVGTGYEFAHNDDYDRTMKEMGDS TLLYDDVTGYEFEEFVKRIKPDLIGSGIKEKFIFQKMGIPFREMHSWDYSGPYHGFDG FAIFARDMDMTLNNPCWKKLQAPWEASEGAEKVAASA
[0045] In some aspects, the at least one modification can increase the number of amino acid side chains in the binding pocket which interact with the iron-sulfur cluster cofactor. For example, in some specific aspects, at least four amino acid side chains in the binding pocket can interact with the iron-sulfur cluster cofactor, and the interactions between the at least four amino acid side chains and an iron-sulfur cluster cofactor molecule bound to the binding pocket can replicate or mimic naturally occurring interactions between an iron-sulfur cluster cofactor molecule and His 442, Cys275, Hisl95, and Glnl91 of SEQ ID NO: 67. In some such aspects, at least one (or at least two, or at least three, or at least four) of these amino acids may be introduced by the at least one modification. It is considered that, in some such aspects, more than four total amino acid side chains in the binding pocket may interact with the iron-sulfur cluster cofactor, but at least these four amino acid side chains must be present.
[0046] As an alternate example, in other specific aspects, at least six amino acid side chains in the binding pocket can interact with the iron-sulfur cluster cofactor, and the interactionsbetween the at least six amino acid side chains and an iron-sulfur cluster cofactor molecule bound to the binding pocket can replicate or mimic naturally occurring interactions between an iron-sulfur cluster cofactor molecule and, for example, His 442, Cys275, Hisl95, Glnl91, Arg96, and Arg359 of SEQ ID NO: 67. In some such aspects, at least one (or at least two, or at least three, or at least four, or at least five, or at least six) of these amino acids may be introduced by the at least one modification. It is considered that, in some such aspects, more than six total amino acid side chains in the binding pocket may interact with the iron-sulfur cluster cofactor, but at least these six amino acid side chains must be present.
[0047] An iron-sulfur cluster cofactor molecule includes two primary components - a metallic cluster core including a metal such a Mo, V, or Fe (which, in the specific example of naturally occurring iron-sulfur clusters, is generally understood to be made of one Fe4Ss (iron(III) sulfide) cluster and one MFe?S? cluster, where M is the metal) and a homocitrate or other organic group bound to the metallic cluster core. As used herein, the term “HCA value” refers to the solvent accessible surface area (i. e. , the area which, if the peptide is submerged in a solvent, would be in contact with said solvent) of the region of the binding pocket which binds the homocitrate or other organic group. As used herein, the term “ICS value” refers to the solvent accessible surface area of the region of the binding pocket which binds the metallic cluster core.
[0048] In some aspects, the peptide can have an HCA value of at least about 50 A2(e g., at least about 55 A2, at least about 60 A2, at least about 65 A2, at least about 70 A2, at least about 75 A2, at least about 80 A2, at least about 85 A2, at least about 90 A2, at least about 95 A2, at least about 100 A2). In some aspects, the peptide can have an HCA value of up to about 100 A2(e.g., up to about 95 A2, up to about 90 A2, up to about 85 A2, up to about 80 A2, up to about 75 A2, up to about 70 A2, up to about 65 A2, up to about 60 A2, up to about 55 A2, up to about 50 A2).
[0049] It is considered that the peptide can have an HCA value ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the peptide can have an HCA value of from about 50 A2to about 100 A2(e.g., from about 55 A2to about 95 A2, from about 60 A2to about 90 A2, from about 65 A2to about 85 A2, from about 70 A2to about 80 A2, from about 50 A2to about 75 A2, from about 55 A2to about 70 A2, from about 60 A2to about 65 A2, from about 75 A2to about 100 A2, from about 80 A2to about 95 A2, from about 85 A2to about 90 A2).
[0050] In some aspects, the peptide can have an ICS value of at least about 0.25 A2(e.g., at least about 0.5 A2, at least about 0.75 A2, at least about 1 A2, at least about 1.5 A2, at leastabout 2 A2, at least about 2.5 A2, at least about 3 A2, at least about 3.5 A2, at least about 4 A2, at least about 4.5 A2, at least about 5 A2, at least about 5.5 A2, at least about 6 A2, at least about6.5 A2, at least about 7 A2, at least about 7.5 A2, at least about 8 A2, at least about 8.5 A2, at least about 9 A2, at least about 9.5 A2, at least about 10 A2). In some aspects, the peptide can have an ICS value of up to about 10 A2(e g., up to about 9.5 A2, up to about 9 A2, up to about8.5 A2, up to about 8 A2, up to about 7.5 A2, up to about 7 A2, up to about 6.5 A2, up to about 6 A2, up to about 5.5 A2, up to about 5 A2, up to about 4.5 A2, up to about 4 A2, up to about 3.5 A2, up to about 3 A2, up to about 2.5 A2, up to about 2 A2, up to about 1.5 A2, up to about 1 A2, up to about 0.75 A2, up to about 0.5 A2, up to about 0.25 A2).
[0051] It is considered that the peptide can have an ICS value ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the peptide can have an ICS value of from about 0.25 A2to about 10 A2(e.g., from about 0.5 A2to about 9.5 A2, from about 0.75 A2to about 9 A2, from about 1 A2to about8.5 A2, from about 1.5 A2to about 8 A2, from about 2 A2to about 7.5 A2, from about 2.5 A2to about 7 A2, from about 3 A2to about 6.5 A2, from about 3.5 A2to about 6 A2, from about 4 A2to about 5.5 A2, from about 4.5 A2to about 5 A2, from about 0.25 A2to about 5 A2, from about 0.5 A2to about 4.5 A2, from about 0.75 A2to about 4 A2, from about 1 A2to about 3.5 A2, from about 1.5 A2to about 3 A2, from about 2 A2to about 2.5 A2, from about 4.5 A2to about 10 A2, from about 5 A2to about 9.5 A2, from about 5.5 A2to about 9 A2, from about 6 A2to about 8.5 A2, from about 6.5 A2to about 8 A2, from about 7 A2to about 7.5 A2).
[0052] In some aspects, the at least one modification can alter the HCA value and / or ICS value of the peptide. For example, in some such aspects, the addition of mutations and / or structural optimizations can alter the HCA value and / or ICS value.
[0053] In some aspects, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be at least about 1.5 A (e.g., at least about 1.6 A, at least about 1.7 A, at least about 1.8 A, at least about 1.9 A, at least about 2 A, at least about 2.1 A, at least about 2.2 A, at least about 2.3 A, at least about 2.4 A. at least about 2.5 A) away from the iron-sulfur cluster cofactor molecule. In some aspects, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be up to about 2.5 A (e.g., up to about 2.4 A, up to about 2.3 A, up to about 2.2 A, up to about 2. 1 A, up to about 2 A, up to about 1.9 A, up to about 1.8 A, up to about 1.7 A, up to about 1.6 A. up to about 1.5 A) away from the iron-sulfur cluster cofactor molecule.
[0054] It is considered that, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be a distance away from the iron-sulfur cluster cofactor molecule ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be from about 1.5 A to about 2.5 A (e.g., from about 1.6 A to about 2.4 A, from about 1.7 A to about 2.3 A, from about 1.8 A to about 2.2 A, from about 1.9 A to about 2.1 A, from about 1.5 A to about 2 A, from about 1.6 A to about 1.9 A, from about 1.7 A to about 1.8 A, from about 2 A to about 2.5 A, from about 2.1 A to about 2.4 A. from about 2.2 A to about 2.3 A) away from the iron-sulfur cluster cofactor molecule.
[0055] In some aspects, the at least one modification can reduce distance between the ironsulfur cluster cofactor molecule and each amino acid side chain which interacts with said ironsulfur cluster cofactor molecule.
[0056] In some aspects, the at least one modification can improve binding complementarity with the iron-sulfur cluster cofactor, such that the native residues in the binding pocket do not disturb the binding of the iron-sulfur cluster cofactor.
[0057] In some aspects, the at least one modification can at least partially eliminate native activity of the peptide chain
[0058] In some aspects, the at least one modification can at least partially eliminate unnecessary disulfide bonds which may interfere with the process of either purification or cofactor binding.
[0059] In some aspects, the peptide can be monomeric.
[0060] In some aspects, the peptide can include up to about 400 amino acids (e.g., up to 390 amino acids, up to about 380 amino acids, up to about 370 amino acids, up to about 360 amino acids, up to about 350 amino acids, up to about 340 amino acids, up to about 330 amino acids, up to about 320 amino acids, up to about 310 amino acids, up to about 300 amino acids, up to about 290 acids, up to about 280 amino acids, up to about 270 amino acids, up to about 260 amino acids, up to about 250 amino acids, up to about 240 amino acids, up to about 230 amino acids, up to about 220 amino acids, up to about 210 amino acids, up to about 200 amino acids, up to about 190 amino acids, up to about 180 amino acids, up to about 170 amino acids, up to about 160 amino acids, up to about 150 amino acids, up to about 140 amino acids, up toabout 130 amino acids, up to about 120 amino acids, up to about 110 amino acids, up to about 100 amino acids).
[0061] In some aspects, the peptide may not have any post-translational modifications.
[0062] In some aspects, the naturally occurring peptide may not be a membrane protein.
[0063] In some aspects, the peptide can reduce N2 to ammonia.
[0064] In some aspects, the peptide may not consume ATP in the reduction of N2 to ammonia.
[0065] In some aspects, the peptide can have a turnover frequency (TOF) of ammonia production of at least about 3 min'1(e.g., at least about 3.5 min'1, at least about 4 min'1, at least about 4.5 min'1, at least about 5 min'1, at least about 5.5 min'1, at least about 6 min'1, at least about 6.5 min1, at least about 7 min'1, at least about 7.5 min1, at least about 8 min'1, at least about 8.5 min'1, at least about 9 min'1, at least about 9.5 min'1, at least about 10 min'1, at least about 10.5 min'1, at least about 11 min'1, at least about 11.5 min'1, at least about 12 min'1, at least about 12.5 min'1, at least about 13 min'1, at least about 13.5 min'1, at least about 14 min' \ at least about 14.5 min1, at least about 15 min1). In some aspects, the peptide can have a TOF of ammonia production of up to about 15 min'1(e.g., up to about 14.5 min'1, up to about 14 min'1, up to about 13.5 min'1, up to about 13 min'1, up to about 12.5 min'1, up to about 12 min'1, up to about 11.5 min'1, up to about 11 min'1, up to about 10.5 min'1, up to about 10 min' \ up to about 9.5 min'1, up to about 9 min'1, up to about 8.5 min'1, up to about 8 min'1, up to about 7.5 min'1, up to about 7 min'1, up to about 6.5 min'1, up to about 6 min'1, up to about 5.5 min'1, up to about 5 min'1, up to about 4.5 min'1, up to about 4 min'1, up to about 3.5 min'1, up to about 3 min'1
[0066] It is considered that the peptide can have a TOF of ammonia production ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the peptide can have a TOF of ammonia production of from about 3 min'1to about 15 min'1(e.g., from about 3.5 min'1to about 14.5 min'1, from about 4 min'1to about 14 min'1, from about 4.5 min'1to about 13.5 min'1, from about 5 min'1to about 13 min'1, from about 5.5 min'1to about 12.5 min'1, from about 6 min'1to about 12 min'1, from about 6.5 min'1to about 11.5 min'1, from about 7 min'1to about 11 min'1, from about 7.5 min'1to about 10.5 min'1, from about 8 min'1to about 10 min'1, from about 8.5 min'1to about 9.5 min'1, from about 3 min'1to about 9 min'1, from about 3.5 min'1to about 8.5 min'1, from about 4 min'1to about 8 min'1, from about 4.5 min'1to about 7.5 min'1, from about 5 min'1to about 7 min'1, from about 5.5 min1to about 6.5 min'1, from about 9 min'1to about 15 min1, from about 9.5 min'1to about 14.5 min'1, from about 10 min'1to about 14 min'1, from about 10.5min'1to about 13.5 min'1, from about 11 min'1to about 13 min'1, from about 11.5 min'1to about 12.5 min1).
[0067] In some aspects, a reductase activity of the peptide can be at least about 10% greater (e.g., at least about 20% greater, at least about 30% greater, at least about 40% greater, at least about 50% greater, at least about 60% greater, at least about 70% greater, at least about 80% greater, at least about 90% greater, at least about 100% greater, at least about 110% greater, at least about 120% greater, at least about 130% greater, at least about 140% greater, at least about 150% greater, at least about 160% greater, at least about 170% greater, at least about 180% greater, at least about 190% greater, at least about 200% greater, at least about 210% greater, at least about 220% greater, at least about 230% greater, at least about 240% greater, at least about 250% greater, at least about 260% greater, at least about 270% greater, at least about 280% greater, at least about 290% greater, at least about 300% greater, at least about 310% greater, at least about 320% greater, at least about 330% greater, at least about 340% greater, at least about 350% greater) than a reductase activity7of a naturally occurring nitrogenase.
[0068] In some aspects, a reductase activity of the peptide can be up to about 350% greater (e.g.. up to about 340% greater, up to about 330% greater, up to about 320% greater, up to about 310% greater, up to about 300% greater, up to about 290% greater, up to about 280% greater, up to about 270% greater, up to about 260% greater, up to about 250% greater, up to about 240% greater, up to about 230% greater, up to about 220% greater, up to about 210% greater, up to about 200% greater, up to about 190% greater, up to about 180% greater, up to about 170% greater, up to about 1 0% greater, up to about 150% greater, up to about 140% greater, up to about 130% greater, up to about 120% greater, up to about 110% greater, up to about 100% greater, up to about 90% greater, up to about 80% greater, up to about 70% greater, up to about 60% greater, up to about 50% greater, up to about 40% greater, up to about 30% greater, up to about 20% greater, up to about 10% greater) than a reductase activity of a naturally occurring nitrogenase.
[0069] It is considered that a reductase activity7of the peptide can be a percent greater than a reductase activity of a naturally occurring nitrogenase ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, a reductase activity of the peptide can be from about 10% to about 350% greater (e.g., from about 20% to about 340% greater, from about 30% to about 330% greater, from about 40% to about 320% greater, from about 50% to about 310% greater, from about 60% to about 300% greater, from about 70% to about 290% greater, from about 80% to about 280% greater, from about 90% to about 270% greater, from about 100% to about 260% greater, from about110% to about 250% greater, from about 120% to about 240% greater, from about 130% to about 230% greater, from about 140% to about 220% greater, from about 150% to about 210% greater, from about 160% to about 200% greater, from about 170% to about 190% greater, from about 10% to about 180% greater, from about 20% to about 170% greater, from about 30% to about 160% greater, from about 40% to about 150% greater, from about 50% to about 140% greater, from about 60% to about 130% greater, from about 70% to about 120% greater, from about 80% to about 110% greater, from about 90% to about 100% greater, from about 180% to about 350% greater, from about 190% to about 340% greater, from about 200% to about 330% greater, from about 210% to about 320% greater, from about 220% to about 310% greater, from about 230% to about 300% greater, from about 240% to about 290% greater, from about 250% to about 280% greater, from about 260% to about 270% greater) than a reductase activity of a naturally occurring nitrogenase.
[0070] In some aspects, the peptide can include about 80% similarity' or more (e.g., about81% similarity or more, about 82% similarity or more, about 83% similarity' or more, about 84% similarity or more, about 85% similarity or more, about 86% similarity or more, about 87% similarity or more, about 88% similarity or more, about 89% similarity or more, about 90% similarity or more, about 91% similarity or more, about 92% similarity or more, about 93% similarity or more, about 94% similarity or more, about 95% similarity or more, about 96% similarity or more, about 97% similarity or more, about 98% similarity or more, about99% similarity or more) to any one of SEQ ID NOs: 1-66. In some aspects, the peptide can include any one of SEQ ID NOs: 1-66.
[0071] In some aspects, the peptide can be expressed by a bacterium or a yeast.
[0072] In some aspects, the peptide can be expressed by E. coli.
[0073] In some aspects, the peptide can produced by any of the disclosed methods of producing a non-naturally occurring peptide for reducing nitrogen, which are discussed below.
[0074] In another aspect, provided is a vector encoding any of the disclosed peptides. As used herein, the term “vector” refers to any moiety which can deliver a nucleic acid sequence into a cell or virus so that the nucleic acid sequence can be replicated and / or expressed by the cell or virus. In some aspects, the vector can be a plasmid. In some aspects, the vector can be a viral vector. In some aspects, the vector can be a cosmid. In some aspects, the vector can be an artificial chromosome.
[0075] In yet another aspect, provided is a cell including any of the disclosed vectors. In some aspects, the cell can have been transfected with any of the disclosed vectors. In someaspects, the cell can be a bacterium, a yeast, an archaeon, or another prokaryotic or singlecelled organism. In some aspects, the cell can be E. coli.METHODS
[0076] In another aspect, provided is a method of producing a non-naturally occurring peptide for reducing nitrogen, the method including: a) identifying a first group of peptides each having a binding pocket similar to a naturally occurring iron-sulfur cluster cofactor binding pocket of a naturally occurring nitrogenase, wherein each of the first group of peptides is not a nitrogenase or a subunit or domain of a nitrogenase; b) selecting, from the first group of peptides, a second group of peptides each having at least one target amino acid (e.g., at least two target amino acids, at least three target amino acids, at least four target amino acids, at least five target amino acids, at least six target amino acids) in the binding pocket; c) selecting, from the second group of peptides, a third group of peptides each having at least one target property; and d) modifying the sequence and / or structure of each of the third group of peptides to improve binding of an iron-sulfur cluster cofactor in the binding pocket.
[0077] In some aspects, the iron-sulfur cluster cofactor can be 7Fe-9S-C-Mo-R- homocitrate (FeMoco). 7Fe-9S-C-V-R-homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof. For example, in some aspects, the iron-sulfur cluster cofactor can have a structure as depicted in SCHEME 1.
[0078] In some aspects, the binding pocket of each of the first group of peptides can have at least about 85% (e.g.. at least about 86%, at least about 87%. at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%) structural overlap (e.g., similarity in shape) with the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase. In some aspects, the binding pocket of each of the first group of peptides can have up to about 95% (e.g., up to about 94%, up to about 93%, up to about 92%, up to about 91%, up to about 90%, up to about 89%, up to about 88%, up to about 87%, up to about 86%, up to about 85%) structural overlap with the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0079] It is considered that the binding pocket of each of the first group of peptides can have a percent structural overlap with the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the binding pocket of each of the first group of peptides can have from about 85% to about 95% (e.g., from about 86% to about 94%, from about 87% to about 93%, from about 88% toabout 92%. from about 89% to about 91%, from about 85% to about 90%, from about 86% to about 89%. from about 87% to about 88%, from about 90% to about 95%, from about 91% to about 94%, from about 92% to about 93%) structural overlap with the naturally occurring ironsulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0080] In some aspects, the binding pocket of each of the first group of peptides can be at least about 115% (e.g., at least about 120%. at least about 125%, at least about 130%, at least about 135%, at least about 140%, at least about 145%, at least about 150%, at least about 155%. at least about 160%, at least about 165%, at least about 170%, at least about 175%, at least about 180%, at least about 185%, at least about 190%, at least about 195%, at least about 200%, at least about 205%, at least about 210%, at least about 215%) of a pocket volume of the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase. In some aspects, the binding pocket of each of the first group of peptides can be up to about 215% (e.g., up to about 210%, up to about 205%, up to about 200%, up to about195%, up to about 190%, up to about 185%, up to about 180%, up to about 175%, up to about170%, up to about 165%, up to about 160%. up to about 155%, up to about 150%, up to about145%, up to about 140%, up to about 135%. up to about 130%. up to about 125%, up to about120%, up to about 115%) of a pocket volume of the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0081] It is considered that the binding pocket of each of the first group of peptides can be a percent of a pocket volume of the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the binding pocket of each of the first group of peptides can be from about 115% to about 215% (e.g., from about 120% to about 210%, from about 125% to about 205%. from about 130% to about 200%, from about 135% to about 195%, from about 140% to about 190%, from about 145% to about 185%, from about 150% to about 180%, from about 155% to about 175%, from about 160% to about 170%, from about 115% to about 175%, from about 120% to about 170%, from about 125% to about 165%, from about 130% to about 160%, from about 135% to about 155%, from about 140% to about 150%, from about 175% to about 215%, from about 180% to about 210%, from about 185% to about 205%, from about 190% to about 200%) of a pocket volume of the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0082] In some aspects, the naturally occurring nitrogenase can be SEQ ID NO: 67.
[0083] In some aspects, the at least one target amino acid can interact with an iron-sulfur cluster cofactor molecule bound to the binding pocket; said interaction can replicate or mimic a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and the at least one naturally occurring amino acid can be His442, Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67. In some such aspects, the at least one naturally occurring amino acid can be His442, Cys275, Hisl95, and / or Glnl91 of SEQ ID NO: 67.
[0084] In some aspects, the at least one target property can be: having up to about 400 amino acids (e.g., up to 390 amino acids, up to about 380 amino acids, up to about 370 amino acids, up to about 360 amino acids, up to about 350 amino acids, up to about 340 amino acids, up to about 330 amino acids, up to about 320 amino acids, up to about 310 amino acids, up to about 300 amino acids, up to about 290 acids, up to about 280 amino acids, up to about 270 amino acids, up to about 260 amino acids, up to about 250 amino acids, up to about 240 amino acids, up to about 230 amino acids, up to about 220 amino acids, up to about 210 amino acids, up to about 200 amino acids, up to about 190 amino acids, up to about 180 amino acids, up to about 170 amino acids, up to about 160 amino acids, up to about 150 amino acids, up to about 140 amino acids, up to about 130 amino acids, up to about 120 amino acids, up to about 110 amino acids, up to about 100 amino acids); being a monomeric peptide; having no post- translational modifications; not being a membrane protein; having a target HCA value and / or a target ICS value; and / or being expressible by a bacterium or yeast.
[0085] In some aspects, the target HCA value can be at least about 50 A2(e.g., at least about 55 A2, at least about 60 A2, at least about 65 A2, at least about 70 A2, at least about 75 A2, at least about 80 A2, at least about 85 A2, at least about 90 A2, at least about 95 A2, at least about 100 A2). In some aspects, the target HCA value can be up to about 100 A2(e.g., up to about 95 A2, up to about 90 A2, up to about 85 A2, up to about 80 A2, up to about 75 A2, up to about 70 A2, up to about 65 A2, up to about 60 A2, up to about 55 A2, up to about 50 A2).
[0086] It is considered that the target HCA value can range from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the target HCA value can be from about 50 A2to about 100 A2(e.g., from about 55 A2to about 95 A2, from about 60 A2to about 90 A2, from about 65 A2to about 85 A2, from about 70 A2to about 80 A2, from about 50 A2to about 75 A2, from about 55 A2to about 70 A2, from about 60 A2to about 65 A2, from about 75 A2to about 100 A2, from about 80 A2to about 95 A2, from about 85 A2to about 90 A2).
[0087] In some aspects, the target ICS value can be at least about 0.25 A2(e.g., at least about 0.5 A2, at least about 0.75 A2, at least about 1 A2, at least about 1.5 A2, at least about 2 A2, at least about 2.5 A2, at least about 3 A2, at least about 3.5 A2, at least about 4 A2, at least about 4.5 A2, at least about 5 A2, at least about 5.5 A2, at least about 6 A2, at least about 6.5 A2, at least about 7 A2, at least about 7.5 A2, at least about 8 A2, at least about 8.5 A2, at least about 9 A2, at least about 9.5 A2, at least about 10 A2). In some aspects, the target ICS value can be up to about 10 A2(e.g., up to about 9.5 A2, up to about 9 A2, up to about 8.5 A2, up to about 8 A2, up to about 7.5 A2, up to about 7 A2, up to about 6.5 A2, up to about 6 A2, up to about 5.5 A2, up to about 5 A2, up to about 4.5 A2, up to about 4 A2, up to about 3.5 A2, up to about 3 A2, up to about 2.5 A2, up to about 2 A2, up to about 1.5 A2, up to about 1 A2, up to about 0.75 A2, up to about 0.5 A2, up to about 0.25 A2).
[0088] It is considered that the target ICS value can range from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the target ICS value can be from about 0.25 A2to about 10 A2(e.g., from about 0.5 A2to about 9.5 A2, from about 0.75 A2to about 9 A2, from about 1 A2to about 8.5 A2, from about 1.5 A2to about 8 A2, from about 2 A2to about 7.5 A2, from about 2.5 A2to about 7 A2, from about 3 A2to about 6.5 A2, from about 3.5 A2to about 6 A2, from about 4 A2to about 5.5 A2, from about 4.5 A2to about 5 A2, from about 0.25 A2to about 5 A2, from about 0.5 A2to about 4.5 A2, from about 0.75 A2to about 4 A2, from about 1 A2to about 3.5 A2, from about 1.5 A2to about 3 A2, from about 2 A2to about 2.5 A2, from about 4.5 A2to about 10 A2, from about 5 A2to about 9.5 A2, from about 5.5 A2to about 9 A2, from about 6 A2to about 8.5 A2, from about 6.5 A2to about 8 A2, from about 7 A2to about 7.5 A2).
[0089] In some aspects, the at least one target property can be being expressible by E. coli.
[0090] In some aspects, step d) can include substitution of at least one amino acid (e.g., at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, at least six amino acids) in the binding pocket.
[0091] In some aspects, the substitution can create at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; the at least one new interaction can replicate or mimic a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and at least one naturally occurring amino acid can be His442, Cys275, Hisl95, Glnl91, Arg96, Arg359. Val70. Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67. In somesuch aspects, the at least one naturally occurring amino acid can be His442, Cys275, Hisl95, and / or Glnl91 of SEQ ID NO: 67.
[0092] In some aspects, step d) can include altering the HCA value and / or ICS value of the peptide. For example, in some such aspects, the addition of mutations and / or structural optimizations can alter the HCA value and / or ICS value.
[0093] In some aspects, step d) can include reducing distance between an iron-sulfur cluster cofactor molecule bound to the binding pocket and each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule.
[0094] In some aspects, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be at least about 1.5 A (e.g., at least about 1.6 A, at least about 1.7 A, at least about 1.8 A, at least about 1.9 A, at least about 2 A, at least about 2.1 A, at least about 2.2 A, at least about 2.3 A, at least about 2.4 A, at least about 2.5 A) away from the iron-sulfur cluster cofactor molecule. In some aspects, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be up to about 2.5 A (e.g., up to about 2.4 A, up to about 2.3 A, up to about 2.2 A, up to about 2. 1 A, up to about 2 A, up to about 1 .9 A, up to about 1.8 A, up to about 1.7 A, up to about 1.6 A, up to about 1.5 A) away from the iron-sulfur cluster cofactor molecule.
[0095] It is considered that, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be a distance away from the iron-sulfur cluster cofactor molecule ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule can independently be from about 1.5 A to about 2.5 A (e.g., from about 1.6 A to about 2.4 A, from about 1.7 A to about 2.3 A, from about 1.8 A to about 2.2 A, from about 1.9 A to about 2.1 A, from about 1.5 A to about 2 A, from about 1.6 A to about 1.9 A, from about 1.7 A to about 1.8 A, from about 2 A to about 2.5 A, from about 2.1 A to about 2.4 A, from about 2.2 A to about 2.3 A) away from the iron-sulfur cluster cofactor molecule.
[0096] In some aspects, the modification in step d) can improve binding complementarity with the iron-sulfur cluster cofactor, such that the native residues in the binding pocket do not disturb the binding of the iron-sulfur cluster cofactor.
[0097] In some aspects, the modification in step d) can at least partially eliminate native activity of the peptide chain
[0098] In some aspects, the modification in step d) can at least partially eliminate unnecessary disulfide bonds which may interfere with the process of either purification or cofactor binding.
[0099] In some aspects, the method can produce a peptide including about 80% similarity or more (e.g.. about 81% similarity or more, about 82% similarity or more, about 83% similarity’ or more, about 84% similarity or more, about 85% similarity or more, about 86% similarity’ or more, about 87% similarity or more, about 88% similarity or more, about 89% similarity or more, about 90% similarity' or more, about 91% similarity or more, about 92% similarity or more, about 93% similarity or more, about 94% similarity or more, about 95% similarity’ or more, about 96% similarity or more, about 97% similarity or more, about 98% similarity’ or more, about 99% similarity or more) to any one of SEQ ID NOs: 1-66. In some aspects, the method can produce a peptide including any one of SEQ ID NOs: 1-66.
[0100] In some aspects, the method can further include: e) computationally analyzing each of the third group of peptides to select a fourth group of peptides having low free energy, high binding affinity for the iron-sulfur cluster cofactor, high geometrical similarity to a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase, and / or a minimal number of mutations.
[0101] In some aspects, each of the fourth group of peptides can be selected from at least a top 5% (e g., at least about a top 4.5%, at least about a top 4%, at least about a top 3.5%, at least about a top 3%, at least about a top 2.5%, at least about a top 2%, at least about a top 1.5%, at least about a top 1%, at least about a top 0.5%) of the third group of peptides when ranked by lowest free energy. In some aspects, each of the fourth group of peptides can be selected from up to about a top 0.5% (e.g., up to about a top 1%, up to about a top 1.5%, up to about a top 2%, up to about a top 2.5%, up to about a top 3%, up to about a top 3.5%, up to about a top 4%, up to about a top 4.5%, up to about a top 5%) of the third group of peptides when ranked by lowest free energy.
[0102] It is considered that each of the fourth group of peptides can be selected from a percentile ranging from any of the minimum values described above to any of the maximum values described above of the third group of peptides when ranked by lowest free energy. For example, in some aspects, each of the fourth group of peptides can be selected from about a top 5% to about a top 0.5% (e.g., from about a top 4.5% to about a top 1%, from about a top 4% to about a top 1.5%, from about a top 3.5% to about a top 2%, from about a top 3% to abouta top 2.5%, from about a top 5% to about a top 2.5%, from about a top 4.5% to about a top 3%, from about a top 4% to about a top 3.5%, from about a top 3% to about a top 0.5%, from about a top 2.5% to about a top 1 %, from about a top 2% to about a top 1.5%) of the third group of peptides when ranked by lowest free energy. For example, the third group of peptides may have a given range of free energies, and the fourth group of peptides can be selected such that each of the fourth group of peptides has a free energy that is in about the lowest 5% to about the lowest 0.5% said range of free energies.
[0103] In another aspect, provided is a method of making any of the disclosed peptides, the method including transfecting any of the disclosed vectors into a cell. In some aspects, the cell can be a bacterium, a yeast, an archaeon, or another prokaryotic or single-celled organism. In some aspects, the cell can be E. coli.
[0104] In yet another aspect, provided is a method of making ammonia, the method including exposing any of the disclosed peptides to nitrogen and an iron-sulfur cluster cofactor
[0105] In some aspects, the nitrogen can be gaseous N2. In some aspects, the gaseous N2 can be part of a bulk gas including at least about 50% N2 (e.g., at least about 55% N2, at least about 60% N2, at least about 65% N2, at least about 70% N2, at least about 75% N2, at least about 80% N2, at least about 85% N2, at least about 90% N2, at least about 95% N2, about 100% N2). In some aspects, the gaseous N2 can be part of a bulk gas including up to about 100% N2 (e.g.. up to about 95% N2, up to about 90% N2, up to about 85% N2. up to about 80% N2. up to about 75% N2, up to about 70% N2, up to about 65% N2, up to about 60% N2, up to about 55% N2, up to about 50% N2).
[0106] It is considered that the gaseous N2 can be part of a bulk gas including N2 in an amount ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the gaseous N2 can be part of a bulk gas including from about 50% N2 to about 100% N2 (e.g., from about 55% N2 to about 95% N2, from about 60% N2 to about 90% N2, from about 65% N2 to about 85% N2, from about 70% N2 to about 80% N2, from about 50% N2 to about 75% N2, from about 55% N2 to about 70% N2. from about 60% N2 to about 65% N2. from about 75% N2 to about 100% N2, from about 80% N2 to about 95% N2, from about 85% N2 to about 90% N2).
[0107] In some aspects, the iron-sulfur cluster cofactor can be 7Fe-9S-C-Mo-R- homocitrate (FeMoco), 7Fe-9S-C-V-R-homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof. For example, in some aspects, the iron-sulfur cluster cofactor can have a structure as depicted in SCHEME 1.
[0108] In some aspects, the method may be carried out at a temperature less than the temperature of the Haber-Bosch process. For example, in some aspects, the method may be carried out a temperature of at least about 25°C (e.g., at least about 30°C, at least about 35°C, at least about 40°C, at least about 45°C, at least about 50°C, at least about 55°C, at least about 60°C, at least about 65°C, at least about 70°C, at least about 75°C, at least about 80°C, at least about 85°C, at least about 90°C, at least about 95°C, at least about 100°C). In some aspects, the method may be carried out at a temperature of up to about 100°C (e.g., up to about 95°C. up to about 90°C, up to about 85°C, up to about 80°C, up to about 75°C, up to about 70°C, up to about 65°C, up to about 60°C, up to about 55°C, up to about 50°C, up to about 45°C, up to about 40°C, up to about 35°C, up to about 30°C, up to about 25°C).
[0109] It is considered that the method may be carried out at a temperature ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the method may be carried out at a temperature of from about 25°C to about 100°C (e.g., from about 30°C to about 95°C, from about 35°C to about 90°C, from about 40°C to about 85°C, from about 45°C to about 80°C, from about 50°C to about 75°C, from about 55°C to about 70°C, from about 60°C to about 65°C. from about 25°C to about 65°C, from about 30°C to about 60°C, from about 35°C to about 65°C, from about 40°C to about 60°C, from about 45°C to about 55°C, from about 60°C to about 100°C, from about 65°C to about 95°C, from about 70°C to about 90°C, from about 75°C to about 85°C).
[0110] In some aspects, the method may be earned out at a pressure less than the pressure of the Haber-Bosch process. For example, in some aspects, the method may be carried out at a pressure of at least about 1 bar (e.g., at least about 1.5 bar, at least about 2 bar, at least about 2.5 bar. at least about 3 bar, at least about 3.5 bar, at least about 4 bar, at least about 4.5 bar, at least about 5 bar. at least about 5.5 bar, at least about 6 bar, at least about 6.5 bar, at least about 7 bar, at least about 7.5 bar, at least about 8 bar, at least about 8.5 bar, at least about 9 bar, at least about 9.5 bar, at least about 10 bar). In some aspects, the method may be carried out at a pressure of up to about 10 bar (e.g., up to about 9.5 bar, up to about 9 bar, up to about 8.5 bar, up to about 8 bar, up to about 7.5 bar, up to about 7 bar. up to about 6.5 bar, up to about 6 bar, up to about 5.5 bar, up to about 5 bar, up to about 4.5 bar. up to about 4 bar, up to about 3.5 bar, up to about 3 bar, up to about 2.5 bar, up to about 2 bar, up to about 1.5 bar, up to about 1 bar).[OHl] It is considered that the method may be carried out at a pressure ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the method may be carried out at a pressure of from about 1 bar toabout 10 bar (e.g., from about 1.5 bar to about 9.5 bar, from about 2 bar to about 9 bar, from about 2.5 bar to about 8.5 bar, from about 3 bar to about 8 bar, from about 3.5 bar to about 7.5 bar, from about 4 bar to about 7 bar, from about 4.5 bar to about 6.5 bar, from about 5 bar to about 6 bar, from about 1 bar to about 5.5 bar, from about 1.5 bar to about 5 bar, from about 2 bar to about 4.5 bar, from about 2.5 bar to about 4 bar, from about 3 bar to about 3.5 bar, from about 5.5 bar to about 10 bar, from about 6 bar to about 9.5 bar, from about 6.5 bar to about 9 bar, from about 7 bar to about 8.5 bar. from about 7.5 bar to about 8 bar).
[0112] In some aspects, the peptide may be activated for nitrogen reduction by incubating said peptide with the iron-sulfur cluster cofactor in a ratio of at least about 1:0.5 (e.g., at least about 1:0.75, at least about 1: 1, at least about 1: 1.25, at least about 1: 1.5, at least about 1: 1.75, at least about 1:2, at least about 1:2.25, at least about 1 :2.5, at least about 1 :2.75, at least about 1 :3, at least about 1 :3.25, at least about 1 :3.5, at least about 1:3.75, at least about 1 :4, at least about 1:4.25, at least about 1:4.5, at least about 1:4.75, at least about 1:5). In some aspects, the peptide may be activated for nitrogen reduction by incubating said peptide with the iron-sulfur cluster cofactor in a ratio of up to about 1:5 (e.g.. up to about 1 :4.75, up to about 1:4.5, up to about 1:4.25. up to about 1 :4, up to about 1 :3.75. up to about 1:3.5, up to about 1:3.25. up to about 1 :3, up to about 1:2.75, up to about 1 :2.5, up to about 1 :2.25, up to about 1 :2, up to about 1: 1.75, up to about 1: 1.5, up to about 1 :1.25, up to about 1: 1, up to about 1:0.75, up to about 1:0.5).
[0113] It is considered that the peptide may be activated for nitrogen reduction by incubating said peptide with the iron-sulfur cluster cofactor in a ratio ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the peptide may be activated for nitrogen reduction by incubating said peptide with the iron-sulfur cluster cofactor in a ratio of from about 1:0.5 to about 1:5 (e.g., from about 1 :0.75 to about 1 :4.75, from about 1: 1 to about 1:4.5, from about 1 : 1.25 to about 1:4.25, from about 1 : 1.5 to about 1:4, from about 1: 1.75 to about 1:3.75, from about 1:2 to about 1 :3.5, from about 1:2.25 to about 1:3.25, from about 1:2.5 to about 1 :3, from about 1:0.5 to about 1:3, from about 1:0.75 to about 1:2.75. from about 1: 1 to about 1 :2.5, from about 1:1.25 to about 1:2.25, from about 1: 1.5 to about 1:2, from about 1 :2.5 to about 1:5. from about 1 :2.75 to about 1:4.75. from about 1 :3 to about 1 :4.5, from about 1:3.25 to about 1 :4.25, from about 1:3.5 to about 1:4).
[0114] In some aspects, the method may take place in an aqueous solvent (e.g., Tris buffer). In other aspects, the method may take place in a mixed solvent.
[0115] In some aspects, the method may not consume ATP. For example, in some such aspects, the method may use a chemical electron donor. In some such aspects, the chemical electron donor can be Eu(II)-DTPA. In some aspects, the peptide and the chemical electron donor can be incubated in a ratio of at least about 1 :0.01 (e.g., at least about 1:0.02, at least about 1:0.03, at least about 1:0.04, at least about 1:0.05, at least about 1 :0.06, at least about 1:0.07, at least about 1:0.08, at least about 1:0.09, at least about 1:0.1, at least about 1 :0.15, at least about 1 :0.2. at least about 1 :0.25. at least about 1:0.3, at least about 1 :0.35, at least about 1 :0.4, at least about 1 :0.45, at least about 1:0.5). In some aspects, the peptide and the chemical electron donor can be incubated in a ratio of up to about 1:0.5 (e.g., up to about 1:0.45, up to about 1:0.4, up to about 1 :0.35, up to about 1:0.3. up to about 1:0.25, up to about 1 :0.2, up to about 1 :0.15. up to about 1 :0.1, up to about 1 :0.09, up to about 1:0.08. up to about 1:0.07, up to about 1 :0.06, up to about 1 :0.05, up to about 1:0.04, up to about 1:0.03, up to about 1 :0.02, up to about 1:0.01).
[0116] It is considered that the peptide and the chemical electron donor can be incubated in a ratio ranging from any of the minimum values described above to any of the maximum values described above. For example, in some aspects, the peptide and the chemical electron donor can be incubated in a ratio of form about 1 :0.01 to about 1:0.5 (e g., from about 1 :0.02 to about 1 :0.45, from about 1 :0.03 to about 1 :0.4, from about 1 :0.04 to about 1 :0.35, from about 1:0.05 to about 1:0.3, from about 1:0.06 to about 1:0.25, from about 1 :0.07 to about 1 :0.2, from about 1 :0.08 to about 1:0.15. from about 1 :0.09 to about 1 :0. 1, from about 1:0.01 to about 1:0.1, from about 1 :0.02 to about 1 :0.09, from about 1 :0.03 to about 1 :0.08, from about 1 :0.04 to about 1 :0.07, from about 1 :0.05 to about 1:0.06, from about 1:0.1 to about 1 :0.5, from about 1:0.15 to about 1:0.45, from about 1:0.2 to about 1:0.4, from about 1:0.25 to about 1 :0.35).
[0117] As another example, in other such aspects, the peptide may be adsorbed onto a photoexcitable surface and illuminated with a light source, thereby delivering photogenerated electrons to the peptide. An example implementation of such a method is provided in Katherine A. Brown et al., Light-driven dinitrogen reduction catalyzed by a CdS:nitrogenase MoFe protein biohybrid. Science352,448-450(2016).DOI: 10. 1126 / science.aaf2091 and Supplementary Materials, which is hereby incorporated by reference in its entirety.EXAMPLESExample 1: Artificial Nitrogenase (Art ase) Enzymes for Biocatalytic Reduction of N2 into Ammonia
[0118] Nitrogen is one of the most important elements in life. Despite it being the most abundant element in the atmosphere in form of dinitrogen, only a select microorganism can use it as their nitrogen source and other life forms depend on them directly or indirectly to obtain their nitrogen. In the early 20th century, the development of the Haber-Bosch process enabled synthetic mass-production of assimilable nitrogen in form of ammonia which was then converted into ammonia-based fertilizers. This invention revolutionized fertilizer production and has since supported the agriculture industry' to the point that half of human’s nitrogen is produced via this process. Ammonia demand will only increase with the growing earth's population, and while the Haber Bosch process has been well-optimized, it suffers from the need for hydrogen gas, which limits it to natural plants and results in high energy consumption and greenhouse gas generation. Disclosed herein is an alternative method of nitrogen fixation which is both decentralized and environmentally friendly. The disclosed artificial enzymes are capable of producing ammonia without the need of hydrogen gas.
[0119] Nitrogen is central in biology and ubiquitous in organic and pharmaceutical molecules, despite its abundance in the form of the predominant N2 gas in the atmosphere, its utilization into biological systems is very difficult energetically due to the inertness of N2. Nature activates dinitrogen and converts it into ammonia via their nitrogenase enzymes in a process called biological nitrogen fixation. Industrially this is achieved via the Haber-Bosch process. The Haber Bosch process, while being highly optimized, is limited by the requirement of hydrogen gas, which amounts to a significant amount of carbon dioxide generation and global energy usage. It also requires the production to be centralized around hydrogen production plants which requires significant logistic infrastructures for the transportation of the produced ammonia. The environmental effects and the ammonia loss during the transportation of it are significant problems that need to be addressed.
[0120] Disclosed herein is an artificial enzyme computationally engineered from native proteins or using state-of-the-art generative models to activate gaseous dinitrogen gas and catalytically reduce it to ammonia. The proteins are referred to as Artificial nitrogenase enz mes or ArtN2ase in short. The proteins function with either the native cluster of the nitrogenase enzymes or synthetic Iron-Sulfur clusters bearing structural resemblance to the native clusters. This artificial enzyme uses the air-abundant N2 gas to generate ammonia independent from the constraint of industrial ammonia synthesis without using harsh reactionconditions. Nitrogen fixation is thus achieved through nitrogenase-like catalysis without using the native enzyme. It can open up a new avenue of decentralized and environmentally friendly ammonia production, compared to present technologies which either use native nitrogenase, which needs special electrochemical setup to be commercially viable, or the Haber-Bosch (HB) process, which uses totally different conditions.
[0121] These artificial enzymes solve the need of centralized ammonia production in natural gas plants and alleviate the need of mass ammonia transportation logistics. Due to the requirements of hydrogen gas for HB process, the ammonia production requires natural gas plants, generates over one percent of global carbon dioxide emissions, and consumes more than two percent of global energy. This number will only increase with the increasing population.
[0122] These artificial enzymes also eliminate the need for additional synthetic steps of substrate protect! on / deprotecti on and ammonia surrogate usage while using the source atom, nitrogen in its cheapest form. They are both efficient and economically more viable. The enzyme can also be further developed to perform nitrogen transfer catalysis into organic substrates or to be used as a general purpose hydrogenation catalyst.Example 2: FeMoco Binding
[0123] Many synthetic models of N2ases have been prepared that either a structurally resemble FeMoco without exhibiting nitrogenase activity or b) functionally reduce N2 while lacking any significant structural resemblance, thus offering limited insights into the function of native N2ases. In addition, most of these models use strong reductants and strong acids to achieve N2 binding and activation. While these achievements are remarkable, the harsh conditions make it difficult for practical applications. Also, strong evidence shows that structural features beyond the PCS, particularly weak and non-covalent electrostatic and hydrogen bonding interactions in SCS. play essential roles in conferring and tuning N2ase activity. It is difficult to incorporate these SCS features into the models.
[0124] To overcome the above limitations of the Haber-Bosch process, native N2ases and their synthetic models, a study was conducted which explored a different approach of designing artificial metalloenzymes (ArMs) that mimic native enzymes by incorporating both the primary and secondary coordination spheres (PCS and SCS) into small, stable and monomeric proteins that can be readily expressed in E. coli with high yields. Using this approach, the study succeeded in designing heme-copper, heme-nonheme iron, and heme-[4Fe-4S] that structurally and functionally mimic heme-copper oxidases, nitric oxide reductase and sulfite reductases, respectively, and used these ArMs to gain deeper insights into native enzymes than studying native ones. Built upon these experiences, the study also reports further development of themethod for designing artificial nitrogenase (ArtlShases) that bind FeMoco and reduce N2 to NHs.
[0125] Previously, successful designs of ArMs have largely relied on either a) reengineering existing metalloproteins for new functions through rational mutagenesis, metal / cofactor substitution, or directed evolution, or b) design of a new7metal-binding site in either known protein scaffolds or de novo designed scaffolds that share similar tertiary structures. Since FeMoco is among the most unique and complex metallocofactors known to nature, neither of these approaches are feasible for the design of Artlkhases. Therefore, the study employed a four-step design approach involving 1) defining the binding pocket containing FeMoco and homocitrate, 2) using this definition to search for proteins in the PDB that are capable of forming a similar pocket, 3) screening candidates for proper cofactor orientation within the pocket, and finally 4) introducing amino acids residues that directly coordinate FeMoco, as w ell as critical SCS residues known to play key roles in N2ase function. The overall design approach is shown in FIG. 1.
[0126] The native FeMoco binding site in the NifD subunit was used to define the target pocket geometry and as reference for evaluation of potential matches. PocketSearch was used to screen the RCSB protein database to both identify matching solvent-exposed pockets and rank them based on the pocket volume and degree of overlap. To simplify the expression, purification, and characterization of Art hases, this search was performed over an abbreviated library of -45000 proteins with unique PDB IDs that are small, monomeric, non-membrane proteins and can be expressed and purified from E. coll. Approximately 14000 PDB structures with comparably large, solvent-exposed pockets were identified, and of these, 562 PDBs exhibited sufficient pocket overlap with the reference binding geometry obtained from NifD.
[0127] Successful implementation of PCS and SCS effects requires both the incorporation and appropriate spatial arrangement of selected amino acid residues. The PCS of Mo N2ase includes His442 and Cys275, which coordinate Fel and Mo of FeMoco. In addition, His 195 and Glnl91 interact with S2B and homocitrate, respectively, through H-bonding, and comprehensive mutagenesis research has shown these residues to be critical in inferring full N2ase activity. Taking this information into account, the study chose to replicate His 195 and Glnl91 from the SCS interactions in addition to the PCS residues.
[0128] RosettaMatcher was employed to evaluate potential scaffolds for their ability to incorporate essential PCS and SCS features in a desired geometry for FeMoco binding, as well as screen the degree of pocket-solvent exposure. Of the initial 562 PDBs displaying appropriate pocket shapes, 428 demonstrated potential FeMoco binding through incorporation of the fourinteractions equivalent to His442, Cys275, Hisl95, and Glnl91 in native MoFe. Multiple mutation combinations are possible for each of these 428 hits, producing a total of 12406 unique PDB-mutation combinations. These combinations were further screened based on cofactor orientation and solvent exposure. The distribution of SASA values is shown in FIG.2.
[0129] By comparing the distribution of SASA values with that of FeMoco in native MoFe, potential mutation-combinations were grouped into four categories: 1) Fully buried cluster core, 2) Highly solvent exposed cluster core, 3) Highly buried homocitrate 4) Adequately exposed cluster core and homocitrate. This process allowed reduction of the 12406 unique PDB-mutation combinations to 45 unique proteins displaying SASA scores in good agreement with the native NifD pocket reference. As a penultimate step, the RosettaScripts interface was used to optimize the 45 filtered structures to ensure compatibility of the protein pocket with FeMoco binding. This step produced 38 final designs with satisfactory coordination environments, constraints, and total scores that were selected for experimental evaluation. Selected designs are shown in FIGS. 3A-3E.
[0130] DNA corresponding to each of the 38 selected scaffold designs was synthesized with codons optimized for expression in E. coli and cloned into 2M-T vector, which includes sequences for N-terminal hexa-histidine tagged maltose binding fusion protein to aid in the expression and purification of these designs in E. coli. Of the 38 constructs, 11 could be expressed and purified in sufficient quantity for further cofactor binding and catalytic assessment.
[0131] Isolated FeMoco was obtained from native Mo N2ase purified from Azobacter Vinelandii and further incorporated into the 11 proteins using previously established protocols. To assess potential FeMoco binding, acetylene reduction activity assays were performed (FIG. 4). While the isolated FeMoco cluster is capable of reducing acetylene to ethylene and ethane, its activity is significantly reduced during turnover (45% activity after 3 hours) due to degradation of the cluster with time. Seven designs display similar time-dependent loss of acetylene reduction activity to the isolated cluster, suggesting either non-specific FeMoco binding at the protein surface, or incomplete incorporation of the cluster into the designed binding pocket. However, constructs His-8, His-11, His-13, and Trp-10 retain activity during turnover, with losses of -15% over the course of 3 hours. The retention of catalytic activity suggests improved cluster stability, indicating FeMoco binding. To veril this observation, the designed ligating cysteine and histidine residues of Trp-10 were mutated to Ala to eliminate potential specific covalent binding. This mutant, termed Trp-10-HC-AA, displayed overalllower acety lene reduction activity, as well a reduced 3 hr activity retention (76%, as compared to 87% in Trp-10). suggesting that the designed coordinating residues play an important role in retaining activity.
[0132] To further characterize FeMoco in Trp-10, the X-band EPR spectrum was measured and compared to native MoFe and isolated FeMoco (FIG. 5). The FeMoco cofactor in native MoFe displays a relatively low rhombicity S = 3 / 2 spectrum with geff = (4.34, 3.66, 2.01), with E / D = 0.056. The free FeMoco in DMF displays two signals, the first a more rhombic S = 3 / 2 system corresponding to the FeMoco cluster (geff = (4.50, 3.58, 2.00), E / D ~ 0.08), and a large isotropic S = 1 / 2 signal at g ~ 2 (-3400 G), which arises from a partially degraded cluster. The Trp-10 sample shows an intermediate S = 3 / 2 signal with geff = (4.44, 3.57. 2.00), E / D ~ 0.07, and lacks any sign of cluster degradation. These results indicate successful binding of FeMoco to the Trp-10 scaffold with no detectable degradation. The EPR signal around g = 4 (-1600 G) suggests that the FeMoco in Trp-10 experienced a greater electronic symmetry than the free FeMoco, probably due to a different orientation of the cluster. The broadening of the EPR signal together with an increase in E / D is also observed in the lyophilized state of FeMoco in native Nzases, suggesting a less well-defined SCS environment.
[0133] Encouraged by findings from acety lene reduction activity assays and EPR measurements, the study further performed a cyanide reduction assay that is commonly used in native IShases. As shown in FIG. 6A, while the Trp-10 predominantly produces methane, similar to the free FeMoco. it does not produce any of the three carbon products that are observed in the free FeMoco, further indicating that the binding pocket of Trp-10 restricts the binding / formation of larger molecules. Finally, the study carried out an N2 reduction assay using15N2 as the starting materials and frequency-selective pulse 'H-NMR spectroscopy to characterize the product of the reduction. As shown in FIG. 6B, the ArtN2ase. 1 reduce N2 to ammonia, similar to the native enzyme.
[0134] In conclusion, the study designed Art bases capable of binding FeMoco and reducing small molecules like native N2ases, including reducing N2 to NH3. This work marks a design of metalloproteins that contain the most complex metallocofactors and most challenging reaction to date. ArtN2ases present simplified, easily expressed / purified systems which lack additional cofactors and accessory' proteins, providing an excellent alternative system to native IShases for studying biological N2 fixation. Further studies are ongoing to establish a secondary coordination sphere that further resembles the native enzy me. The methods developed in this work also provide a steppingstone for the design more sophisticatedproteins capable of harnessing FeMoco and other high-nucleanty clusters for small molecule activation and reduction.Computational Methods
[0135] The goal of the computational simulations was to come up with repurposed protein scaffolds capable of binding the FeMoco. The main assumption of these simulations was that the replication of MoFe protein's binding pocket, its primary coordination sphere, and some if not all of its secondary coordination sphere in another protein is sufficient to bind the FeMoco.
[0136] The majority of computational work was performed using the Carl R. Woese Institute for Genomic Biology Biocluster. Final computational validation of the first generation of the designs is being completed via the stampede supercomputer of the Texas Advanced Computing Center.
[0137] Selection of Initial Protein Pool for Filtering'. From the Protein Data Bank (PDB), using their advanced search method, four filters were applied to the search: first, the search was restricted to proteins only; second, the protein should have a size less than or equal to 400 residues; third, the protein should be deposited from an X-Ray diffraction structure; and forth, the structure should only be an asymmetric unit, which means that only monomers are selected. At the time of the query in October 2020, these filters yielded around 45000 structures. The PDB resolution was 2 Angstroms or higher.
[0138] Filtering based on Pocket Shape Complementarity. The Imln pdb structure was picked up for this purpose and prepared for pocket generation on pymol. All cofactors, ions, and water molecules were removed from the structure. Then, from the heterotetramer, all sequences but one alpha subunit (NifD) were removed. This structure was used as the input structure for the pocket generation via fpocket. All the default parameters of the code were used for the pocket generation and the structure only generated one pocket, which is at the site of the FeMoco.
[0139] For fpocket, one main parameter was called alpha sphere, which was defined as "a contact sphere, that touches 4 atoms in 3D space without having any internal atoms.'’ The minimum and maximum size of the sphere can be modified in the run script and was 3.4 A and 6.3 A by default. The more alpha sphere a pocket has, the larger its size becomes. Hence, this can be used as a rough filter; pockets too large or too small compared to the FeMoco native binding site can be filtered out from the pool after a simple alpha sphere number calculation. The FeMoco binding site pocket has 108 alpha spheres with the default minimum and maximum sizes. After filtering out pockets over 150 and under 80 alpha spheres from the pool, 14054 unique PDB IDs were left.
[0140] Here, the PocketSearch main code was used to measure the matching of the FeMoco pockets with the candidates. Any pdb structure that had an overlap over 85% was passed through the filtering, which in this case was 562 structures.
[0141] The output files of PocketSearch contain a pos file for each protein which includes the list of amino acid residue numbers that encapsulate the pockets. This list was used alongside the protein structure as the input for the next step.
[0142] Finding Residue Interaction with the Cofactor'. In this step, the residues surrounding the pocket were probed for binding the cofactor in the appropriate geometry resembling the native, holo, alpha subunit. Rosetta Matcher, developed by the Baker lab, was utilized for this purpose. In short, Rosetta Matcher attempts to “place’' the given input molecule into the protein by mutating a specific set of protein residues given in another input list to the desired residue, and mix-and-matching the target molecule to those residues until at least one conformational combination of those residues are found which fulfill the given geometrical constraint file. For FeMoco, to make the match more accurate, a set of homocitrate conformations in the FeMoco were generated and accounted for. The list of residues for matching has been generated by PocketSearch is described below. For the first generation of designs, the constraint file included four interactions, namely the His442, Cys274, H195, Q191 interactions, and was benchmarked with the Imln protein structure to ensure the native binding geomctiy can be replicated. The constraint file also included interactions for the two ligation residues of His442, Cys 275, and two constraints for the non-covalent interactions of His 195 and Glnl91. The constraint file is depicted in FIG. 7.
[0143] For the histidine ligation to the molybdenum atom, the two nitrogen of the histidine were considered for the binding. This uses the Rosetta atom ty pe N-His and accounts for both delta and epsilon nitrogen. Further explanation of the exact atom type considerations for the deprotonated nitrogen in histidine can be found in the rosetta matcher source code. The nomenclature is internal in Rosetta and can be found on their w ebsite.
[0144] Filtering based on Solvent Accessible Surface Area. To narrow' down the number of proteins, another filter was implemented. In this filtering, the aim was to sort out structures that have the homocitrate surface exposed and the cluster side of the cofactor buried inside the protein. This was done by SASA calculation of each fraction of the cofactor separately and comparing it to the native cofactor. The SASA calculation was done via a TCL script in VMD and needed each part of the cofactor to have a separate label. To standardize all the scores compared to Imln, solvation waters for Imln through pymol and all hydrogen atoms were removed via a python script.
[0145] After the SASA scoring, structures were sorted based on their ICS (cluster core) or HCA (homocitrate) SASA scores and manually inspected. The cutoff was set such that the structures above it would be certainly deemed un-fit to bind FeMoco at the given binding position. The goal was to eliminate false negatives; false positives were dealt with later with the HRR score, manual analysis, and / or post-enzyme design. Analysis of the post-enzyme design SASA scores was done by benchmarking the PDB itself and the Imln match benchmarks with the experimental PCS and SCS residues. Values were 10.8 A2for ICS and 76.2 A2for HCA in the Imln native structure.
[0146] Structure Refinement Using Rosetta Scripts'. The non-mutational enzyme design rosetta script was run on the generated structures for 11 rounds, in which the optimization output was used as the input of the next round. The first eleven rounds were done without any rotamer consideration for the homocitrate, but the last optimization round included rotamers.
[0147] The outputs of these runs were analyzed via the filtering code described below, and the best candidates were optimized for the mutational runs. The mutational runs were the same as the non-mutational ones except they allow mutations around the active site biased by the PSSM files. These optimizations were also done in 10 rounds in the same manner as the non- mutational runs.
[0148] Manual Analysis of the high throughput results'. Structures scored by the high throughput method were ranked based on their score and grouped together based on their protein groups. Then the structures were overlayed in pymol and assessed by mutations and residue positionings. For some structures, the additional introduced mutations were changed or reverted and then re-optimized with the minimal rosetta scripts to ascertain the intended effect of the mutation reversion (for example, if a mutation was unnecessary to the cofactor binding or if a hydrogen-bonding interaction worked the way expected).Experimental Methods
[0149] Plasmid Design and Ordering
[0150] First Generation designs'. The 2M-T Cloning vector, which contains hexa-histidine and MBP tags on the N-terminus, was chosen for the protein constructs. pET His6 MBP TEV LIC cloning vector (2M-T) was a gift from Scott Gradia (Addgene plasmid #29708; RRID:Addgene_29708). This vector has the ampicillin resistance gene and contains the T7 promoter. The expressed sequences were as follows: N-terminus-Histidine Tag-MBP- ENLYFQ / S-IGSG-Protein of interest-C -terminus. The ENLYFQ / S is a sequence that is recognizable by the Tabaco Etch Virus Protease (TEV -Protease) which is a highly specific cysteine protease. The cleavage site is between the glutamine and the serine. The IGSGsequence was added between the TEV site and the protein as a spacer to make the cleavage site more accessible to the protease.
[0151] Genes were ordered from Genscript. The BL21DE3 cell line was chosen as the host cell since it contains the T7 polymerase gene controlled by the Lac Repressor. It was ordered from NEB and transformation was done via Heat Shock. 1 pL of the plasmids ranging from 10 to 100 nanogram per microliter was used for each 15 pL of the cell line and transformed and plated based on NEB protocols. In order to confirm the successful transform and absence of spurious mutations, small volume cultures of 5 mL for each protein was grown from the grown colonies overnight at 37°C in triplicates. Plasmids were extracted using the Qiagen miniprep kit according to their protocols and sent for sanger sequencing from the C-terminus with a custom reverse primer of: CAACTCAGCTTCCTTTCG (SEQ ID NO: 68).
[0152] Out of all the plasmids, except his-BlAR, every' other protein sequence was successfully transformed and had the correct sequence.
[0153] For each protein, a colony was grown in 5mL of LB-media containing 50 pg / mL of ampicillin at 37°C overnight and 600 pL of the cell media was mixed with 400 pL of autoclaved 60% glycerol solution in an Eppendorf 1.5 mL tube and flash frozen using liquid nitrogen. The cell stocks w ere kept in a -80°C freezer for further use.
[0154] Large-Scale Protein Expression and Purification
[0155] Protein Expression using IPTG induction'. A 5 mL overnight culture containing 50 pg / mL ampicillin and the protein was made from its colony, and added into either a 6-liter flask containing 2 liter of media or a 4-liter flask containing 1.5 liter of media. Media was a variant of the LB-Lennox including 10 g / L tr ptone. 8 g / L yeast extract, 5 g / L NaCl in water. Water was filtered via Milli-Q IQ water purification system from Millipore Sigma. All material was from Fisher Scientific. The media was then autoclaved at 120°C for 45 minutes and let out to cool down to room temperature. All the 5 mL of cell culture was then added to the media flasks, then ampicillin sodium was added to a concentration of 50 mg / L and the media flasks were shaken for at least four hours at 37°C and 200 RPM or until its optical density with respect to ungrown media exceeded 0.4. Afterwards, the shaker was cooled down to 18°C and IPTG was added, initially, for a final concentration of 500 pM. After several rounds of expressions, it was found that IPTG concentration beyond 100 pM does not benefit the protein expression, hence the IPTG concentration was fixed to 100 pM. The flasks were shaken at 18°C and then taken for cell harvesting.
[0156] Protein Expression using Autoinduction'. The protocol was taken from Studier. Preparation of the small culture media was the same as above. The media was as follows: 20g / L tryptone, 5 g / L yeast extract, 5 g / L NaCl, 6 g / L Na2HPC>4, and 3 g / L KH2PO4. All chemicals were from Fisher Scientific. Water was filtered via Milli-Q IQ water purification system from Millipore Sigma. The media was then autoclaved at 120°C for 45 minutes and let out to cool down to room temperature. All the 5mL of cell culture was then added to the media flasks, then ampicillin sodium was added to a concentration of 50mg / L and the media flasks were shaken at 30°C and 200 RPM for 3 days. ArtlShase-lO-HC-AA was expressed in high amounts this way.
[0157] Cell Harvesting'. Cells were harvested from media via centrifugation at 8000 RPM (13800 RCF) for 20 minutes at 4°C, using a Fiberlite™ F9-6 x 1000 LEX fixed angle rotor in a Sorvall LYNX 6000 centrifuge. The supernatant solution was discarded, and the pellet was flash frozen with liquid nitrogen and stored in a minus 20°C freezer until further use.
[0158] Lysis: The cell-to-ly sis buffer ratio was set to at least 1 to 9, lysis buffer was 50mM KPi, 250mM KC1, pH=7.7, 5mM BME, 25mM Imidazole. Solution was added to the pellet, and the pellet was resuspended. Afterwards, the following chemicals were added: 1 mg / mL of Egg White Lysozyme (Sigma- Aldrich), 0.174 mg / mL (ImM) phenylmethylsulfonyl fluoride (PMSF) from a fresh-made 100 mM Ethanol Stock Solution. 1 pL per 10 mL total volume up to 10 pL of Thermofisher’s Pierce™ Universal Nuclease, and Triton X-100 to final concentration of 0.1%V / V. The solutions were stirred at 4°C for one hour and the sonicated with the Misonix S-4000 ultrasonic sonicator cell disruptor as follows: If the total volume was less than 100 mL, solution was lysed under microtip conditions. 100% amplitude 2 seconds on 8 seconds off and 5 minutes total on time. If the volume was above 100 mL and below 250 mL, a normal tip was used with 70% amplitude and same time conditions. If the volume was above 250mL, the batch was split into volumes less than 250 mL. Then they were centrifuged in the Fiberlite™ F21-8x50y fixed angle rotor in a Sorvall LYNX 6000 centrifuge at 18000 rpm for 30 minutes at 4°C. Then the supernatant solution was filtered via Millex™ 33 mm PES syringe filters, first with a 0.45 micron filter, then a 0.22 micron filter. Afterwards, the solution was loaded into aNickel IMAC column for further purification not later than 12 hours after filtering, usually almost immediately.
[0159] Protein Purification Protocol: Using the AKTA-go system, the solution was loaded into the Ni-IMAC column. The column was assembled from a Cytiva Cl 6 / 20 with a plunger and filled with a Sepharose HP Ni IMAC resin. First, it was washed using buffer K-A with 25 mM imidazole until a stable UV-Vis and conductivity baseline was reached. After the sample load, the column was washed with 5-15 column volumes of 15% Buffer K-B (which is K-A wi th a final concentration of 500mM imidazole) followed by a stepwise elution where thebuffer imidazole concentration was increased to 50% Buffer K-B for three column volumes. Then a wash with only K-A for 1-2 column volumes was done, followed by a last isocratic wash with 100% K-B for two column volumes. The peaks in the 50% and 100% K-B wash steps were pooled and buffer exchanged to K-A so that the imidazole content reached below 15% at least. Then the column run was repeated. Fractions containing UV 280 absorption peaks were then selected for SDS-PAGE analysis. The purest samples were then either buffer exchanged via dialysis into storage buffer of 50mM Tris. 150mM NaCl, 5mM Maltose, 5mM BME, 20% V / V Glycerol at pH=8.00 and flash frozen.
[0160] Sample Submission for MS analysis
[0161] ESI-MS'. Mass Spectrometry was performed using an ESI-TOF mass spectrometer, with samples needed in a final volume of 500 pL and concentration of about 50 pM or more in -50% v / v acetonitrile, -50% v / v water, 1% v / v formic acid. The sample solutions were buffer exchanged to w ater via a desalting column, then after the final dilution or concentration, water w as added to 250 pL. 250 pL of HPLC grade acetonitrile and 5 pL of formic acid was then added to the solution and the sample submitted for analysis. From the raw mass spectrum, any peaks above a minimum threshold were digitized and analyzed using UniDec. Recommended protein concentration was at least 0.05 mM.
[0162] MALD-TOF'. The sample preparation was the same as above with the difference that the final solution did not have formic acid, but instead used 0.1% of Trifluoro Acetic Acid. Sinapic acid was used as a matrix. The protocol of the sample mixing with the matrix solution and the matrix solvent was taken from the Bruker protocol. The droplet volumes were either 0.5 or 1.0 pL.
[0163] SDS-PAGE'. The solutions used were as follows:
[0164] 1) Buffer solution lOx: 30.3 g tris-base, 144 g glycine, 10g SDS, in 1000 mL H2O.
[0165] 2) Running solution in gel (15%): 3.45 mL H2O, 3.75 mL acrylamide (40%), 2.6 mL tris 1.5M pH=8.8, 100 pL SDS, 50 pL APS ((NfG^Os) 25% W / V), 10 pL TEMED (T etramethylethylenediamine).
[0166] 3) Stacking solution in gel: 3.167 mL H2O, 1.25mL tris 0.5 M pH=6.8. 0.50 mL AA40%. 50 pL APS, 50 pL SDS. 5 pL TEMED.
[0167] Gels were either made in 15 or 10 wells with a 0.75 mm thickness. Samples were loaded in 12-13 pL for the 15 well and 20 pL for the 10 well unless otherwise stated. The ladders w ere loaded in 5 pL aliquots. For the samples. 5 pL of 4x laemmli sample buffer, nonreducing. was used. 5 pL of IM DTT solution was added, then water and protein samples were added to a final volume of 20 pL in a 1.5 mL Eppendorf tube. The amount of added proteinwas dependent on its concentration, roughly 50 mAu of a FPLC run in a final concentration. The solution was then centrifuged for at least 30 seconds to mix the solutions and collect all the droplets at the bottom. Occasionally, the solution was boiled at 90°C 500 RPM for 10 minutes. The voltage and wattage were as follows: 180-185 mV, 7-15 Watts for 55 minutes to one hour.
[0168] Site-directed mutagenesis of the mutants'. Three approaches for trp-dFABP-5gkb H75A and CI25A knockout mutations were attempted. First. SDM via rationally-designed primers. Second was a computationally generated primer sequence via Agilent for SDM. Last was a Gibson Assembly protocol. Only the Gibson Assembly was successful. Mutagenesis was performed via the NEB protocol, and sequences were ordered by IDT. All protocols done according to NEB via their enzymes for the SDM. The sequences are given in TABLE 1.TABLE 1. Sequences.
[0169] Expressions'. The expression group of the first generation design proteins are shown in TABLE 2TABLE 2. The expression group of the first generation design proteins. Not all of them were expressed, and his-blAr was only successfully transformed to the DH5 Alpha cell line (not BL21DE3 cell line).
[0170] Binding Assays: The assays were done in different manners. First, the study conducted the assays in varying concentrations of the components, followed by assaying the cofactor incubated protein at different time points for acetylene reduction, then for the productdistribution of the acetylene reduction, and finally for cyanide reduction. The reductant was 20mM EuDTPA, which is strong compared to the MoFe protein redox partner. The TON overtime measurement was used as a probe for the cofactor’s stabilization by the protein; the more stable the activity is, the more protection the protein is providing, so it is assumed.
[0171] The results of the three rounds of assays are as shown in TABLE 3: the blank activity of the cofactor was higher than all the designs and also had a fluctuation of around ten percent between different batches. It dropped down to less than half in all the assays. The higher activity of the cofactor showcases that the proteins do not enhance the catalytic capabilities of the cofactor. However, in some designs such as Trp-10, which is trp-dFABP-5gkb, the activity was much more stable compared to the naked cofactor without the protein’s protection. This is an indicator of the protein binding; however, it is not sufficient. The study also attempted to measure EPR signals for the Trp-10 (FIG. 8), but the incubation yield was very low, and the signal hence w as very w eak and unreliable.TABLE 3. Activity assay of selected proteins.
[0172] The theory is that the SASA window of a successful design is much narrower than initially thought and the designed binding pockets might be too large or too small. The finalized design sequences were initially optimized with a modified rosetta enzyme design script. Then a SASA analysis was performed, and the best structures were additionally analyzed based on their binding geometries. Four candidates were then selected for assaying. Those four proteins were his-NP7 (his- 12), his-3HAO (his- 15), and trp-PDE9A (trp-7), and min-hMetAP2 (min- 2). Aside of min-hMetAP2 which was not successfully expressed, the three others were assayed and only his-NP7 showed improved cofactor stability. This means that the theory on the correct criteria is either wrong or incomplete. Further experiments are needed to distinguish between these two possibilities.Example 3: Ammonia Production
[0173] To prepare an active catalyst. lOOuM of TrplO peptides were incubated with 200uM of FeMoco, followed by repurification on aNi-NTA column. In a 10 ml vial, 60uM of FeMoco- containing peptides were subjected to turnover by mixing 10 mM Eu(II)-DTPA in 100 mM Tris (pH 8.0) with a total volume of 1.5 ml. The reaction mixture was incubated at 30°C for 10 min under a gas atmosphere of 100%15N2. Formation of15NHs is then confirmed by frequency selective pulse NMR analysis. Samples are first cleaned by removing the peptides via filtration, then acidified and combined with locking reagent. They are then measured by1H NMR using a Bruker AVANCE600 spectrometer equipped with a CBBFO cryoprobe employing water suppression. The15NH3 peaks were then quantified through referencing to standards.Example 4: Nitrogenase Models
[0174] The following tables summarize example nitrogenase models.TABLE 4. Example protein designs.TABLE 5. First generation protein designs.TABLE 6. Second generation protein designs.Example 5: Design of an Artificial Nitrogenase that Reduces Dinitrogen to Ammonia
[0175] Nitrogen is present in almost all biomolecules, natural products, and synthetic compounds, and its source can almost always be traced back to atmospheric N2. Despite its abundance, activating the 945 kJ / mol triple bond of N2 efficiently under mild conditions has been a major challenge facing scientists and engineers for over a century [1], While the Haber- Bosch process has been highly effective in reducing N2 to NH3. which is the key ingredient of fertilizers that have supported human population growth throughout the 20th century [2], the process requires large quantities of pure H2 generated by natural gas steam reforming. As a result, the process not only consumes 2% of global energy but also generates CO2 at an equally colossal scale (-450 Mtons annually, -1.2% of global CO2 emissions [3]). These environmental concerns, along with the centralized nature of industrial ammonia production, have intensified efforts to develop alternative methods of N2 fixation [4],
[0176] In contrast to the Haber-Bosch process, bacteria and archaea use nitrogenases (N2ases) to fix N2 at ambient conditions, independent of H2 [5,6], Despite the potential and efforts spanning >50 years, translating N2ases into practical catalysts remains elusive due to several major challenges. First, while enzy me engineering is a powerful tool for understanding the reaction mechanisms of enzymes, it has proven difficult to apply this approach to N2ases, as it is hindered by limited expression vectors and phenotypic screening methods for N ase mutants [7], So far. most Nzase mutants have been generated through homologous recombination - an approach that is cumbersome and prone to selection of mutants that revert back to the wild-type N2ase phenot pe due to the essential role N2ase in the survival of the native host. One promising strategy involves heterologous expression of N2ase in genetically accessible organisms such as E. coli [8,9] which facilitates engineering efforts. Recent breakthroughs in the heterologous expression and purification of Nzase in E. coli mark a significant step forward
[0010] , especially given the challenge to piece together a viable pathway for the heterologous synthesis of a metalloci uster-replete N2ase in a non-diazotrophic host
[0011] . However, Nzases are include intricate subunits and metallocofactors, e.g.. [Fe4S4] and the [Mo:Fe?:S9:C]:homocitrate Iron-molybdenum cofactor (FeMoco) [12-14], As a result, the biosynthesis and assembly of the N2ase requires a complex maturation process that demands multiple proteins, making the production of the N2ases, both from native and heterologous host organisms to be exceptionally difficult. Furthermore, the presence of other metallocofactors has made it difficult to focus on studying the catalytic activity of the FeMoco that binds and reduces N2 to NH3 using biochemical and biophysical tools, including many spectroscopictechniques [15-18], Additionally, N2ases rely on ATP for activity
[0019] , posing an additional economic barrier to practical implementation
[0020] ,
[0177] To address the aforementioned issues, many synthetic models of N2ases have been prepared, including structural models mimicking FeMoco in IShases but displaying no N2ase activity and functional models that do not structurally resemble FeMoco [21-25], While these synthetic systems represent significant strides toward enzyme-inspired catalysis, not being able to mimic N2ase both structurally and functionally made it difficult to offer insights into how N2ase works. In addition, most rely on strong reducing reagents for N2 activation, making their practical applications difficult and costly. Extensive evidence indicate that structural features beyond the primary' coordination spheres (PCS), particularly weak and non-covalent electrostatic and hydrogen bonding interactions in the secondary coordination spheres (SCS), play essential roles in conferring and tuning N2ase activity
[0026] , Despite the importance of these interactions, it has been challenging to incorporate elaborate SCS features into model systems due to the requirement of precisely placing them in a rigid scaffold.
[0178] To overcome the limitations of the Haber-Bosch process, native IShases, and synthetic models of the N2ase active site, a study was conducted which explores the use of a different approach of designing artificial metalloenzymes (ArMs). These ArMs mimic native enzy mes by incorporating both the primary and secondary' coordination spheres (PCS and SCS) into small, stable, and monomeric proteins that can be readily expressed in E. coll with high yields [28-33], Small and robust ArMs have allowed not only’ to reveal the minimal structural features for the activity, but also to gain deeper insights into native enzymes that are otherwise difficult to obtain [34,35], Despite the success in designing ArMs for other metalloenzymes, the complexity of the metallocofactor and the challenge of the reactions these ArMs cannot match those in N2ases.Materials and Methods
[0179] Computational Methods:
[0180] The goal of computational simulations is to come up with repurposed protein scaffolds capable of binding the FeMoco. The computational pipeline is an adaptation from earlier work done by Mirts et al. and is also inspired by two computational protein design approaches [32,42,57],
[0181] FeMoco geometry was taken from the 3U7Q PDB ID12. Homocitrate rotamers were considered by generating a library' of 27 different fixed geometries and the bond dihedrals were left to be optimized during the Rosetta runs.
[0182] Selection of Initial Protein Pool for Filtering: From the Protein DataBank (PDB), using their advanced search method, four filters were applied to the search: first, the search is restricted to E.coli-expressed proteins only, second, the protein will have a size less than or equal to 400 residues, third, the protein should be deposited from an X-Ray diffraction structure, and fourth, the structure should only be an asymmetric unit, which means that only monomers are selected. This at the time of the query, in October 2020. yielded around 46000 structures. The PDB resolution is 2 Angstroms or better.
[0183] Filtering based on Pocket Shape Complementarity: The 3u7q PDB structure was selected as the target active site to be engineered into a different scaffold. All cofactors, ions and water were stripped from the model before pocket detection was performed. Of the heterotetramer, only chain A, the NifD subunit which contains FeMoco, was selected to represent the scenario of a singular peptide chain binding the cofactor. This structure was then used as the input structure for pocket generation via fpocket
[0058] , All the default parameters were used with the exception of having 80 minimum alpha spheres for pocket clustering in order to reproducibly select the catalytic pocket at the binding site of FeMoco.
[0184] Fpocket contains several parameters which can be tuned to generate different sized pockets, including the number of alpha spheres to be considered, and how to cluster those alpha spheres to ultimately generate the pocket. In the case of fpocket, an alpha sphere is simply a dummy atom which contacts at least 4 other atoms within its cutoff radius, including other alpha spheres. The radius is tunable and variable amongst dummy particles which allows for better volume estimation, and the defaults for minimum and maximum radii, 3.4 A and 6.3 A respectively, were used in this work. This also allowed the use of the number of alpha spheres as a rough filter, removing any potential scaffolds which have pockets too small or too large to accommodate the target active site. The FeMoco binding site pocket has 108 alpha spheres, and it was reasoned that it would be easier to fill in a scaffold pocket that was too large than to open up a scaffold pocket too small, so the study settled on a range of acceptable sizes from 80 to 150 alpha spheres as a cutoff. After filtering out pockets over 150 and under 80 alpha spheres from the pool, 14053 unique PDB ids were left.
[0185] Finally, the PocketSearch code, utilizing both VASP and Fpocket [58.59]. was used to measure the matching of the FeMoco pockets with the candidates by volume and shape complementarity. Any PDB structure that has a volume overlap over 85% was passed through the filtering which in this case, were 562 structures.
[0186] The output files of PocketSearch contain a pos file for each protein which is the list of amino acid residue numbers that encapsulate the pockets. This list has been used alongside the protein structure as the input for the Rosetta Matcher step.
[0187] Finding Residue Interaction with the Cofactor: In this step, the residues surrounding the pocket are probed for binding the cofactor in the appropriate geometry resembling the native, holo, alpha subunit. Rosetta Matcher
[0047] is utilized for this purpose. For FeMoco, to make the match more accurate, a set of homocitrate conformations in the FeMoco are generated and accounted for. The list of residues for matching has been generated by PocketSearch. The constraint file contains four interactions, namely the His442, Cys275, H195, Q191 interactions, and has been benchmarked with the 3u7q protein structure to ensure the native binding geometry can be replicated. The constraint file has interactions for the two ligation residues of His442, Cys 275 and two constraints for the non-covalent interactions of Hisl95 and Glnl91.
[0188] For the histidine ligation to the molybdenum atom and the His 191 hydrogen bonding interactions, the two nitrogens of the histidine have been considered for binding through the Rosetta atom types N-His and N-Trp, which accounts for both delta and epsilon nitrogens. Further explanation of the exact atom type considerations for the deprotonated nitrogen in histidine can be found in the Rosetta Matcher source code. The nomenclature is internal to Rosetta and can be found on their website
[0060] , In addition to glutamine, asparagine was also used for homocitrate hydrogen bonding.
[0189] Rosetta Matcher outputs have several mutation combinations possible for each of these 428 hits, producing a total of 12406 unique PDB-mutation combinations.
[0190] Filtering based on Solvent Accessible Surface Area: To narrow down the number of proteins, another filter was implemented to sort out structures that have the homocitrate surface exposed and the cluster side of the cofactor buried inside the protein. This is done by solvent-accessible surface area (SASA) calculation of each fraction of the cofactor separately and comparing it to the native cofactor. The SASA calculation is done via a TCL script in VMD and needs each part of the cofactor to have a separate label. To standardize all the scores compared to 3U7Q, solvation waters for 3U7Q through PyMol and all hydrogen atoms, via a Python script, were removed.
[0191] After the SASA scoring, structures were sorted based on their ICS (cluster core) or HCA (homocitrate) SASA scores and edge cases heuristically inspected. The cutoff was set at values such that structures beyond would be certainly deemed unfit to bind FeMoco at the given binding position. Analysis of the post enzyme design SASA scores was done by benchmarkingthe PDB itself and the 3u7q match benchmarks with the experimental PCS and SCS residues. Values are 8.48 square angstrom for ICS and 70.1 square angstrom for HCA in the 3u7q native structure. Values for selected Rosetta Matcher outputs are 3-15 square angstrom for ICS and 30-90 square angstrom for HCA according to the crystal structure SASA and the SASA of Rosetta Matcher benchmarks of the FeMoco-removed NifD.
[0192] By comparing the solvent exposure values with that of FeMoco in native N2ases, potential combinations of mutation can be grouped into four categories: 1) Fully buried cluster core, which prevents substrate access to the cluster and will not be practical for cofactor incubation, 2) High solvent exposure of cluster core, which defeats the purpose of having the protein stabilize the cofactor as no residues will surround it, 3) Highly buried homocitrate, which would render the homocitrate unable for proton transfer to or from the cluster, and 4) Adequate solvent exposure statistically close to native NifD solvent exposure of both for the cluster core and homocitrate. In the last category, similar to N2ase, the cluster core is buried within the protein while retaining a small, and quantifiable, degree of solvent access while as the homocitrate moiety is solvent exposed and can freely exchange protons with it.
[0193] Structure Refinement Using Rosetta Scripts: A non-mutational enzyme design Rosetta script was run on the structures passed in 1-4 for 10 rounds in which the optimization output is used as the input of the next round.
[0194] The mutational runs are the same as the non-mutational ones except they allow mutations around the active site defined by a distance cutoff from the cluster. These optimizations were also done in 10 rounds in the same manner as the non-mutational runs.
[0195] Manual Analysis of the high throughput results: Structures scored by the high throughput method have been ranked based on their ligand score and grouped by their protein scaffolds. Then the structures were overlayed and visually assessed by mutations and residue positionings. For some structures, the additional introduced mutations were changed or reverted and then re-optimized with the minimal Rosetta script pipeline to ascertain the intended effect of the mutation reversion.
[0196] Experimental Methods:
[0197] Plasmid Design and Ordering for ArtNzase: The 2M-T Cloning vector, which contains hexa-histidine and MBP tags on the N-terminus, has been chosen for the protein constructs. pET His6 MBP TEV LIC cloning vector (2M-T) was a gift from Scott Gradia (Addgene plasmid # 29708; RRID:Addgene_29708). The expressed sequences will be as follows: N-terminus-Histidine Tag-MBP-ENLYFQ / S-IGSG- Protein of interest- C-terminus.
[0198] Codon-optimized genes were ordered from Genscript. The plasmids were transferred into E.coli BL21DE3 cells, purchased from NEB and following their protocols. Cell stocks were prepared from 5mL overnight growth cultures in LB media and stored in -80C freezers for further use.
[0199] Large-Scale Protein Expression and Purification for ArtNzase:
[0200] Protein Expression and Harvesting'. 5 mL cultures of the protein with 50 pg / mL ampicillin is made from its colony or frozen cells stocks and grown between 10-16 hours at 37C with 250rpm shaking rate. Large growth was done via either 6-liter flask containing 2 liters of media or a 4.5-liter flask containing 1.5 liters of autoclaved media. Media is a variant of the LB-Lennox including 10 g / L try ptone, 8 g / L yeast extract, 5 g / L NaCl in water filtered via Milli-Q IQ water purification system from Millipore Sigma. All material is from Fisher Scientific. All the 5 mL of cell culture is then added to per media flask, ampicillin sodium is added to the concentration of 50 mg / L and shaken for at least four hours at 37°C and 200 RPM or until its optical density7with respect to ungrown media reaches the range of 0.4-1.2. Afterwards, the shaker is cooled down to 18°C and then aqueous IPTG stock solution is added to a final concentration of 300 pM. The flasks are shaken for 18-24 hours.
[0201] Cells were harvested from media via centrifugation at 8000 RPM (13800 RCF) for 20 minutes at 4 C, using a Fiberlite™ F9-6 x 1000 LEX fixed angle rotor in a Sorvall LYNX 6000 centrifuge. The supernatant solution was discarded, and the pellet was flash frozen with liquid nitrogen and stored in a -20°C freezer until further use.
[0202] Lysis'. The cell is suspended with the lysis buffer to a ratio of at least 9 volumes of buffer per 1 volume of pellet. Lysis buffer contains 50 mM KPi, 250 mM KC1 pH = 7.7, 25 mM Imidazole. 5 mM BME, 1 mg / mL of Egg White Lysozyme (Sigma- Aldrich), 0. 174 mg / mL (1 mM) phenylmethylsulfonyl fluoride (PMSF) from an 100 mM Ethanol stock solution. 1 pL per 10 mL total volume up to 10 pL of Thermofisher’s Pierce™ Universal Nuclease, and 0. 1%V / V Triton X- The solutions have been stirred at 4 C for one hour and the sonicated with the Misonix S-4000 ultrasonic sonicator cell disruptor as follows: for volumes below below 250 mL, a normal tip is used with 70% amplitude. 2s on time and 8s off time for total of 5 minutes of sonication. If the volume is above 250 mL. the batch will be split into volumes less than 250 mL then sonicated. Sonicated solution is then centrifuged in the Fiberlite™ F21- 8x50y fixed angle rotor in a Sorvall LYNX 6000 centrifuge at 18000 rpm for 30 minutes in 4 C. Supernatant is filtered via a Millex™ 33 mm, 0.45-micron PES syringe filter.
[0203] Protein Purification'. Protein is purified from the clarified cell lysate via a Ni- IMAC chromatography using the AKTA-go FPLC system and a 25mL Cytiva Sepharose HPNi IMAC resin packed in a Cl 6-20 column. Used buffers are as follows: 50mM KPi, 250mM KC1 pH=7.7. 5mM BME, and either 25mM Imidazole for binding buffer or 500mM imidazole for elution buffer. After loading the protein into an equilibrated column, it is washed using with 5-10 column volumes of 15% elution buffer followed by a stepwise protein elution where the buffer imidazole concentration is increased to 50% elution buffer for three column volumes. Then a wash with only lysis buffer for 1 column volumes is done followed by a last isocratic wash with 100% elution buffer for two column volumes. The protein peaks in the 50% and 100% elution buffer wash steps are pooled and double dialyzed into a buffer of 50mM Tris 150mM NaCl 5mM Maltose, 5mM BME, 20% V / V Glycerol at pH=8.00 after concentrating the protein solution, it was further purified via size exclusion chromatography using a Cytiva S-200 HR resin equilibrated to the same buffer. Purified proteins were confirmed by either ESI-MS or MALDI. Deconvolution of ESI-MS spectra was done via the UniDec software suite61.
[0204] SDM for ArtNiase-lO-AA: Mutagenesis for ArtN2ase-10 residue H75A and C 125 A knockout mutations was done via a Gibson Assembly via the NEB protocol, sequences were ordered through IDT. The relevant sequences are as follows:TABLE 7. Sequences for mutagenesis.
[0205] Purification of MoFe ase and extraction of FeMoco: This was done as reported previously [62,63],
[0206] Synthesis of CdS / ZnS quantum dots: Synthesis protocol has been devised by a combination of several protocols [56,64,65,65], The details are as follows: in a three neck round bottom flask, 51.2mg of CdO (99.998%, Thermo Scientific), 1.004g Oleic Acid (technical grade, Sigma Aldrich), and 12g Octadecene (ODE, technical grade, Sigma Aldrich) are added. On neck is covered by a septum while, a second one contains the thermometer submerged into the ODE solution, and the RBF is connected to the Schlenk line through a distillation column. The solution is under stirring via a magnetic stirrer and heated via a heating mantle. First, it is heated to 90°C and then vacuum degassed three times to argon, each time including 5 minutes of purging and evacuation. The solution is then heated to 285°C and kept for 5 minutes or until the solution becomes clear which indicates the CdO being fully dissolved. Afterwards the temperature is reduced to 260°C. Two grams of an aerobic and room temperature solution of 3.2mg / g of dissolved sulfur (flakes, 99.99%, Sigma Aldrich) in technical grade ODE is then added via syringe in less than 5 seconds to the solution. The reaction vessel is then taken off the heating mantle after 150 seconds to cool down. Once the solution is close to RT, unreacted components will be extracted with MeOH three times. The UV-vis of the CdS quantum dot is then measured and the based on their absorption maxima, the QD size and concentration is calculated for the addition of a monolayer of ZnS66. One millimeter of technical grade oleylamine (technical grade, Sigma Adri ch) per each lOOnmol of CdS is added to the solution, and then the solution was heated to 80°C and evacuated for 30 minutes. Afterwards, leq of technical grade zinc stearate powder (Sigma Aldrich) and 3.2mg S / lg ODE was added to the ODE solution and the solution was heated to 130°C and vacuum degassed three times into argon. Then the solution was heated at 220°C for 20 minutes and the heating source is removed afterwards. After cooling down, the impurities were extracted with MeOH and the quantum dot is precipitated by addition of acetone and pelleted by centrifuging. The CdS / ZnS quantum dot is then dissolved in chloroform and stored under dinitrogen at 4°C.
[0207] Ligand exchange is done aerobically as follows: QD solution is diluted until the QD main UV-Vis peak absorption roughly equals 2. It is then rigorously stirred overnight with 1.5x volume of aqueous lOOmM BME (Sigma Aldrich) at pH 10 in a water bath at around 55°C. The QD will transfer to the aqueous solution with BME as the surface ligand.
[0208] Cofactor incubations and assays for ArtNiases: Proteins w ere transformed into an anaerobic glovebox and buffer exchanged into an anaerobic tris buffer. NMF extracted FeMoco was added to the apoprotein in anaerobic Tris buffer in an anaerobic chamber such that the final NMF concentration did not exceed 1% N / N. Subsequently, the incubated solutionwas concentrated using a stirred cell in the glovebox to a volume of less than 2 mL and then passed through a G25 column to separate it from non-bound FeMoco [16,36.67],
[0209] Assays were done in triplicates following previously published protocols [36,56,68], Protein concentration for the15N2 reduction assays were 13.6 pM and the15Nammonia standards had a concentration range of 0-250 pM. Product calculation was done via integrating the corresponding NMR peaks.
[0210] EPR spectra were recorded by an ESP 300E spectrometer (Bruker) interfaced with an ESR-9002 liquid-helium continuous-flow cryostat (Oxford Instruments). The measurements were taken using a microwave power of 50 mW, a gain factor of 5 - 104, a modulation frequency of 100 kHz, and a modulation amplitude of 5 G. The spectra were recorded for each sample at 10 K using a microwave frequency of 9.62 GHz and using five scans for each sample
[0069] ,
[0211] Design Information: The full protein sequences of the designs including the 6His- MBP-TevP tags are given in TABLE 8.TABLE 8. Design information.Results
[0213] Computational design of ArtNiases: Previously, successful designs of ArMs have largely relied on a) reengineering existing metalloproteins for new functions through rational mutagenesis, metallocofactor substitution, or directed evolution [36-41], b) introduction of known metal-binding motifs in well-characterized protein scaffolds [32,42], c) de novo design of metalloproteins considering known ternary structures and cofactor binding requirements [31,33,43,44], and d) using dual-purpose protein binders and cofactor ligands to anchor the metal center to the protein or within its active site [30,45], Since FeMoco is among the most unique and complex metallocofactors found in nature and catalyzes one of the most challenging reactions, none of these approaches alone is sufficient for the design of ArtN2ases. Consequently, additional considerations on the pocket shape and cofactor orientation within the protein need to be taken into account. Therefore, this study has employed a five-step design approach involving 1) geometrically defining the binding pocket containing FeMoco and homocitrate, 2) using this definition to search for proteins in the PDB that contain a similar pocket. 3) introducing amino acids residues that coordinate FeMoco and are essential to N2ase function 4) screening candidates for proper cofactor orientation within the pocket, and finally 5) structure and sequence optimization of the passed candidates to probe cofactor binding and stability (FIG. 10)
[0214] The native FeMoco binding site of the NifD subunit of the MoFe protein component of Mo N2ase (called native N2ase hereafter) was used to define the target pocket geometry and as a reference for evaluation of potential matches. PocketSearch
[0046] was then used to screen the RCSB protein database to identify matching solvent-exposed pockets and rank them based on volume and degree of overlap with the reference FeMoco pocket obtained from NifD (FIG. 11). To simplify the expression, purification, and characterization of ArtN2ases, this search was performed across a curated library of -46,000 unique PDB IDs of small, monomeric nonmembrane proteins that could be expressed and purified from E. coli. From this search, -14,000 PDB structures with comparably large, solvent-exposed pockets were identified, with 562 PDBs exhibiting -85% overlap with the reference binding pocket.
[0215] After finding the candidate proteins with a pocket to accommodate the FeMoco, the study used Rosetta Matcher to design residues that spatially occupy similar positions corresponding to His442 and Cys275 in native N2ases, which coordinate Fei and Mo of FeMoco. respectively (FIG. 10) [47,48], In addition, the study introduced residues that correspond to Hisl95 and Glnl91 in the SCS of native N2ases that form hydrogen bondinginteractions with S2B and homocitrate, respectively. Mutagenesis studies have shown that these residues are critical in conferring N2ase activity [26,49,50], Among the 562 proteins in the PDB from the pocket search, 428 had at least one mutational combination matching the intended residue interactions.
[0216] The above combinations were further filtered based on the solvent exposure of the cluster core and homocitrate ligand of the FeMoco, as it was hypothesized that the solvent exposure may play important roles in the N2ase activity. Specifically, current evidence suggests that the FeMoco core needs to be buried within the protein to shield it from the solvent, modulate its activity via the SCS interactions, and establish substrate and proton channels, while the solvent exposure of the homocitrate is necessary to facilitate proton transfer to and from the protein environment
[0051] , Based on this evidence, the study calculated the solvent exposure of the cluster core and the homocitrate moiety of the design candidates and compared them with that of the native N2ase pocket reference. This process narrowed the 12,406 unique PDB-mutation combinations down to 45 scaffolds that display solvent exposures comparable to that of the native N2ase. As the last step, an optimization calculation using the RosettaScripts
[0052] was performed to optimize the protein structure and sequence of the 45 filtered candidates to ensure compatibility of the protein pocket with FeMoco binding. Assessing the optimized structures in this stage, the study settled on 38 final designs with satisfactory7coordination environments, constraints, and total scores that were selected for experimental evaluations.
[0217] Preparation, biochemical and biophysical characterization of ArtNzases: The codon-optimized genes encoding the above 38 designs were cloned into a 2M-T vector including sequences for a hexa-histidine-tagged maltose-binding fusion protein at the N- terminal to facilitate protein expression and purification. Of the 38 constructs, 11 were successfully expressed in E. coll and purified in sufficient quantity to further assess their cofactor binding and catalytic activity. The FeMoco was isolated from native N2ase purified from Azotobacter vinelandii into NMF and then incorporated into the 11 proteins using previously established protocols [36,53], To assess potential FeMoco binding, acety lene reduction activity assays were performed (FIG. 12). While the isolated FeMoco cluster is capable of reducing acetylene to ethylene and ethane on its own, its activity is significantly reduced during turnover over a period of time due to degradation of the cluster, retaining just 48% activity after 3 hours. In contrast, native Ikhase retains 83% of its activity under the same conditions over the same period. Seven of the designs display a time-dependent loss of acetylene reduction activity similar to that displayed by the isolated cluster, suggesting either non-specific FeMoco binding at the protein surface or incomplete incorporation of the clusterinto the designed binding pocket. In contrast, four designs retain activity' during turnover, with modest losses of -25% over the course of 3 hours. Notably, ArtN2ase-10 retains 74% activity over this time period, similar to the 83% activity retention displayed by native Mo N2ase. The retention of catalytic activity suggests improved cluster stability, an important indicator of FeMoco binding. To verify this observation, the designed coordinating cysteine and histidine residues of the ArtN2ase-10 were mutated to Ala to eliminate potential specific covalent binding. This mutant, termed ArtN2ase-10-AA, retained only 55% of activity over 3 hours, less than 79% of ArtN2ase-10 and similar to 48% retention of activity of the isolated FeMoco over the same period, thereby providing support that the designed coordinating residues play an important role in retaining activity.
[0218] To further characterize the FeMoco-bound holo-ArtN2ase-10. the study performed X-band electron paramagnetic resonance (EPR) spectroscopy to compare this construct with isolated FeMoco and FeMoco bound to the native Mo N2ase (FIG. 13). The FeMoco of the resting-state Mo N2ase displays a relatively low rhombicity S = 3 / 2 spectrum with geff = [4.34, 3.66, 2.01] and E / D = 0.056. In comparison, the isolated FeMoco in NMF displays two signals, the first a more rhombic, broadened S = 3 / 2 spectrum arising from the FeMoco (geff = [4.50, 3.58, 2.00], E / D - 0.08), and a large isotropic S = 1 / 2 signal at g - 2 (-3400 G), which arises from a partially degraded cluster. The EPR spectrum of ArtN2ase-10 reveals a single S = 3 / 2 signal with geff = [4.46, 3.57. 2.00] and an intermediate rhombicity of E / D - 0.073, larger than in those of the native Mo N2ase but smaller than those of the isolated FeMoco. It is noted that this spectral broadening is similar to that of lyophilized native Mo N2ase, where desolvation leads to a less well-defined SCS environment54. These results indicate the successful binding of FeMoco to the ArtN2ase-10 scaffold with minimal detectable degradation and a greater degree of electronic symmetry than isolated FeMoco. Finally, the ArtN2ase-10-AA displayed a much weaker (-40%) EPR spectrum than that of ArtN2ase-10, suggesting the importance of the two coordinating ligands designed into ArtN2ase-10.
[0219] ArtNiase-lO reduces15N2 to15NH3: Encouraged by the findings from acety lene reduction activity’ assays and EPR measurements, the study carried out15N2 reduction assays using a photoreduction approach that has previously been successfully applied to native N2ases
[0055] , The study followed the same protocol using ZnS / CdS quantum dots (QDs) with betamercaptoethanol as a capping ligand and dithionite as a sacrificial reductant
[0056] , Excitingly, frequency-selective pulse 1H-NMR spectroscopy revealed15NH4+production by holo- ArtN2ase-10. with a turnover number (TON) of 11 ± 2. In contrast, QD photoreduction system alone or the apo-ArtN2ase-10 did not produce any observable15NH4+under the sameconditions. When the two coordinating residues were mutated to alanine., the holo-ArtN2ase- 10-AA produced less15NH4+(6.6 ± 0.2) than holo-ArtN2ase-10, indicating that the Arthbase- 10-AA pocket is still able to accommodate and protect FeMoco, but the protection provided is not as robust as in Arthbase- 10 due to the absence of the two coordinating residues.Conclusion
[0220] This study has successfully designed ArtNhases, engineered to bind FeMoco and reduce small-molecule substrates of native N2ases, including N2 to NH3. Unlike their natural counterparts, the ArtlShases are small, stable, easily expressed monomeric proteins, lacking additional cofactors and accessory' components. These properties make them invaluable tools for studying biological nitrogen fixation and serve as a foundation for the next generation of catalysts designed to overcome the limitations of the Haber-Bosch process. The ‘"bottom-up” approach, starting from protein scaffolds completely unrelated to leases to gain the N2ase function, allowed the study to both identify the minimal structural features for biological N2 fixation and have unprecedented control of activity'. Beyond advancing understanding of biological N2 fixation, this work represents a major advance in metalloprotein design, demonstrating the feasibility of engineering proteins that can accommodate one of the most complex metallocofactors to facilitate one of the most challenging reactions. The methodologies developed here pave the way for the creation of next-generation artificial enzymes, capable of harnessing FeMoco and other high-nuclearity clusters for diverse applications in small-molecule activation.EXAMPLE ASPECTS
[0221] Example 1: A method of producing a non-naturally occurring peptide for reducing nitrogen, the method comprising: a) identifying a first group of peptides each having a binding pocket similar to a naturally occurring iron-sulfur cluster cofactor binding pocket of a naturally occurring nitrogenase, wherein each of the first group of peptides is not a nitrogenase or a subunit or domain of a nitrogenase; b) selecting, from the first group of peptides, a second group of peptides each having at least one target amino acid in the binding pocket; c) selecting, from the second group of peptides, a third group of peptides each having at least one target property; and d) modifying the sequence and / or structure of each of the third group of peptides to improve binding of the iron-sulfur cluster cofactor in the binding pocket.
[0222] Example 2: The method of any examples herein, particularly Example 1, wherein the iron-sulfur cluster cofactor is 7Fe-9S-C-Mo-R-homocitrate (FeMoco). 7Fe-9S-C-V-R- homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof.
[0223] Example 3: The method of any examples herein, particularly Examples 1-2, wherein the binding pocket of each of the first group of peptides has from about 85% to about 95% structural overlap with the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0224] Example 4: The method of any examples herein, particularly Examples 1-3, wherein the binding pocket of each of the first group of peptides is from about 115% to about 215% of a pocket volume of the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0225] Example 5: The method of any examples herein, particularly Examples 1-4, wherein the naturally occurring nitrogenase is SEQ ID NO: 67.
[0226] Example 6: The method of any examples herein, particularly Examples 1-5, wherein the at least one target amino acid interacts with an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein said interaction replicates or mimics a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442. Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
[0227] Example 7: The method of any examples herein, particularly Examples 1-6, wherein the at least one target property is: having up to about 400 amino acids; being a monomeric peptide; having no post-translational modifications; not being a membrane protein; having a target HCA value and / or a target ICS value; and / or being expressible by a bacterium or yeast.
[0228] Example 8: The method of any examples herein, particularly Example 7, wherein the target HCA value is from about 50 A2to about 100 A2.
[0229] Example 9: The method of any examples herein, particularly Examples 7-8, wherein the target ICS value is from about 0.25 A2to about 10 A2.
[0230] Example 10: The method of any examples herein, particularly Examples 7-9, wherein the at least one target property is being expressible by E. coli.
[0231] Example 11: The method of any examples herein, particularly Examples 1-10, wherein step d) comprises substitution of at least one amino acid in the binding pocket.
[0232] Example 12: The method of any examples herein, particularly Example 11, wherein the substitution creates at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein the at leastone new interaction replicates or mimics a naturally occurring interaction between an ironsulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442, Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
[0233] Example 13: The method of any examples herein, particularly Examples 1-12, wherein step d) comprises altering the HCA value and / or ICS value of the peptide.
[0234] Example 14: The method of any examples herein, particularly Examples 1-13, wherein step d) comprises modifying distance between an iron-sulfur cluster cofactor molecule bound to the binding pocket and each amino acid side chain which interacts with said ironsulfur cluster cofactor molecule.
[0235] Example 15: The method of any examples herein, particularly Example 14, wherein, when the iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule is independently from about 1.5 A to about 2.5 A away from the iron-sulfur cluster cofactor molecule.
[0236] Example 16: The method of any examples herein, particularly Examples 1-15, wherein the method produces a peptide comprising about 80% similarity or more to any one of SEQ ID NOs: 1-66.
[0237] Example 17: The method of any examples herein, particularly Examples 1-16, wherein the method produces a peptide comprising about 90% similarity or more to any one of SEQ ID NOs: 1-66.
[0238] Example 18: The method of any examples herein, particularly Examples 1-17, wherein the method produces a peptide comprising any one of SEQ ID NOs: 1-66.
[0239] Example 19: The method of any examples herein, particularly Examples 1-18, further comprising: e) computationally analyzing each of the third group of peptides to select a fourth group of peptides having low free energy, high binding affinity for the iron-sulfur cluster cofactor, high geometrical similarity to a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase, and / or a minimal number of mutations.
[0240] Example 20: The method of any examples herein, particularly Example 19, wherein each of the fourth group of peptides is selected from about a top 5% to about a top 0.5% of the third group of peptides when ranked by lowest free energy.
[0241] Example 21: A non-naturally occurring peptide for reducing nitrogen, comprising: at least one modification to a naturally occurring peptide; and a binding pocket for binding aniron-sulfur cluster cofactor; wherein the naturally occurring peptide is not a nitrogenase or a subunit or domain of a nitrogenase; and wherein the at least one modification improves binding of the iron-sulfur cluster cofactor in the binding pocket.
[0242] Example 22: The peptide of any examples herein, particularly Example 21, wherein the iron-sulfur cluster cofactor is 7Fe-9S-C-Mo-R-homocitrate (FeMoco), 7Fe-9S-C-V-R- homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof.
[0243] Example 23: The peptide of any examples herein, particularly Examples 21-22. wherein the at least one modification comprises a substitution of at least one amino acid in the binding pocket.
[0244] Example 24: The peptide of any examples herein, particularly Example 23, wherein the substitution creates at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein the at least one new interaction replicates or mimics a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring ironsulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442, Cys275, Hisl95, Glnl91, Arg96, Arg359. Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
[0245] Example 25: The peptide of any examples herein, particularly Examples 21-24, wherein the at least one modification increases the number of amino acid side chains in the binding pocket which interact with the iron-sulfur cluster cofactor.
[0246] Example 26: The peptide of any examples herein, particularly Examples 21 -25, wherein the peptide has an HCA value of from about 50 A2to about 100 A2.
[0247] Example 27: The peptide of any examples herein, particularly Examples 21-26, wherein the peptide has an ICS value of from about 0.25 A2to about 10 A2.
[0248] Example 28: The peptide of any examples herein, particularly Examples26-27, wherein the at least one modification alters the HCA value and / or ICS value of the peptide.
[0249] Example 29: The peptide of any examples herein, particularly Examples 21-28, wherein, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule is independently from about 1.5 A to about 2.5 A away from the iron-sulfur cluster cofactor molecule.
[0250] Example 30: The peptide of any examples herein, particularly Example 29, wherein the at least one modification modifies distance between the iron-sulfur cluster cofactormolecule and each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule.
[0251] Example 31: The peptide of any examples herein, particularly Examples 21-30, wherein the peptide is monomeric.
[0252] Example 32: The peptide of any examples herein, particularly Examples 21-31, wherein the peptide comprises up to about 400 amino acids.
[0253] Example 33: The peptide of any examples herein, particularly Examples 21-32. wherein the peptide does not have any post-translational modifications.
[0254] Example 34: The peptide of any examples herein, particularly Examples 21-33, wherein the naturally occurring peptide is not a membrane protein.
[0255] Example 35: The peptide of any examples herein, particularly Examples 21-34, wherein the peptide reduces N2 to ammonia.
[0256] Example 36: The peptide of any examples herein, particularly Example 35, wherein the peptide does not consume ATP in the reduction of N2 to ammonia.
[0257] Example 37: The peptide of any examples herein, particularly Examples 35-36, wherein the peptide has a turnover frequency (TOF) of ammonia of from about 3 min’1to about 15 min’1.
[0258] Example 38: The peptide of any examples herein, particularly Examples 35-37, wherein a reductase activity of the peptide is up to about 350% greater than a reductase activity of a naturally occurring nitrogenase.
[0259] Example 39: The peptide of any examples herein, particularly Examples 21 -38, wherein the peptide comprises about 80% similarity or more to any one of SEQ ID NOs: 1-66.
[0260] Example 40: The peptide of any examples herein, particularly Examples 21-39, wherein the peptide comprises about 90% similarity or more to any one of SEQ ID NOs: 1-66.
[0261] Example 41: The peptide of any examples herein, particularly Examples 21-40, wherein the peptide comprises any one of SEQ ID NOs: 1-66.
[0262] Example 42: The peptide of any examples herein, particularly Examples 21-41, wherein the peptide can be expressed by a bacterium or a yeast.
[0263] Example 43: The peptide of any examples herein, particularly Example 42, wherein the peptide can be expressed by E. coll.
[0264] Example 44: The peptide of any examples herein, particularly Examples 21-43, wherein the peptide is produced by the method of any examples herein, particularly Examples 1-20.
[0265] Example 45: A vector encoding the peptide of any examples herein, particularly Examples 21-44.
[0266] Example 46: The vector of any examples herein, particularly Example 45, wherein the vector is a plasmid.
[0267] Example 47: A cell comprising the vector of any examples herein, particularly Examples 45-46.
[0268] Example 48: The cell of any examples herein, particularly Example 47, wherein the cell is a bacterium or a yeast.
[0269] Example 49: The cell of any examples herein, particularly Example 48, wherein the cell is E. coli.
[0270] Example 50: A method of making the peptide of any examples herein, particularly Examples 18-40, the method comprising transfecting the vector of any examples herein, particularly Examples 45-46 into a cell.
[0271] Example 51: The method of any examples herein, particularly Example 50, wherein the cell is a bacterium or yeast.
[0272] Example 52: The method of any examples herein, particularly Example 51. wherein the cell is E. coli.
[0273] Example 53: A method of making ammonia, the method comprising exposing the peptide of any examples herein, particularly Examples 21-44 to nitrogen and an iron-sulfur cluster cofactor.
[0274] Example 54: The method of any examples herein, particularly Example 53, wherein the nitrogen is gaseous N2.
[0275] Example 55: The method of any examples herein, particularly Examples 53-54, wherein the iron-sulfur cluster cofactor is 7Fe-9S-C-Mo-R-homocitrate (FeMoco). 7Fe-9S-C- V-R-homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof.
[0276] Example 56: The method of any examples herein, particularly Examples 53-55, wherein the method does not consume ATP.
[0277] Example 57: The method of any examples herein, particularly Example 56. wherein the method uses a chemical electron donor or is carried out on a photoexciteable surface.
[0278] Example 58: The method of any examples herein, particularly Examples 53-57, wherein the method is carried out at a temperature and / or pressure that is less than a temperature and / or pressure of the Haber-Bosch process.
[0279] Example 59: A method of producing a non-naturally occurring peptide for reducing nitrogen, the method comprising: a) identifying a first group of peptides each having a binding pocket similar to a naturally occurring iron-sulfur cluster cofactor binding pocket of a naturally occurring nitrogenase, wherein each of the first group of peptides is not a nitrogenase or a subunit or domain of a nitrogenase; b) selecting, from the first group of peptides, a second group of peptides each having at least one target amino acid in the binding pocket; c) selecting, from the second group of peptides, a third group of peptides each having at least one target property; and d) modifying the sequence and / or structure of each of the third group of peptides to improve binding of the iron-sulfur cluster cofactor in the binding pocket.
[0280] Example 60: The method of any examples herein, particularly Example 59, wherein the iron-sulfur cluster cofactor is 7Fe-9S-C-Mo-R-homocitrate (FeMoco). 7Fe-9S-C- V-R-homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof.
[0281] Example 61: The method of any examples herein, particularly Example 59, wherein the binding pocket of each of the first group of peptides has from about 85% to about 95% structural overlap with the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0282] Example 62: The method of any examples herein, particularly Example 59, wherein the binding pocket of each of the first group of peptides is from about 115% to about 215% of a pocket volume of the naturally occurring iron-sulfur cluster cofactor binding pocket of the naturally occurring nitrogenase.
[0283] Example 63: The method of any examples herein, particularly Example 59, wherein the naturally occurring nitrogenase is SEQ ID NO: 67.
[0284] Example 64: The method of any examples herein, particularly Example 59, w herein the at least one target amino acid interacts with an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein said interaction replicates or mimics a naturally occurring interaction betw een an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442. Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
[0285] Example 65: The method of any examples herein, particularly Example 59, wherein the at least one target property is: having up to about 400 amino acids; being a monomeric peptide; having no post-translational modifications; not being a membrane protein;having a target HCA value and / or a target ICS value; and / or being expressible by a bacterium or veast.
[0286] Example 66: The method of any examples herein, particularly Example 65, wherein the target HCA value is from about 50 A2to about 100 A2.
[0287] Example 67: The method of any examples herein, particularly Example 65, wherein the target ICS value is from about 0.25 A2to about 10 A2.
[0288] Example 68: The method of any examples herein, particularly Example 65. wherein the at least one target property is being expressible by E. coll.
[0289] Example 69: The method of any examples herein, particularly Example 59, wherein step d) comprises one or more of: substitution of at least one amino acid in the binding pocket; altering the HCA value and / or ICS value of the peptide; or modifying distance between an iron-sulfur cluster cofactor molecule bound to the binding pocket and each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule.
[0290] Example 70: The method of any examples herein, particularly Example 69, wherein the substitution creates at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein the at least one new interaction replicates or mimics a naturally occurring interaction between an ironsulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442. Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
[0291] Example 71: The method of any examples herein, particularly Example 69, wherein, when the iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule is independently from about 1.5 A to about 2.5 A away from the iron-sulfur cluster cofactor molecule.
[0292] Example 72: The method of any examples herein, particularly Example 59, further comprising: e) computationally analyzing each of the third group of peptides to select a fourth group of peptides having low free energy, high binding affinity for the iron-sulfur cluster cofactor, high geometrical similarity to a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase, and / or a minimal number of mutations.
[0293] Example 73: The method of any examples herein, particularly Example 72, wherein each of the fourth group of peptides is selected from about a top 5% to about a top 0.5% of the third group of peptides when ranked by lowest free energy.
[0294] Example 74: A non-naturally occurring peptide for reducing nitrogen produced by the method of any examples herein, particularly Example 59.
[0295] Example 75: The peptide of any examples herein, particularly Example 74, wherein the peptide comprises about 80% similarity or more to any one of SEQ ID NOs: 1-66.
[0296] Example 76: A method of making ammonia, the method comprising exposing the peptide of any examples herein, particularly Example 74, to nitrogen and an iron-sulfur cluster cofactor.
[0297] Example 77: The method of any examples herein, particularly Example 76, wherein the method uses a chemical electron donor or is carried out on a photoexciteable surface.
[0298] Example 78: The method of any examples herein, particularly Example 76, wherein the method is carried out at a temperature and / or pressure that is less than a temperature and / or pressure of the Haber-Bosch process.
[0299] Example 79: A non-naturally occurring peptide for reducing nitrogen, comprising: at least one modification to a naturally occurring peptide; and a binding pocket for binding an iron-sulfur cluster cofactor; wherein the naturally occurring peptide is not a nitrogenase or a subunit or domain of a nitrogenase: and wherein the at least one modification improves binding of the iron-sulfur cluster cofactor in the binding pocket.
[0300] Example 80: The peptide of any examples herein, particularly Example 79, wherein the iron-sulfur cluster cofactor is 7Fe-9S-C-Mo-R-homocitrate (FeMoco). 7Fe-9S-C-V-R- homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof.
[0301] Example 81: The peptide of any examples herein, particularly Example 79, wherein the at least one modification comprises a substitution of at least one amino acid in the binding pocket.
[0302] Example 82: The peptide of any examples herein, particularly Example 81, wherein the substitution creates at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein the at least one new interaction replicates or mimics a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring ironsulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442, Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
[0303] Example 83: The peptide of any examples herein, particularly Example 79, wherein the at least one modification increases the number of amino acid side chains in the binding pocket which interact with the iron-sulfur cluster cofactor.
[0304] Example 84: The peptide of any examples herein, particularly Example 79, wherein the peptide has an HCA value of from about 50 A2to about 100 A2and / or an ICS value of from about 0.25 A2to about 10 A2.
[0305] Example 85: The peptide of any examples herein, particularly Example 84, wherein the at least one modification alters the HCA value and / or ICS value of the peptide.
[0306] Example 86: The peptide of any examples herein, particularly Example 79, wherein, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule is independently from about 1.5 A to about 2.5 A away from the iron-sulfur cluster cofactor molecule.
[0307] Example 87: The peptide of any examples herein, particularly Example 86, wherein the at least one modification modifies distance between the iron-sulfur cluster cofactor molecule and each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule.
[0308] Example 88: The peptide of any examples herein, particularly Example 79, wherein the peptide is monomeric.
[0309] Example 89: The peptide of any examples herein, particularly Example 79, wherein the peptide does not have any post-translational modifications.
[0310] Example 90: The peptide of any examples herein, particularly Example 79, wherein the naturally occurring peptide is not a membrane protein.
[0311] Example 91: The peptide of any examples herein, particularly Example 79, wherein the peptide reduces N2 to ammonia.
[0312] Example 92: The peptide of any examples herein, particularly Example 91, wherein the peptide does not consume ATP in the reduction of N2 to ammonia.
[0313] Example 93: The peptide of any examples herein, particularly Example 91, wherein the peptide has a turnover frequency (TOF) of ammonia of from about 3 min’1to about 15 min’ and wherein a reductase activity of the peptide is up to about 350% greater than a reductase activity of a naturally occurring nitrogenase.
[0314] Example 94: The peptide of any examples herein, particularly Example 79, wherein the peptide comprises about 80% similarity or more to any one of SEQ ID NOs: 1-66.
[0315] Example 95: The peptide of any examples herein, particularly Example 79, wherein the peptide can be expressed by a bacterium or a yeast.
[0316] Example 96: A method of making ammonia, the method comprising exposing the peptide of any examples herein, particularly Example 79, to nitrogen and an iron-sulfur cluster cofactor.
[0317] Example 97: The method of any examples herein, particularly Example 96, wherein the method uses a chemical electron donor or is carried out on a photoexciteable surface.
[0318] Example 98: The method of any examples herein, particularly Example 96, wherein the method is carried out at a temperature and / or pressure that is less than a temperature and / or pressure of the Haber-Bosch process.
[0319] The following patents, applications and publications as listed below and throughout this document are hereby incorporated by reference in their entirety herein.Reference List1. Zhang. X., et al. Global Nitrogen Cycle: Critical Enzymes. Organisms, and Processes for Nitrogen Budgets and Dynamics. Chem. Rev. 120, 5308-5351 (2020).2. Erisman, J. W., et al. How a century of ammonia synthesis changed the world. Nat. Geosci. 1, 636-639 (2008).3. Executive Summary - Ammonia Technology Roadmap - Analysis. IEA.4. Foster, S. L. et al. Catalysts for nitrogen reduction to ammonia. Nat. Catal. 1, 490-500 (2018).5. Westhead, O. et al. Near ambient N2 fixation on solid electrodes versus enzymes and homogeneous catalysts. Nat. Rev. Chem. 7, 184-201 (2023).6. Chen, J. G. et al. Beyond fossil fuel-driven nitrogen transformations. Science 360, eaar6611 (2018).7. Dos Santos, P. C. Genomic Manipulations of the Diazotroph Azotobacter vinelandii. in Metalloproteins: Methods and Protocols (ed. Hu, Y.) 91-109 (Springer, New York, NY, 2019). doi : 10. 1007 / 978- 1 -4939-8864-8_6.8. Ryu, M.-H. et al. Control of nitrogen fixation in bacteria that associate with cereals. Nat. Microbiol. 5, 314-330 (2020).9. Li, X -X., et al. Using synthetic biology to increase nitrogenase activity. Microb. Cell Factories 15, 43 (2016).10. Solomon. J. B. et al. Ammonia synthesis via an engineered nitrogenase assembly pathway in Escherichia coli. Nat. Catal. 7, 1130-1141 (2024).11. Bennett, E. M., et al. Engineering Nitrogenases for Synthetic Nitrogen Fixation: From Pathway Engineering to Directed Evolution. Biodesign Res. 5, 0005.12. Spatzal, T. et al. Evidence for Interstitial Carbon in Nitrogenase FeMo Cofactor. Science 334, 940-940 (2011).13. Lancaster, K. M. et al. X-ray Emission Spectroscopy Evidences a Central Carbon in the Nitrogenase Iron-Molybdenum Cofactor. Science 334, 974-977 (201 1).14. Einsle, O. & Rees, D. C. Structural Enzymology of Nitrogenase Enzymes. Chem. Rev. 120, 4969-5004 (2020).15. Van Stappen, C. et al. The Spectroscopy of Nitrogenases. Chem. Rev. 120, 5005-5081 (2020).16. Badding, E. D., et al. Connecting the geometric and electronic structures of the nitrogenase iron-molybdenum cofactor through site-selective57Fe labelling. Nat. Chem. 15, 658-665 (2023).17. Owens, C. P., et al. Evidence for Functionally Relevant Encounter Complexes in Nitrogenase Catalysis. J. Am. Chem. Soc. 137, 12704-12712 (2015).18. Sengupta, K. et al. Investigating the Molybdenum Nitrogenase Mechanistic Cycle Using Spectroelectrochemistry. J. Am. Chem. Soc. 147, 2099-2114 (2025).19. Mortenson. L. E. Ferredoxin and atp. requirements for nitrogen fixation in cell-free extracts of Clostridium pasteurianum*. Proc. Natl. Acad. Sci. 52, 272-279 (1964).20. Hoffman, B. M., et al. Mechanism of Nitrogen Fixation by Nitrogenase: The Next Stage. Chem. Rev. 114, 4041-4062 (2014).21. Lee, S. C., et al. Developments in the Biomimetic Chemistry of Cubane-Type and Higher Nuclearity Iron-Sulfur Clusters. Chem. Rev. 114, 3579-3600 (2014).22. Chalkley. M. J., et al. Catalytic N240-NH3 (or -N2H4) Conversion by Well-Defined Molecular Coordination Complexes. Chem. Rev. 120, 5582-5636 (2020).23. Tanifuji, K. & Ohki, Y. Metal-Sulfur Compounds in N2 Reduction and Nitrogenase- Related Chemistry. Chem. Rev. 120, 5194-5251 (2020).24. Wang, C.-H. & DeBeer, S. Structure, reactivity, and spectroscopy of nitrogenase- related synthetic and biological clusters. Chem. Soc. Rev. 50, 8743-8761 (2021).25. Wilson, D. W. N. & Holland, P. L. 15.03 - Nitrogenases and Model Complexes in Bioorganometallic Chemistry, in Comprehensive Organometallic Chemistry' IV (eds. Parkin, G„ Meyer, K. & O’hare, D.) 41-72 (Elsevier, Oxford, 2022). doi: 10.1016 / B978-0-12-820206- 7.00035-4.26. Stripp, S. T. et al. Second and Outer Coordination Sphere Effects in Nitrogenase. Hydrogenase, Formate Dehydrogenase, and CO Dehydrogenase. Chem. Rev. 122, 11900- 11973 (2022).27. Pettersen, E. F. et al. UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci. Publ. Protein Soc. 30, 70-82 (2021).28. Yeung, N. et al. Rational design of a structural and functional nitric oxide reductase. Nature 462, 1079-1082 (2009).29. Siegel, J. B. et al. Computational Design of an Enzyme Catalyst for a Stereoselective Bimolecular Diels-Alder Reaction. Science 329, 309-313 (2010).30. Olshansky', L. et al. Artificial Metalloproteins Containing CO4O4 Cubane Active Sites. J. Am. Chem. Soc. 140, 2739-2742 (2018).31. Chino, M. et al. Spectroscopic and metal binding properties of a de novo metalloprotein binding a tetrazinc cluster. Biopolymers 109, e23339 (2018).32. Mirts, E. N., et al. A designed heme-[4Fe-4S] metalloenzyme catalyzes sulfite reduction like the native enzyme. Science 361, 1098-1101 (2018).33. Kalvet, I. et al. Design of Heme Enzymes with a Tunable Substrate Binding Pocket Adjacent to an Open Metal Coordination Site. J. Am. Chem. Soc. 145, 14307-14315 (2023).34. Yu. Y. et al. Direct EPR Observation of a Tyrosyl Radical in a Functional Oxidase Model in Myoglobin during both H202 and O2 Reactions. J. Am. Chem. Soc. 136, 1174-1177 (2014).35. Bhagi-Damodaran, A. et al. Why copper is preferred over iron for oxygen activation and reduction in haem-copper oxidases. Nat. Chem. 9, 257-263 (2017).36. Tanifuji, K. et al. Combining a Nitrogenase Scaffold and a Synthetic Compound into an Artificial Enzyme. Angew. Chem. 127, 14228-14231 (2015).37. Dydio, P. et al. An artificial metalloenzyme with the kinetics of native enzymes. Science 354, 102-106 (2016).38. Hosseinzadeh, P. et al. Design of a single protein that spans the entire 2-V range of physiological redox potentials. Proc. Natl. Acad. Sci. U. S. A. 113, 262-267 (2016).39. Slater, J. W. et al. Power of the Secondary Sphere: Modulating Hydrogenase Activity in Nickel-Substituted Rubredoxin. Acs Catal. 9, 8928-8942 (2019).40. Yang, Y. & Arnold, F. H. Navigatingthe Unnatural Reaction Space: Directed Evolution of Heme Proteins for Selective Carbene and Nitrene Transfer. Acc. Chem. Res. 54, 1209-1225 (2021).41. Tanifuji, K. et al. Incorporation of an Asymmetric Mo-Fe-S Cluster as an Artificial Cofactor into Nitrogenase. ChemBioChem 23, e202200384 (2022).42. Heinisch, T. et al. Improving the Catalytic Performance of an Artificial Metalloenzyme by Computational Design. J. Am. Chem. Soc. 137, 10414-10419 (2015).43. Krishna, R. et al. Generalized biomolecular modeling and design with RoseTTAFold All-Atom. Science 384, eadl2528 (2024).44. Chalkley. M. J., et al. De novo metalloprotein design. Nat. Rev. Chem. 6, 31-50 (2022).45. Waser, V., et al. An Artificial [Fe4S4]-Containing Metalloenzyme for the Reduction of CO2 to Hydrocarbons. J. Am. Chem. Soc. 145, 14823-14830 (2023).46. Sinclair, M. pocketSearch. (2021).47. Zanghellini, A. et al. New algorithms and an in silico benchmark for computational enzyme design. Protein Sci. 15, 2785-2794 (2006).48. Richter, F., et al. De Novo Enzyme Design Using Rosetta3. PLOS ONE 6, el9230 (2011).49. Kim, C. H , et al. Role of the MoFe protein alpha-subunit histidine-195 residue in FeMo-cofactor binding and nitrogenase catalysis. Biochemistry 34. 2798-2808 (1995).50. Scott. D. J., et al. Role for the nitrogenase MoFe protein a-subunit in FeMo-cofactor binding and catalysis. Nature 343, 188-190 (1990).51. Seefeldt, L. C. et al. Reduction of Substrates by Nitrogenases. Chem. Rev. 120, 5082- 5106 (2020).52. Bender. B. J. et al. Protocols for Molecular Modeling with Rosetta3 and RosettaS cripts. Biochemistry 55, 4748-4763 (2016).53. Rebelein. J. G., et al. Characterization of an M-Cluster-Substituted Nitrogenase VFe Protein. mBio 9, 10.1128 / mbio.00310-18 (2018).54. Van Stappen, C., et al. Preparation and spectroscopic characterization of lyophilized Mo nitrogenase. JBIC J. Biol. Inorg. Chem. (2020) doi: 10.1007 / s00775-020-01838-4.55. Brown, K. A. et al. Light-driven dinitrogen reduction catalyzed by a CdS nitrogenase MoFe protein biohybrid. Science 352, 448-450 (2016).56. Ding. Y. et al. Light-driven Transformation of Carbon Monoxide into Hydrocarbons using CdS@ZnS : VFe Protein Biohybrids. ChemSusChem 16, e202300981 (2023).57. Mills, J. H. et al. Computational Design of an Unnatural Amino Acid Dependent Metalloprotein with Atomic Level Accuracy. J. Am. Chem. Soc. 135, 13393-13399 (2013).58. Le Guilloux, V., et al. Fpocket: An open source platform for ligand pocket detection. BMC Bioinformatics 10. 168 (2009).59. Chen, B. Y. & Honig, B. VASP: a volumetric analysis of surface properties yields insights into protein-ligand binding specificity. PLoS Comput. Biol. 6, el000881 (2010).60. RosettaAtomTypes.61. Marty, M. T. et al. Bayesian Deconvolution of Mass and Ion Mobility Spectra: From Binary Interactions to Polydisperse Ensembles. Anal. Chem. 87, 4370-4376 (2015).62. Lee, C.-C., et al. Purification of Nitrogenase Proteins, in Metalloproteins: Methods and Protocols (ed. Hu, Y.) 111-124 (Springer, New York, NY, 2019). doi: 10. 1007 / 978-1-4939- 8864-8 7.63. Yang, S. S. et al. Iron-molybdenum cofactor from nitrogenase. Modified extraction methods as probes for composition. J. Biol. Chem. 257. 8042-8048 (1982).64. Yu. W. W. & Peng. X. Formation of High-Quality CdS and Other II-VI Semiconductor Nanocrystals in Noncoordinating Solvents: Tunable Reactivity of Monomers. Angew. Chem. Int. Ed. 41, 2368-2371 (2002).65. Chen, D., et al. Bright and Stable Purple / Blue Emitting CdS / ZnS Core / Shell Nanocrystals Grown by Thermal Cycling Using a Single-Source Precursor. Chem. Mater. 22, 1437-1444 (2010).66. Yu. W. W., et al. Experimental Determination of the Extinction Coefficient of CdTe. CdSe, and CdS Nanocrystals. Chem. Mater. 15, 2854-2860 (2003).67. H. Phillips, A. et al. Environment and coordination of FeMo-co in the nitrogenase metallochaperone NafY. RSC Chem. Biol. 2, 1462-1465 (2021).68. Lee, C. C., et al. ATP -Independent Formation of Hydrocarbons Catalyzed by Isolated Nitrogenase Cofactors. Angew. Chem. 124, 1983-1985 (2012).69. Solomon. J. B. et al. Heterologous expression of a fully active Azotobacter vinelandii nitrogenase Fe protein in Escherichia coli. mBio 14, e02572-23 (2023).
Claims
1. What is claimed is:
1. A non-naturally occurring peptide for reducing nitrogen, comprising: at least one modification to a naturally occurring peptide; and a binding pocket for binding an iron-sulfur cluster cofactor; wherein the naturally occurring peptide is not a nitrogenase or a subunit or domain of a nitrogenase; and wherein the at least one modification improves binding of the iron-sulfur cluster cofactor in the binding pocket.
2. The peptide of claim 1, wherein the iron-sulfur cluster cofactor is 7Fe-9S-C-Mo-R- homocitrate (FeMoco), 7Fe-9S-C-V-R-homocitrate (FeVco), 7Fe-9S-C-Fe-R-homocitrate (FeFeco), or a synthetic analog thereof.
3. The peptide of claim 1, wherein the at least one modification comprises a substitution of at least one ammo acid in the binding pocket.
4. The peptide of claim 3, wherein the substitution creates at least one new interaction between a substituted amino acid and an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein the at least one new interaction replicates or mimics a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442, Cys275, Hisl95, Glnl91, Arg96, Arg359, Val70, Ser278, Tyr229, and / or Phe381 of SEQ ID NO: 67.
5. The peptide of claim 1, wherein the at least one modification increases the number of amino acid side chains in the binding pocket which interact with the iron-sulfur cluster cofactor.
6. The peptide of claim 1, wherein the peptide has an HCA value of from about 50 A2to about 100 A2and / or an ICS value of from about 0.25 A2to about 10 A2.
7. The peptide of claim 6. wherein the at least one modification alters the HCA value and / or ICS value of the peptide.
8. The peptide of claim 1, wherein, when an iron-sulfur cluster cofactor molecule is bound to the binding pocket, each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule is independently from about 1.5 A to about 2.5 A away from the iron-sulfur cluster cofactor molecule.
9. The peptide of claim 8, wherein the at least one modification modifies distance between the iron-sulfur cluster cofactor molecule and each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule.
10. The peptide of claim 1, wherein the peptide is monomeric.
11. The peptide of claim 1, wherein the peptide does not have any post-translational modifications.
12. The peptide of claim 1, wherein the naturally occurring peptide is not a membrane protein.
13. The peptide of claim 1 , wherein the peptide reduces N2 to ammonia.
14. The peptide of claim 13, wherein the peptide does not consume ATP in the reduction of N2 to ammonia.
15. The peptide of claim 13, wherein the peptide has a turnover frequency (TOF) of ammonia of from about 3 min'1to about 15 min'1; and wherein a reductase activity of the peptide is up to about 350% greater than a reductase activity of a naturally occurring nitrogenase.
16. The peptide of claim 1, wherein the peptide comprises about 80% similarity or more to any one of SEQ ID NOs: 1-66.
17. The peptide of claim 1. wherein the peptide can be expressed by a bacterium or a yeast.
18. A method of making ammonia, the method comprising exposing the peptide of claim 1 to nitrogen and an iron-sulfur cluster cofactor.
19. The method of claim 18. wherein the method uses a chemical electron donor or is carried out on a photoexciteable surface.
20. The method of claim 18, wherein the method is carried out at a temperature and / or pressure that is less than a temperature and / or pressure of the Haber-Bosch process.
21. A method of producing a non-naturally occurring peptide for reducing nitrogen, the method comprising: a) identifying a first group of peptides each having a binding pocket similar to a naturally occurring iron-sulfur cluster cofactor binding pocket of a naturally occurring nitrogenase, wherein each of the first group of peptides is not a nitrogenase or a subunit or domain of a nitrogenase; b) selecting, from the first group of peptides, a second group of peptides each having at least one target amino acid in the binding pocket; c) selecting, from the second group of peptides, a third group of peptides each having at least one target property; and d) modify ing the sequence and / or structure of each of the third group of peptides to improve binding of the iron-sulfur cluster cofactor in the binding pocket.
22. The method of claim 21, wherein the at least one target amino acid interacts with an iron-sulfur cluster cofactor molecule bound to the binding pocket; wherein said interaction replicates or mimics a naturally occurring interaction between an iron-sulfur cluster cofactor molecule and at least one naturally occurring amino acid in a naturally occurring iron-sulfur cluster cofactor binding pocket in a naturally occurring nitrogenase; and wherein the at least one naturally occurring amino acid is His442, Cys275, His 195, Glnl91, Arg96, Arg359. Val70, Ser278. Tyr229, and / or Phe381 of SEQ ID NO: 67.
23. The method of claim 21, wherein the at least one target property is: having up to about 400 amino acids; being a monomeric peptide; having no post-translational modifications; not being a membrane protein; having a target HCA value and / or a target ICS value; and / or being expressible by a bacterium or yeast.
24. The method of claim 21, wherein step d) comprises one or more of: substitution of at least one amino acid in the binding pocket; altering the HCA value and / or ICS value of the peptide; or modifying distance between an iron-sulfur cluster cofactor molecule bound to the binding pocket and each amino acid side chain which interacts with said iron-sulfur cluster cofactor molecule.
25. A non-naturally occurring peptide for reducing nitrogen produced by the method of claim 21.
Citation Information
Patent Citations
Nitrogenase protein variant
CN113755459A
Compositions and methods for bioremediation
US20040038382A1
Compositions and methods for expression of nitrogenase in plant cells
US20160304842A1
N2-reduction by simplified nifen-based nitrogenase systems
WO2025240869A1