Plant enzymes with improved activity
By replacing specific amino acids in GCC, the activity and energy efficiency of hydroxyacetyl-CoA carboxylase were improved, solving the problem of low carbon dioxide fixation efficiency in photosynthesis and promoting the improvement of biomass production and carbon yield.
Patent Information
- Application Number
- CN202480049177.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-02
- Filing Date
- 2024-05-31
- Publication Date
- 2026-02-27
AI Technical Summary
In the prior art, RuBisCO enzymes cannot reliably distinguish between oxygen and carbon dioxide, which weakens the photosynthetic efficiency of the Calvin-Benson cycle and results in undesirable carbon loss during photorespiration. Therefore, it is necessary to improve the activity of carbon dioxide fixation enzymes to increase the growth rate and carbon production of photosynthetic organisms.
By replacing amino acids in hydroxyacetyl-CoA carboxylase (GCC), particularly at positions 20 and 100, its specific activity for hydroxyacetyl-CoA carboxylation is increased and ATP consumption is reduced, thus optimizing the tartrate-CoA pathway to reduce carbon loss.
It has achieved an increase in GCC activity, specifically a 2-3 fold increase in carboxylation activity, a reduction of more than 50% in ATP consumption, an increase in catalytic rate, an improvement in energy efficiency, and a promotion of carbon dioxide fixation and biomass production.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention provides a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized in that it comprises an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 1, and has one or more amino acid substitutions, deletions or insertions at a position selected from the group consisting of the 20th and 100th positions of the amino acid sequence shown in SEQ ID NO: 1 or at a position corresponding to any of these positions, preferably wherein the GCC has improved activity relative to a reference GCC, preferably wherein the reference GCC comprises the amino acid sequence of SEQ ID NO: 1. Background Technology
[0002] Photosynthesis plays a crucial role in the global carbon cycle. It converts carbon dioxide into organic compounds, which are used as feedstock by heterotrophic organisms. However, the efficiency of photosynthesis in the Calvin-Benson cycle is compromised because RuBisCO enzymes cannot reliably distinguish between oxygen and carbon dioxide. Photorespiration is the pathway required to recover 2-phosphoglycolic acid, a byproduct of RuBisCO oxidation, but it results in undesirable carbon loss. The inventors recently developed the tartronyl-CoA (TaCo) pathway, a synthetic carboxylation module that circumvents this carbon loss and further fixes additional carbon dioxide (see WO 2016 / 207219, etc.). This pathway requires less adenosine triphosphate (ATP) and fewer reducing equivalents to produce biomass, thus promising to improve the growth rate and carbon yield of photosynthetic organisms. Using structure-guided enzyme engineering combined with large-scale mutagenesis library screening, the key enzyme in this pathway—hydroxyacetyl-CoA carboxylase (GCC M5)—was evolved. In the iterative engineering process, five mutations were introduced to convert the previous propionyl-CoA carboxylase (PCC) into GCC M5, enabling it to... (The sentence is incomplete and requires more context to translate accurately.) -1 The catalytic rate of GCCM5 for carboxylation of hydroxyacetyl-CoA is comparable to that of the natural carboxylase (Scheffen, Nat. Catal. 4, 105-115, 2021). However, each carboxylation by GCCM5 requires approximately 4 ATP molecules, compared to a theoretical ratio of 1.
[0003] Therefore, there is a need to provide means and methods to improve the activity of carbon dioxide fixative enzymes (especially hydroxyacetyl-CoA carboxylase). Summary of the Invention
[0004] The technical problem upon which this invention is based is to provide a hydroxyacetyl-CoA carboxylase with improved activity.
[0005] This technical problem is solved by providing embodiments as described in the claims and provided below. Specifically, the above-mentioned technical problem and difficulties are overcome by providing a hydroxyacetyl-CoA carboxylase (GCC) with specific amino acid substitutions. The inventors have determined that the substitution of amino acids at positions 20 and 100 in the GCC of SEQ ID NO: 1 results in improved activity.
[0006] Therefore, the present invention provides a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized in that: it comprises an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 1, and has one or more amino acid substitutions, deletions or insertions at a position selected from the group consisting of the 20th and 100th positions of the amino acid sequence shown in SEQ ID NO: 1 or at a position corresponding to any of these positions, preferably wherein the GCC has improved activity relative to a reference GCC, preferably wherein the reference GCC comprises the amino acid sequence of SEQ ID NO: 1.
[0007] The above content will be described in the following examples.
[0008] The present invention will now be described in more detail.
[0009] The inventors have discovered that certain amino acid substitutions in GCC increase the specific activity of GCC for hydroxyacetyl-CoA carboxylation and the ATP / CO2 ratio. The inventors can demonstrate that certain amino acid substitutions increase the specific activity of GCC for hydroxyacetyl-CoA carboxylation and the ATP / CO2 ratio (Table 6). Specifically, the inventors have found that the amino acid substitution G20R increases the carboxylation activity (i.e., the conversion of hydroxyacetyl-CoA to tartratel-CoA) by 2-3 times. Figure 1 Furthermore, the inventors have discovered that amino acid substitution S100N (relative to GCC M5; SEQ ID NO: 1) reduces ATP consumption by more than 50%. Figure 2 In other words, the substitution improves energy efficiency.
[0010] The present invention is characterized by the following:
[0011] 1. A hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0012] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0013] At a position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions, there is one or more amino acid substitutions, deletions, or insertions.
[0014] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC comprises the amino acid sequence of SEQ ID NO: 1.
[0015] 2. The GCC according to item 1, wherein
[0016] (1) The amino acid at position 20 in the amino acid sequence shown in SEQ ID NO: 1, or the amino acid at the corresponding position, is replaced by alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine, preferably by arginine, histidine, lysine, or tyrosine, and more preferably by arginine.
[0017] Preferably, the substituted amino acid is glycine; and / or
[0018] (2) The amino acid at position 100 in the amino acid sequence shown in SEQ ID NO: 1, or the amino acid at the corresponding position, is replaced by alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, threonine, tryptophan, tyrosine, or valine, preferably by asparagine.
[0019] Preferably, the amino acid being replaced is serine.
[0020] 3. The GCC according to item 1 or 2, wherein the improved activity is improved energy efficiency.
[0021] 4. The GCC according to item 3, wherein the improved energy efficiency is a reduction in ATP consumption, preferably a reduction of at least 50%.
[0022] 5. The GCC according to any one of items 1 to 4, wherein the ATP / CO2 ratio of said GCC is less than about 4.00 ± 0.02.
[0023] 6. The GCC according to any one of items 1 to 5, wherein the ATP / CO2 ratio of the GCC is about 1.71 ± 0.11.
[0024] 7. The GCC according to any one of items 1 to 6, wherein the improved activity is an increase in carboxylation activity, preferably at least 2 to 3 times.
[0025] 8. The GCC according to item 7, wherein the increase in carboxylation activity is an increase in the conversion of hydroxyacetyl-CoA to tartrate-CoA.
[0026] 9. The GCC according to any one of claims 1 to 8, wherein the GCC catalyzes the carboxylation of hydroxyacetyl-CoA (conversion of hydroxyacetyl-CoA to tartrate-CoA) at a rate greater than about 5.6 ± 0.3 s. -1 .
[0027] 10. The GCC according to any one of claims 1 to 9, wherein the specific activity of said GCC is higher than about 937 ± 40 nmol hydroxyacetyl-CoA min. -1 mg -1 .
[0028] 11. The GCC according to claim 10, wherein the specific activity of said GCC is about 2602 ± 430 nmol hydroxyacetyl-CoAmin. -1 mg -1 .
[0029] 12. A nucleic acid molecule encoding a GCC according to any one of items 1 to 11.
[0030] 13. A carrier comprising the nucleic acid molecule according to claim 12.
[0031] 14. An organelle comprising a nucleic acid molecule according to claim 12 or a carrier according to claim 13.
[0032] 15. A host cell comprising a nucleic acid molecule according to claim 12, a vector according to claim 13, or an organelle according to claim 14.
[0033] 16. A tissue comprising the host cells according to claim 15.
[0034] 17. An organism comprising a nucleic acid molecule according to claim 12, a carrier according to claim 13, or an organelle according to claim 14.
[0035] 18. The organism according to claim 17, wherein the organism is a plant, algae or microorganism.
[0036] 19. The organism according to claim 18, wherein the microorganism is bacteria.
[0037] 20. An organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19, wherein the organelle, the host cell, the tissue, or the organism converts hydroxyacetyl-CoA to tartrate-CoA at a higher rate than the corresponding organelle, host cell, tissue, or organism comprising a reference GCC, preferably, wherein the reference GCC comprises SEQ ID NO: 1.
[0038] 21. An organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19, wherein the growth rate and / or carbon production rate of the organelle, the host cell, the tissue, or the organism is higher than that of the corresponding organelle, host cell, tissue, or organism comprising a reference GCC, preferably, wherein the reference GCC comprises SEQ ID NO: 1.
[0039] 22. Use of the GCC according to any one of items 1 to 11, the organelle according to item 14, the host cell according to item 15, the tissue according to item 16, or the organism according to any one of items 17 to 19 for the process of CO2 fixation and / or photosynthesis.
[0040] 23. Use of the GCC according to any one of items 1 to 11, the organelle according to item 14, the host cell according to item 15, the tissue according to item 16, or the organism according to any one of items 17 to 19 for the decomposition of synthetic materials.
[0041] 24. The use according to item 23, wherein the synthetic material is polyester.
[0042] 25. The use according to item 24, wherein the polyester is polyethylene terephthalate.
[0043] 26. Use of the GCC according to any one of items 1 to 11, the organelle according to item 14, the host cell according to item 15, the tissue according to item 16, or the organism according to any one of items 17 to 19 for the decomposition of ethylene glycol and / or glycolic acid.
[0044] 27. A method for producing biomass using a GCC according to any one of claims 1 to 11, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
[0045] 28. A method for decomposing synthetic materials using a GCC according to any one of claims 1 to 11, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
[0046] 29. The method according to item 28, wherein the synthetic material is polyester.
[0047] 30. The method according to claim 29, wherein the polyester is polyethylene terephthalate.
[0048] 31. A method for decomposing ethylene glycol and / or glycolic acid using a GCC according to any one of claims 1 to 11, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
[0049] 32. A composition comprising a GCC according to any one of claims 1 to 11, a nucleic acid according to claim 12, a vector according to claim 13, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
[0050] The conversion of C2 compounds glycolic acid and glyoxylic acid into C3 metabolites plays a central role in many carbon metabolism processes. However, very few natural metabolic pathways directly convert these C2 intermediates into C3 metabolites, and all of them result in carbon loss. The glyoxylic acid cycle (Kornberg, Nature 179, 988–991, 1957) and the recently described β-hydroxyaspartate cycle (Borzyskowski, Nature 575, 500–504, 2019) ultimately produce C4 compounds, which require decarboxylation to generate C3 metabolites. Similarly, photorespiration and the glyceric acid pathway convert two glyoxylic acid molecules into glyceric acid by releasing CO2 (Krakow, J. Bacteriol. 81, 509–518, 1961; Bauwe, TrendsPlant Sci. 15, 330–336, 2010). The unavoidable CO2 loss in all these pathways severely limits their carbon efficiency, especially in photorespiration. It is estimated that crop yields can be reduced by up to 50% due to carbon loss from photorespiration in hot, dry climates (Walker, Annu. Rev. Plant Biol. 67, 107–129, 2016). Therefore, avoiding the energy and carbon losses during glycolate assimilation via synthetic pathways holds promise for significantly improving productivity.
[0051] The tartrate-CoA pathway has recently been proposed as a direct route for glycolate assimilation into central carbon metabolism (Trudeau, Proc. Natl Acad. Sci. USA 115, E11455–E11464, 2018). This previously hypothesized pathway aims to fix CO2 rather than release it, and its performance is expected to be superior to all naturally evolved glycolate assimilation pathways. However, this pathway remained theoretical until the inventors developed the key enzyme of the tartrate-CoA pathway, hydroxyacetyl-CoA carboxylase (GCC), in previous work. To develop the GCC, the inventors previously introduced five mutations to convert *Tricholoma matsutake* (…). Methylorubrum extorquens The naturally occurring propionyl-CoA carboxylase (PCC) is converted into GCC M5, which enables it to carboxylate hydroxyacetyl-CoA.
[0052] The PCC of *Tricholoma mater* is a two-subunit protein comprising an α subunit (protein: SEQ ID NO: 13; DNA: SEQ ID NO: 26) and a β subunit (protein: SEQ ID NO: 12; DNA: SEQ ID NO: 25). It will be apparent to those skilled in the art that the “actual wild-type” variant / version of the GCC variant described herein is the PCC naturally present in *Tricholoma mater*. Therefore, the “actual wild-type” sequence should be the sequence of the *Tricholoma mater* PCC (α subunit (protein: SEQ ID NO: 13; DNA: SEQ ID NO: 26) and β subunit (protein: SEQ ID NO: 12; DNA: SEQ ID NO: 25)).
[0053] However, it will be further apparent to those skilled in the art that, depending on the specific circumstances, GCC M5 may also be considered as the wild-type or reference sequence of the variant described herein. For the GCC variant of the present invention described herein, GCC M5 may be considered as the wild-type or reference version, as the inventors have optimized the GCC starting with GCC M5. The amino acid substitutions described herein are located in the β subunit. Therefore, this document primarily refers to the sequence of the β subunit of PCC or GCC, or the sequence encoding the β subunit. The amino acid sequence of the β subunit of GCC M5 is SEQ ID NO: 1 (the corresponding nucleotide sequence is shown as SEQ ID NO: 14). All amino acid changes in GCC M5 relative to PCC (“actual wild-type”) are located in the β subunit. Therefore, the GCCM5 α subunit is the α subunit of PCC (“actual wild-type”; protein: SEQ ID NO: 13; DNA: SEQ ID NO: 26). All amino acid changes in the GCC of the present invention described herein relative to GCC M5 (and relative to PCC (“actual wild-type”)) are also present in the β subunit. Therefore, it is envisioned that the GCC of the present invention described herein (the GCC variant of the present invention) has / contains the α subunit of PCC (“actual wild type”; protein: SEQ ID NO: 13; DNA: SEQ ID NO: 26). Thus, all disclosures herein that define only the β subunit can be associated with the α subunit, preferably, the α subunit being composed of or containing SEQ ID NO: 13. In the context of the present invention, those skilled in the art will understand that the GCC provided herein (i.e., a protein complex having / containing GCC enzymatic activity) comprises the GCC β subunit of the present invention provided herein (e.g., containing G20R and / or S100N amino acid substitutions) and a suitable α subunit (e.g., the α subunit of PCC).
[0054] As described above, those skilled in the art know that carboxylases (i.e., enzymes that catalyze the carboxylation of one or more substrates) comprise two subunits (i.e., an α subunit and a β subunit), each with its own active site. Specifically, in biotin-dependent carboxylases, the α subunit has biotin carboxylase activity (i.e., cleaving ATP and carboxylating biotin), while the β subunit has carboxyltransferase activity (i.e., catalyzing the transcarboxylation reaction from biotin to a receptor molecule, such as hydroxyacetyl-CoA). Although the substrate specificity of a carboxylase is primarily determined by the β subunit, the α subunit may be highly conserved among different carboxylases (with different substrate specificities). For example, previous studies have shown that the α and β subunits of different carboxylases from the same organism or even different organisms can form functional (chimeric) complexes (Huang et al., 2010, Nature, 466(7309):1001-1005, and Lombard and Moreira, 2011, BMC Evol Biol 11, 232, respectively). Therefore, in the context of this invention, it is contemplated that the GCC provided herein comprises the β subunit of this invention (such as a β subunit comprising G20R and / or S100N substitution) and also comprises an α subunit. Its properties are not particularly limited as long as the α subunit has biotinylate carboxylase activity and can form a catalytically active complex with the β subunit provided herein. Those skilled in the art also know that such α subunits may themselves comprise two subunits that together form an active enzyme (complex) with biotinylate carboxylase activity (e.g., see Cronan and Waldrop, 2002, Progress in Lipid Research, 41, 407-435). Exemplary biotinylate carboxylases comprising two subunits (collectively referred to herein as the α subunit) include acetyl-CoA carboxylase from *Escherichia coli* and acetyl-CoA carboxylases from most other bacteria, wherein the α subunit is divided into two subunits (i.e., the biotinylate carboxylase domain and the biotinylate carboxyl carrier protein domain are divided into two subunits; family 1.1 in Tong, 2013, Cell.Mol. Life Sci. 70, 863–891; see also Cronan, 2021, Microbiol Mol Biol Rev 85:10; see also Cronan, 2021, Microbiol Mol Biol Rev 85:10). Therefore, as used herein, the term "α subunit" generally refers to an enzyme (or enzyme complex) having biotinylate carboxylase activity in the context of this invention. Exemplary biotinylate carboxylases containing one subunit (i.e., the α subunit) are: propionyl-CoA carboxylase from Homo sapiens and most other propionyl-CoA carboxylases (family 1.5 in Tong, 2013, ibid.; see also Huang, 2010, ibid.).Therefore, in the context of this invention, the α subunit (i.e., biotinylate carboxylase) may optionally be selected from the biotinylate carboxylases listed above. However, preferably, the α subunit that forms a complex with the β subunit provided herein is the α subunit of PCC (or a variant thereof with enzymatic activity).
[0055] Therefore, the present invention relates to a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0056] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0057] At a position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions, there is one or more amino acid substitutions, deletions, or insertions, and
[0058] It (further) comprises an α subunit, preferably wherein the α subunit is a biotinylate carboxylase or has biotinylate carboxylase activity, more preferably wherein the α subunit comprises / has at least 60% sequence identity with SEQ ID NO: 13, more preferably wherein the α subunit comprises / has at least 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with SEQ ID NO: 13, and even more preferably wherein the α subunit comprises / has about 100% sequence identity with SEQ ID NO: 13, and
[0059] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC comprises the amino acid sequence of SEQ ID NO:1 and optionally the amino acid sequence of SEQ ID NO:13.
[0060] Therefore, the present invention relates to a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0061] It contains a β subunit, wherein the β subunit:
[0062] Contains / has an amino acid sequence that has at least 60% sequence identity with SEQ ID NO:1, and
[0063] At a position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions, there is one or more amino acid substitutions, deletions, or insertions, and
[0064] The GCC (further) comprises an α subunit, preferably wherein the α subunit is a biotinylate carboxylase or has biotinylate carboxylase activity, more preferably wherein the α subunit comprises / has at least 60% sequence identity with SEQ ID NO: 13, more preferably wherein the α subunit comprises / has at least 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 13, and even more preferably wherein the α subunit comprises / has about 100% sequence identity with SEQ ID NO: 13.
[0065] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC comprises the amino acid sequence of SEQ ID NO:1 and optionally the amino acid sequence of SEQ ID NO:13.
[0066] In this context, the β subunit and α subunit can be defined as described above or anywhere below.
[0067] In the context of this invention, and as described in the appended embodiments and Figure 6 As exemplarily demonstrated, the GCC provided herein may comprise 6 β subunits and 6 α subunits (i.e., forming a β6α6 complex). However, in the context of this application, complexes comprising fewer than 6 α subunits are also contemplated. Therefore, the GCC provided herein may comprise 6 β subunits and 1 to 6 α subunits, preferably 6 α subunits. These are as follows... Figure 6 As shown.
[0068] Therefore, the present invention relates to a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0069] It contains six β subunits, wherein each of the β subunits:
[0070] Contains / has an amino acid sequence that has at least 60% sequence identity with SEQ ID NO:1, and
[0071] At a position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions, there is one or more amino acid substitutions, deletions, or insertions, and
[0072] The GCC (further) comprises 1 to 6 α subunits, preferably wherein each α subunit is a biotinylate carboxylase or has biotinylate carboxylase activity, more preferably wherein each α subunit contains / has at least 60% sequence identity with SEQ ID NO: 13, more preferably wherein each α subunit contains / has at least 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 13, and even more preferably wherein each α subunit contains / has about 100% sequence identity with SEQ ID NO: 13, and
[0073] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC comprises the amino acid sequence of SEQ ID NO:1 and optionally the amino acid sequence of SEQ ID NO:13.
[0074] In this context, the β subunit and α subunit can be defined as described above or anywhere below.
[0075] This paper further envisions that the GCC provided herein contains one or more β subunits and α subunits, each having / containing different amino acid sequences; however, it is preferred that the GCC provided herein contains only one type of β subunits and α subunits (i.e., preferably all β subunits contain the same amino acid sequence and all α subunits contain the same amino acid sequence).
[0076] Therefore, amino acid substitutions can be represented relative to the “actual wild-type” sequence (which could be the β subunit of PCC (protein: SEQ ID NO: 12; DNA: SEQ ID NO: 25)). However, amino acid substitutions can also be represented relative to the GCC M5 sequence (i.e., relative to SEQ ID NO: 1). Figure 3The sequence alignments of PCC, GCC M5, and the two GCCs of this invention are shown. To construct GCC M5, the inventors introduced five mutations (L100S, Y143H, D407I, I450V, and W502R) into PCC. Therefore, GCC M5 has a serine residue at position 100. However, when referring herein to an amino acid substitution at position 100 of GCC M5 (SEQ ID NO: 1), this substitution can still be represented relative to the "actual wild-type" sequence. For example, when referring herein to an amino acid substitution at position 100 of GCC M5, such as arginine, the substitution can be represented as L100N. In this case, the substitution is represented relative to the "actual wild-type" sequence. However, it is obvious to those skilled in the art that an amino acid substitution relative to GCC M5 (and therefore relative to SEQ ID NO: 1) is S100N. In other words, it is obvious to those skilled in the art that when the amino acid at position 100 of GCC M5 is replaced by arginine, serine is replaced by arginine.
[0077] Since GCC M5 is an enzyme in the prior art, SEQ ID NO: 1 is used as a reference sequence to describe variants of the present invention. Accordingly, the present invention provides positions in GCC M5 (SEQ ID NO: 1) where amino acid substitutions improve its activity. In other words, the inventors have identified certain amino acid positions where substitutions enhance the activity of GCC. In other words, the inventors have identified certain amino acid residues in GCC M5 that can be substituted with other amino acid residues, thereby altering the activity of GCCM5 (as described elsewhere herein, changes in energy efficiency and / or carboxylation activity are observed). In this document, the terms “amino acid” and “amino acid residue” are used interchangeably. The terms “substituted,” “replaced,” and “mutated” are used interchangeably. Accordingly, the terms “substitution,” “replacement,” and “mutation” are also used interchangeably. An example of substitution could be the substitution of glycine at position 20 or the position corresponding to that position in SEQ ID NO: 1 with arginine. Another example of substitution could be the substitution of serine at position 100 or the position corresponding to that position in SEQ ID NO: 1 with asparagine. Those skilled in the art are well aware of how to use molecular biology techniques, such as site-directed mutagenesis as described herein, to substitute amino acids. Furthermore, the insertion and / or deletion of amino acids at the positions described herein are also envisioned.
[0078] In the context of this invention, "missing" or "deleted" means the deletion or removal of an amino acid at the corresponding position.
[0079] In the context of this invention, "inserted" or "inserted" means the insertion of one or more, preferably one, amino acid residue at the corresponding position.
[0080] It should be noted that the GCC of the present invention can also be referred to as a GCC variant of the present invention. Therefore, "GCC" and "GCC variant" can be used interchangeably. In addition, "GCC mutation variant" or "GCC version" can also be used.
[0081] Those skilled in the art are well aware of how to introduce substitutions, insertions, or deletions. The site-directed mutagenesis described below can be used to generate the GCC described herein. Furthermore, a randomized mutagenesis library can be generated to produce a dataset for input into artificial intelligence (AI) algorithms. The randomized mutagenesis library can be generated as described below.
[0082] A plasmid library for random mutagenesis of GCC M5 can be created using macroprimer-based full plasmid PCR (MEGAWHOP) (Miyazaki, Methods Enzymol. 498, 399-406, 2011). To generate, for example, a random fragment of the β subunit of GCC M5 (pTE3101), error-prone PCR can be performed in a 50 µL reaction system using 2.5 U Taq polymerase and magnesium-free buffer (New England Biolabs; M0320), 7 mM MgCl2, 0.4 mM each of dGTP and dATP, 2 mM each of dCTP and dTTP, 0.4 µM each of primers PccB_fw_P1 and PccB_rv_P1, 10% (v / v) dimethyl sulfoxide, 50 ng pTE3101 template DNA, and 200–500 µM MnCl2. Random fragments can be digested with DpnI (NEB, R0176), purified by agarose gel electrophoresis, and used as macroprimers for whole plasmid PCR (MEGAWHOP), as described in other literature (Miyazaki, see above); or a second error-prone PCR reaction can be performed to further increase the mutation rate. The MEGAWHOP reaction system (50 µL) may contain 1× KOD Hot Start reaction buffer (Novagen), 0.2 mM dNTPs, 1.5 mM MgSO4, 500 ng macroprimers, 50 ng template plasmid (GCC M5; pTE3101), and 2.5 U KOD Hot Start DNA polymerase (Novagen). MEGAWHOP products can be purified, digested with DpnI, and transformed into ElectroMAX DH5α (ThermoFisher Scientific) to ensure a high number of transformants in the resulting library. To estimate the mutation rate of different concentrations of MnCl2 used in error-prone PCR, plasmids from 10 randomly selected clones after MEGAWHOP purification were sequenced and analyzed for nucleotide exchange.
[0083] To develop (other) GCC variants with improved kinetic properties, artificial intelligence (AI) algorithms can be used. The training dataset for the AI model can be obtained by generating a random mutagenesis library of pTE3101 (GCC M5) and transforming it into chemically competent *E. coli* BL21- biraColonies were then picked and inoculated into 96-well deep-well plates (PlateOne) containing lysate (Miller) at a concentration of 100 µg / mL ampicillin and 50 µg / mL spectinomycin. Expression, lysis, and screening of GCC samples were performed as previously described (Scheffen, see above). Up to 2100 mutant variants were screened, and a representative subset of candidates was selected for sequencing. A machine learning model for Exazyme was trained using 161 samples to predict beneficial mutations in GCC M5. The resulting list of all possible single mutations, sorted by their efficiency, was used as a template to identify suitable candidates for biochemical characterization through homology modeling and structural studies of promising mutations.
[0084] To evaluate the mutations predicted by the artificial intelligence algorithm, homology modeling can be performed on each promising mutant variant using SWISS-MODEL. As a template for GCC mutation homology modeling, the structure of engineered GCC M5 (PDB ID 6YBQ) from *Tricholoma fusca* can be used. Structural analysis of the model can be performed using PyMOL (PyMOL Molecular Graphics System; version 1.8; Schrödinger). Modeling hydroxyacetyl-CoA to the active site of GCC can be based on the position of CoA in the GCC M5 structure and the structure of *Propionibacterium fischeri* (…). Propionibacterium freudenreichii The position of methylmalonyl-CoA in the structure of methylmalonyl-CoA carboxyltransferase (PDB ID 1ON3; 52% amino acid identity) can be determined. CoA thioesters can be manually fitted and tuned using COOT and PyMOL to reflect differences in the active site architecture.
[0085] Site-directed mutagenesis can be used to construct GCC mutant variants previously predicted by artificial intelligence algorithms and selected through homology modeling and structural analysis. Introducing new mutations can be achieved through single-site mutagenesis oligonucleotide PCR, as described in other literature (ShenoyAnal. Biochem. 319, 335-336, 2003). These techniques can also be used for the GCC described in this paper. A 25 µL reaction mixture containing 0.5 µM primers, 3% (v / v) dimethyl sulfoxide, 50 ng template DNA (pTE3101), and Phusion high-fidelity PCR premix (NEB, M0531) can be used for PCR. Subsequently, 20 U DpnI (NEB, R0176) is added to the reaction mixture, and the mixture is incubated at 37 °C for 2 hours for digestion. 5 µL of the digest can be transformed into chemically competent *E. coli* NEB Turbo cells and streaked onto Miller agar plates containing 50 µg / mL streptomycin. Three to six colonies can be selected and cultured in 10 mL of lysate broth (Miller) containing 50 µg / mL streptomycin at 37 °C and 180 rpm for 12 h. The plasmid can then be isolated and sequenced to verify the mutagenesis effect. Primers / oligonucleotides from Table 4 can be used for mutagenesis and sequencing. Strains from Table 2 can be used as strains for mutagenesis and cloning. *E. coli* ElectroMAX DH5α can be used to construct a random mutagenesis library, which may be needed to generate a dataset of randomly mutated GCC variants to train an artificial intelligence model for predicting beneficial mutations. *E. coli* NEB Turbo can be used to construct and maintain plasmids with site-specific mutations in genes / nucleic acids encoding GCC. *E. coli* BL21- beer It can be obtained from *E. coli* BL21DE3 by introducing a vector carrying the biotin ligase gene from *Tricholoma mater* required for GCC activation. *E. coli* BL21- beer It can be used for protein overexpression of GCC variants.
[0086] Those skilled in the art will know that some carboxylases are biotin-dependent. For example, PCC belongs to the family of biotin-dependent carboxylases. Biotin can covalently bind to the lysine residue in the α subunit of PCC and can act as a cofactor in the carboxylation reaction. Therefore, those skilled in the art will know that the PCC of *Tricholoma mater* is also biotin-dependent (Scheffen, see above). Therefore, those skilled in the art will know that GCC M5 is also biotin-dependent. It is therefore contemplated that the GCC of the present invention is also biotin-dependent. Therefore, it will be apparent to those skilled in the art that the GCC described herein can be activated by contact with a biotin ligase at some point in time. In other words, it is contemplated that the GCC described herein can be contacted with a biotin ligase at some point in time, thereby enabling it to catalyze the carboxylation reaction. For example, the biotin ligase can be co-expressed / co-produced with the GCC described herein. Preferably, the biotin ligase of *Tricholoma mater* is co-expressed / co-produced with the GCC described herein. It is also contemplated that the biotin ligase gene be introduced into a host used to produce the GCC described herein. Preferably, the biotin ligase gene from *Thiazoma truncatum* is introduced into the host used to produce the GCC described herein.
[0087] Furthermore, it is apparent that when the host cell, organelle, tissue, and / or organism described herein contains the GCC described herein, it may also contain a suitable biotin ligase. A suitable biotin ligase may be an endogenous biotin ligase or a recombinantly introduced biotin ligase. Those skilled in the art will know that the biotinylation of GCC can be determined, for example, by an avidin gel migration assay (Scheffen, see above). Therefore, those skilled in the art can test whether a given biotin ligase is capable of biotinylating GCC, and thus determine its suitability. Furthermore, those skilled in the art can test whether the GCC is fully biotinylated or only partially biotinylated. In a preferred aspect, the host cell, organelle, tissue, and / or organism described herein may contain a biotin ligase gene from *Tricholoma mater*. In other words, in a preferred aspect, the host cell, organelle, tissue, and / or organism described herein may produce a biotin ligase (in addition to the GCC described herein) from *Tricholoma mater*. Biotin ligases from *Thizotomyces truncatum* have an amino acid sequence as shown in SEQ ID NO: 27 and / or a nucleic acid sequence as shown in SEQ ID NO: 28.
[0088] The chemical reagents required for the techniques described herein are available from Sigma-Aldrich, Carl Roth GmbH + Co. KG, Santa Cruz Biotechnology Inc., and Merck. Biochemical substances and materials for cloning and protein expression are available from Thermo Fisher Scientific, New England Biolabs GmbH, and Macherey-Nagel GmbH. Coenzyme A is available from Roche Diagnostics. Materials and equipment for protein purification are available from GE Healthcare, BioRad, and Merck Millipore GmbH. Pyruvate kinase / lactate dehydrogenase, malate dehydrogenase, glucose-6-phosphate dehydrogenase, glucose dehydrogenase, and phosphoenolpyruvate carboxylase are available from Sigma-Aldrich.
[0089] This invention provides a GCC, wherein the GCC is characterized by:
[0090] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0091] It has one or more amino acid substitutions, deletions or insertions at a position selected from the group consisting of the 20th and 100th positions in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0092] It is envisioned that the GCC described herein has improved activity compared to a reference GCC that does not contain the amino acid substitutions, deletions, or insertions described herein.
[0093] This invention provides a GCC, wherein the GCC is characterized by:
[0094] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0095] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0096] The GCC has improved activity compared to the reference GCC.
[0097] It is envisioned that the GCC described herein (e.g., the GCC having the amino acid substitutions of the present invention) has improved activity compared to the reference GCC containing SEQ ID NO: 1.
[0098] This invention provides a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0099] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0100] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0101] The GCC exhibits improved activity compared to the reference GCC containing the amino acid sequence of SEQ ID NO: 1.
[0102] As described above, this document envisions that GCC comprises two subunits (e.g., the GCC variants provided herein comprise an α subunit and a β subunit). Therefore, it will be apparent to those skilled in the art that the reference GCC also comprises two subunits (an α subunit and a β subunit). Preferably, the β subunit of the reference GCC consists of or comprises SEQ ID NO: 1. Preferably, the α subunit of the reference GCC consists of or comprises SEQ ID NO: 13. Therefore, it is preferred to compare the activity of the GCC of the present invention with the activity of the reference GCC, wherein the reference GCC comprises an α subunit consisting of or comprising SEQ ID NO: 13 and a β subunit consisting of or comprising SEQ ID NO: 1. As described above, the prior art enzyme GCC M5 consists of an α subunit consisting of SEQ ID NO: 13 and a β subunit consisting of SEQ ID NO: 1. Therefore, it is preferred to compare the activity of the GCC of the present invention with the activity of the reference GCC, wherein the reference GCC is GCC M5.
[0103] In the context of this invention, improved activity of GCC or improved enzymatic activity of GCC may refer to improved energy efficiency and / or improved carboxylation activity. Preferably, the GCC variant with improved energy efficiency is a variant having an amino acid substitution, deletion, or insertion at position 100 as defined herein.
[0104] In this document, improved energy efficiency can mean that less energy input is required to carry out a particular reaction. In other words, improved energy efficiency can mean that the GCC requires less energy equivalent to carry out a particular reaction. In other words, improved energy efficiency can mean that, relative to a reference GCC (preferably containing SEQ ID NO: 1), the GCC of the present invention requires less energy input / energy equivalent (i.e., ATP) to convert hydroxyacetyl-CoA to tartratel-CoA. The term "tartratel-CoA" as used herein specifically refers to "(S)-tartratel-CoA".
[0105] Therefore, it is assumed that the improved activity of the GCC described herein is an improved energy efficiency. In other words, the present invention provides a GCC with improved activity, wherein the improved activity is an improved energy efficiency.
[0106] In the context of this invention, "improved energy efficiency" may refer to a reduction in ATP consumption. ATP is a well-known energy equivalent for biological systems. Therefore, as used herein, improved energy efficiency may refer to a GCC that, relative to a reference GCC (preferably comprising SEQ ID NO: 1), requires less ATP to convert hydroxyacetyl-CoA to tartrate-CoA. Thus, the improved energy efficiency of the GCC described herein is presumably a reduction in ATP consumption. In other words, this invention provides a GCC with improved energy efficiency, wherein the improved energy efficiency is a reduction in ATP consumption.
[0107] It is envisioned that ATP consumption is reduced by at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%, preferably by at least 50%. Therefore, the ATP consumption of the GCC provided by the present invention is reduced by at least 50% relative to a reference GCC (preferably containing the reference GCC of SEQ ID NO: 1).
[0108] By adding CO2, GCC converts hydroxyacetyl-CoA to tartrate-CoA by consuming ATP (hydroxyacetyl-CoA + HCO3-). -+ ATP → Tartratel-CoA + ADP + Pi). As mentioned above, the ATP / carboxylation ratio of GCC M5 is approximately 4, although the theoretical ratio is 1. That is, the theoretical ratio of ATP molecules consumed per CO2 molecule added is 1. However, the measured ATP consumption per CO2 molecule added is approximately 4 ATP molecules. As demonstrated in the accompanying examples, the ATP / CO2 ratio of the GCC can be determined. Those skilled in the art can easily determine the ATP / CO2 ratio of any GCC variant by first assessing the ATP consumption and carboxylation rate of a given GCC variant and then calculating the ATP / CO2 ratio. An ATP / CO2 ratio of 4 means that four ATP molecules are required to fix one CO2 molecule (i.e., to convert one hydroxyacetyl-CoA molecule to a tartratel-CoA molecule).
[0109] Therefore, it is envisioned that the GCC described herein has / shown an ATP / CO2 ratio below about 4.00 ± 0.02. Thus, the present invention provides a GCC wherein the ATP / CO2 ratio of said GCC is below about 4.00 ± 0.02. In other words, the improved activity and / or improved energy efficiency of the GCC described herein can be seen from the fact that the ATP / CO2 ratio of said GCC is below about 4.00 ± 0.02. It is evident from the appended examples that the ATP / CO2 ratio of GCC M5 is 4.00 ± 0.02. Therefore, when the ATP / CO2 ratio of the GCC of the present invention is below about 4.00 ± 0.02, it has improved energy efficiency relative to a reference GCC (e.g., GCC M5). It is also envisioned that the ATP / CO2 ratio of the GCC described herein is below about 4, below about 3, below about 2, or below about 1, or any value between the two, such as 2.5.
[0110] Therefore, it is envisioned that the ATP / CO2 ratio of the GCC described herein is about 1.0, about 1.1, about 1.2, about 1.3, about 1.4, about 1.5, about 1.6, about 1.7, about 1.8, about 1.9, about 2.0, about 2.1, about 2.2, about 2.3, about 2.4, about 2.5, about 2.6, about 2.7, about 2.8, about 2.9, about 3.0, about 3.1, about 3.2, about 3.3, about 3.4, about 3.5, about 3.6, about 3.7, about 3.8, or about 3.9, or any value between these values, such as 2.23. Preferably, the ATP / CO2 ratio of the GCC is about 1.71 ± 0.11. Therefore, the present invention provides a GCC wherein the ATP / CO2 ratio of the GCC is about 1.71 ± 0.11. The GCC comprising SEQ ID NO:7 of the present invention can have an ATP / CO2 ratio of about 1.71 ± 0.11. In a preferred aspect, the present invention relates to a GCC comprising SEQ ID NO: 7, wherein the ATP / CO2 ratio of the GCC is about 1.71 ± 0.11.
[0111] The explanations in this article regarding “improved energy efficiency,” “reduced ATP consumption,” “ATP / CO2 ratio below about 4.00 ± 0.02,” and “ATP / CO2 ratio of about 1.71 ± 0.11” are particularly applicable to hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by containing an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 1, and having (only having) one amino acid substitution, deletion, or insertion (preferably substitution) at position 100 or the position corresponding to that position (preferably asparagine at position 100) in the amino acid sequence shown in SEQ ID NO: 1. In this context, “only” means that the GCC has no amino acid substitution, deletion, or insertion (preferably substitution) at position 20 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1. In this context, it is preferred that GCC has only one amino acid substitution, deletion, or insertion (preferably substitution) at position 100 of the amino acid sequence shown in SEQ ID NO: 1 or at any of these positions, i.e., no amino acid substitution, deletion, or insertion (preferably substitution) at position 20 as defined above.
[0112] Therefore, in one aspect, the present invention relates to a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0113] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0114] It has (only has) one amino acid substitution, deletion or insertion at position 100 of the amino acid sequence shown in SEQ ID NO: 1 or at a position corresponding to any of these positions, preferably wherein the GCC has improved activity relative to the reference GCC, preferably wherein the reference GCC comprises the amino acid sequence of SEQ ID NO: 1, and preferably wherein the GCC has improved energy efficiency as defined herein.
[0115] The catalytic rates of "increased carboxylation activity," "increased conversion of hydroxyacetyl-CoA to tartratel-CoA," and "hydroxyacetyl-CoA carboxylation (conversion of hydroxyacetyl-CoA to tartratel-CoA)" in this article are higher than approximately 5.6 ± 0.3 s. -1 "The specific activity of GCC is higher than approximately 937 ± 40 nmol of hydroxyacetyl-CoA min." -1 mg -1 "The specific activity of GCC is approximately 2602 ± 430 nmol hydroxyacetyl-CoA min." -1 mg -1 The interpretation of "etc." applies particularly to hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by containing an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 1, and having (only having) one amino acid substitution, deletion, or insertion (preferably substitution) at position 20 or the position corresponding to that position (preferably arginine at position 20) in the amino acid sequence shown in SEQ ID NO: 1. In this context, "only" means that the GCC has no amino acid substitution, deletion, or insertion (preferably substitution) at position 100 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1. In this context, it is preferred that the GCC has only one amino acid substitution, deletion, or insertion (preferably substitution) at position 20 or the position corresponding to any of these positions in the amino acid sequence shown in SEQ ID NO: 1, i.e., no amino acid substitution, deletion, or insertion (preferably substitution) at position 100 as defined above.
[0116] Therefore, in one aspect, the present invention relates to a hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by:
[0117] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0118] It has (only has) one amino acid substitution, deletion or insertion at position 20 of the amino acid sequence shown in SEQ ID NO: 1 or at a position corresponding to any of these positions, preferably wherein the GCC has improved activity relative to the reference GCC, preferably wherein the reference GCC comprises the amino acid sequence of SEQ ID NO: 1, preferably wherein the improved activity is an increase in carboxylation activity, as defined herein.
[0119] In one respect, the GCC of this article has no amino acid substitutions, deletions, or insertions (preferably substitutions) at positions 20 and 100 as defined herein, for example, no amino acid substitutions, deletions, or insertions (preferably substitutions) at either position 20 of the amino acid sequence shown in SEQ ID NO: 1 or at position 100 of the asparagine, or at positions corresponding to either of these positions.
[0120] Those skilled in the art are well aware of how to measure ATP consumption and CO2 fixation (Scheffen, see above). In particular, ATP hydrolysis can be measured spectrophotometrically. The ATP / CO2 ratio is also known as the "ATP consumption per carboxylation". The determination of the ATP consumption per carboxylation ratio can be achieved spectrophotometrically using purified enzymes.
[0121] For overexpression of GCC M5 and its mutant variants, the corresponding plasmids can be transformed into chemically competent Escherichia coli BL21- beer Cells can be grown overnight at 25 °C on Miller lysate agar plates containing 100 µg / mL ampicillin and 50 µg / mL spectinomycin. 8 μL of Golden Miller lysate containing 5 g / L yeast extract, 10 g / L tryptone, 10 g / L NaCl, 17 mM KH₂PO₄, 72 mM K₂HPO₄, and 0.4% glycerol can be inoculated from the agar plate and incubated at 37 °C and 140 rpm. When OD... 600When the pH is 0.4–0.6, protein expression can be induced with 500 µM IPTG, and cells can be incubated overnight at 25 °C. Cells can be collected by centrifugation at 8,000 g at 4 °C for 12 min, followed by lysis by French pressure filtration and His-Trap purification using an ÄktaStart (GE Healthcare) column connected to a HisTrap FF column. The purification buffer contains 50 mM HEPES pH 7.8 and 500 mM KCl, and elution can be performed using 500 mM imidazole. Protein desalting can be performed by gel filtration chromatography using a HiLoad 16 / 600 Superdex 200 pg column (GE Healthcare) and a buffer containing 50 mM HEPES pH 7.8 and 150 mM KCl. Protein quantification can be performed by measuring absorbance at 280 nm. Protein purity can be verified by SDS-PAGE, using 15 µg of purified protein on a 4-20% Mini-Protean TGX pre-fabricated protein gel (Biorad).
[0122] Hydroxyacetyl-CoA can be synthesized and purified as previously described (Trudeau, see above; Scheffen, see above). The concentration of coenzyme A ester can be quantified by measuring the absorbance at 260 nm (ε = 16.4 mM). -1 cm -1 ).
[0123] To measure the ATP consumption to carboxylation ratio (ATP / CO2 ratio) of GCC, a CaMCR-coupled enzyme assay can be performed under ATP-limited conditions. The assay consisted of 100 mM MOPS at pH 7.8, 50 mM KHCO3, 0.15 mM ATP, 0.5 mM NADPH, 5 mM MgCl2, and 1.8 mg / mL of *Flexibrio orangeense*. Chloroflexus aurantiacus CaMCR and 0.05–3 mg / mL GCC can be mixed in a cuvette and incubated at 37 °C for 2 minutes. The reaction can be initiated with 0.5 mM hydroxyacetyl-CoA, and absorbance can be measured over time at λ = 340 nm. The ATP consumption ratio for each carboxylation is calculated as the ratio between the amount of ATP in the reaction mixture and the amount of NADPH consumed as reflected by the decrease in absorbance during the reaction.
[0124] Hydroxyacetyl-CoA + HCO3 - + ATP → Tartrate acyl-CoA + ADP + Pi (GCC)
[0125] Tartrate-CoA + 2 NADPH → Glyceric acid + CoA + 2 NADP + (CaMCR)
[0126] As described above, in the context of this invention, improved activity or improved enzyme activity of GCC can refer to improved carboxylation activity. Preferably, a GCC variant having improved carboxylation activity is a variant having an amino acid substitution, deletion, or insertion at position 20 as defined herein. Carboxylation is a chemical reaction that can include incorporating carbon dioxide into an organic compound to generate a carboxylic acid. As used herein, "improved carboxylation activity" can refer to an increased conversion of hydroxyacetyl-CoA to tartratel-CoA. Therefore, it is contemplated that the GCC described herein has / exhibits an increased carboxylation activity. It is also contemplated that the GCC described herein has / exhibits an increased conversion of hydroxyacetyl-CoA to tartratel-CoA. Therefore, this invention provides a GCC having improved activity, wherein the improved activity is an increased carboxylation activity. This invention provides a GCC having / exhibiting improved activity, wherein the improved activity is an increased carboxylation activity.
[0127] The present invention also provides GCCs with increased carboxylation activity, wherein the increase in carboxylation activity is due to an increased conversion of hydroxyacetyl-CoA to tartrate-CoA. The present invention also provides GCCs having / exhibiting increased carboxylation activity, wherein the increase in carboxylation activity is due to an increased conversion of hydroxyacetyl-CoA to tartrate-CoA.
[0128] The carboxylation activity is envisioned to increase by at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 times. Preferably, the carboxylation activity increases by at least 2 or 3 times.
[0129] Carboxylation activity can be measured by the catalytic rate of hydroxyacetyl-CoA carboxylation (the conversion of hydroxyacetyl-CoA to tartrate-CoA) (Scheffen, see above). Catalytic rate (k cat Here it is defined as hydroxyacetyl-CoA and HCO3-. - The maximum formation rate of tartrate-CoA. It describes the number of moles of substrate converted per mole of enzyme per unit time. Since 100% substrate saturation of the active site is required, only approximate calculations can be made. Therefore, the turnover rate at different initial substrate concentrations can be measured, and the catalytic rate can be calculated using the Michealis-Menton equation. It is assumed that the catalytic rate of hydroxyacetyl-CoA carboxylation by the GCC described in this paper is greater than approximately 5.6 ± 0.3 s. -1 .
[0130] It is assumed that the catalytic rate of hydroxyacetyl-CoA carboxylation by the GCC described in this paper is higher than approximately 5.3 s. -¹, approximately 5.4 s - ¹, Approximately 5.5 s - ¹, approximately 5.6 s - ¹, approximately 5.7 s - ¹, approximately 5.8 s - ¹, approximately 5.9 s - ¹, approximately 6 s - ¹, approximately 7 seconds - ¹, approximately 8 s - ¹, approximately 9 s - ¹, approximately 10 s - ¹, approximately 11 s - ¹, Approximately 12 s - ¹, approximately 13 s - ¹, approximately 14 s - ¹, Approximately 15 seconds - ¹, Approximately 16 s - ¹, approximately 17 s - ¹, Approximately 18 s - ¹, approximately 19 s - ¹, Approximately 20 s - ¹, approximately 21 s - ¹, Approximately 22 s - ¹, approximately 23 s - ¹, Approximately 24 s - ¹, Approximately 25 s - ¹Or any value between these values, such as 8.3 s - ¹. Preferably, the GCC described herein exhibits a catalytic rate greater than approximately 5.6 ± 0.3 s for the carboxylation of hydroxyacetyl-CoA (converting hydroxyacetyl-CoA to tartrate-CoA). - ¹. In a preferred aspect, the present invention relates to a GCC comprising SEQ ID NO: 3, wherein the GCC exhibits a catalytic rate greater than about 5.6 ± 0.3 s for the carboxylation of hydroxyacetyl-CoA (converting hydroxyacetyl-CoA to tartrate-CoA). - ¹.
[0131] Carboxylation activity can also be measured using the specific activity of GCC. Specific activity describes the number of moles of substrate converted per unit amount of enzyme per unit time. Maximum specific activity (v0) max The specific activity (v) is proportional to the catalytic rate, but since 100% substrate saturation of the active sites is required, it can only be approximated. Therefore, the specific activity at different initial substrate concentrations can be measured to calculate v. max Specifically, specific activity described herein refers to the number of nanomoles of hydroxyacetyl-CoA converted per minute per mg of GCC. As will be apparent from the accompanying examples, the preferred reference GCC, GCC M5, has a specific activity of approximately 937 ± 40 nmol of hydroxyacetyl-CoA per minute.- ¹ mg - ¹. Therefore, it is envisioned that the specific activity of the GCC of the present invention is higher than about 937 ± 40 nmol hydroxyacetyl-CoA min. - ¹ mg - ¹. This invention provides a GCC, wherein the specific activity of the GCC is higher than about 937 ± 40 nmol hydroxyacetyl-CoA min. - ¹mg - ¹. In a preferred aspect, the present invention relates to a GCC comprising SEQ ID NO: 3, wherein the specific activity of said GCC is higher than about 937 ± 40 nmol hydroxyacetyl-CoA min. - ¹ mg - ¹. It is also hypothesized that the GCC described herein has a specific activity higher than the following value: approximately 940 nmol hydroxyacetyl-CoA min. -1 mg -1 Approximately 950 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 960 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 970 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 980 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 990 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1000 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1100 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1200 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1300 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1400 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1500 nmol of hydroxyacetyl-CoAmin -1 mg -1 Approximately 1600 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1700 nmol of hydroxyacetyl-CoA min -1 mg -1Approximately 1800 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 1900 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2000 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2100 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2200 nmol of hydroxyacetyl-CoAmin -1 mg -1 Approximately 2300 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2400 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2500 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2600 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2700 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2800 nmol of hydroxyacetyl-CoA min -1 mg -1 Approximately 2900 nmol of hydroxyacetyl-CoAmin -1 mg -1 Approximately 3000 nmol of hydroxyacetyl-CoA min -1 mg -1 Or any value in between, such as 2339 nmol hydroxyacetyl-CoA min -1 mg -1 .
[0132] It is also hypothesized that the specific activity of the GCC described in this paper is approximately 940 nmol of hydroxyacetyl-CoA min. - ¹ mg - ¹, Approximately 950 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 960 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 970 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 980 nmol of hydroxyacetyl-CoA min - ¹ mg- ¹, Approximately 990 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1000 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1100 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1200 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1300 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1400 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1500 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1600 nmol of hydroxyacetyl-CoA min - ¹mg - ¹, Approximately 1700 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1800 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 1900 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2000 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2100 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2200 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2300 nmol of hydroxyacetyl-CoA min - ¹mg - ¹, Approximately 2400 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2500 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2600 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2700 nmol of hydroxyacetyl-CoA min - ¹ mg -¹, Approximately 2800 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 2900 nmol of hydroxyacetyl-CoA min - ¹ mg - ¹, Approximately 3000 nmol of hydroxyacetyl-CoA min - ¹mg - ¹, or any value in between, such as 2339 nmol of hydroxyacetyl-CoA min. - ¹ mg - ¹. Preferably, the specific activity of the GCC described herein is approximately 2602 ± 430 nmol hydroxyacetyl-CoA min. - ¹ mg - ¹. Therefore, the present invention provides a GCC, wherein the specific activity of the GCC is about 2602 ± 430 nmol hydroxyacetyl-CoA min. - ¹ mg - ¹. In a preferred method, the present invention relates to a GCC comprising SEQ ID NO: 3, wherein the specific activity of said GCC is about 2602 ± 430 nmol hydroxyacetyl-CoA min. -1 mg -1 In another preferred aspect, the present invention relates to a GCC comprising SEQ ID NO: 3, wherein the specific activity of said GCC is about 2602 ± 430 nmol hydroxyacetyl-CoA min. -1 mg -1 Furthermore, the catalytic rate of the GCC for the carboxylation of hydroxyacetyl-CoA (conversion of hydroxyacetyl-CoA to tartrate-CoA) is greater than approximately 5.6 ± 0.3 s. -1 In another preferred aspect, the present invention relates to a GCC comprising SEQ ID NO: 3, wherein the specific activity of said GCC is greater than about 937 ± 40 nmol hydroxyacetyl-CoA min. -1 mg -1 Furthermore, the catalytic rate of the GCC for the carboxylation of hydroxyacetyl-CoA (conversion of hydroxyacetyl-CoA to tartrate-CoA) is greater than approximately 5.6 ± 0.3 s. -1 .
[0133] Those skilled in the art will readily understand how to determine carboxylation activity, for example, through the determination methods described in the appended examples.
[0134] Carboxylation activity can be determined by measuring lysate-based enzyme-linked immunosorbent assay (ELISA) or by spectrophotometry using purified enzyme (Scheffen, see above).
[0135] Lysis-based measurements can be performed as described below. Constructs encoding GCC or GCC random mutagenesis libraries can be transformed into *E. coli* BL21_birA (see above). Eight colonies from each construct can be picked and inoculated into 96-well PlateOne plates containing Miller broth (100 μg / mL ampicillin and 50 μg / mL streptomycin). The culture plates can be incubated overnight at 37 °C and then transferred to new 96-well plates containing Miller broth, 100 μg / mL ampicillin, 50 μg / mL spectinomycin, and 2 μg / mL biotin to allow OD to be measured. 600 The value reaches 0.1. When OD 600 When the value reaches 0.4-0.6, protein expression can be induced with 0.25 mM isopropyl-β-d-1-thiogalactopyranoside, and the cells should be incubated overnight at 25°C. Cells can be lysed using CelLytic B (Sigma–Aldrich) and stored in 20% glycerol at -80°C. Enzyme activity can be measured in a microplate reader using a coupled enzyme assay with purified malonyl-CoA reductase (also known as CaMCR) from *Flexibrio cirrhosa*, as described previously (Scheffen, see above). A small-volume 384-well plate (Greiner Bio-One) can be used, containing 2 µL of cell extract, 100 mM 3-(N-morpholino)propanesulfonic acid (MOPS) at pH 7.8, 1 mM ATP, 50 mM KHCO3, 500 μg / mL CaMCR, 1 mM NADPH, 10 mM MgCl2, and 1 mM hydroxyacetyl-CoA in a 10 µL reaction volume. The absorbance of NADPH can be measured every 47 seconds at 340 nm at 37 °C using a Tecan Infinite M Plex microplate reader for 5 hours.
[0136] The carboxylation activity can be determined spectrophotometrically using the purified enzyme as described below. Enzyme production and purification, as well as the synthesis of CoA esters, have been described elsewhere in this document. To determine the carboxylation rate of GCC, a CaMCR-coupled enzyme assay can be performed. 100 mM MOPS pH 7.8, 50 mM KHCO3, 2 mM ATP, 0.3 mM NADPH, 5 mM MgCl2, 1.8 mg / mL CaMCR from *Flexibrio citrinum*, and 0.01–1 mg / mL GCC can be mixed in a cuvette and incubated at 37 °C for 2 minutes. The reaction can be initiated with 0.5 mM hydroxyacetyl-CoA, and absorbance can be measured over time at λ = 340 nm.
[0137] Hydroxyacetyl-CoA + HCO3- + ATP → Tartrate acyl-CoA + ADP + P i (GCC)
[0138] Tartrate-CoA + 2 NADPH → Glyceric acid + CoA + 2 NADP + (CaMCR)
[0139] It will be apparent to those skilled in the art that the present invention also covers GCCs that have a specific sequence identity with the specific sequence described herein.
[0140] For example, the present invention provides a GCC, wherein the GCC is characterized by:
[0141] It contains an amino acid sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.8% sequence identity with SEQ ID NO: 1, and
[0142] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0143] Preferably, the GCC has improved activity relative to the reference GCC, and preferably the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0144] Those skilled in the art are well aware of how to use, for example, algorithms such as those based on the CLUSTALW computer program (Thompson, Nucl. Acids Res. 2, 4673-4680, 1994), CLUSTAL Omega (Sievers, Curr. Protoc. Bioinformatics 48, 3.13.1-3.13.16, 2014), or FASTDB (Brutlag, Comp. App. Biosci. 6, 237-245, 1990) to determine the percentage of identity between sequences. Furthermore, those skilled in the art can also use BLAST (which stands for Basic Local Alignment Search Tool) and the BLAST 2.0 algorithm (Altschul, Nucl. Acids Res. 25:3389-3402, 1997; Altschul, J. Mol. Biol. 215, 403-410, 1990) and related tools.
[0145] It is assumed that a BLAST comparison of the sequence of SEQ ID NO: 1 with a SEQ ID NO: 1 having one amino acid substitution would yield 99.8% sequence identity.
[0146] This invention provides a GCC, wherein the GCC is characterized by:
[0147] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0148] Arginine is present at position 20 or the corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0149] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0150] In a preferred aspect, the present invention provides a GCC, wherein the GCC is characterized by:
[0151] It contains the amino acid sequence of SEQ ID NO: 3, and
[0152] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC comprises the amino acid sequence of SEQ ID NO: 1.
[0153] In another preferred aspect, the present invention provides a GCC, wherein the GCC is characterized by:
[0154] It comprises an α subunit, which comprises or is composed of the amino acid sequence of SEQ ID NO: 13, and / or comprises a β subunit, which comprises or is composed of the amino acid sequence of SEQ ID NO: 3.
[0155] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0156] This invention provides a GCC, wherein the GCC is characterized by:
[0157] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0158] The amino acid sequence shown in SEQ ID NO: 1 contains asparagine at position 100 or the position corresponding to that position.
[0159] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0160] In a preferred aspect, the present invention provides a GCC, wherein the GCC is characterized by:
[0161] It contains the amino acid sequence of SEQ ID NO: 7, and
[0162] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0163] In another preferred aspect, the present invention provides a GCC, wherein the GCC is characterized by:
[0164] It comprises an α subunit, which comprises or is composed of the amino acid sequence of SEQ ID NO: 13, and / or comprises a β subunit, which comprises or is composed of the amino acid sequence of SEQ ID NO: 7.
[0165] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0166] As detailed herein, the inventors constructed GCC M5 from PCC by introducing five mutations (L100S, Y143H, D407I, I450V, and W502R) into naturally occurring PCC. Therefore, preferably, the GCC described herein may have the same amino acids as GCC M5 (as shown in SEQ ID NO:1) at one or more of the five positions (e.g., one, two, three, four, or five positions). It is conceivable that the GCC described herein has one or more (e.g., one, two, three, four, or five) amino acids at the following positions in the amino acid sequence shown in SEQ ID NO: 1: serine at position 100 (if the GCC variant has more mutations at positions other than those mutated in GCC M5 compared to PCC), histidine at position 143, isoleucine at position 407, valine at position 450, and / or arginine at position 502, or one or more of these amino acids at positions corresponding to any of these positions. It will be apparent to those skilled in the art that when the 100th amino acid is replaced by, for example, asparagine, the amino acid at the 100th position is different from serine, for example, asparagine. For instance, when GCC has an amino acid substitution (preferably asparagine) at the 100th position, it is conceivable that GCC has histidine at position 143, isoleucine at position 407, valine at position 450, and / or arginine at position 502 in the amino acid sequence shown in SEQ ID NO: 1, or has the aforementioned amino acid at any of these corresponding positions. However, it is conceivable that in a GCC in which position 20 in SEQ ID NO: 1 or at the position corresponding to that position is substituted, the GCC has serine at position 100, histidine at position 143, isoleucine at position 407, valine at position 450, and / or arginine at position 502 in the amino acid sequence shown in SEQ ID NO: 1, or has the aforementioned amino acid at any of these corresponding positions.
[0167] Based on the above, the GCC disclosed herein may also have certain variations in addition to the mutations described herein, such as further modifications as described below. This is indicated and implied by expressions such as "at least 60% identity," as used herein.
[0168] To understand, the GCC disclosed herein is characterized by containing an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 1, and is disclosed to have one or more disclosed mutations / substitutions, for example, one or more of the following: arginine at position 20, serine at position 100, histidine at position 143, isoleucine at position 407, valine at position 450, and / or arginine at position 502 in the amino acid sequence shown in SEQ ID NO: 1, or having one or more of the above amino acids at positions corresponding to any of these positions. This means that the GCC has these mutations / substitutions as constituent elements, i.e., they need to be present in the amino acid sequence of the GCC. However, due to having at least 60% sequence identity with SEQ ID NO: 1, the GCC may have further mutations / substitutions at other positions within the threshold range of having at least 60% sequence identity with SEQ ID NO: 1. For example, in addition to one or more mutations as described above, the amino acid sequence of GCC may have further modifications, such as substitutions, additions, and / or deletions (preferably substitutions) of one or more amino acids, preferably 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more amino acids, such as 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids. For example, the amino acid sequence of GCC may have more such modifications at positions corresponding to positions 1 to 19, 20, 21-99, 100, 101-142, 143, 144-406, 407, 408-449, 450, 451-501, 502, 503-510. For example, if GCC has (only) the mutation described above at position 20, it may have one or more further mutations at any other position, for example, at one or more positions 1-19 and / or 21-510. As another example, if GCC has one or more disclosed mutations / substitutions, for example, one or more of the following: in the amino acid sequence shown in SEQ ID NO: 1, position 20 is arginine, position 143 is histidine, position 407 is isoleucine, position 450 is valine and / or position 502 is arginine, or has one or more of the above amino acids at positions corresponding to any of these positions, then GCC may have one or more mutations / substitutions, for example at positions corresponding to positions 1-19, 21-142, 144-406, 408-449, 451-501 and / or 503-510.
[0169] Specifically, it is envisioned that such further modifications (mutations / substitutions) do not affect the activity of the GCC as defined herein, or do not substantially affect the activity of the GCC as defined herein. "Not substantially affect" in this context may mean that the activity of the GCC is reduced by no more than 10% compared to a GCC without / with such modifications. Specifically, a GCC with such further modifications (mutations / substitutions) still exhibits higher activity compared to a reference GCC, preferably wherein the reference GCC comprises the amino acid sequence of SEQ ID NO:1.
[0170] To understand the GCC provided and used herein, it is necessary to understand that it contains the mutations defined and interpreted herein, specifically, one or more amino acid substitutions at position 20 or 100 of the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to either of these positions. Apart from these mutations, it is contemplated that the sequence of the mutated GCC at all other positions may comprise or consist of the same amino acid sequence as its counterpart (GCC-M5, SEQ ID NO. 1). For example, if a GCC contains the mutations defined and interpreted herein, it may comprise or consist of the same amino acid sequence as GCC-M5, SEQ ID NO. 1, except for those mutations.
[0171] Preferably, any further modifications of this kind (e.g., deletion, insertion, addition, and / or substitution (especially substitution in this context)) are conservative, i.e., the amino acid is substituted with an amino acid having the same or similar properties. For example, a hydrophobic amino acid is preferably substituted with another hydrophobic amino acid, and so on.
[0172] This invention provides a GCC, wherein the GCC is characterized by:
[0173] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0174] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0175] The 502nd amino acid is arginine.
[0176] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0177] The present invention also provides a GCC, wherein the GCC is characterized by:
[0178] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0179] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0180] The 450th amino acid is valine.
[0181] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0182] The present invention also provides a GCC, wherein the GCC is characterized by:
[0183] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0184] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0185] The 407th amino acid is isoleucine.
[0186] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0187] The present invention also provides a GCC, wherein the GCC is characterized by:
[0188] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0189] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0190] The 143rd amino acid is histidine.
[0191] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0192] The present invention also provides a GCC, wherein the GCC is characterized by:
[0193] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0194] It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions.
[0195] The amino acid at position 143 is histidine, the amino acid at position 407 is isoleucine, the amino acid at position 450 is valine, and / or the amino acid at position 502 is arginine.
[0196] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0197] The present invention also provides a GCC, wherein the GCC is characterized by:
[0198] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0199] It has an amino acid substitution, deletion, or insertion at position 20 or the corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0200] The 100th amino acid is serine, the 143rd amino acid is histidine, the 407th amino acid is isoleucine, the 450th amino acid is valine, and / or the 502nd amino acid is arginine.
[0201] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0202] The present invention also provides a GCC, wherein the GCC is characterized by:
[0203] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0204] It contains arginine at position 20 or the corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0205] The 100th amino acid is serine, the 143rd amino acid is histidine, the 407th amino acid is isoleucine, the 450th amino acid is valine, and / or the 502nd amino acid is arginine.
[0206] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0207] The present invention also provides a GCC, wherein the GCC is characterized by:
[0208] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0209] It contains arginine at position 20 or any corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0210] The 100th amino acid is alanine, arginine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine; the 143rd amino acid is histidine; the 407th amino acid is isoleucine; the 450th amino acid is valine; and / or the 502nd amino acid is arginine.
[0211] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0212] The present invention also provides a GCC, wherein the GCC is characterized by:
[0213] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0214] It has an amino acid substitution, deletion, or insertion at position 100 or the corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0215] The amino acid at position 143 is histidine, the amino acid at position 407 is isoleucine, the amino acid at position 450 is valine, and / or the amino acid at position 502 is arginine.
[0216] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0217] The present invention also provides a GCC, wherein the GCC is characterized by:
[0218] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0219] It contains asparagine at position 100 or the corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0220] The amino acid at position 143 is histidine, the amino acid at position 407 is isoleucine, the amino acid at position 450 is valine, and / or the amino acid at position 502 is arginine.
[0221] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0222] The present invention also provides a GCC, wherein the GCC is characterized by:
[0223] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0224] It contains asparagine at position 100 or any corresponding position in the amino acid sequence shown in SEQ ID NO: 1.
[0225] The 20th amino acid is alanine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine; the 143rd amino acid is histidine; the 407th amino acid is isoleucine; the 450th amino acid is valine; and / or the 502nd amino acid is arginine.
[0226] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0227] From the above aspects, it is clear that the stated sequence identity with SEQ ID NO: 1 may not refer to all positions in SEQ ID NO: 1. In other words, amino acids at certain positions should not be substituted, even though the resulting GCC has, for example, at least 60% sequence identity with SEQ ID NO: 1. Specifically, the amino acid substitutions that lead from PCC to GCC M5 may not require further substitution.
[0228] As explained in detail, the inventors have discovered that amino acid substitutions at certain positions enhance the activity of GCC M5. Therefore, the present invention provides a GCC having specific amino acid substitutions for glycine and serine at positions 20 and 100, respectively, relative to the reference amino acid sequence of SEQ ID NO: 1.
[0229] Therefore, the present invention provides a GCC, wherein the GCC is characterized by:
[0230] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0231] (1) The amino acid at position 20 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine, preferably by arginine, histidine, lysine, or tyrosine, and more preferably by arginine.
[0232] Preferably, the amino acid being replaced is glycine; and / or
[0233] (2) The amino acid at position 100 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, threonine, tryptophan, tyrosine, or valine, preferably by asparagine.
[0234] Preferably, the amino acid being replaced is serine.
[0235] The present invention also provides a GCC, wherein the GCC is characterized by:
[0236] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0237] (1) The amino acid at position 20 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine, preferably by arginine, histidine, lysine, or tyrosine, and more preferably by arginine.
[0238] Preferably, the amino acid being replaced is glycine; and / or
[0239] (2) The amino acid at position 100 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, threonine, tryptophan, tyrosine, or valine, preferably by asparagine.
[0240] Preferably, the amino acid being replaced is serine;
[0241] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0242] The present invention also provides a GCC, wherein the GCC is characterized by:
[0243] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0244] (1) The amino acid at position 20 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by arginine.
[0245] Preferably, the amino acid being replaced is glycine; and / or
[0246] (2) The amino acid at position 100 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by asparagine.
[0247] Preferably, the amino acid being replaced is serine.
[0248] The present invention also provides a GCC, wherein the GCC is characterized by:
[0249] It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and
[0250] (1) The amino acid at position 20 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by arginine.
[0251] Preferably, the amino acid being replaced is glycine; and / or
[0252] (2) The amino acid at position 100 or the position corresponding to that position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by asparagine.
[0253] Preferably, the amino acid being replaced is serine;
[0254] Preferably, the GCC has improved activity relative to the reference GCC, and preferably, the reference GCC has the amino acid sequence of SEQ ID NO: 1.
[0255] This invention also relates to nucleic acid molecules encoding the GCC described herein. In the context of this invention, the nucleic acid molecule may be a DNA molecule or an RNA molecule.
[0256] In a preferred aspect, the present invention relates to a nucleic acid molecule as shown in SEQ ID NO: 16. In a preferred aspect, the α subunit of the GCC of the present invention is encoded by a nucleic acid molecule as shown in SEQ ID NO: 26, and the β subunit is encoded by a nucleic acid molecule as shown in SEQ ID NO: 16.
[0257] In another preferred aspect, the present invention relates to a nucleic acid molecule as shown in SEQ ID NO: 20. In a preferred aspect, the α subunit of the GCC of the present invention is encoded by a nucleic acid molecule as shown in SEQ ID NO: 26, and the β subunit is encoded by a nucleic acid molecule as shown in SEQ ID NO: 20.
[0258] This invention further relates to vectors comprising the nucleic acid molecules described herein. Vectors that can be used according to the invention are known in the art. These vectors may also contain expression regulatory sequences operatively linked to the nucleic acid molecules of the invention contained in the vector. These expression regulatory sequences can be adapted to ensure transcription and synthesis of translatable RNA in bacteria or fungi. The expression regulatory sequence may, for example, be a promoter. The promoter used in conjunction with the nucleic acid molecules of the invention may be homologous or heterologous, depending on its origin and / or the gene to be expressed. Suitable promoters are, for example, promoters that constitutively express themselves. However, promoters that are activated only at a time point determined by external influences may also be used. In this case, artificial and / or chemically inducible promoters may be used.
[0259] Preferably, the vector of the present invention is an expression vector. Expression vectors have been extensively described in the literature. Typically, they contain not only a selection marker gene and a replication origin to ensure replication in a selected host, but also a bacterial or viral promoter, and in most cases, a transcription termination signal. There is usually at least one restriction enzyme site or polylinker between the promoter and the termination signal, which allows the insertion of a coding DNA sequence. If this sequence is active in the selected host organism, a naturally occurring DNA sequence that controls the transcription of the corresponding gene can be used as the promoter sequence. However, this sequence can also be replaced by other promoter sequences. Promoters that ensure constitutive gene expression and inducible promoters that allow intentional control of gene expression can be used. Bacterial and viral promoter sequences with these characteristics have been described in detail in the literature. Regulatory sequences for expression in microorganisms (e.g., *Escherichia coli*, *Saccharomyces cerevisiae*) have been well described in the literature. Promoters that allow for particularly efficient expression of downstream sequences include, for example, the T7 promoter (Studier, Methods in Enzymology 185, 60-89, 1990), lacUV5, trp, trp-lacUV5 (DeBoer, Promoters, Structure and Function, Praeger, 462-481, 1982; DeBoer, Proc. Natl. Acad. Sci. USA 80, 21-25, 1983), lp1, or rac (Boros, Gene 42, 97-100, 1986). Inducible promoters are preferred for downstream sequence expression. These promoters typically produce higher polypeptide yields than constitutive promoters. To obtain optimal polypeptide yields, a two-stage process is usually employed. First, host cells are cultured under optimal conditions until a relatively high cell density is achieved. In the second step, transcription is induced according to the type of promoter used. In this respect, the tac promoter is particularly suitable, as it can be induced by lactose or IPTG (= isopropyl-β-D-thiogalactopyranoside) (DeBoer, see above). Transcription termination signals are also described in the literature.
[0260] The present invention further provides an organelle comprising the nucleic acid molecule or vector described herein.
[0261] Therefore, the present invention provides an organelle comprising a nucleic acid molecule encoding the GCC of the present invention or a vector comprising a nucleic acid molecule encoding the GCC of the present invention. Thus, the present invention provides an organelle for producing the GCC of the present invention. Organelles are specialized subunits within a cell (typically a eukaryotic cell) that perform specific functions. Examples of organelles are mitochondria, the nucleus, the Golgi apparatus, the endoplasmic reticulum, and plastids. Preferably, in the context of the present invention, the organelle is a plastid. Plastids can be chloroplasts, cyanoplasts, erythroplasts, or phaneroplasts, etc., preferably chloroplasts.
[0262] This invention also provides a host cell comprising the nucleic acid molecules, vectors, or organelles described herein. Those skilled in the art can easily select a suitable host cell, such as a host cell for protein expression, based on various factors, including the desired protein, desired expression level, downstream processing requirements, and process costs. Commonly used host cells for protein expression include bacteria (such as *Escherichia coli*), yeast (such as *Saccharomyces cerevisiae*), insect cells (such as Sf9 and Sf21), and mammalian cells (such as CHO and HEK293). Each host cell has its advantages and limitations. For example, bacterial systems are relatively easy to grow and produce high yields of proteins, but they may not be suitable for expressing complex eukaryotic proteins with post-translational modifications. Yeast systems are also easy to grow and capable of some post-translational modifications, but they may not be suitable for certain types of proteins. Insect cells are capable of some post-translational modifications. They are suitable for expressing large, complex proteins, but their culture costs are higher than those of bacterial and yeast systems. Mammalian cells are generally preferred for producing biologically active and properly folded proteins with complex post-translational modifications. However, their culture is more difficult and costly than that of bacterial and yeast systems. Ultimately, the choice of host cell for protein expression will depend on the specific requirements of the protein and its downstream applications.
[0263] Escherichia coli is the preferred organism if the host cell is used as a chassis for cloning and constructing carbon dioxide fixation modules or for expressing any carbon dioxide fixation or photosynthesis-related proteins independent of post-translational modifications, due to its genetic accessibility, ease of culture, and rapid growth. The cyanobacterial strain *Synechococcus slenderus* (…) Synechococcus elongatus ) and species of the genus Synechococcus ( Synechocystis sp. PCC 6803 is a preferred cell type for protein expression in a photosynthetic metabolic context because these organisms are easy to genetically manipulate, easy to culture, grow rapidly, and have a simpler structure than eukaryotic phototrophs. Single-celled algae, such as *Chlamydomonas reinhardtii* (…), are also suitable. Chlamydomonas reinhardtii), are preferred organisms for expressing proteins in eukaryotic cells or chloroplasts under photosynthetic metabolic backgrounds because their metabolic and physiological characteristics are similar to plants, but they grow faster and are easier to cultivate. Among plants, Arabidopsis thaliana is a preferred organism for expressing proteins under photosynthetic backgrounds. Arabidopsis thaliana It grows rapidly compared to other plants or any crops; moreover, it has industrial value.
[0264] The present invention also provides a tissue comprising the host cells described herein. Therefore, the present invention provides a tissue comprising host cells containing the nucleic acid molecules, vectors, or organelles described herein. A tissue can be defined as a collection of similar or different host cells performing a specific function as described herein. The host cells described herein can be seeded onto a scaffold, which serves as a three-dimensional support structure for cell growth and, for example, the generation of GCCs described herein.
[0265] This invention further provides an organism comprising the nucleic acid molecules, vectors, or organelles described herein. Furthermore, this invention relates to an organism or other related biological system comprising or encoding the GCC described herein. Preferably, in the context of this invention, the organism is a plant, algae, or microorganism.
[0266] Algae are not particularly restricted, but can be eukaryotic single-celled organisms (diatoms, yellow-green algae, dinoflagellates, etc.) or multicellular organisms, such as seaweed (e.g., red algae, brown algae, green algae, etc.). They exist in a variety of environments, including oceans, freshwater bodies, soil, and even on other organisms. Technicians can easily select the appropriate algae according to the required application.
[0267] Algae containing the nucleic acid molecules, carriers, or organelles described herein are commonly used in biofuel production, food production, bioplastics production, or carbon sequestration. One algae-based carbon sequestration method involves cultivating algae in large, closed systems, such as photobioreactors or ponds. In these systems, algae grow under controlled conditions and are provided with the nutrients and light required for photosynthesis. As the algae grow, they absorb carbon dioxide from the atmosphere and convert it into biomass. Once mature, the algae can be harvested and processed for the production of biofuels, food, or other products, while the remaining biomass can be used as a soil conditioner or stored for long-term carbon sequestration. Another approach utilizes algae to capture and recover carbon dioxide emissions from industrial sources, such as power plants or factories. In this process, industrial flue gas is introduced into algae ponds or bioreactors, where the algae absorb carbon dioxide and convert it into biomass. This method promises to reduce greenhouse gas emissions from industrial sources and produce renewable, carbon-neutral, or carbon-negative products.
[0268] The plants that can be used in the context of this invention are not limited and can belong to either dicotyledons or monocotyledons. Non-limiting examples of plants that can be used in the context of this invention are provided below.
[0269] Brassicaceae: Arabidopsis thaliana ( Arabidopsis thaliana ),turnip( Turnip cabbage , Cabbage turnip ),cabbage( Brassica oleracea var. Capitalized ), Chinese cabbage ( Brassica rapa var. of Beijing ), rosewood ( Brassica rapa var. hakabura ), water vegetables ( Brassica rapa var. lancinium ), Komatsuna ( Brassica rapa var. perviridis ), Japanese radish ( Radish agricultural Wasabi Japanese wasabi ) etc. Solanaceae: Tobacco ( Nicotiana tabacum ),eggplant( Nightshade eggplant ),potato( Solanum tuberosum ),tomato( Lycopersicon lycopersicum ),chili( Red pepper Petunias, etc. Legumes: Soybeans ( Glycine max ),pea( Pea ),broad bean( Vetches beans Wisteria ( Wisteria floribunda ),peanut( Peanut ), lotus root ( Lotus corniculatus var. Japonicus ),kidney bean( Common bean ), adzuki beans (Angular vine ), Acacia ( Acacia) Etc. Asteraceae family: Chrysanthemum, Sunflower ( Sunflower ) etc. Palm family: Oil palm ( Olives Guinean, Elaeis oleifera ),coconut( Cocos nucifera ),date palm( Phoenix dactylifera ), wax palm ( Copernicus ) etc. Anacardiaceae: Forget-me-not ( Rhus succedanea ),cashew( Cashew western ), lacquer tree ( Toxicodendron vernicifluus ),mango( Mangosteen ),pistachio( Pistachio tree ) etc. Cucurbitaceae: Pumpkin ( Cucurbita maxima, Cucurbita moschata, Cucurbita watermelon ),cucumber( Cucumber ),gooseberry( Trichosanthes cucumeroides ),gourd( Lagenaria siceraria var. Gourda ) etc. Rosaceae: Almond ( Common almond ),Rose( Rose ),strawberry( Strawberries ),cherry( Prunus ),apple( Malus pumila var. Domestica ) etc. Salicaceae: Poplar ( Populus trichocarpa, Populus nigra, Populus tremula ) etc. Poaceae: *Brucea*, *Zea* (… Zea mays ), rice Oryza sativa ),barley( Hordeum vulgare ),wheat( Triticum aestivum ),bamboo( Phyllostachys ),sugar cane( Saccharum officinarum ), elephant grass ( Pennisetum putureum ), Miscanthus sinensis ( Miscanthus virgatum ), sorghum ( Sorghum ), willow branch millet ( Panicum ) etc. Liliaceae: Tulip ( Tulipa ),lily( Lilium ) etc. Myrtaceae: Eucalyptus ( Eucalyptus camaldulensis, Eucalyptus grandis )wait.
[0270] In the context of this invention, the plant can also be a moss. Mosses are sometimes also referred to as bryophytes, including three classes of non-vascular terrestrial plants (liverwort, hornwort, and moss). In the context of this invention, the plant can also be a fern.
[0271] As described above, this invention provides a microorganism comprising the nucleic acid molecules, vectors, or organelles described herein. As used herein, microorganisms are not limited to these categories, but refer to microscopic organisms that can exist as single cells or as cell communities. Microorganisms include, in particular, bacteria, fungi, archaea, and protozoa.
[0272] Non-limiting examples of microorganisms include Aspergillus ( Aspergillus ), such as Aspergillus echinosporum ( Aspergillus aculeatus ) and Aspergillus oryzae ( Aspergillus oryzae Various known yeasts have also been envisioned, such as yeasts of the genus *Saccharomyces*, including *Saccharomyces cerevisiae*. Saccharomyces cerevisiae Yeasts of the genus *Schizosaccharomyces*, such as *Schizosaccharomyces* var. *saccharomyces*. Schizosaccharomyces pombe Candida ( ) ; Candida genus ( Candida Yeasts, such as Candida hujatae ( Candida shehatae ); Pichia pastoris ( Pichia Yeasts, such as Pichia pastoris ( Pichia stipitis ); Hansenula genus ( Hansenula ) yeast; Klerk yeast ( Klocekera Yeasts of the genus *Swaniomyces*; yeasts of the genus *Yersinia*. Yarrowia yeasts; Trichosporium genus ( Trichosporon ) yeast; Brett's yeast ( Brettanomyces Yeasts of the genus *Trichostomum*; and yeasts of the genus *Trichostomum* Pachysolen ) yeast.
[0273] In the context of this invention, microorganisms can refer to bacteria. Therefore, this invention provides a bacterium comprising the nucleic acid molecule or vector described herein. Those skilled in the art are familiar with the types of bacteria suitable for the desired application. Non-limiting examples include species of the genera *Escherichia coli* (such as *Escherichia coli*), *Bacillus* (such as *Bacillus subtilis*), *Streptomyces*, and *Pseudomonas* (such as *Pseudomonas putida*).
[0274] While it will be apparent to those skilled in the art, it should be noted that when an organelle, host cell, tissue, or organism contains the nucleic acid described herein, it is contemplated that the organelle, host cell, tissue, or organism expresses the nucleic acid, i.e., the organelle, host cell, tissue, or organism produces a protein / peptide encoded by the nucleic acid. Therefore, it is contemplated that organelles, host cells, tissues, or organisms produce the GCC described herein. In a preferred aspect of the invention, the organelle, host cell, tissue, or organism produces the protein of SEQ ID NO: 3 or 7. Therefore, in a preferred aspect of the invention, the organelle, host cell, tissue, or organism contains the nucleic acid of SEQ ID NO: 16 or 20. As previously stated, GCC contains an α subunit and a β subunit. Therefore, it is contemplated that organelles, host cells, tissues, or organisms produce the α and β subunits of GCC. Therefore, in a preferred aspect of the invention, the organelle, host cell, tissue, or organism contains the nucleic acid of SEQ ID NO: 26 and the nucleic acid of SEQ ID NO: 16 or 20.
[0275] It is envisioned that, compared to organelles, host cells, tissues, or organisms that produce GCC in the prior art, the organelles, host cells, tissues, or organisms that produce GCC of the present invention may be characterized by a higher conversion rate of hydroxyacetyl-CoA to tartrate-CoA.
[0276] Therefore, the present invention provides an organelle, host cell, tissue or organism that produces the GCC of the present invention, wherein the organelle, the host cell, the tissue or the organism converts hydroxyacetyl-CoA to tartrate-CoA at a higher rate than the corresponding organelle, host cell, tissue or organism that produces a reference GCC (preferably the reference GCC comprising SEQ ID NO: 1).
[0277] Therefore, the present invention provides an organelle, host cell, tissue or organism comprising the GCC of the present invention, wherein the organelle, the host cell, the tissue or the organism converts hydroxyacetyl-CoA to tartrate-CoA at a higher rate than the corresponding organelle, host cell, tissue or organism comprising a reference GCC (preferably comprising the reference GCC of SEQ ID NO: 1).
[0278] The production of GCC leads to a higher conversion rate of hydroxyacetyl-CoA to tartrate-CoA, which may in turn result in a higher carboxylation rate and thus a higher carbon yield. This can potentially affect the growth rate of, for example, organelles, host cells, tissues, or organisms. Growth rate may be directly related to biomass accumulation and can be quantified, particularly as an indicator such as the proliferation, amplification, or reproduction of organelles, host cells, tissues, or organisms. Methods for quantifying carbon yield and improved growth / higher growth rate are known in the art.
[0279] Therefore, the present invention provides an organelle, host cell, tissue or organism that produces the GCC of the present invention, wherein the growth rate and / or carbon production rate of the organelle, the host cell, the tissue or the organism is higher than that of the corresponding organelle, host cell, tissue or organism that produces a reference GCC (preferably comprising the reference GCC of SEQ ID NO: 1).
[0280] Therefore, the present invention provides an organelle, host cell, tissue or organism comprising the GCC of the present invention, wherein the growth rate and / or carbon production rate of the organelle, the host cell, the tissue or the organism is higher than that of the corresponding organelle, host cell, tissue or organism comprising a reference GCC (preferably comprising the reference GCC of SEQ ID NO: 1).
[0281] Imagine that the GCCs, organelles, host cells, tissues, or organisms described herein are used in certain methods and processes.
[0282] Therefore, this invention provides uses of the GCCs, organelles, host cells, tissues, or organisms described herein. It should be noted that if the nucleic acids and vectors can be used in a manner that generates a polypeptide encoding a peptide to, for example, catalyze the carboxylation of hydroxyacetyl-CoA to tartrate-CoA, then all disclosed uses and methods of the GCCs, organelles, host cells, tissues, or organisms described herein are also disclosed herein for use with the nucleic acids and vectors.
[0283] The GCCs, organelles, host cells, tissues, or organisms mentioned can be used for carbon dioxide fixation because they can catalyze the conversion of hydroxyacetyl-CoA to tartrate-CoA via carboxylation. Similarly, the GCCs, organelles, host cells, tissues, or organisms can also be used in photosynthesis, a process that also relies on the incorporation of carbon dioxide into organic molecules. Carbon dioxide fixation and photosynthesis are defined as the incorporation of carbon dioxide from inorganic gaseous or aqueous solutions into organic molecules. These processes can rely on the use of energy equivalents assimilated by means of light energy. However, other methods of providing energy equivalents, such as ATP, are also conceivable.
[0284] Therefore, this invention provides the use of the GCCs, organelles, host cells, tissues or organisms described herein for carbon dioxide fixation and / or photosynthesis.
[0285] The GCC, organelles, host cells, tissues, or organisms mentioned above can also be used to decompose certain substances, such as hazardous waste. For example, it is conceivable that the GCC, organelles, host cells, tissues, or organisms could be used to decompose synthetic materials that pollute the environment.
[0286] Therefore, the use of the GCC, organelles, host cells, tissues, or organisms may involve the decomposition of synthetic materials. Synthetic materials refer to substances formulated or manufactured through chemical processes, or substances formulated or manufactured through processes that chemically alter substances extracted from natural plant, animal, or mineral sources.
[0287] Therefore, the present invention provides a method for decomposing synthetic materials from the GCC, organelles, host cells, tissues or organisms.
[0288] In the context of this invention, the synthetic material may be, for example, a polyester, prepared from monomers derived from mineral oil. Therefore, the GCC, organelles, host cells, tissues, or organisms can be used to decompose polyesters (such as synthetic resins), wherein the monomer units are linked by ester groups. Polyesters and methods for their decomposition are well known to those skilled in the art.
[0289] Therefore, the present invention provides a use for the decomposition of synthetic materials from the GCC, organelles, host cells, tissues, or organisms, wherein the synthetic material is a polyester. In other words, the present invention provides a use for the decomposition of polyester from the GCC, organelles, host cells, tissues, or organisms.
[0290] Polyethylene terephthalate (PET) is a widely used polyester, primarily derived from mineral oil, and a major environmental pollutant. PET consists of ethylene glycol and terephthalic acid monomers linked together by ester bonds to form polymer structures of varying lengths. Degradation of PET can include the hydrolytic cleavage of these ester bonds, yielding the monomers ethylene glycol and terephthalic acid. As is well known to those skilled in the art, ethylene glycol can be readily enzymatically modified to produce hydroxyacetyl-CoA, which can then be carboxylated by said GCC, organelles, host cells, tissues, or organisms to produce tartrate-CoA.
[0291] Therefore, the present invention provides the use of the GCC, organelles, host cells, tissues, or organisms for the decomposition of polyester, wherein the polyester is PET. In other words, the present invention provides the use of the GCC, organelles, host cells, tissues, or organisms for the decomposition of PET.
[0292] Ethylene glycol is also a significant environmental pollutant because it is widely used as a de-icing fluid, for example, in aircraft. Furthermore, one of the most abundant organic compounds in the ocean is glycolic acid, as it is secreted by marine algae (Wright, Mar. Biol. 43, 257-263, 1977).
[0293] Those skilled in the art will understand that ethylene glycol and / or glycolic acid can be precursors of hydroxyacetyl-CoA. In other words, under appropriate catalytic conditions or in the presence of sufficient enzymes, ethylene glycol and / or glycolic acid can be converted to hydroxyacetyl-CoA. Hydroxyacetyl-CoA can then be converted to tartrate-CoA by the GCCs, organelles, host cells, tissues, or organisms described herein. Therefore, the GCCs, organelles, host cells, tissues, or organisms described herein can be used for the breakdown of ethylene glycol and / or glycolic acid. Suitable enzymes for breaking down ethylene glycol into glycolic acid are derived from *Glucosobacterium oxysporum* (Glucosobacterium oxysporum). Gluconobacter oxydans Gox0313 from *E. coli*, FucO, or any homologous enzyme, is suitable for activating glycolic acid to hydroxyacetyl-CoA. The recently developed hydroxyacetyl-CoA synthase GCS, derived from the *Rhodotorula* species NAP1, or any homologous enzyme (Scheffen, see above) is also considered. The use of enzymes such as those from *Clostridium gamma* (Gamma-aminobutyric acid bacteria) is also envisioned. Clostridium aminobutyricum ) of the coenzyme A transferase of AbfT.
[0294] Therefore, this invention provides the use of the GCCs, organelles, host cells, tissues, or organisms described herein for the degradation of ethylene glycol and / or glycolic acid. While it is obvious, it should be noted that those skilled in the art can readily identify other substances or materials that can be degraded by hydroxyacetyl-CoA or its precursors. Therefore, those skilled in the art can identify other substances or materials whose degradation can be assisted by the GCCs, organelles, host cells, tissues, or organisms described herein.
[0295] The disclosures made in the context of the methods described herein, after appropriate modifications, are considered disclosures for the corresponding uses. The disclosures made in the context of the uses described herein, after appropriate modifications, are considered disclosures of the corresponding methods.
[0296] Therefore, the present invention provides methods for catabolism and anabolism using the GCCs, organelles, host cells, tissues or organisms described herein.
[0297] Therefore, the present invention provides a method for producing biomass using the GCCs, organelles, host cells, tissues, or organisms described herein. Those skilled in the art will understand that methods for producing biomass may include incorporating inorganic carbon (such as carbon dioxide) into organic molecules, thereby resulting, in particular, the proliferation, amplification, or reproduction of organelles, host cells, tissues, or organisms.
[0298] Furthermore, the present invention provides a method for degrading synthetic materials (preferably polyester, more preferably PET) using the GCC, organelles, host cells, tissues or organisms described herein.
[0299] Therefore, the present invention provides a method for decomposing and synthesizing materials using GCC, organelles, host cells, tissues or organisms as described herein.
[0300] The present invention also provides a method for decomposing polyesters using the GCC, organelles, host cells, tissues or organisms described herein.
[0301] This invention also provides a method for decomposing PET using the GCC, organelles, host cells, tissues or organisms described herein.
[0302] The present invention also provides a method for decomposing ethylene glycol and / or glycolic acid using the GCC, organelles, host cells, tissues or organisms described herein.
[0303] It is also envisioned that the nucleic acids, vectors, GCCs, organelles, host cells, tissues and / or organisms described herein are included in the composition, which may, for example, contain other reagents for a particular method or use.
[0304] The present invention also provides a composition comprising the nucleic acids, vectors, GCCs, organelles, host cells, tissues and / or organisms described herein.
[0305] The nucleic acids, vectors, GCCs, organelles, host cells, tissues or organisms described herein, when modified accordingly, shall be considered as disclosures of the corresponding compositions.
[0306] Unless otherwise defined, all technical terms, symbols, and other scientific terms used herein are intended to have the meaning commonly understood by one of ordinary skill in the art to which this invention pertains. In some cases, for clarity and / or ease of reference, terms with commonly understood meanings are defined herein, and the inclusion of these definitions herein is not necessarily to be construed as differing from their commonly understood meaning in the art. The techniques and procedures described or referenced herein are generally well known to those skilled in the art and are typically used using conventional methods. Where appropriate, procedures involving the use of commercially available kits and reagents are generally performed according to the manufacturer's prescribed operating procedures and conditions, unless otherwise stated.
[0307] As used herein, the singular forms “a,” “an,” and “the” include plural references unless the context clearly indicates otherwise. The terms “including,” “such as,” and “etc.” are intended to mean that they include, but are not limited to, the listed contents unless otherwise clearly stated.
[0308] As used herein, the term "or" generally takes its common meaning, including "and / or," unless the context clearly specifies otherwise. The term "and / or" refers to one or all of the listed elements, or any combination of two or more of the listed elements.
[0309] As used herein, the term “comprising” also specifically refers to the elements listed in “consisting of” and “substantially consisting of”, unless otherwise explicitly stated.
[0310] As used herein, the term "about" indicates and covers a specified value as well as a range above and below that value. The term "about" may mean a specified value ± 10%, ± 5%, or ± 1%. In some embodiments, if applicable, the term "about" means a specified value ± one standard deviation of that value.
[0311] As used herein, the terms “comprising,” “including,” “having,” or grammatical variations thereof should be understood to specify the stated feature, integer, step, or component, but do not preclude the addition of one or more additional features, integers, steps, components, or combinations thereof. The terms “comprising,” “including,” or “having” encompass the terms “consisting of,” and “substantially consisting of.” Therefore, whenever the terms “comprising,” “including,” or “having” are used herein, they may be replaced with “substantially consisting of,” or preferably “consisting of,”.
[0312] The terms “comprise” / “include” / “have” mean that any other component (or similar feature, integer, step, etc.) may exist.
[0313] The term "composed of" means that no other components (or similar features, integers, steps, etc.) may be present.
[0314] The term “consistently of” or its grammatical variations as used herein shall be understood to specify the stated feature, integer, step or component, but does not preclude the addition of one or more additional features, integers, steps, components or groups thereof, provided that such additional features, integers, steps, components or groups thereof do not materially alter the essential and novel characteristics of the claimed product, composition, use or method, etc.
[0315] The term "method" refers to the manner, means, techniques, and procedures for accomplishing a particular task, including but not limited to those manner, means, techniques, and procedures known to practitioners in the fields of chemistry, biology, and biophysics, or that can be readily developed from known manner, means, techniques, and procedures.
[0316] All sequences disclosed in this article are listed in the table below.
[0317] Table 1: Sequences
[0318]
[0319]
[0320]
[0321] This invention relates to the following nucleotide and amino acid sequences:
[0322] amino acid sequence
[0323] SEQ ID NO: 1
[0324] The amino acid sequence of GCC M5 (Typhobic methylbacterium) β-subunit L100S, Y143H, D407I, I450V, W502R mutations.
[0325] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0326] SEQ ID NO: 2
[0327] Amino acid sequence of the GCC M5_L8F (Methylobacterium extorquens) β-subunit with L8F, L100S, Y143H, D407I, I450V, W502R mutations
[0328] MKDILEKFEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0329] SEQ ID NO: 3
[0330] Amino acid sequence of the β-subunit G20R, L100S, Y143H, D407I, I450V, W502R mutants of GCC M5_G20R (Methylobacterium extorquens)
[0331] MKDILEKLEERRAQARLGGREKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0332] SEQ ID NO: 4
[0333] Amino acid sequence of the β-subunit M53E, L100S, Y143H, D407I, I450V, W502R mutants of GCC M5_M53E (Methylobacterium extorquens)
[0334] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDEFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0335] SEQ ID NO: 5
[0336] Amino acid sequence of the β-subunit M53Q, L100S, Y143H, D407I, I450V, W502R mutants of GCC M5_M53Q (Methylobacillus extorquens)
[0337] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDQFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0338] SEQ ID NO: 6
[0339] Amino acid sequence of GCC M5_M64R (Methylobacterium extorquens) β-subunit with M64R, L100S, Y143H, D407I, I450V, W502R mutations
[0340] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGREKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0341] SEQ ID NO: 7
[0342] Amino acid sequence of the β-subunit L100N, Y143H, D407I, I450V, W502R mutations of GCC M5_S100N (Methylobacterium extorquens)
[0343] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSNSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0344] SEQ ID NO: 8
[0345] Amino acid sequence of the β-subunit of GCC M5_D407V (Methylobacterium extorquens) with mutations L100S, Y143H, D407V, I450V, W502R
[0346] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYVVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPL
[0347] SEQ ID NO: 9
[0348] Amino acid sequence of the β-subunit of GCC M5_T495N (Methylobacterium extorquens) with mutations L100S, Y143H, D407I, I450V, T495N, W502R
[0349] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRNKEMEQPRKKHDNIPL
[0350] SEQ ID NO: 10
[0351] Amino acid sequence of the β-subunit L100S, Y143H, D407I, I450V, W502H mutations of GCC M5_W502H (Methylobacillus extorquens)
[0352] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPHKKHDNIPL
[0353] SEQ ID NO: 11
[0354] Amino acid sequence of the β-subunit of GCC M5_L510F (Methylobacillus thiooxidans) with mutations L100S, Y143H, D407I, I450V, W502R, L510F
[0355] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSSSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGHGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYIVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKVAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPRKKHDNIPF
[0356] SEQ ID NO: 12
[0357] Amino acid sequence of the β-subunit of PCC WT (Methylobacterium extorquens)
[0358] MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEEFDMFVQHRSTDFGMEKQKIPGDGVVTGWGTVNGRTVFLFSKDFTVFGGSLSEAHAAKIVKVQDMALKMRAPIIGIFDAGGARIQEGVAALGGYGEVFRRNVAASGVIPQISVIMGPCAGGDVYSPAMTDFIFMVRDTSYMFVTGPDVVKTVTNEVVTAEELGGAKVHTSKSSIADGSFENDVEAILQIRRLLDFLPANNIEGVPEIESFDDVNRLDKSLDTLIPDNPNKPYDMGELIRRVVDEGDFFEIQAAYARNIITGFGRVEGRTVGFVANQPLVLAGVLDSDASRKAARFVRFCNAFSIPIVTFVDVPGFLPGTAQEYGGLIKHGAKLLFAYSQATVPLVTIITRKAFGGAYDVMASKHVGADLNYAWPTAQIAVMGAKGAVEIIFRAEIGDADKIAERTKEYEDRFLSPFVAAERGYIDEVIMPHSTRKRIARALGMLRTKEMEQPWKKHDNIPL
[0359] SEQ ID NO: 13
[0360] Amino acid sequence of the α-subunit of PCC WT (Methylobacterium extorquens)
[0361] MFDKILIANRGEIACRIIKTAQKMGIKTVAVYSDADRDAVHVAMADEAVHIGPAPAAQSYLLIEKIIDACKQTGAQAVHPGYGFLSERESFPKALAEAGIVFIGPNPGAIAAMGDKIESKKAAAAAEVSTVPGFLGVIESPEHAVTIADEIGYPVMIKASAGGGGKGMRIAESADEVAEGFARAKSEASSSFGDDRVFVEKFITDPRHIEIQVIGDKHGNVIYLGERECSIQRRNQKVIEEAPSPLLDEETRRKMGEQAVALAKAVNYDSAGTVEFVAGQDKSFYFLEMNTRLQVEHPVTEMITGLDLVELMIRVAAGEKLPLSQDQVKLDGWAVESRVYAEDPTRNFLPSIGRLTTYQPPEEGPLGGAIVRNDTGVEEGGEIAIHYDPMIAKLVTWAPTRLEAIEAQATALDAFAIEGIRHNIPFLATLMAHPRWRDGRLSTGFIKEEFPEGFIAPEPEGPVAHRLAAVAAAIDHKLNIRKRGISGQMRDPSLLTFQRERVVVLSGQRFNVTVDPDGDDLLVTFDDGTTAPVRSAWRPGAPVWSGTVGDQSVAIQVRPLLNGVFLQHAGAAAEARVFTRREAELADLMPVKENAGSGKQLLCPMPGLVKQIMVSEGQEVKNGEPLAIVEAMKMENVLRAERDGTISKIAAKEGDSLAVDAVILEFA
[0362] SEQ ID NO: 27
[0363] Amino acid sequence of biotin ligase BirA (Bacillus methylotrophicus)
[0364] MQFRLSQAARSEGHRLHSHDRLDSTNSEAMRLAQGGETGPLWVTTQRQEAGRGRRGNAWTSPEGNLAASLLMPVAGVAPEMVATLGFVAGVALVDALRDACRLALSPRAEMGDAPLPQGIAPSAAAIHLKWPNDVLADG QKLAGILLEAETLPGGRRAVVVGFGVNVAAAPDGLPYPAAALAAFSAADAPMLLEFLSERFVEAVRIWNKGRGFSNIRRRWLERAAGVGAPVSVRMAEVTLTGIFETIDEGGRLVILAPDGTRRTVTAGEVHFGSAATAA
[0365] nucleotide sequence
[0366] SEQ ID NO: 14
[0367] Nucleotide sequences encoding the L100S, Y143H, D407I, I450V, and W502R mutations in the β-subunit of GCC M5 (Thunb.)
[0368]
[0369] SEQ ID NO: 15
[0370] The nucleotide sequence encoding the mutations in the β-subunits L8F, L100S, Y143H, D407I, I450V, and W502R of GCC M5_L8F (Thunb.)
[0371]
[0372] SEQ ID NO: 16
[0373] Nucleotide sequences encoding the GCC M5_G20R (Twisted Methylbacterium) β-subunit G20R, L100S, Y143H, D407I, I450V, W502R mutations.
[0374]
[0375] SEQ ID NO: 17
[0376] Nucleotide sequences encoding the mutations in the β-subunits M53E, L100S, Y143H, D407I, I450V, and W502R of GCC M5_M53E (Thunb.)
[0377]
[0378] SEQ ID NO: 18
[0379] The nucleotide sequence encoding the mutations in the β-subunits M53Q, L100S, Y143H, D407I, I450V, and W502R of GCC M5_M53Q (Thunb.)
[0380]
[0381] SEQ ID NO: 19
[0382] Nucleotide sequences encoding the mutations in the β-subunits M64R, L100S, Y143H, D407I, I450V, and W502R of GCC M5_M64R (Twisted Methylbacterium).
[0383]
[0384] SEQ ID NO: 20
[0385] The nucleotide sequence encoding the mutations in the β-subunits L100N, Y143H, D407I, I450V, and W502R of GCC M5_S100N (Twisted Methylbacterium).
[0386]
[0387] SEQ ID NO: 21
[0388] The nucleotide sequence encoding the mutations in the β-subunits L100S, Y143H, D407V, I450V, and W502R of GCC M5_D407V (Thunb.)
[0389]
[0390] SEQ ID NO: 22
[0391] The nucleotide sequence encoding the mutations in the β-subunits L100S, Y143H, D407I, I450V, T495N, and W502R of GCC M5_T495N (Tremyla truncatella).
[0392]
[0393] SEQ ID NO: 23
[0394] Nucleotide sequence encoding the mutations in the β-subunits L100S, Y143H, D407I, I450V, and W502H of GCC M5_W502H (Twisted Methylbacterium).
[0395]
[0396] SEQ ID NO: 24
[0397] The nucleotide sequence encoding the mutations in the β-subunits L100S, Y143H, D407I, I450V, W502R, and L510F of GCC M5_L510F (Thunb. methylbacterium).
[0398]
[0399] SEQ ID NO: 25
[0400] Nucleotide sequence encoding the β-subunit of PCC WT (Methylobacterium twistutum)
[0401]
[0402] SEQ ID NO: 26
[0403] Nucleotide sequence encoding the α-subunit of PCC WT (methylbacterium twistutum)
[0404]
[0405] SEQ ID NO: 28
[0406] Nucleotide sequence encoding biotin ligase BirA (Methylobacterium extorquens)
[0407] ATGCAGTTCCGGCTAAGTCAGGCGGCCCGGTCCGAGGGGCATCGGCTCCACAGCCACGACCGGCTCGACTCGACCAACAGCGAGGCCATGCGCCTCGCCCAGGGCGGCGAGACGGGCCCGCTCTGGGTCACGACCCAACGCCAGGAGGCGGGCCGCGGCCGACGCGGCAACGCCTGGACCTCGCCGGAGGGCAACCTCGCCGCGAGCCTGCTGATGCCGGTGGCCGGCGTGGCGCCGGAGATGGTGGCCACGCTGGGCTTCGTGGCGGGCGTGGCGTTGGTGGACGCTCTGCGCGATGCGTGCCGTTTAGCCCTCTCCCCGCGTGCGGAGATGGGGGATGCACCGCTGCCGCAGGGCATTGCTCCTAGCGCTGCCGCCATCCATCTGAAATGGCCCAACGACGTCCTCGCCGACGGCCAAAAACTCGCCGGCATCCTGCTGGAGGCCGAGACGCTGCCGGGCGGGCGGCGGGCGGTGGTGGTCGGTTTCGGCGTCAACGTCGCGGCGGCGCCGGACGGTTTGCCCTATCCGGCGGCGGCGCTTGCGGCTTTTTCCGCGGCGGATGCGCCGATGCTGCTCGAATTCCTGTCGGAACGGTTTGTCGAAGCCGTCCGGATCTGGAACAAGGGGCGCGGATTCTCGAATATCCGGCGGCGGTGGCTGGAGCGCGCGGCGGGGGTAGGGGCCCCCGTGTCGGTTCGCATGGCCGAAGTCACGCTGACGGGCATCTTTGAAACGATCGACGAGGGGGGGCGGCTCGTGATCCTCGCCCCGGACGGAACCCGACGAACCGTGACGGCGGGCGAGGTGCATTTCGGAAGTGCCGCGACGGCGGCCTGA
[0408] The invention will be further described with reference to the following non-limiting drawings and embodiments. Brief description of the attached diagram
[0410] Figure 1 Mutant activity: hydroxyacetyl-CoA carboxylation
[0411] The rate of hydroxyacetyl-CoA carboxylation was measured spectrophotometrically using CaMCR as the coupling enzyme. Error bars represent standard deviation. All measurements were repeated three times. M5: GCC M5, G20R: mutant G20R, S100N: mutant S100N (identical to L100N used in this paper).
[0412] Figure 2 Mutant activity: ATP / CO2 ratio
[0413] CaMCR was used as the coupling enzyme, and the ratio of ATP consumption to hydroxyacetyl-CoA carboxylation was measured spectrophotometrically. Error bars represent standard deviation. All measurements were repeated three times. M5: GCC M5, G20R: mutant G20R, S100N: mutant S100N (identical to L100N used in this paper).
[0414] Figure 3 Sequence alignment between PCC, GCC-M5 and its mutant variants G20R and S100N
[0415] Sequence alignments were generated using Clustal Omega (version 1.2.4) between PCC, GCC M5, and its mutant variants G20R and S100N (identical to L100N used in this paper). GCC_M5 has the following mutations relative to PCC: L100S (identical to the term S100S), Y143H, D407I, I450V, and W502R (highlighted). GCC_M5_G20R has the M5 mutation (highlighted) and an additional G20R mutation (bold highlighted). GCC_M5_S100N has the M5 mutation (highlighted) and an additional S100N mutation (bold highlighted). (Astro) () indicates a position with completely conserved residues, and a colon (:) indicates a position with non-conserved but highly similar residues (score > 0.5 in the Gonnet PAM 250 matrix).
[0416] Figure 4Site-saturated library screening based on lysis buffer. Library M5-G20X contains a GCC M5 variant with a substitution at position 20, and library M5-S100X contains a GCC M5 variant with a substitution at position 100 (detailed description of library construction is in the Methods section). Solid dots represent samples with insignificant activity, which were not considered in subsequent analyses. Hollow triangles, squares, and circles represent samples that are active in the GCC M5 background, carrying G20R (triangle), S100N (square), or S100S (circle) residues, respectively. Therefore, S100S corresponds to GCC M5. The y-axis represents the initial slope of the decrease in absorbance at 340 nm within the first 500 seconds of the reaction, which indicates the carboxylation rate. A total of 367 samples were measured, including 324 mutagenized variants, 20 positive controls with GCC M5-G20R (no mutations added except for the G20R mutation compared to GCC M5), 15 positive controls with GCC M5-S100N (no mutations added except for the S100N mutation compared to GCC M5), and 8 negative controls with lysis buffer (CelLytic B; Sigma Aldrich). Because this assay is based on cell lysates and the amount of GCC M5 protein loaded was not quantified / normalized, it primarily provides qualitative results. The assay showed that only S100S, S100N, and G20R contained the associated GCC activity levels; all other assays showed comparable substitutions at positions 20 and 100 to the negative controls.
[0417] Figure 5Structural analysis of the novel GCC M5 variants G20R and S100N. A) and B) show surface representations of cryo-electron microscopy electron density maps of the G20R (EMD-17777) and S100N (EMD-17778) variants, respectively. The β subunit is shown in dark gray, while the partial electron density of the α subunit is shown in light gray. C) The location of the G20R substitution (PDB 8PN7) is shown on the surface of the β subunit core (left panel), and a magnified view shows the location of Arg20, close to the binding site of the adenosine moiety of hydroxyacetyl-CoA (right panel). D) A magnified view of the active site (PDB 8PN8) of the S100N variant shows the environment of His143 and its probable interaction with hydroxyacetyl-CoA. The modeling of hydroxyacetyl-CoA corresponds to methylmalonyl-CoA in PDB 1ON3, with additional manual fitting to reflect the binding of coenzyme A in the cryo-electron microscopy structure. Manually fitted carboxybiotin is shown at its most likely position for carboxyl transfer to the substrate. His143 interacts polarly with hydroxyacetyl-CoA and Asp171. The amide group of S100N is positioned parallel to the imidazole ring of His143 at a distance of 3.9 Å.
[0418] Figure 6 Mass spectrophotometric analysis of novel GCC variants. Mass spectrophotometric (MP) data for GCC M5 (A), GCC M5 G20R variant (B), GCC M5 S100N variant (C), and PCC from *Tricholoma mater* (D) are shown. Each peak represents the absolute number of protein complexes of a specific size, thus the distribution of different protein complexes in each sample is shown in the figure. All variants show a broad distribution of complexes with varying numbers of α subunits attached to the β6 core, highlighting transient interactions between subunits. The G20R variant appears to favor the formation of the β6α6 complex, as the peak formed by this complex is larger than the others. The S100N variant appears to slightly favor the formation of the β6α6 complex compared to GCC M5. The PCC wild-type did not show a significant preference for any oligomeric state. All samples show peaks at ~70 kDa, representing inactive monomeric β subunits.
[0419] Figure 7The efficiency of the tartrate-CoA (TaCo) pathway relative to natural plant photorespiration using different GCC variants. Relative theoretical yields were calculated by comparing the amount of NADPH and ATP required to generate 3-phosphoglycerate (3PG) from three CO2 molecules (see Methods section, Table 9). It was assumed that 25% of the reaction in the RuBisCO reaction involved oxygenation (Fu et al., 2023, Nature Plants, 9(1), 169-178; Walker, see above). NPR = natural plant photorespiration (occurring during oxygen-producing photosynthesis); TaCo (M5) = TaCo pathway using GCC M5; TaCo (M5 S100N) = TaCo pathway using GCC M5S100N; TaCo (theoret.) = theoretically optimal TaCo pathway, where all ATP hydrolysis is used for 3PG generation.
[0420] Example
[0421] Materials and methods
[0422] Material
[0423] Chemical reagents were purchased from Sigma-Aldrich, Carl Roth GmbH + Co. KG, Santa Cruz Biotechnology Inc., and Merck. Biochemical reagents and materials used for cloning and protein expression were purchased from Thermo Fisher Scientific, New England Biolabs GmbH, and Macherey-Nagel GmbH. Coenzyme A was purchased from Roche Diagnostics. Materials and equipment used for protein purification were purchased from GE Healthcare, BioRad, and Merck Millipore GmbH. Pyruvate kinase / lactate dehydrogenase, malate dehydrogenase, glucose-6-phosphate dehydrogenase, glucose dehydrogenase, and phosphoenolpyruvate carboxylase were purchased from Sigma-Aldrich.
[0424] strain
[0425] All strains used in this study are listed in the table below.
[0426] Table 2 Strains
[0427]
[0428] *E. coli* ElectroMAX DH5α was used to create a randomized mutagenesis library required to generate datasets of randomly mutated GCC variants for training an AI model to predict beneficial mutations. *E. coli* NEB Turbo was used to construct and maintain plasmids with GCC gene site-specific mutations. *E. coli* BL21-birA was derived from *E. coli* BL21 DE3 by introducing a vector carrying the biotin ligase gene from *Tricholoma materia fusiforme*, required for GCC activation. *E. coli* BL21-birA was used for protein overexpression of GCC variants.
[0429] plasmid
[0430] All plasmids used in this study are listed in the table below.
[0431] Table 3 Plasmids
[0432]
[0433] Oligonucleotides
[0434] All oligonucleotides used in this study are listed in the table below.
[0435] Table 4 Oligonucleotides
[0436]
[0437]
[0438] Synthesis of Coenzyme A Ester
[0439] GCC catalyzes the carboxylation of hydroxyacetyl-CoA to tartrate-CoA. To measure this reaction, hydroxyacetyl-CoA was synthesized and purified as previously described (Trudeau, see above; Scheffen, see above). The absorbance at 260 nm (ε = 16.4 mM) was measured. -1 cm -1 To quantify the concentration of coenzyme A ester.
[0440] Construction of random mutation library
[0441] To generate a dataset for artificial intelligence algorithms, a random mutagenesis library of GCC M5 was constructed. A plasmid library of GCC M5 was constructed using macroprimer-based full plasmid PCR (MEGAWHOP) (Miyazaki, see above). To generate random fragments of the GCC M5 (pTE3101) β subunit, error-prone PCR was performed in a 50 µL reaction volume using 2.5 U Taq polymerase and magnesium-free buffer (New England Biolabs; M0320), 7 mM MgCl2, 0.4 mM each of dGTP and dATP, 2 mM each of dCTP and dTTP, 0.4 µM each of primers PccB_fw_P1 and PccB_rv_P1, 10% (v / v) dimethyl sulfoxide, 50 ng of template DNA from pTE3101, and 200–500 µM MnCl2. Random fragments were digested with DpnI (NEB, R0176), purified by agarose gel electrophoresis, and used as macroprimers for whole-plasmid PCR (MEGAWHOP), as described elsewhere (Miyazaki, see above), or for another error-prone PCR reaction to further improve the mutation rate. The MEGAWHOP reaction system (50 µL) contained 1× KOD Hot Start reaction buffer (Novagen), 0.2 mM dNTPs, 1.5 mM MgSO4, 500 ng macroprimers, 50 ng template plasmid (GCC M5; pTE3101), and 2.5 U KOD Hot Start DNA polymerase (Novagen). The MEGAWHOP product was purified, digested with DpnI, and transformed into ElectroMAX DH5α (ThermoFisher Scientific) to ensure an adequate number of transformants in the final library. To estimate the mutation rate of different concentrations of MnCl2 used in error-prone PCR, plasmids from 10 clones randomly selected after MEGAWHOP were purified, sequenced, and subjected to nucleotide exchange analysis.
[0442] Protein expression and purification (including SDS-PAGE)
[0443] For overexpression of GCC M5 and its mutant variants, the corresponding plasmids were transformed into chemically competent Escherichia coli BL21- birACells were grown overnight at 25 °C on Miller lysate agar plates containing 100 µg / mL ampicillin and 50 µg / mL spectinomycin. 8 μL of Golden Miller lysate containing 5 g / L yeast extract, 10 g / L tryptone, 10 g / L NaCl, 17 mM KH₂PO₄, 72 mM K₂HPO₄, and 0.4% glycerol was inoculated from the agar plates and incubated at 37 °C at 140 rpm. When OD... 600 At a pH of 0.4–0.6, protein expression was induced with 500 µM IPTG, and cells were incubated overnight at 25 °C. Cells were collected by centrifugation at 8,000 g at 4 °C for 12 min and lysed by French pressure filtration. His-Trap purification was then performed using an Äkta Start (GE Healthcare) connected to a HisTrap FF column. The purification buffer contained 50 mM HEPES pH 7.8 and 500 mM KCl, and elution was performed using 500 mM imidazole. Protein desalting was performed by gel filtration chromatography using a HiLoad 16 / 600 Superdex 200 pg column (GE Healthcare) and a buffer containing 50 mM HEPES pH 7.8 and 150 mM KCl. Protein quantification was performed by measuring absorbance at 280 nm. Protein purity was verified by SDS-PAGE, with 15 µg of purified protein electrophoresed on a 4–20% Mini-Protean TGX pre-fabricated protein gel (Biorad).
[0444] Enzyme assay
[0445] Enzyme activity assays were performed using three different methods. Randomly mutated GCCs were screened to generate a dataset for training an artificial intelligence algorithm; this screening was conducted on a microplate reader using lysis-based measurements. Mutant variants predicted by the AI algorithm and selected through homology modeling and structural analysis were pre-screened using the same assays. Carboxylation rate and ATP / carboxylation ratio were determined spectrophotometrically using purified enzymes.
[0446] Buffer-based measurements of carboxylation rate and ATP hydrolysis
[0447] GCC-encoded constructs or random mutagenesis libraries of GCC were transformed into *E. coli* BL21_birA (see above), and eight colonies from each construct were picked and inoculated into 96-well PlateOne plates containing lysate (Miller) containing 100 μg / mL ampicillin and 50 μg / mL streptomycin. The plates were incubated overnight at 37 °C, and then transferred to new 96-well PlateOne plates containing lysate (Miller), 100 μg / mL ampicillin, 50 μg / mL spectinomycin, and 2 μg / mL biotin to allow OD to reach the target concentration. 600 The value reaches 0.1. When OD 600 When the pH reached 0.4–0.6, protein expression was induced with 0.25 mM isopropyl β-d-1-thiogalactopyranoside, and cells were incubated overnight at 25°C. Cells were lysed using CelLytic B (Sigma-Aldrich) and stored in 20% glycerol at -80°C. Enzyme activity was measured using a coupled enzyme assay with purified malonyl-CoA reductase (also known as CaMCR) from *Flexibrio citrinum*, as described previously (Scheffen, see above). We used small-volume 384-well plates (Greiner Bio-One) containing 2 µL of cell extract, 100 mM 3-(N-morpholino)propanesulfonic acid (MOPS) pH 7.8, 1 mM ATP, 50 mM KHCO3, 500 μg / mL CaMCR, 1 mM NADPH, 10 mM MgCl2, and 1 mM hydroxyacetyl-CoA in a 10 µL reaction volume. The absorbance of NADPH was measured every 47 seconds at 340 nm in a ELISA reader (TecanInfinite M Plex) at 37 °C for 5 hours.
[0448] Spectrophotometric determination of carboxylation rate
[0449] To determine the carboxylation rate of GCC, a CaMCR coupling assay was performed. 100 mM MOPS (pH 7.8), 50 mM KHCO3, 2 mM ATP, 0.3 mM NADPH, 5 mM MgCl2, 1.8 mg / mL CaMCR from *Flexibrio citrinum*, and 0.01–1 mg / mL GCC were mixed in a cuvette and incubated at 37 °C for 2 min. The reaction was initiated with 0.5 mM hydroxyacetyl-CoA, and absorbance was measured over time at λ = 340 nm.
[0450] Hydroxyacetyl-CoA + HCO3 - + ATP → Tartrate acyl-CoA + ADP + P i(GCC)
[0451] Tartrate-CoA + 2 NADPH → Glyceric acid + CoA + 2 NADP + (CaMCR)
[0452] Spectrophotometric determination of ATP hydrolysis
[0453] To measure the ratio of ATP consumption to carboxylation in GCC, a CaMCR-coupled enzyme assay was performed under ATP-limited conditions. 100 mM MOPS (pH 7.8), 50 mM KHCO3, 0.15 mM ATP, 0.5 mM NADPH, 5 mM MgCl2, 1.8 mg / mL CaMCR from *Flexibrio citrinum*, and 0.05–3 mg / mL GCC were mixed in a cuvette and incubated at 37 °C for 2 min. The reaction was initiated with 0.5 mM hydroxyacetyl-CoA, and absorbance was measured over time at λ = 340 nm. The ATP consumption ratio for each carboxylation reaction was calculated as the ratio of the amount of ATP in the reaction mixture to the amount of NADPH consumed, as reflected by the decrease in absorbance during the reaction.
[0454] Hydroxyacetyl-CoA + HCO3 - + ATP → Tartrate acyl-CoA + ADP + Pi (GCC)
[0455] Tartrate-CoA + 2 NADPH → Glyceric acid + CoA + 2 NADP + (CaMCR)
[0456] Artificial intelligence prediction
[0457] Initially, 5,000 mutant variants were screened, covering approximately 35% of all possible single mutations, but no significant improvement in carboxylation rate or ATP / CO2 ratio was detected.
[0458] To develop GCC variants with improved kinetic properties, an artificial intelligence (AI) algorithm was employed. The AI algorithm for predicting beneficial mutations was performed by Exazyme (Berlin, Germany). The training dataset for the AI model was generated as follows: a random mutagenesis library of pTE3101 (GCC M5) was constructed, transformed into chemically competent *E. coli* BL21-birA, and colonies were picked and inoculated into 96-well PlateOne plates containing lysate broth (Miller) containing 100 µg / mL ampicillin and 50 µg / mL spectinomycin. Expression, lysis, and screening of GCC samples were performed as previously described (Scheffen, see above). 2100 mutant variants were screened, and a representative subset of candidate mutant variants was selected for sequencing. Exazyme's machine learning model was trained using 161 samples to predict beneficial mutations in GCC M5. The resulting list of all possible single mutations, sorted by their efficiency, was used as a template to identify suitable candidates for biochemical characterization through homology modeling and structural studies of promising mutations.
[0459] Structural modeling and analysis
[0460] To evaluate the mutations predicted by the AI algorithm, homology modeling was performed for each promising mutant variant using SWISS-MODEL. The structure of engineered GCC M5 (PDBID 6YBQ) from *Methylobacterium twirlii* was used as a template for GCC mutation homology modeling. Structural analysis of the models was performed using PyMOL (PyMOL Molecular Graphics System; version 1.8; Schrödinger). The active site of hydroxyacetyl-CoA was modeled to GCC based on the position of CoA in the GCC M5 structure and the position of methylmalonyl-CoA in the structure of *Propionibacterium fischeri* methylmalonyl-CoA carboxyltransferase (PDB ID 1ON3; 52% amino acid identity). CoA thioesters were manually fitted and adjusted using COOT and PyMOL to reflect differences in the active site architecture.
[0461] Targeted mutagenesis
[0462] Site-directed mutagenesis was used to construct GCC mutant variants previously predicted by artificial intelligence algorithms and screened through homology modeling and structural analysis. New mutations were introduced via site-directed mutagenesis oligonucleotide PCR, as described in other literature (ShenoyAnal. Biochem. 319, 335-336, 2003). A 25 µL reaction mixture containing 0.5 µM primers, 3% (v / v) dimethyl sulfoxide, 50 ng template DNA (pTE3101), and Phusion high-fidelity PCR premix (NEB, M0531) was used for PCR. 20 U DpnI (NEB, R0176) was then added to the reaction mixture, and the mixture was incubated at 37 °C for 2 h for digestion. 5 µL of the digested product was transformed into chemically competent *E. coli* NEB Turbo cells and streaked onto Miller agar plates containing 50 µg / mL streptomycin. Three to six colonies were selected and cultured in 10 mL of lysate broth (Miller) containing 50 µg / mL streptomycin at 37 °C and 180 rpm for 12 h. The plasmid was then isolated and sequenced to verify the mutagenesis effect.
[0463] Site-saturated mutagenesis library generation:
[0464] GCC M5 plasmid libraries saturated with all amino acids at residues 20 or 100 were constructed by full plasmid PCR using a mixture of primers containing different editing bases. Primers were designed using the 22c-trick to reduce codon redundancy (Kille et al., 2013, ACS Synthetic Biology 2(2), 83-92). For full plasmid PCR, G20X primers (oDM0164, oDM0165, oDM0166, and oDM0170) or L100X primers (oDM0167, oDM0168, oDM0169, and oDM0171) were mixed in a ratio of 12:9:1:22, according to Kille et al. 2013 (Kille et al., ACS synthetic biology, 2(2), 83-92), to ensure that the amount of each primer was equal (Table 4). PCR was performed using a 50 µL reaction mixture containing a 0.5 µM primer mixture, 3% (v / v) dimethyl sulfoxide, 100 ng template plasmid DNA encoding GCC M5 (pTE3101), and Phusion high-fidelity PCR premix (NEB, M0531). The mixture was then digested with 20 U DpnI (NEB, R0176) and incubated at 37 °C for 2 h. After PCR purification using a Macherey-Nagel NucleoSpin gel and PCR purification kit (REF 740609) according to the appropriate experimental protocol, 5 µL of the purified PCR solution was transformed into chemically competent *E. coli* NEB Turbo cells and streaked onto Miller agar plates containing 50 µg / mL streptomycin. Colonies were collected from the plates, and plasmids were isolated. To ensure coverage of all plasmid variants in the library, at least 1300 colonies were collected, equivalent to 65-fold oversampling (meaning that theoretically each protein amino acid is covered approximately 65-fold at either position 20 or 100 of GCC M5). Codon diversity was verified by sequencing the library.
[0465] Cryo-electron microscopy sample preparation and data acquisition
[0466] Add 3 µL of protein solution (1 mg / mL) in 50 mM HEPES pH 7.8 and 150 mM KCl containing 2 mM MgCl2, 1 mM ATP and 4 mM hydroxyacetyl-CoA to a QUANTIFOIL that has been glow-discharged for 90 seconds immediately before use. ®The mesh was placed on an R2 / 1 300 copper mesh and aspirated for 3.5 seconds using a Vitrobot Mark IV (Thermo Scientific) at 100% humidity and 4°C with a suction power of 4. The mesh was then rapidly immersed in liquid ethane cooled by liquid nitrogen for quick freezing and immediately used for data acquisition.
[0467] Cryo-electron microscopy data were acquired using a Titan Krios G3i electron microscope (Thermo Scientific) operating at an accelerating voltage of 300 kV and equipped with a BioQuantum-K3 imaging filter (Gatan). Data were measured in electron counting mode at a nominal magnification of 105,000x (0.837 Å / pixel) and a total dose of 55 e. - / A 2 (55 frames) and aberration-free image shift (AFIS) correction was performed using the EPU (Thermo Scientific). Five images were acquired per foil aperture, with a nominal defocus range of -0.5 to -2.0 µm.
[0468] Cryo-electron microscopy data processing
[0469] All data were processed in CryoSPARC (version 4.1 or 4.2; Punjani et al., 2017, Nat Methods 14, 290-296). For all datasets, dose-segmented videos were normalized, registered, and dose-weighted using Patch Motion correction. The contrast transfer function (CTF) was determined using the Patch CTF procedure. Information regarding cryo-electron microscopy data acquisition, model refinement, and statistics is listed in Table 8.
[0470] Processing of GCC M5 G20R
[0471] Using a speckle picker and manual particle inspection, 837,101 particles (box size 256 pixels) were initially extracted, which were used to construct 2D categories. 2D categories with protein-like features were used to initialize template selection. After inspection and extraction with a 256-pixel box size, a total of 3,439,715 particles were obtained, which were used to construct 2D categories. From the 2D categories, 1,845,969 candidate particles were selected for de novo reconstruction and divided into 4 categories. The particles in the best-aligned category (647,870 particles) underwent non-uniform refinement, along with single-particle defocus optimization, grouped CTF parameter optimization, and EWS correction. This process yielded a global resolution of 2.08 Å and a temperature factor of 66.1 Å. 2The map was then refined locally to achieve a global resolution of 2.03 Å and a temperature factor of 58.9 Å. 2 The resulting map. The map was obtained at -40 Å. 2 B-factor sharpening was applied. Further classification failed to achieve higher resolution.
[0472] Processing of GCC M5 S100N
[0473] Using a speckle picker and manual particle inspection, 620,147 particles (500-pixel box size) were initially extracted, which were used to construct 50 2D categories. 2D categories with protein-like features were used to initialize template selection. After inspection and extraction with a 500-pixel box size, 2,511,911 particles were obtained, which were used to construct 2D categories. A total of 324,129 candidate particles were selected and used for de novo reconstruction and divided into 5 categories. The particles with the best aligned categories (113,824 particles) underwent non-uniform refinement, along with single-particle defocus optimization, grouped CTF parameter optimization, and EWS correction. This process yielded a global resolution of 2.36 Å and a temperature factor of 66.9 Å. 2 The map was then refined locally to achieve a global resolution of 2.31 Å and a temperature factor of 60.6 Å. 2 The resulting map. The map is oriented at -50 Å. 2 B-factor sharpening was applied. Further classification failed to achieve higher resolution.
[0474] Model building and refinement
[0475] Initially, cryo-electron microscopy (cryo-EM) patterns were fitted using GCC M5 (PDB 6YBQ) as a template in UCSF-ChimeraX (v1.6; Pettersen et al., 2020, Protein Science, 301, 70-82). The resulting model was then manually constructed in Coot (v0.9.8.3; Emsley et al., 2010, Structural Biology, 66 4, 486-501). Automatic refinement of the structure was performed using phenix.real_space_refine in the Phenix (v1.20.1) software suite (Liebschner et al., 2019, Structural Biology, 75 10, 861-877). Manual refinement and aqueous phase selection were performed in Coot. Model statistics are listed in Table 8.
[0476] Mass photometry
[0477] Mass spectrophotometry (MP) measurements were performed using a CultureWell™ reusable pad (CW-50R-1.0, 50-3 mm diameter x 1 mm depth) on a microscope coverslip (1.5 H, 24 x 50 mm, Carl Roth). The pad was rinsed three times consecutively with water and isopropanol and then dried under compressed air. The pad was assembled onto the microscope coverslip and placed on the stage of a TwoMP mass spectrometer (MP, Refeyn Ltd, Oxford, UK) after being soaked in oil. Measurements were performed in 1× phosphate-buffered saline (PBS, 10 mM Na2HPO4, 1.8 mM KH2PO4, 137 mM NaCl, 2.7 mM KCl (pH 7.4)). For this purpose, 18 µL of 1× PBS was focused into the MP, followed by the addition of 2 µL of sample (1 µM protein), rapid mixing, and measurement. Shortly before measurement, samples were prepared by diluting purified proteins to a monomer concentration of 1 µM in buffer (50 mM HEPES, pH 7.8, 150 mM KCl), and measurements were taken by absorbance at 280 nm. Data were acquired for 60 seconds at 100 frames per second using AcquireMP (Refeyn Ltd, Oxford, UK). MP contrast was molecularly calibrated using a 50 nM in-house purified protein mixture containing citrate synthase complexes with known molecular weights ranging from 86 to 430 kDa. MP datasets were processed and analyzed using DiscoverMP (Refeyn Ltd, Oxford, UK). Details of MP image analysis have been previously described (Sonn-Segev et al., 2020, Nat Commun 11, 1772).
[0478] Flux balance analysis
[0479] To compare the differences between different versions of the tartrate-CoA pathway and natural plant photorespiration (which occurs during oxygen-producing photosynthesis), we used COBRApy (v0.20.0) (Ebrahim et al., 2013, BMC systemsbiology, 7, 74) with flux balance analysis (FBA) to perform chemometric modeling. We used the same framework previously used to compare the tartrate-CoA (TaCo) pathway (including GCC M5) with other pathways (Scheffen above) and extended it with additional TaCo pathway variants including the GCC M5 S100N provided in this paper. The chemometric models used primarily considered ATP consumption and reducing equivalents for each pathway. This is a widely used standard method in the field of computational analysis of metabolic pathways, providing an indication of the energy efficiency of the pathway being evaluated (Wilbert et al., 2012, Metabolic Engineering, 14 3, 270-280; Berkvens et al., Essays Biochem, April 30, 2024, 68 1, 41-51). Therefore, this analysis focuses on whether the GCC M5 S100N presented in this paper (with its surprisingly low ATP / CO2 consumption, especially as shown in Table 6) is also superior to GCC M5 in the context of the complete TaCo pathway. Therefore, for each pathway, we calculated the consumption of ATP, NAD(P)H, and reduced ferricredoxin, as well as the number of CBB cycles (including RuBisCO) required to produce one unit of 3-phosphoglycerate (3PG). To this end, we constructed a simplified metabolic model consisting of the Calvin-Benson-Barthum (CBB) cycle, specific reactions for each photorespiratory pathway considered, and universal / artificial cofactor regeneration and interconversion reactions (e.g., ADP + inorganic phosphate → ATP; NAD+ → NADH and NADH + NADP+ → NAD+ + NADPH; ATP + AMP → 2 ADP; oxidized ferricredo → reduced ferricredo). We note that in this model, electron transfer from NADH to NADPH and vice versa can occur freely via universal transhydrogenase reactions without any ATP input (simulating NADPH production in photosynthesis). We assume a RuBisCO carboxylation to oxidation ratio of 3:1 (i.e., 25% of all RuBisCO reactions are oxidation reactions) (Walker, see above; Fu, see above; Sharkey, 1988, Physiologia Pantarum, 73 1; 147-152).To compare the yields of all pathways, we calculated the total “ATP equivalent” required to generate 3-phosphoglycerate (3PG) from CO2 using the conversion relationships of 1 NAD(P)H = 2.5 ATP (Ferguson, 1986, Trends in Biochemical Sciences, 11 9, 351-353; Hinkle, 2005, Biochimica et Biophysica Acta (BBA) – Bioenergetics, 1706 1-2, 1-11) and 2 reduced ferroredoxins = 1 NAD(P)H.
[0480] result
[0481] The enzyme GCC M5 was developed for the synthesis of carbon dioxide fixation pathway. It can catalyze the carboxylation of hydroxyacetyl-CoA to tartrate-CoA (Scheffen, see above).
[0482] However, GCC M5 requires approximately 4 ATP units for each carboxylation, although the theoretical ratio is 1. Therefore, further enzyme engineering is needed to provide candidate enzymes with reduced ATP consumption and improved energy efficiency.
[0483] As previously mentioned, by combining random mutagenesis, rational design, directed evolution, and high-throughput screening of mutant variants of GCC M3, improved variants GCC M4 and GCC M5 were generated.
[0484] To continue this approach in new iterations, a random mutagenesis library of GCC M5 was constructed, containing the carboxyltransferase encoding gene of the carboxylase. pccB_M5 Targeting was achieved via error-prone PCR. Large-scale screening of the library was performed as previously described (Scheffen, see above). 5000 mutant variants were screened, covering approximately 35% of all possible single mutations, but no significant increase in carboxylation rate or ATP / CO2 ratio was detected. Therefore, conventional methods for identifying variants with improved kinetic properties failed to yield results.
[0485] Further attempts were made to develop improved GCCs (based on M5), but without any improvement. For example, we tried the following approaches: (i) searching for naturally occurring homologous enzymes carrying certain M5 mutations; (ii) using molecular dynamics simulations to model the reaction mechanism and predict new mutations; (iii) developing an in vivo screening system for GCCs; and (iv) repeating the directed evolution process previously used to generate M5 (Scheffen, see above).
[0486] Although these attempts failed, it was subsequently decided to use the generated screening data to train an artificial intelligence (AI) model to predict beneficial mutations in GCC M5. That is, without the screening data generated in our lab, the AI model could not be developed and the AI could not produce any useful results.
[0487] The AI model ranked all possible single mutations (approximately 10,000 mutations) based on efficiency. However, we decided not to perform biochemical characterization solely on the top-ranked hits. Instead, we further investigated and evaluated the proposed mutation list to identify candidates we believed might truly demonstrate potential for improved kinetic properties for biochemical characterization. From the top 1% of mutations (covering 105 predictions), a homology model based on the PDB structure 6YBQ was constructed, and structural studies were performed on the mutant variants.
[0488] Although the mutation at position 20 is not located at or near the active site, and therefore we initially expected it to have no effect on enzyme activity, we still selected mutations at that position. However, the mutation at position 20 appeared multiple times in the top 1% of predictions, suggesting a potential relevance. We also selected other mutations that we assumed might affect enzyme activity and / or had a high frequency in the top 1% of predictions (the mutations at positions 53 and 510 are examples of the latter). Based on these considerations, we selected seven mutant variants from the top 1% for in vitro testing (Table 5).
[0489] In addition, we selected three mutant variants from the top 5% of the mutations. Mutations at these positions had been tested previously; however, some mutations (such as S100N) failed to show improvements in kinetic properties, and other amino acid substitutions at positions 502 and 407 had also been tested previously (see Scheffen, above). Despite the failure of these previous tests, we decided to test these three selected variants. Previous variants were known not to reduce ATP consumption.
[0490] The analyzed mutations are shown in Table 5 below. Some mutations are in Figure 3 The sequence alignment will be explained separately.
[0491]
[0492]
[0493] In spectrophotometric measurements of hydroxyacetyl-CoA carboxylation activity using cell lysates, all mutant variants except M64R showed significant activity and were purified for subsequent analysis. The lack of activity in M64R indicates that the AI model can only provide a pre-screening of potentially beneficial amino acid substitution sites, but cannot reliably predict whether mutant variants possess activity, let alone whether their activity is actually enhanced. Furthermore, the selection of positions from the AI-generated pre-screening list for biochemical characterization needs to consider factors that AI cannot model.
[0494] The specific activity of hydroxyacetyl-CoA carboxylation and the ATP / CO2 ratio were measured spectrophotometrically (Table 6). The data show that we developed two evolutionary variants of GCC M5. The candidate variant G20R showed a 2-3 fold increase in carboxylation activity, while the candidate variant L100N showed a more than 50% reduction in ATP consumption (see Tables 6 and 6 respectively). Figure 1 and Figure 2 Structural analysis of these mutant variants was performed using cryo-electron microscopy to investigate the roles of these residues in the overall complex and their impact on catalysis. Furthermore, site-directed mutagenesis was used to construct mutant variants G20A, G20H, G20K, and G20Y, which were expected to affect catalysis. Saturation mutagenesis was also performed at positions 20 and 100 to cover all possible mutations at these positions and to further understand the roles of these positions and the residue properties required to improve hydroxyacetyl-CoA carboxylation.
[0495]
[0496] The results show that even among the top 1%, most AI-predicted mutant variants did not exhibit improved kinetic properties, despite these positions having a high frequency among the top 1% of mutations (e.g., positions 53 and 510). Mutations at positions 53 and 510 did not even show detectable activity. Other predicted mutant variants showed activity, but less than the reference GCC M5. Furthermore, mutations at previously occurring positions (e.g., positions 407 and 502) also did not provide improvements in kinetic properties. Therefore, we are surprised that mutations at positions 20 and 100 showed improvements in activity / kinetic properties.
[0497] Using LC-MS, we performed a more detailed biochemical characterization of this enzyme variant and determined its... V max 、k cat and K M The values indicated that its catalytic activity was improved compared to GCC M5 (Table 7). Despite the apparent efficacy of the G20R variant against hydroxyacetyl-CoA... KM The value increased slightly, but V max It is 1.8 times higher than GCC M5, therefore, k cat The G20R variant, with a concentration of 9.8 ± 0.2, shows great promise for GCC-based applications, such as, for example, for the tartrate-CoA pathway at higher production rates (Scheffen, see above). These results further confirm the unexpectedly favorable high enzyme activity of the G20R variant, particularly superior to GCC M5, as shown in Table 6. In this variant, the substitution of arginine for glycine at position 20 on the surface ring of the β-core significantly increases the carboxylation rate, while the ATP ratio required for each carboxylation changes only slightly ( Figure 1 and Figure 2 For example, as shown in Table 6, compared to GCC M5 (which requires 4.0 ± 0.0 ATP per hydroxyacetyl-CoA carboxylation), the S100N variant had a significantly lower ATP / carboxylation ratio of 1.7 ± 0.1 ATP, representing a 60% reduction in reaction energy requirements. To assess whether other mutations at positions 20 and 100 would have a beneficial effect on the catalytic properties of GCC, we tested site-saturated libraries at these two positions and applied a lysis-based screening method. Figure 4 However, apart from the G20R and S100N or S100S (i.e., GCC M5) variants, we did not detect any other active substitutions. These findings strongly support the view that substitutions at positions 20 and / or 100 of GCC M5 (especially with Arg and Asn, respectively) could not have been predicted to produce enzymatically active GCC variants (let alone GCC variants with even improved properties, particularly compared to GCC M5, such as enzyme activity and ATP / CO2 ratio).
[0498] To identify the structural changes that may lead to improved catalytic activity of the G20R or S100N variants, we resolved their cryo-electron microscopy structures at 2.05 Å and 2.31 Å, respectively. Figure 5 A and 5B). The G20R substitution is located approximately 8 Å from the 3'-phosphate group of coenzyme A, and surprisingly, it does not appear to interact with other residues or substrates. Figure 5 C). The cryo-electron microscopy structure lacks well-defined contact points, reflected in the weak electron density observed in the Arg20 side chain, indicating its high flexibility. Arg20 is located in a flexible ring, preceded by two glycine residues, further enhancing its flexibility. Based on the position of G20R at the upper edge of the β6 core of GCC, we speculate that it may contribute to stabilizing the interaction with the α subunit and indirectly promote the localization of coenzyme A. Figure 5C). In fact, in mass spectrophotometric (MP) measurements, the G20R variant formed more and higher quality complexes compared to GCC M5, indicating that its complex formation is more stable. Figure 6 (A and 6B). Recent studies have shown that the formation of the propionyl-CoA carboxylase complex is a dynamic process, with an equilibrium between complex assembly and dissociation, and that β6α6 is the most active (Lee et al., 2023, J Struct Biol X, 7, 100088). Therefore, a higher proportion of α subunits bind to the β6 core may explain the higher in vitro activity of the G20R mutant. Thus, it seems plausible that the use of the GCC M5 G20R (or its variants) presented herein in metabolic pathways such as the tartrate-CoA pathway would result in higher product formation or a reduced metabolic burden in living organisms (compared to the same metabolic pathway containing GCC M5 but not the GCC M5 G20R variant), due to the smaller amount of enzyme required to achieve comparable reaction rates.
[0499] The S100N substitution is located on the periphery of the active site, a position previously targeted in GCC M5 engineering efforts where this site was engineered to be a serine residue. The Asn100 substitution is closely adjacent to His143, which is believed to coordinate with the hydroxyl group of hydroxyacetyl-CoA. Figure 5 D). Although the active site superposition diagrams of the S100N variant and GCC M5 are almost identical, we hypothesize that Asn100 forces His143 to form a more favorable rotatoric isomer conformation for substrate binding and catalysis, thereby making the carboxylation reaction more efficient. This hypothesis is supported by the fact that His143 is classified as a rotatoric isomer outlier in the cryo-electron microscopy structures of all subunits of S100N, while this is not the case in the structures of GCC M5 or G20R variants. This slight shift of the His143 side chain toward hydroxyacetyl-CoA may help improve substrate orientation / localization, thereby reducing ineffective decarboxylation of carboxybiotin (i.e., CO2 is released from carboxybiotin but not transferred to the substrate). This, in turn, reduces the energy requirement of the reaction on ATP. Although the S100N variant showed a slightly increased proportion of higher molecular weight oligomeric complexes in MP experiments ( Figure 6 C), but its complex distribution is still similar to that of GCC M5 and PCC (C). Figure 6 A and D). All variants studied, including the PCC wild-type ( Figure 6In Mp measurements, stable β6 cores are formed and bound with varying numbers of α subunits. However, as mentioned above, the formation of the GCC enzyme complexes disclosed herein is a dynamic process (meaning that subunits may dissociate). Therefore, further improvements in the stability of such complexes may be of significant importance (particularly for their efficient purification and / or in vitro application). To further improve the formation of active GCC complexes with fully bound α subunits (β6α6), for example, the expression of α subunits relative to β subunits could be increased to saturate the β core.
[0500] Table 7. Kinetic characteristics of GCC M5 variants G20R and S100N.
[0501]
[0502] The data are represented as mean ± standard deviation, determined by nonlinear regression from n=18 independent measurements.
[0503] Table 8. Cryo-electron microscopy data collection, refinement, and model statistics.
[0504]
[0505]
[0506] After demonstrating improvements in carboxylation rate and ATP consumption in GCC M5 G20R and S100N in in vitro experiments, we further analyzed their potential to enhance natural CO2 fixation in plants via the tartrate-CoA pathway (Scheffen, see above). Using flux balance analysis (Orth et al., 2010, Nat Biotechnol 28, 245-248), we assessed the energy requirements of natural plant photorespiration and the combination of different versions of the tartrate-CoA (TaCo) pathway (Scheffen, see above) with the Calvin-Benson-Bassam (CBB) cycle of photosynthesis. Figure 7 (Table 9). This model considers the demand for ATP and NADPH and predicts the theoretical yield of 3PG, a product of the CBB cycle generated from three CO2 molecules. Since GCC M5 G20R leads to improved kinetics, but flux balance analysis (FBA) only considers stoichiometric data, we limited the analysis to tartrate-CoA pathway versions of theoretical GCCs containing existing technology GCC M5, novel variant GCC M5 S100N, or those without ineffective ATP hydrolysis. The model reveals an improvement in the theoretical 3PG yield for all TaCo versions, with a 20% improvement in the TaCo pathway containing GCC M5 and a 28% improvement in the version containing GCC S100N, very close to the theoretical maximum of a 30% improvement. Figure 7 (Table 9). The tartrate-CoA pathway transforms photorespiration into a process of assimilation rather than CO2 release, thereby doubling its carbon use efficiency from 75% to 150% (Bar-Even, 2018, Plant Science, 273, 71-83; Scheffen, see above; Trudeau, see above; Marchal et al., 2023, ACS Synth. Biol. 12 12, 3521-3530). In terms of energy efficiency, coupling the CBB cycle with the TaCo pathway with GCC S100N reduces ATP by 18% and reducing equivalent by 24% compared to natural plant photorespiration; compared to the TaCo pathway with GCC M5, it reduces ATP by 13% ( Figure 7 (See Table 9). Therefore, the GCC M5 S100N variant (or GCC variants carrying S100N substitution) presented in this paper not only exhibits a surprisingly favorable ATP / CO2 ratio (i.e., less ATP required per carboxylation reaction), but computer simulations show that its carboxylation rate is effectively increased by approximately 28% when applied in the TaCo pathway. Therefore, the energy efficiency of plants containing the TaCo pathway (which replaces or supplements the plant's natural photorespiration pathway) with the GCC M5S100N presented in this paper should be increased by 28%, thereby increasing biomass yield by approximately 28%—an increase of significant importance for fields such as agriculture and many others.
[0507] Table 9. Comparison of natural photorespiration and the TaCo pathway, which is modified from (Scheffen, see above).
[0508]
[0509] a The conversion of 2-phosphoglycolic acid to 3-phosphoglyceric acid
[0510] b The two reactions of TCR (tartrate acyl-CoA reductase) are considered as two independent enzymes.
[0511] c The required "ATP equivalent" is calculated using the following formula: 1 NAD(P)H = 2.5 ATP (Ferguson, 2010, Proceedings of the National Academy of Sciences of the United States of America, 107 39, 16755-16756; Hinkle, see above), 2 reduced ferricoxanes (1 Fd... 2- = 1 NAD(P)H
[0512] The bold figures are the results of flux balance analysis, calculated by converting 3x CO2 net to 1 unit of 3-phosphoglycerate through the CBB cycle and the corresponding photorespiration pathway (see the Methods section above for more details).
[0513] Abbreviation: Fd 2- = 2 reduced ferricoxins; NPR = plant natural photorespiration (glycine decarboxylation pathway); TaCo: tartrate coenzyme A pathway; M5 = including the ineffective ATP hydrolysis of the enzyme "GCC(M5)" in the TaCo pathway (generating 4.0 ATP per carboxylation reaction); M5+S100N = including the ineffective ATP hydrolysis of the optimized enzyme "GCC M5 S100N" in the TaCo pathway (generating 1.7 ATP per carboxylation reaction).
[0514] All references cited herein are incorporated herein by reference in their entirety. This invention has now been fully described, and those skilled in the art will understand that it can be practiced under broad and equivalent conditions, parameters, etc., without affecting the spirit or scope of the invention or any of its embodiments.
Claims
1. A hydroxyacetyl-CoA carboxylase (GCC), wherein the GCC is characterized by: It contains an amino acid sequence that has at least 60% sequence identity with SEQ ID NO: 1, and It has one or more amino acid substitutions, deletions, or insertions at the position selected from the group consisting of positions 20 and 100 in the amino acid sequence shown in SEQ ID NO: 1, or at a position corresponding to any of these positions. Preferably, the GCC has improved activity compared to the reference GCC. Preferably, the reference GCC comprises the amino acid sequence of SEQ ID NO:
1.
2. The GCC according to claim 1, wherein (1) The amino acid at position 20 or at the corresponding position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by arginine. Preferably, the substituted amino acid is glycine; and / or (2) The amino acid at position 100 or at the corresponding position in the amino acid sequence shown in SEQ ID NO: 1 is replaced by asparagine. Preferably, the amino acid being replaced is serine.
3. The GCC according to claim 1 or 2, wherein the improved activity is improved energy efficiency.
4. The GCC according to claim 3, wherein the improved energy efficiency is a reduction in ATP consumption, preferably a reduction of at least 50%.
5. The GCC according to any one of claims 1 to 4, wherein the ATP / CO2 ratio of the GCC is less than about 4.00 ± 0.
02.
6. The GCC according to any one of claims 1 to 5, wherein the ATP / CO2 ratio of the GCC is about 1.71 ± 0.
11.
7. The GCC according to any one of claims 1 to 6, wherein the improved activity is an increase in carboxylation activity, preferably at least 2 to 3 times.
8. The GCC according to claim 7, wherein the increase in carboxylation activity is due to an increase in the conversion of hydroxyacetyl-CoA to tartrate-CoA.
9. The GCC according to any one of claims 1 to 8, wherein the GCC catalyzes the carboxylation of hydroxyacetyl-CoA (conversion of hydroxyacetyl-CoA to tartrate-CoA) at a rate greater than about 5.6 ± 0.3 s. -1 .
10. The GCC according to any one of claims 1 to 9, wherein the specific activity of said GCC is higher than about 937 ± 40 nmol hydroxyacetyl-CoA min. -1 mg -1 .
11. The GCC according to claim 10, wherein the specific activity of the GCC is about 2602 ± 430 nmol hydroxyacetyl-CoA min. -1 mg -1 .
12. A nucleic acid molecule encoding a GCC according to any one of claims 1 to 11.
13. A carrier comprising the nucleic acid molecule according to claim 12.
14. An organelle comprising a nucleic acid molecule according to claim 12 or a carrier according to claim 13.
15. A host cell comprising a nucleic acid molecule according to claim 12, a vector according to claim 13, or an organelle according to claim 14.
16. A tissue comprising the host cells according to claim 15.
17. An organism comprising a nucleic acid molecule according to claim 12, a carrier according to claim 13, or an organelle according to claim 14.
18. The organism according to claim 17, wherein the organism is a plant, algae, or microorganism.
19. The organism according to claim 18, wherein the microorganism is bacteria.
20. The organelle of claim 14, the host cell of claim 15, the tissue of claim 16, or the organism of any one of claims 17 to 19, wherein the organelle, the host cell, the tissue, or the organism converts hydroxyacetyl-CoA to tartrate-CoA at a higher rate than the corresponding organelle, host cell, tissue, or organism containing a reference GCC, preferably, wherein the reference GCC contains SEQ ID NO:
1.
21. The organelle of claim 14, the host cell of claim 15, the tissue of claim 16, or the organism of any one of claims 17 to 19, wherein the growth rate and / or carbon production rate of the organelle, the host cell, the tissue, or the organism is higher than that of the corresponding organelle, host cell, tissue, or organism containing a reference GCC, preferably, wherein the reference GCC contains SEQ ID NO:
1.
22. Use of the GCC according to any one of claims 1 to 11, the organelle according to claim 14, the host cell according to claim 15, the tissue according to claim 16, or the organism according to any one of claims 17 to 19 for the process of CO2 fixation and / or photosynthesis.
23. Use of the GCC according to any one of claims 1 to 11, the organelle according to claim 14, the host cell according to claim 15, the tissue according to claim 16, or the organism according to any one of claims 17 to 19 for the decomposition of synthetic materials.
24. The use according to claim 23, wherein the synthetic material is polyester.
25. The use according to claim 24, wherein the polyester is polyethylene terephthalate.
26. Use of the GCC according to any one of claims 1 to 11, the organelle according to claim 14, the host cell according to claim 15, the tissue according to claim 16, or the organism according to any one of claims 17 to 19 for the decomposition of ethylene glycol and / or glycolic acid.
27. A method for producing biomass using a GCC according to any one of claims 1 to 11, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
28. A method for decomposing synthetic materials using a GCC according to any one of claims 1 to 11, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
29. The method according to claim 28, wherein the synthetic material is polyester.
30. The method of claim 29, wherein the polyester is polyethylene terephthalate.
31. A method for decomposing ethylene glycol and / or glycolic acid using a GCC according to any one of claims 1 to 11, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
32. A composition comprising a GCC according to any one of claims 1 to 11, a nucleic acid according to claim 12, a vector according to claim 13, an organelle according to claim 14, a host cell according to claim 15, a tissue according to claim 16, or an organism according to any one of claims 17 to 19.
Citation Information
Patent Citations
Carbon-neutral and carbon-positive photorespiration bypass routes supporting higher photosynthetic rate and yield
WO2016207219A1