Enzyme selection systems

WO2026114812A1PCT designated stage Publication Date: 2026-06-04DANMARKS TEKNISKE UNIV

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DANMARKS TEKNISKE UNIV
Filing Date
2025-11-24
Publication Date
2026-06-04

Smart Images

  • Figure EP2025084022_04062026_PF_FP_ABST
    Figure EP2025084022_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to enzyme selection systems for glycine amidinotransferase and guanidinoacetate methyltransferase, and the use of these optimized enzyme in fermentation processes. In particular, the present invention relates to the production of creatine, using such enzymes, as well as cells optimized for creatine production.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 83737PC01

[0002] 1

[0003] Enzyme selection systems

[0004] Technical field of the invention

[0005] The present invention relates to enzyme selection systems, and the use of these optimized enzyme in fermentation processes. In particular, the present invention relates to the production of creatine, using such enzymes, as well as cells optimized for creatine production.

[0006] Background of the invention

[0007] Creatine is a molecule that is used to store energy in vertebrates. It plays important functions in muscle growth and performance, and is one of the most popular sports supplements. It has also been found to enhance neurophysiological performance, especially in elderly people. The daily requirement for creatine is ~2 g / d. In sports nutrition, creatine supplementation is recommended at 3-5 g / d. Humans derive creatine from both endogenous synthesis (~1 g / d) and creatine- containing diet. Importantly, creatine is not found in plants and plant-based food unless supplemented, heightening the benefit / need for creatine supplementation for people with vegan and vegetarian diets.

[0008] The creatine currently available on the market is produced through chemical synthesis, mostly from cyanamide and sarcosine, both of which are derived from petroleum-based compounds and are thus not environmentally friendly. Cyanamide is further a toxic compound. In a 2011 study, derivatives of cyanamide (dicyanamide, dihydrotriazine) have been detected in creatine supplements on the market, in some cases exceeding the EFSA (the European Food Safety Authority) recommendation for maximum intake.

[0009] In living organisms, creatine is synthesized from arginine, glycine, and S- adenosylmethionine (SAM) with two enzymatic steps: 1) glycine a midi notransferase (GAT) to convert arginine and glycine into ornithine and guanidinoacetate (GAA); 2) guanidinoacetate methyltransferase (GAMT) to add a methyl group from SAM to GAA, forming creatine. Enzymatic synthesis of creatine, either based on microbial cell factories or enzymatic catalysis, can replace the problematic chemical synthesis methods with the potential for higher safety and lower environmental cost. 83737PC01

[0010] 2

[0011] Growth coupling, a high-throughput screening approach, aims to couple the activity of the target enzyme to cell growth. This is achieved through metabolic coupling, where the metabolic network is rewired such that in a specific growth media, the energy, biomass, or some essential metabolite relies on a certain pathway, which the target enzyme is part of. This allows an arbitrarily large library to be "screened" with minimal effort and cost.

[0012] Hence, improved variants of enzymes applicable for use in bacterial creatine production would be advantageous, and in particular a more efficient and / or reliable process of screening such enzymes would be advantageous.

[0013] Summary of the invention

[0014] Thus, an object of the present invention relates to selection systems for selecting optimized enzymes, for use in fermentation processes such as creatine production, as well as the provision of such optimized enzymes.

[0015] In particular, it is an object of the present invention to provide enzyme mutations that solves the above mentioned problems of the prior art with inefficient fermentation of creatine.

[0016] The Examples of the present disclosure shows ways of selecting optimized variants, and the use of such variants to produce creatine.

[0017] Example 1 provides an overview of the materials and methods used in the following examples.

[0018] Firstly, example 2, provides that it is possible to generate creatine in E. coli by transforming the cells with GAT and GAMT coding sequences, however that the production is far from optimal.

[0019] Example 3 and 4 provide a growth-coupled selection system for GAT, and shows optimized variants of OCD.

[0020] Example 5 and 6 provides a GAT selection using auxotrophy, and improved variants of GAT, derived through the auxotrophy based selection system Example 7 provides an optimized fermentation of creatine, wherein the novel GAT variants are co-expressed with GAMT and shows 58% improvement compared to 83737PC01

[0021] 3

[0022] GAT8_wt. Moreover, we utilized an adaptive evolution to alleviate product inhibition of the cells.

[0023] In summary, an optimized GAT growth-coupled strategy to improve the efficiency of GAT was initially designed. Additionally, the strain was engineered for improved creatine tolerance. The resulting strain reached more than 50% titre increase compared to the starting strain.

[0024] Thus, in one aspect of the present disclosure a bacterial cell for selecting optimized variants capable of catalysing the reaction arginine < = > ornithine is provided, in such an aspect the bacterial cell comprises:

[0025] • a nucleic acid sequence encoding an enzyme, such as Glycine a midi notransferase (GAT) or a functional variant thereof, capable of catalysing the reaction arginine < = > ornithine, wherein the enzyme is an enzyme selected from any of the EC numbers EC 2.1.4.1, EC 2.1.4.3, or EC 3.5.3.1; and

[0026] • a nucleic acid encoding o an enzyme, such as Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of catalysing the reaction ornithine < = > proline, wherein the enzyme is an enzyme according to EC 4.3.1.12, or o an enzyme, such as Ornithine aminotransferase (OAT) or a functional variant thereof, capable of catalysing the reaction ornithine < = > glutamate+ semialdehyde, wherein the enzyme is an enzyme according to EC 2.6.1.13, and wherein pathways for generating proline and ornithine are downregulated or deleted.

[0027] Another aspect relates to a process for selecting mutants of GAT and / or OCD / OAT, the process comprising: i. Providing a bacterial host cell comprising

[0028] ■ a nucleic acid sequence encoding an enzyme, such as Glycine a midi notransferase (GAT) or a functional variant thereof, capable of catalysing the reaction arginine < = > ornithine, wherein the enzyme is an enzyme selected from any of the EC numbers EC 2.1.4.1, EC 2.1.4.3, or EC 3.5.3.1; and a nucleic acid encoding 83737PC01

[0029] 4

[0030] • an enzyme, such as Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of catalysing the reaction ornithine < = > proline, wherein the enzyme is an enzyme according to EC 4.3.1.12, or

[0031] • an enzyme, such as Ornithine aminotransferase (OAT) or a functional variant thereof, capable of catalysing the reaction ornithine < = > glutamate+ semialdehyde, wherein the enzyme is an enzyme according to EC 2.6.1.13, ii. optionally, the GAT as described herein and 1) a nucleic acid encoding an OCD as described herein, or 2) a nucleic acid encoding an OAT as described herein; iii. Downregulating or deleting pathways for generating proline or ornithine, wherein the pathways for generating proline or ornithine as described herein, or optionally if only mutants of OCD are selected the pathways may be encoded by astABCDE and speAB; iv. Growing the cells for a time sufficient to obtain improved variants of GAT and / or OCD / OAT; v. Identifying said improved variants.

[0032] Another aspect relates to the production of creatine, and thus in another aspect is provided a bacterial cell, such as e. coli, comprising one or more of the following features; i. a nucleic acid encoding a GAT as described herein; ii. a nucleic acid encoding a GAMT as described herein; iii. A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; iv. A nusA mutation, such as A234E R104L, and / or R228P; v. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; vi. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or vii. a combination of the above. 83737PC01

[0033] 5

[0034] Another aspect of the disclosure relates to a process for selecting bacterial strains with increased tolerance to creatine, the process comprising: i. providing an E. coli with a rewired SAM-cycle and being deficient of the natural cysteine synthesis; ii. introducing into the bacterial cell: o a nucleic acid sequence encoding GAT as described herein; and o a nucleic acid sequence encoding GAMT as described herein; iii. optionally, the e. coli may comprise mutations selected from the group consisting of; o A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; o A nusA mutation, such as A234E R104L, and / or R228P; o Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; o One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or o a combination of the above. iv. Growing the bacterial cell in conditions with substantially no cysteine being present, for a time sufficient to obtain strains with increased tolerance to creatine; and v. optionally, isolating and / or identifying said strains with increased tolerance to creatine.

[0035] Other aspects disclosed herein relates to the optimized enzyme variants of GAT, OCD, and OAT, nucleic acids encoding such, vectors, and cells comprising such. Furthermore the use of these is provided in a process for producing creatine and in a process for selecting improved variants of GAT.

[0036] Another aspect relates to a novel E. coli, and the use of E.coH, such as the strain BW25113, comprising mutations selected from the group consisting of;

[0037] • A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y;

[0038] • A nusA mutation, such as A234E R104L, and / or R228P; 83737PC01

[0039] 6

[0040] Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308;

[0041] • One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or

[0042] • a combination of the above; in a process for producing creatine.

[0043] Another aspect relates to a method of producing creatine, the method comprising: i. providing a cell, such as a bacterial cell, such as E. coir, ii. introducing into the cell: o a nucleic acid sequence encoding GAT as described herein; and o a nucleic acid sequence encoding GAMT as described herein; iii. optionally, wherein the cell is e. coli and the e. coli comprises genomic mutations as described herein; iv. growing the cell under conditions suitable for the production of creatine; and v. optionally, purifying creatine from the supernatant.

[0044] Brief description of the figures

[0045] Figure 1: Creatine production in Escherichia coli and GAT growth coupling.

[0046] A) Biosynthesis of creatine from glycine and L-arginine. B) Creatine production over time for SDT631, which is E. coli BW25113 transformed with a plasmid expressing GAMT4 and GAT8 (pSD168). Cultures were grown in M9 with glucose, with or without supplementation of 0.2% arginine and 0.2% glycine. Concentrations greater than the maximum concentration used for calibration of the creatine assay are shaded grey. C) Endpoint creatine concentrations after 25 h of cultivation of BW25113 with GAMT4 and different GAT variants (pSD163-166, pSD168, pSD169). The strains were cultured in M9 with 0.2% glucose, 0.2% arginine and 0.2% glycine. Concentrations outside the concertation range used for 83737PC01

[0047] 7 calibration of the creatine assay are shaded gray D) A computationally designed growth-coupling design for GAT, based on the conversion of L-ornithine to L- proline and subsequently biomass and energy. E) Feasible GAT activity versus growth rate planes for the GAT selection system using arginine as carbon source (left diagonal) and co-fed with glycine (right diagonal). Minimum required fluxes for GAT are indicated for a growth rate of 0.2 / h. GAT : glycine a midi notransferase, GAMT: guanidinoacetate N-methyltransferase, SAM: S- Adenosyl-L-Methionine, SAH: S-adenosyl homocysteine, Arg: L-arginine, Gly: L- glycine, OCD: ornithine cyclodeaminase, SPE: arginine decarboxylase / agmatinase pathway, AST: L-arginine succinyltransferase pathway.

[0048] Figure 2: The growth of background strain BW25113 in M9 glucose media with different concentrations of GAA and creatine supplemented. No significant growth inhibition was observed in the tested concentrations.

[0049] Figure 3: Improvement of the conversion from proline to biomass. A) The activity of two OCD variants was tested by assessing their ability to rescue proline auxotrophy. Glucose, proline, and ornithine were supplemented at 0.2%. B) Growth rate during the proline-ALE experiment using BW25113Zlast4 (SDT854), after the helper substrate glycerol was completely removed from the experiment. C) Growth rate and lag time for different strains growing on M9 with ornithine as the sole carbon and energy source. See Table 2 and Table 3 for the genomic backgrounds, OCD versions and expression backbones.

[0050] Figure 4: Inhibition of growth by arginine. A) Growth rates of strain SDT969 and SDT981 that can grow on ornithine via OCD for various arginine and glycine concentrations. Biomass is either generated from glucose (I) or ornithine (II). Increasing the arginine concentration leads to stronger inhibition of ornithine- based growth, but not glucose-based growth. B) Growth rates of a ! proB strain expressing OCD (SDT941), and the same strain with GAT expression (SDT989)for various arginine and glycine concentrations. Proline is either directly provided (I) or generated endogenously from glucose (II). Increasing the arginine concentration leads to stronger inhibition growth when proline needs to be produced from endogenous ornithine. C) Growth rates for proline auxotrophic strains expressing three additional OCD and two OAT variants (SDT1034-38, 83737PC01

[0051] 8

[0052] SDT1042, SDT1052) in the presence and absence of arginine (0.2%). All variants showed clear growth inhibition by arginine.

[0053] Figure 5: Proline auxotrophy-based selection system. A) Design of the auxotrophy-based selection system. The deletion of arg A eliminates intracellular synthesis of ornithine so ornithine can only be produced from arginine via GAT. B) Feasible GAT activity versus growth rate planes for the GAT selection system using arginine as carbon source (left diagonal) and co-fed with glycine (right diagonal). Minimum required fluxes for GAT are indicated for a growth rate of 0.2 / h. C) Inhibition effects of OCDv4 by arginine for a concentration gradient. D) Growth rates for the auxotrophy-based selection system with (SDT1127) and without (SDT1130) GAT8 for a range of arginine concentrations. E) Addition of ornithine to the selection system with GAT8 (SDT1127) and 0.025% or 0.05% arginine leads to an increase in growth rate. F) Growth rates for the expression of GAT at different levels in the selection system (SDT1127, SDT1130, SDT1152- 58), using combinations of different promoters and ribosome binding sites.

[0054] Figure 6. Determination of operating window for proline auxotrophy-based selection system. The requirement of ornithine (A) and arginine (B) were determined by cultivating the selection system strain SDT1066 or SDT1018 in different concentrations of the substrates.

[0055] Figure 7: A) Selection experiment using the selection system based on proline- auxotrophy. Three different GAT8 variants were isolated from this experiment. B) Growth rates of the selection system with the three re-cloned GAT8 variants (SDT1366-68) and GAT8WT (SDT1156). C) GAA concentration after bioconversion of glycine and arginine using the novel GAT variants.

[0056] Figure 8. Creatine inhibits growth when produced intracellularly from GAA. Growth test in a SAM-dependent methylation growth-coupled strain SDT1004 with different concentrations of GAA fed. When no GAA was added, no growth-coupling allowed, thus there was no growth. Generally, the more GAA was fed the less biomass accumulated, suggesting creatine inhibition to growth. The experiment was repeated twice showing similar trend. 83737PC01

[0057] 9

[0058] Figure 9. Creatine production using different natural GAMT variants. GAMT3 showed the best activity.

[0059] Figure 10. Creatine production E. coli. ALE isolates showed increased creatine tolerance compared to the original strain SDT1004 (A). We assembled the full creatine synthesis pathway from glycine and arginine in a creatine tolerant strain SDT1243 (Materials and methods). One of the variants GAT8_vl (SDT1369) showed more than 50% enhancement in creatine production compared to GAT8_wt (SDT1373) (B).

[0060] Figure 11. Characterization of gxhP KO. The growth profiles of KEIO collection strain GxhP KO (SDT1450, circle) and wt (SDT1448, square) has no significant difference when lOmM GAA was fed (a). The growth rates of the two strains GxhP KO (SDT1450, black) and wt (SDT1448, gray) have not significant difference at different concentrations of GAA feeding. A plasmid containing GAMT (pSD158) was transformed into both strains so that creatine can be produced intracellularly.

[0061] Figure 12. Sequence alignment GAT. Consensus sequences are indicated in gray.

[0062] The present invention will now be described in more detail in the following.

[0063] Detailed description of the invention

[0064] Definitions

[0065] Prior to discussing the present invention in further details, the following terms and conventions will first be defined :

[0066] An enzyme capable of catalysing the reaction arginine < = > ornithine

[0067] In the present context, the enzyme capable of catalysing the reaction arginine < = > ornithine is understood as any enzyme the skilled person may identify, that is able to increase the reaction rate of conversion between arginine and ornithine. Preferably, the enzyme is an enzyme selected from any of the EC numbers EC 2. 1.4.1, EC 2.1.4.3, or EC 3.5.3.1. 83737PC01

[0068] 10

[0069] Glycine amidinotransferase (GAT)

[0070] In the present context, the terms "Glycine amidinotransferase" or "GAT" are to be understood as a protein capable of converting arginine and glycine into L-ornithine and guanidinoacetate (GAA).

[0071] The enzyme catalyses the reaction: glycine + L-arginine < = > guanidinoacetate + L-ornithine

[0072] A functional variant will be able to catalyse said reaction.

[0073] The GAT is preferably an enzyme classified according to EC 2.1.4.1.

[0074] Ornithine cyclodeaminase (OCD)

[0075] In the present context, the terms "ornithine cyclodeaminase" or "OCD" are to be understood as a protein capable of converting L-ornithine into L-proline.

[0076] The enzyme catalyses the reaction:

[0077] L-ornithine < = > L-proline + NH4(+)

[0078] A functional variant will be able to catalyse said reaction.

[0079] The OCD is preferably an enzyme classified according to EC 4.3.1.12.

[0080] Ornithine aminotransferase (OAT)

[0081] In the present context, the terms "ornithine aminotransferase" or "OAT" are to be understood as a protein capable of converting L-ornithine into L-glutamate and 5- semialdehyde. Following spontaneous reaction, proline is generated through 1- pyrroline-5 carboxylate.

[0082] The enzyme catalyses the reaction: a 2-oxocarboxylate + L-ornithine = an L-alpha-amino acid + L-glutamate 5- semialdehyde

[0083] A functional variant will be able to catalyse said reaction.

[0084] The OAT is preferably an enzyme classified according to EC 2.6.1.13.

[0085] Guanidinoacetate methyltransferase (GAMT)

[0086] In the present context, the terms Guanidinoacetate methyltransferase or "GAMT" are understood as a protein capable of adding a methyl group from S-Adenosyl methionine (SAM) to GAA.

[0087] The enzyme catalyses the reaction: 83737PC01

[0088] 11 guanidinoacetate + S-adenosyl-L-methionine < = > creatine + H(+) + S-adenosyl- L-homocysteine

[0089] A functional variant will be able to catalyse said reaction.

[0090] The GAMT is preferably an enzyme classified according to EC 2.1.1.2.

[0091] Variants

[0092] The application discloses variants of nucleic acid and protein sequences. In general, variants may be referred to as "V" or Var" followed by a number, indication the specific variant.

[0093] EC classification

[0094] Enzymes referred to herein may be classified on the basis of the handbook Enzyme Nomenclature from NC-IUBMB, 1992, see also the ENZYME site at the internet: http: / / www.expasy.ch / enzyme / . This is a repository of information relative to the nomenclature of enzymes, and is primarily based on the recommendations of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (IUB-MB). It describes each type of characterized enzyme for which an EC (Enzyme Commission) number has been provided. The IUBMB Enzyme nomenclature is based on the substrate specificity and occasionally on their molecular mechanism.

[0095] Pathway for generating proline

[0096] In the present context, the term "Pathway for generating proline" is to be understood as one or more genes of a bacterial cell, which is, at least partly, responsible generating proline or metabolites in relation their to.

[0097] Pathway for generating ornithine

[0098] In the present context, the term "Pathway for generating proline" is to be understood as one or more genes of a bacterial cell, which is, at least partly, responsible generating ornithine or metabolites in relation their to. 83737PC01

[0099] 12 astABCDE

[0100] In the present context, the terms "astABCDE" and "astABCDE operon" are to be understood as the operon comprising the astA, astB, astD, astC, astE genes, which are responsible for arginine degradation.

[0101] Promoter

[0102] In the present context, the terms "promoter" or "promoter sequence" are to be understood as a region of DNA that initiates transcription of a particular gene.

[0103] Ribosomal binding site / RBS

[0104] In the present context, the terms "ribosomal binding site" or "RBS" are to be understood as a sequence of nucleotides upstream of the start codon of an mRNA transcript that is responsible for the recruitment of a ribosome during the initiation of protein translation.

[0105] Degenerate sequence

[0106] In the present context, a degenerate sequence is understood as a nucleic acid sequence encoding the same protein sequence, with one or more different codons, wherein the one or more different codons code for the same amino acid.

[0107] Nucleic acid sequence, mutation and coding sequence

[0108] In some aspects of the present invention are mentioned nucleic acid sequences, as well as proteins with optimized sequences, by result of mutations. It is thus understood by the skilled person, that when the disclosure refers to a nucleic acid sequence as comprising an amino acid mutation, the nucleic acid sequence does not comprise amino acids, the nucleic acid sequence comprises a sequence, coding for the respective amino acid mutation.

[0109] The present disclosure does not relate to the provision of sequences comprising a mix of nucleotides and amino acids.

[0110] As commonly known within the art a nucleic acid sequence is a succession of bases within the nucleotides forming alleles within a DNA (using GACT) or RNA (GACU) molecule. This sequence represents genetic information and is crucial for the functions of an organism. 83737PC01

[0111] 13

[0112] A nucleic acid sequence refers to an oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA), preferably DNA. This term includes molecules composed of naturally-occurring nucleobases, sugars, and covalent internucleoside (backbone) phosphodiester bond linkages, as well as molecules with non-naturally occurring nucleobases, sugars, and covalent internucleoside linkages that function similarly. Nucleic acids can be composed entirely of deoxyribonucleotides, ribonucleotides, nucleic acid mimics or analogues, or chimeric mixtures thereof. The preferred nucleic acid sequences as discloses herein, are composed of naturally occurring deoxyribonucleotides, or ribonucleotides.

[0113] A protein sequence refers to the specific order of amino acids that are linked together by peptide bonds to form a protein. This sequence determines the protein's three-dimensional structure and its biological function. Proteins are composed of 20 different amino acids. The sequence can be represented using either one-letter or three-letter codes for the amino acids.

[0114] While the application is written in the sense of the nucleic acid sequences encoding a protein, the skilled person will easily be able to understand that embodiments relating to nucleic acid sequences encoding, may also relate to the protein sequence itself. For the ease of understanding, table 1 provides an overview of the comparable nucleic acid- and protein sequences.

[0115] Table 1: Overview of nucleic acid and protein sequences as used herein. Registry entries to uniport or NCBI. 83737PC01

[0116] 14

[0117] Genetic Modification

[0118] In the present context, the term "genetic modification" is defined as the introduction of a genetically inherited change in the host cell genome. Specifically, changes can include mutations in genes and regulatory sequences, coding and non-coding DNA sequences. "Mutations" include deletions, substitutions and insertions of one or more nucleotides or nucleic acid sequences in the genome. Other genetic modifications include the introduction of heterologous genes or coding DNA sequences by recombinant techniques.

[0119] Downregulating

[0120] In the context of this application, "downregulating" an endogenous gene is defined as reducing, optionally eliminating, disrupting or deleting, the transcription or translation of an endogenous gene so that the levels of functional protein, such as an enzyme, encoded by the gene are significantly reduced in the host cell, typically by at least 50%, such as at least 75%, such as at least 90%, such as at least 95%, as compared to a control. Specifically, when the reduced expression is obtained by a genetic modification in the host cell, the control is the unmodified host cell. In some cases, such as gene deletion, the level of native mRNA and 83737PC01

[0121] 15 functional protein encoded by the gene is further reduced, effectively eliminated, by more than 95%, such as 99% or greater.

[0122] Sequence identity

[0123] In the present context, the term "sequence identity" is here defined as the sequence identity between genes or proteins at the nucleotide, base or amino acid level, respectively. Specifically, a DNA and an RNA sequence are considered identical if the transcript of the DNA sequence can be transcribed to the corresponding RNA sequence.

[0124] Thus, in the present context, "sequence identity" is a measure of identity between proteins at the amino acid level and a measure of identity between nucleic acids at nucleotide level. The protein sequence identity may be determined by comparing the amino acid sequence in a given position in each sequence when the sequences are aligned. Similarly, the nucleic acid sequence identity may be determined by comparing the nucleotide sequence in a given position in each sequence when the sequences are aligned.

[0125] To determine the percent identity of two amino acid sequences or of two nucleic acids, the sequences are aligned for optimal comparison purposes (e.g., gaps may be introduced in the sequence of a first amino acid or nucleic acid sequence for optimal alignment with a second amino or nucleic acid sequence). The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide at the corresponding position in the second sequence, then the molecules are identical in that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity = # of identical positions / total # of positions (e.g., overlapping positions) x 100). In one embodiment, the two sequences are the same length.

[0126] In another embodiment, the two sequences are of different length and gaps are seen as different positions. One may manually align the sequences and count the number of identical amino acids. Alternatively, alignment of two sequences for the determination of percent identity may be accomplished using a mathematical algorithm. Such an algorithm is incorporated into the BLASTN and BLASTX programs of (Altschul et al. 1990). BLAST nucleotide searches may be performed 83737PC01

[0127] 16 with the NBLAST program, to obtain nucleotide sequences homologous to a nucleic acid molecule of the invention. BLAST protein searches may be performed with the BLASTX program, to obtain amino acid sequences homologous to a protein molecule of the invention.

[0128] To obtain gapped alignments for comparison purposes, Gapped BLAST may be utilized. Alternatively, PSI-Blast may be used to perform an iterated search that detects distant relationships between molecules. When utilizing the BLASTN, BLASTX, and Gapped BLAST programs, the default parameters of the respective programs may be used. See http: / / www.ncbi.nlm.nih.gov. Alternatively, sequence identity may be calculated after the sequences have been aligned e.g. by the BLAST program in the EMBL database (www.ncbi.nlm.gov / cgi-bin / BLAST). Generally, the default settings with respect to e.g. "scoring matrix" and "gap penalty" may be used for alignment. In the context of the present invention, the BLASTN and PSI BLAST default settings may be advantageous.

[0129] The percent identity between two sequences may be determined using techniques similar to those described above, with or without allowing gaps. In calculating percent identity, only exact matches are counted. An embodiment of the present invention thus relates to sequences of the present invention that has some degree of sequence variation.

[0130] Bacterial selection systems

[0131] Provision of optimized enzymes, for use in fermentation processes, such as creatine production, may be done through selection systems for selecting such enzymes. In silico modelling identified a way to use rewire the metabolic network such that in a specific growth media, the energy, biomass, or some essential metabolite relies on a certain pathway, which the target enzyme is part of. This allows an arbitrarily large library to be "screened" with minimal effort and cost. It was thus identified that introducing enzymes capable of performing the reaction arginine < = > ornithine, and ornithine < = > proline, and removing the endogenous pathways for generating proline and ornithine, would allow for screening for optimized enzymes (figure Id). 83737PC01

[0132] 17

[0133] Thus, in one aspect of the present disclosure a bacterial cell is provided, the bacterial cell comprising:

[0134] • a nucleic acid sequence encoding an enzyme, such as Glycine a midi notransferase (GAT) or a functional variant thereof, capable of catalysing the reaction arginine < = > ornithine, wherein the enzyme is an enzyme selected from any of the EC numbers EC 2.1.4.1, EC 2.1.4.3, or EC 3.5.3.1; and

[0135] • a nucleic acid encoding o an enzyme, such as Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of catalysing the reaction ornithine < = > proline, wherein the enzyme is an enzyme according to EC 4.3.1.12, or o an enzyme, such as Ornithine aminotransferase (OAT) or a functional variant thereof, capable of catalysing the reaction ornithine < = > glutamate+ semialdehyde, wherein the enzyme is an enzyme according to EC 2.6.1.13, and wherein one or more pathways for generating proline and ornithine are downregulated or deleted.

[0136] In preferred embodiments of the present disclosure, the bacterial cell is a genetically modified bacterial cell.

[0137] In some embodiments of the present disclosure, it is only the one or more pathways for generating ornithine that are downregulated or deleted.

[0138] In some embodiments of the present disclosure, one or more genetic modifications are introduced for deleting or downregulating the one or more pathways generating proline and ornithine.

[0139] In preferred embodiments of the present disclosure, the growth of the cell is dependent on the production of ornithine from the enzyme capable of catalysing the reaction arginine < = > ornithine.

[0140] The pathways for generating proline or ornithine may be downregulated or deleted in different ways, in some embodiments of the present disclosure the pathways for generating proline are selected from the group consisting of:

[0141] • proB (N-acetylglutamate synthase);

[0142] • proA (glutamate-5-semialdehyde dehydrogenase); and 83737PC01

[0143] 18

[0144] • proC (pyrroline-5-carboxylate reductase), preferably proB is the pathway for generating proline. In some embodiments of the present disclosure the pathways for generating ornithine are selected from the group consisting of:

[0145] • argA (N-acetylglutamate synthase);

[0146] • argB (acetylglutamate kinase);

[0147] • argC (N-acetylglutamylphosphate reductase);

[0148] • argD, astC, and gabT (N-acetylornithine aminotransferase); and

[0149] • argE (acetylornithine deacetylase) preferably, argA is the pathway for generating ornithine.

[0150] In a preferred embodiment, proB is the pathway for generating proline and argA is the pathway for generating ornithine.

[0151] In a particular favourable variant, the pathway encoded by the operon astABCDE is also downregulated or deleted, and thus in some embodiments of the present disclosure at least one enzyme encoded by the pathway encoded by the operon astABCDE is downregulated or deleted.

[0152] In some embodiments of the present disclosure the downregulated or deleted pathways for generating proline or ornithine are astABCDE, proB, and argA.

[0153] As shown by the experiments, several GAT enzymes can be used, and mutations are favourably introduced into them, allowing for a better performance of the cell.

[0154] A favourable protein was identified as the CyrA from Cylindrospermopsis raciborskii AWT205 (UniProt B0LI36_9CYAN), and thus in some embodiments of the present disclosure the nucleic acid sequence encoding an enzyme capable of catalysing the reaction arginine < = > ornithine is a nucleic acid sequence encoding glycine amidinotransferase (GAT) having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 8, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity. 83737PC01

[0155] 19

[0156] The mutations:

[0157] • M175T, Q211L and R232G (GAT8-V1);

[0158] • G25E, S181N and R232G (GAT8-V2); and

[0159] • M182L and G281D (GAT8-V3), was identified as particularly favourable and identified in three individual variants.

[0160] In some embodiments of the present disclosure the GAT comprises a combination of mutations selected from the group of combinations consisting of:

[0161] • M175T, Q211L and R232G (GAT8-V1);

[0162] • G25E, S181N and R232G (GAT8-V2); and

[0163] • M182L and G281D (GAT8-V3).

[0164] Thus, it is clearly shown that a number of mutations may specifically be introduced, and thus in some embodiments of the present disclosure the encoded GAT has between 1 and 4 mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 8, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 8.

[0165] Further GAT proteins are also shown to work herein, and thus the skilled person may select from a number of specific variants.

[0166] In some embodiments of the present disclosure GAT is selected from the group consisting of GAT3 as encoded by the sequence according to SEQ ID NO: 3, GAT8 as encoded by the sequence according to SEQ ID NO: 8, GAT9 as encoded by the sequence according to SEQ ID NO: 9, preferably GAT8 as encoded by the sequence according to SEQ ID NO: 8. In some embodiments of the present disclosure GAT is selected from the group consisting of GAT1 as encoded by the sequence according to SEQ ID NO: 1, GAT2 as encoded by the sequence according to SEQ ID NO: 2, GAT3 as encoded by the sequence according to SEQ ID NO: 3, GAT4 as encoded by the sequence according to SEQ ID NO: 4, GAT5 as encoded by the sequence according to SEQ ID NO: 5, GAT6 as encoded by the sequence according to SEQ ID NO: 6, GAT7 as encoded by the sequence according to SEQ ID NO: 7, GAT8 as encoded by the sequence according to SEQ ID NO: 8, GAT9 as encoded by the sequence according to SEQ ID NO: 9. 83737PC01

[0167] 20

[0168] The skilled person will find that each mutation may be used separately, and thus in some embodiments of the present disclosure the bacterial cell comprises mutations in GAT, such as one or more of the mutations M175T, Q211L, R232G, G25E, S181N, M182L and G281D, preferably at least R232G.

[0169] In some embodiments of the present disclosure GAT is selected from the group consisting of GAT8-V1 as encoded by the sequence according to SEQ ID NO: 23, GAT8-V2 as encoded by the sequence according to SEQ ID NO: 24, GAT8-V3 as encoded by the sequence according to SEQ ID NO: 25 and wherein the encoded GAT has between 1 and 4 additional mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 23, 24, or 25, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 23, 24, or 25. In some embodiments of the present disclosure GAT is selected from the group consisting of GAT8-V1 as encoded by the sequence according to SEQ ID NO: 23, GAT8-V2 as encoded by the sequence according to SEQ ID NO: 24, GAT8-V3 as encoded by the sequence according to SEQ ID NO: 25.

[0170] As shown in the examples, a number of different promoters and RBS sequences can be used to regulate the expression. In some embodiments of the present disclosure the expression of GAT is regulated by a promoter selected from J23110, J23109, and J23114, preferably the J23110 promoter. In some embodiments of the present disclosure the expression of GAT is regulated by RBSU058 as the RBS.

[0171] In some embodiments of the present disclosure the expression of GAT is regulated by a promoter selected from J23110, J23109, and J23114, preferably the J23110 promoter and RBSU058 as the RBS. In some embodiments of the present disclosure the expression of GAT is regulated by a promoter with an expression corresponding to J23110.

[0172] In some embodiments of the present disclosure, the promoter J23109 has the nucleic acid according to SEQ ID NO: 31. In some embodiments of the present disclosure, the promoter J23114 has the nucleic acid according to SEQ ID NO: 32. In some embodiments of the present disclosure, the promoter J23110 has the 83737PC01

[0173] 21 nucleic acid according to SEQ ID NO: 34. In some embodiments of the present disclosure, the RBS RBSU058 has the nucleic acid sequence according to SEQ ID NO: 19.

[0174] As shown by the experiments, several OCD enzymes can be used, and mutations are favourably introduced into them, allowing for a better performance of the cell.

[0175] A favourable protein was identified as protein referred to as OCDWT herein, more particularly the Ornithine cyclodeaminase from Pseudomonas putida (UniProt Q88H32), and thus in some embodiments of the present disclosure the nucleic acid sequence encoding the enzyme capable of catalysing the reaction ornithine < = > proline is a nucleic acid sequence Ornithine cyclodeaminase (OCD) having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 176, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[0176] Different mutations were found to be favourable in the OCD. With respect to the OCDWT, the following variants was identified:

[0177] • OCDvl - OCDK205G / M86K / T162A;

[0178] • OCDv2 - OCDK205G / M86T / T162A ;

[0179] • OCDv3 - OCDK205G / M86M / T162A ; a nd

[0180] • OCDv4 - OCDR65H / R184H / A185V / R244H .

[0181] Thus, in some embodiments of the present disclosure OCD is selected from the group consisting of OCDvl, OCDv2, OCDv3, OCDv4, OCDWT. The best performing enzyme was found to be OCDv4. In some embodiments of the present disclosure OCD is selected from the group consisting of OCDvl, OCDv2, OCDv3, OCDv4, OCDWT, OCD_AGRT4, and OCD_AGRFC, wherein OCDv4 is preferred.

[0182] A particularly preferred embodiment is thus a nucleic acid sequence encoding a protein having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 176, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such 83737PC01

[0183] 22 as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity, wherein the nucleic acid sequence comprise a sequence encoding the following mutations R65H, R184H, A185V, and / or R244H.

[0184] In some embodiments of the present disclosure the encoded OCD has between 1 and 4 mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 176, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 176.

[0185] In some embodiments of the present disclosure the OCD comprises a methionine at position 86, with respect to the amino acid sequence encoded by OCDWT.

[0186] In some embodiments of the present disclosure the OCD comprises mutations in OCD, such as one or more of the mutations K205G, T162A, R65H, R184H, A185V, or R244H.

[0187] In some embodiments of the present disclosure, the bacterial cell comprises a combination of mutations in OCD, wherein a combination is selected from the group of combinations consisting of:

[0188] • K205G, M86K, T162A;

[0189] • K205G, M86T, T162A;

[0190] • K205G, M86M, T162A; and

[0191] • R65H, R184H, A185V, R244H.

[0192] Instead of OCD, OAT may be used, since the protein is capable of converting + L- ornithine into L-glutamate and 5-semialdehyde. Following spontaneous reaction, proline is generated through l-pyrroline-5 carboxylate. In some embodiments of the present disclosure OAT is selected from OAT_HUMAN, or OAT_BACSU.

[0193] In some embodiments of the present disclosure the nucleic acid sequence encoding the enzyme capable of catalysing the reaction ornithine< = > glutamate+ semialdehyde is a nucleic acid sequence encoding Ornithine aminotransferase (OAT) having at least 80 % sequence identity to the protein encoded by the 83737PC01

[0194] 23 sequence according to SEQ ID NO: 21, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[0195] The in silica analysis leading the inventors to the present invention was further used to identify other organisms wherein the invention will function. Thus, in some embodiments of the present disclosure the bacterial cell is selected from the group consisting of:

[0196] • Escherichia coli;

[0197] • Geobacter metallireducens;

[0198] • Synechococcus elongatus;

[0199] • Synechocystis sp;

[0200] • Saccharomyces cerevisiae;

[0201] • Shigella dysenteriae; and

[0202] • Phaeodactylum tricornutum : preferably Escherichia coli.

[0203] These strains are predicted feasible for GAT selection based on genome-scale models available in BIGG database until 17 July 2024.

[0204] In some embodiments of the present disclosure the bacterial cell is Escherichia coli and the strain is selected from the group consisting of BW25113 and Escherichia coli str. K-12 substr. MG1655.

[0205] Bacterial selection system for GAMT

[0206] Another important enzyme for use in fermentation processes, such as creatine production, as the guanidinoacetate methyltransferase (GAMT) enzyme.

[0207] Preferably the provision of optimized mutants thereof are carried out in e. coli cells, and thus in another aspect of the present disclosure is provided an e. coli, wherein the e. coli comprises a nucleic acid sequence encoding an enzyme capable of catalysing the reaction guanidinoacetate + S-adenosyl-L-methionine < = > creatine + H(+) + S-adenosyl-L-homocysteine, or a functional variant thereof, preferably an enzyme classified according to EC 2.1.1.2. A variant of such an aspect may be an e. coli, wherein the e. coli comprises a nucleic acid sequence 83737PC01

[0208] 24 encoding guanidinoacetate methyltransferase (GAMT) or a functional variant thereof such as an enzyme capable of catalysing the reaction guanidinoacetate + S-adenosyl-L-methionine < = > creatine + H(+) + S-adenosyl-L-homocysteine. As shown in the examples, GAMT sequences in general may be optimized, however a preferred GAMT sequence may for instance be the GAMT3 or the GAMT4 protein.

[0209] In some embodiments of the present disclosure GAMT has at least 80 % sequence identity to the protein, GAMT3, encoded by the sequence according to SEQ ID NO:

[0210] 12, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[0211] In some embodiments of the present disclosure GAMT has at least 80 % sequence identity to the protein, GAMT4, encoded by the sequence according to SEQ ID NO:

[0212] 13, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[0213] In some embodiments of the present disclosure GAMT is selected from the group consisting of:

[0214] • GAMT1 as encoded by the sequence according to SEQ ID NO: 10;

[0215] • GAMT2 as encoded by the sequence according to SEQ ID NO: 11;

[0216] • GAMT3 as encoded by the sequence according to SEQ ID NO: 12;

[0217] • GAMT4 as encoded by the sequence according to SEQ ID NO: 13; or

[0218] • GAMT5 as encoded by the sequence according to SEQ ID NO: 14.

[0219] To ensure efficiency of the engineered cells, the SAM-cycles may be rewired, thus in some embodiments of the present disclosure the bacterial cell has rewired SAM-cycle and is deficient of the natural cysteine synthesis. In some embodiments of the present disclosure cysteine is provided by the SAM cycle and SAM-dependent methylation. In some embodiments of the present disclosure the rewired SAM-cycle is rewired by the introduction of a nucleic acid encoding sacB, CYS3, CYS4. 83737PC01

[0220] 25

[0221] As identified in example 7, several e.coli ALE mutants was identified, each having individual mutations, leading to improved efficiency in generating creatine. Without being bound by theory, the skilled person will understand that these mutations each contribute to the improved function, and may thus be combined individually.

[0222] In some embodiments of the present disclosure the bacterial strain comprises one or more of the following features; i. A rpoA D305X mutation, wherein X may be any amino acid, preferably H or Y; ii. A nusA mutation, such as A234E R104L, and / or R228P; iii. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; iv. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or v. a combination of the above.

[0223] The specific ALE strains, SDT1273, SDT1274, SDT1275, SDT1276, and SDT1277 are listed in example 7 with a specific combination of these mutations. It is thus clear that in some specific embodiments of the disclosure, e. coli are provided with the specific combination of mutations. In some embodiments of the present disclosure the E. coli is the strain BW25113, preferably comprising one or more of the features as described above. In particular, such features may be combined as originally identified, thus in some embodiments of the present disclosure the bacterial strain comprises a combination of mutations selected from the group consisting of:

[0224] • metC G341D, and rpoA D305Y;

[0225] • metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt);

[0226] • trpA Q243P, metC G241C, and nusA A234E; 83737PC01

[0227] 26 folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt).

[0228] Bacteria for producing creatine

[0229] As introduced in the introductory parts of the present disclosure, a main object of the present disclosure is to provide optimized enzymes for use in producing creatine in bacterial cells. Such cells will now be introduced, and the skilled person will easily understand that the previous disclosure relating to the various enzymes, or optimized variants thereof, capable of overall performing the reaction as introduced in figure la will now be used for producing creatine.

[0230] A wide range of bacterial cells may be used for the production of creatine, and thus in a first instance, the disclosure is not limited to a specific set of bacterial cells, however introduced with e. coli as an exemplary cell.

[0231] An aspect of the present disclosure relates to a bacterial cell, such as e. coli, comprising; i. a nucleic acid encoding a GAT as described herein; and ii. a nucleic acid encoding a GAMT as described herein.

[0232] However, as introduced in example 2, this bacterial cell may not be optimal, and thus the skilled person will preferably also introduce some of the mutations identified in E. coli, into the bacterial cell. Without being bound by theory, the skilled person will understand that these mutations each contribute to the improved function and may thus be combined individually.

[0233] Another aspect of the present disclosure relates to a bacterial cell, such as E. coli, comprising one or more of the following features; i. a nucleic acid encoding a GAT as described herein; ii. a nucleic acid encoding a GAMT as described herein; iii. A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; iv. A nusA mutation, such as A234E R104L, and / or R228P; v. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; 83737PC01

[0234] 27 vi. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or vii. a combination of the above.

[0235] Another aspect of the present disclosure relates to a bacterial cell, such as E. coli, comprising one or more of the following features; i. a nucleic acid sequence encoding an enzyme capable of catalysing the reaction arginine < = > ornithine as described herein, wherein the enzyme is an enzyme selected from any of the EC numbers EC

[0236] 2. 1.4.1, EC 2.1.4.3, or EC 3.5.3.1, such an enzyme may for instance be a GAT as described herein; ii. a nucleic acid encoding a GAMT as described herein; iii. A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; iv. A nusA mutation, such as A234E R104L, and / or R228P; v. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; vi. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or vii. a combination of the above.

[0237] Another aspect of the present disclosure relates to a bacterial cell, such as e. coli, comprising; i. a nucleic acid encoding a GAT as described herein; ii. a nucleic acid encoding a GAMT as described herein; and optionally one or more of the following features: i. A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; ii. A nusA mutation, such as A234E R104L, and / or R228P; 83737PC01

[0238] 28 iii. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; iv. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or v. a combination of the above.

[0239] Another aspect of the present disclosure relates to a bacterial cell, such as E. coli, comprising; i. a nucleic acid sequence encoding an enzyme capable of catalysing the reaction arginine < = > ornithine as described herein, wherein the enzyme is an enzyme selected from any of the EC numbers EC

[0240] 2. 1.4.1, EC 2.1.4.3, or EC 3.5.3.1, such an enzyme may for instance be a GAT as described herein; ii. a nucleic acid encoding a GAMT as described herein; and optionally one or more of the following features: i. A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; ii. A nusA mutation, such as A234E R104L, and / or R228P; iii. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; iv. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or v. a combination of the above.

[0241] The specific ALE strains, SDT1273, SDT1274, SDT1275, SDT1276, and SDT1277 are listed in example 7 with a specific combination of these mutations. It is thus clear that in some specific embodiments of the disclosure, E. coli are provided with the specific combination of mutations. In some embodiments of the present disclosure the E. coli is the strain BW25113, preferably comprising one or more of the features as described above. In particular, such features may be combined as 83737PC01

[0242] 29 originally identified, thus in some embodiments of the present disclosure the bacterial strain comprises a combination of mutations selected from the group consisting of:

[0243] • metC G341D, and rpoA D305Y;

[0244] • metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt);

[0245] • trpA Q243P, metC G241C, and nusA A234E;

[0246] • folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and

[0247] • metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt).

[0248] The examples of the present invention, are primarily carried out in e.coli, thus by a further in silica analysis further bacteria, in which the optimized enzymes may be used, have been identified. The bacteria are frequently used engineering hosts. Thus, in some embodiments of the present disclosure the bacterial cell is selected from the group consisting of

[0249] • E. coli;

[0250] • Pseudomonas putida;

[0251] • Vibrio natriegens;

[0252] • Bacillus subtilis;

[0253] • Corynebacterium glutamicum;

[0254] • Cyanobacteria, such as Synechocystis sp. PCC 6803 and / or

[0255] Synechococcus elongatus PCC 7942;

[0256] • Streptomyces spp;

[0257] • Lactobacillus spp;

[0258] • Saccharomyces cerevisiae;

[0259] • Clostridium butyclicum;

[0260] • Aspergillus niger and

[0261] • Bacillus licheniformis.

[0262] In particular, the disclosure provides an e. coli, comprising a combination of mutations selected from the group consisting of:

[0263] • metC G341D, and rpoA D305Y; 83737PC01

[0264] 30

[0265] • metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt);

[0266] • trpA Q243P, metC G241C, and nusA A234E;

[0267] • folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and

[0268] • metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt), preferably, wherein the e. coli comprises a gene coding for GAT and / or GAMT.

[0269] An aspect of the present disclosure relates to the use, of the bacterial cell as described above, in a fermentation process, such as a fermentation process for producing creatine.

[0270] Process for optimized selection systems

[0271] The present disclosure also provides the specific processes for identifying optimized enzyme variants. Thus, in one aspect of the present disclosure is provided a process for selecting mutants of GAT and / or OCD / OAT, the process comprising: i. Providing a bacterial host cell comprising

[0272] ■ a nucleic acid sequence encoding an enzyme, such as Glycine a midi notransferase (GAT) or a functional variant thereof, capable of catalysing the reaction arginine < = > ornithine, wherein the enzyme is an enzyme selected from any of the EC numbers EC 2.1.4.1, EC 2.1.4.3, or EC 3.5.3.1; and

[0273] ■ a nucleic acid encoding

[0274] • an enzyme, such as Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of catalysing the reaction ornithine < = > proline, wherein the enzyme is an enzyme according to EC 4.3.1.12, or

[0275] • an enzyme, such as Ornithine aminotransferase (OAT) or a functional variant thereof, capable of catalysing the reaction ornithine < = > glutamate+ semialdehyde, wherein the enzyme is an enzyme according to EC 2.6.1.13, 83737PC01

[0276] 31 ii. optionally, the GAT as described herein and a. a nucleic acid encoding an OCD as described herein, or b. a nucleic acid encoding an OAT as described herein; iii. Downregulating or deleting pathways for generating proline or ornithine, wherein the pathways for generating proline or ornithine as described herein, or optionally if only mutants of OCD are selected the pathways may be encoded by astABCDE and speAB; iv. Growing the cells for a time sufficient to obtain improved variants of GAT and / or OCD / OAT; v. Identifying said improved variants.

[0277] The functioning of process is preferably carried out in the presence of arginine. The specific concentration of arginine may depend on the specific setup, however the skilled person will easily be able to adjust the arginine concentration, when applying the rest of the teaching as provided herein. Thus, in preferable embodiments of the present disclosure, the cells are in the presence of arginine. In some further embodiments, ornithine may also favourably be supplemented, and this in some embodiments of the present disclosure, the cells are grown in the presence of both arginine and ornithine.

[0278] In some embodiments of the present disclosure the cells are grown in a medium comprising arginine between 0.002 % and 0.020 %, such as between 0.002 % and 0.010 %, such as between 0.003 % and 0.010 %, preferably at about 0.005 %.

[0279] In one embodiment of the present disclosure, the cells are grown in a medium comprising:

[0280] • ornithine supplemented at least 0.005 %, such as at least 0.01, such as at least 0.02%; and arginine between 0.002 % and 0.020 %, such as between 0.002 % and 0.010 %, such as between 0.003 % and 0.010 %, preferably at about 0.005 %. 83737PC01

[0281] 32

[0282] The disclosure also provides how to adapt cells already capable of producing creatine. The skilled person will know that also the optimized enzyme variants as provided herein, may be inserted for further optimization in the creatine production.

[0283] Another aspect of the disclosure relates to a process for selecting bacterial strains with increased tolerance to creatine, the process comprising: i. providing an E. coli with a rewired SAM-cycle and being deficient of the natural cysteine synthesis; ii. introducing into the bacterial cell: o a nucleic acid sequence encoding GAT as described herein; and o a nucleic acid sequence encoding GAMT as described herein; iii. optionally, the E. coli may comprise mutations selected from the group consisting of; o A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; o A nusA mutation, such as A234E R104L, and / or R228P; o Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; o One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or o a combination of the above. iv. Growing the bacterial cell in conditions with substantially no cysteine being present, for a time sufficient to obtain strains with increased tolerance to creatine; and v. optionally, isolating and / or identifying said strains with increased tolerance to creatine.

[0284] In some embodiments of the present disclosure the bacterial strain comprises a combination of mutations selected from the group consisting of: 83737PC01

[0285] 33

[0286] • metC G341D, and rpoA D305Y;

[0287] • metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt);

[0288] • trpA Q243P, metC G241C, and nusA A234E;

[0289] • folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and

[0290] • metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt).

[0291] Enzymes with optimized mutations

[0292] The optimized enzymes as identified in the examples, and introduced herein above, are also provided specifically. The skilled person will know that despite the wildtype proteins being novel, the individual mutations themselves provide optimal benefits, and may also be transferred to homologous proteins. The proteins will be introduced individually, as well as the nucleic acids, vectors and host cells comprising such.

[0293] Thus, one aspect of the present disclosure relates to a Glycine amidinotransferase (GAT) as described herein, wherein the GAT comprises one or more of the mutations M175T, Q211L, R232G, G25E, S181N, M182L and / or G281D, preferably at least R232G.

[0294] In some embodiments of the present disclosure the GAT comprises a combination of mutations selected from the group of combinations consisting of:

[0295] • M175T, Q211L and R232G;

[0296] • G25E, S181N and R232G; and

[0297] • M182L, and G281D.

[0298] Since these mutations were originally identified in the protein encoded by the nucleic acid sequence according to SEQ ID NO: 8, a preferred embodiment may be the GAT according to SEQ ID NO: 8, comprising one or more of the mutations as described above. Thus, in some embodiments of the present disclosure the GAT comprising one or more of these mutations will be a GAT having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID 83737PC01

[0299] 34

[0300] NO: 8, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[0301] Since sequence variance may cause an ambiguity in the actual amount of possible mutations, in some embodiments of the present disclosure the encoded GAT has between 1 and 4 mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 23, 24, or 25, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 23, 24, or 25. In some embodiments, the GAT has between 1 and 4 additional mutations with respect to the protein encoded by the sequence according to SEQ ID NO: NO: 23, 24, or 25, such as between 1 and 3 additional mutations, such as between 1 and 2 additional mutations, such as 1 additional mutation with respect to the protein encoded by the sequence according to SEQ ID NO: NO: 23, 24, or 25.

[0302] Another aspect relates to a nucleic acid sequence encoding the GAT comprising said mutations. Such specific nucleic acid sequences may be any of SEQ ID NO: NO: 23, 24, or 25, or any degenerate sequence thereof.

[0303] Further GAT proteins are also shown to work herein, and thus the skilled person may select from a number of specific variants.

[0304] As shown in example 8 some mutations show consensus in other GAT sequences too. For example, G25 in GAT8 has correspondence in other GATs as G, N or K. E is similar to N in terms of amino acid chemistry.

[0305] Thus, in one aspect of the present disclosure the GAT is a GAT selected from the group consisting of:

[0306] • GAT8 comprising the mutation G25E;

[0307] • GAT3 comprising the mutation N25E;

[0308] • GAT4 comprising the mutation N26E;

[0309] • GAT5 comprising the mutation K26E; 83737PC01

[0310] 35

[0311] GAT6 comprising the mutation G23E; and

[0312] GAT9 comprising the mutation N22E.

[0313] It is preferred that GAT8, GAT3, GAT4, GAT6, and / or GAT9 comprises an E at the indicated positions. Even more preferred that GAT6 and / or GAT8 have an E at the indicated position.

[0314] Thus, in one aspect of the present disclosure the GAT is a GAT selected from the group consisting of:

[0315] • GAT8 comprising the mutation Q211L;

[0316] • GAT3 comprising the mutation Q208L;

[0317] • GAT4 comprising the mutation Q209L;

[0318] • GAT5 comprising the mutation Q188L;

[0319] • GAT6 comprising the amino acid L193; and

[0320] • GAT9 comprising the mutation N223L.

[0321] It is preferred that GAT8, GAT3, GAT4, GAT5, and / or GAT6 comprises an L at the indicated positions.

[0322] Thus, in one aspect of the present disclosure the GAT is a GAT selected from the group consisting of:

[0323] • GAT8 comprising the mutation M182L;

[0324] • GAT3 comprising the mutation A179L;

[0325] • GAT4 comprising the mutation A180L;

[0326] • GAT5 comprising the mutation N159L; and

[0327] • GAT9 comprising the mutation S148L.

[0328] Thus, in one aspect of the present disclosure the GAT is a GAT selected from the group consisting of:

[0329] GAT8 comprising the mutation G281D;

[0330] GAT3 comprising the mutation G274D;

[0331] GAT4 comprising the mutation G275D; 83737PC01

[0332] 36

[0333] GAT5 comprising the amino acid D256;

[0334] GAT6 comprising the mutation S258D; and

[0335] GAT9 comprising the mutation K286D.

[0336] It is preferred that GAT8, GAT3, GAT4, and / or GAT5 comprises a D at the indicated positions.

[0337] Thus, one aspect of the present disclosure relates to a GAT3 as described herein, wherein the GAT3 comprises one or more of the mutations N26E, V173T, L197N, A180L, Q209L, and / or G275D. Thus, in some embodiments of the present disclosure, the GAT3 comprising one or more of these mutations will be a GAT3 having at least 80% sequence identity to the protein encoded by the sequence according to SEQ ID NO: 3, or a degenerate sequence thereof, such as at least

[0338] 85% sequence identity, such as at least 90% sequence identity, such as at least

[0339] 95% sequence identity, such as at least 96% sequence identity, such as at least

[0340] 97% sequence identity, such as at least 98% sequence identity, such as at least

[0341] 99% sequence identity.

[0342] Thus, one aspect of the present disclosure relates to a GAT4 as described herein, wherein the GAT4 comprises one or more of the mutations N26E, V173T, L197N, A180L, Q209L, and / or G275D. Thus, in some embodiments of the present disclosure, the GAT4 comprising one or more of these mutations will be a GAT4 having at least 80% sequence identity to the protein encoded by the sequence according to SEQ ID NO: 4, or a degenerate sequence thereof, such as at least

[0343] 85% sequence identity, such as at least 90% sequence identity, such as at least

[0344] 95% sequence identity, such as at least 96% sequence identity, such as at least

[0345] 97% sequence identity, such as at least 98% sequence identity, such as at least

[0346] 99% sequence identity.

[0347] Thus, one aspect of the present disclosure relates to a GAT5 as described herein, wherein the GAT5 comprises one or more of the mutations K26E, Q188L, and / or D256D. Thus, in some embodiments of the present disclosure, the GAT5 comprising one or more of these mutations will be a GAT5 having at least 80% sequence identity to the protein encoded by the sequence according to SEQ ID NO: 5, or a degenerate sequence thereof, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, 83737PC01

[0348] 37 such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity.

[0349] Thus, one aspect of the present disclosure relates to a GAT6 as described herein, wherein the GAT6 comprises one or more of the mutations G23E, L193L, and / or S258D. Thus, in some embodiments of the present disclosure, the GAT6 comprising one or more of these mutations will be a GAT6 having at least 80% sequence identity to the protein encoded by the sequence according to SEQ ID NO: 6, or a degenerate sequence thereof, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity.

[0350] Thus, one aspect of the present disclosure relates to a GAT9 as described herein, wherein the GAT9 comprises one or more of the mutations N22E, S148L, N223L, and / or K286D. Thus, in some embodiments of the present disclosure, the GAT9 comprising one or more of these mutations will be a GAT9 having at least 80% sequence identity to the protein encoded by the sequence according to SEQ ID NO: 9, or a degenerate sequence thereof, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity.

[0351] Another aspect relates to a vector, such as a plasmid, comprising the nucleic acid sequence described above.

[0352] Another aspect a host cell, such as a bacterial host cell, comprising the nucleic acid or vector as described above, optionally wherein the host cell is a host cell for reproduction of the vector.

[0353] Lastly, the GAT is also provided for the use in producing creatine. Thus another aspect of the present disclosure relates to the use of GAT, the nucleic acid, the vector, or the cell as described above, in a process for producing creatine.

[0354] Another aspect of the present disclosure relates to an ornithine cyclodeaminase (OCD) comprising one or more of the mutations R65H; M86T or M86M; R184H;

[0355] A185V; and / or R244H. In a more specific embodiment, the one or more mutations 83737PC01

[0356] 38 are selected from the group consisting of R65H, R184H, A185V, and / or R244H. In an even more specific embodiment, the OCD comprises the mutations R65H, R184H, A185V, and R244H.

[0357] As shown by the experiments, several OCD enzymes can be used, and mutations are favourably introduced into them, allowing for a better performance of the cell.

[0358] Since these mutations were originally identified in the protein encoded by the nucleic acid sequence according to SEQ ID NO: 176, a preferred embodiment may be the OCD according to SEQ ID NO: 176, comprising one or more of the mutations as described above. Thus, in some embodiments of the present disclosure the OCD comprising one or more of these mutations will be an OCD having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 176, or a degenerate sequence thereof, such as at least

[0359] 85 % sequence identity, such as at least 90 % sequence identity, such as at least

[0360] 95 % sequence identity, such as at least 96 % sequence identity, such as at least

[0361] 97 % sequence identity, such as at least 98 % sequence identity, such as at least

[0362] 99 % sequence identity.

[0363] Since sequence variance may cause an ambiguity in the actual amount of possible mutations, in some embodiments of the present disclosure the encoded OCD has between 1 and 4 mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 176, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 176. In some embodiments, the OCD has between 1 and 4 additional mutations with respect to the protein encoded by the sequence according to SEQ ID NO: NO: 22, such as between 1 and 3 additional mutations, such as between 1 and 2 additional mutations, such as 1 additional mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 176, wherein additional mutations are additional in relation to the one or more of the mutations R65H; M86T or M86M; R184H;

[0364] A185V; and / or R244H. 83737PC01

[0365] 39

[0366] In some embodiments of the present disclosure the OCD comprises a methionine at position 86, with respect to the amino acid sequence encoded by SEQ ID NO: 176.

[0367] Another aspect relates to a nucleic acid sequence encoding the OCD comprising said mutations.

[0368] Another aspect relates to a vector, such as a plasmid, comprising the nucleic acid sequence described above.

[0369] Another aspect a host cell, such as a bacterial host cell, comprising the nucleic acid or vector as described above, optionally wherein the host cell is a host cell for reproduction of the vector.

[0370] Since these optimized OCD variants may be used in the GAT selection process as disclosed herein, the specific use is also provided. Thus, in another aspect of the present disclosure is provided the use of OCD, the nucleic acid, the vector, or the cell as described above in a process for selecting improved variants of GAT, such as a process as described herein.

[0371] Uses

[0372] Use of an E.coli, such as the strain BW25113, comprising mutations selected from the group consisting of;

[0373] • A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y;

[0374] • A nusA mutation, such as A234E R104L, and / or R228P;

[0375] • Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308;

[0376] • One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or

[0377] • a combination of the above; in a process for producing creatine.

[0378] In some embodiments of the present disclosure the e. coli comprises a combination of mutations selected from the group of combinations consisting of: 83737PC01

[0379] 40

[0380] • metC G341D, and rpoA D305Y;

[0381] • metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt);

[0382] • trpA Q243P, metC G241C, and nusA A234E;

[0383] • folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and

[0384] • metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt).

[0385] In some embodiments of the present disclosure the e. coli comprises a gene encoding GAMT, such as GAMT as described herein.

[0386] Creatine produced with the optimized enzymes

[0387] Favourably, the optimized enzymes as provided herein, are also used in methods of producing creatine. Thus, the present disclosure also provides such an aspect.

[0388] Thus another aspect of the present disclosure relates to a method of producing creatine, the method comprising: i. providing a cell, such as a bacterial cell, such as e. coli; ii. introducing into the cell: o a nucleic acid sequence encoding GAT as described herein; and o a nucleic acid sequence encoding GAMT as described herein; iii. optionally, wherein the cell is e. coli and the e. coli comprises genomic mutations as described herein; iv. growing the cell under conditions suitable for the production of creatine; and v. optionally, purifying creatine from the supernatant.

[0389] In some embodiments of the present disclosure the cell is grown in the presence of arginine and optionally glycine.

[0390] In some embodiments of the present disclosure the cell is grown in a fermentation broth comprising L-arginine and optionally glycine. 83737PC01

[0391] 41

[0392] It should be noted that embodiments and features described in the context of one of the aspects of the present invention also apply to the other aspects of the invention.

[0393] All patent and non-patent references cited in the present application, are hereby incorporated by reference in their entirety.

[0394] The invention will now be described in further details in the following non-limiting examples.

[0395] Examples

[0396] Example 1 - Materials and Methods Metabolic model-based analyses

[0397] Metabolic model analyses were carried out to verify the enzyme selection system design for glycine a midi notransferase (GAT) in E. coli. With iML1515 (Monk et al. 2017), the latest version of a continuously curated metabolic reconstruction of E. coli was employed and adapted accordingly as described in the Results section. The COBRApy metabolic modeling package toolbox (version 0.20.0) (Ebrahim et al. 2013) in combination with the Gurobi Optimizer suite (version 9.0.0) were used for model handling and simulations. 2D flux space projections were computed using the Growth-coupling-suite, which is freely accessible on Github (https: / / github.com / biosustain / Growth-coupling-suite). All scripts and workflows were implemented in Python 3.9 and run on a Windows 10 machine with 32 GB of RAM and an AMD Ryzen 5900x processor (12 cores at 3.7 GHz).

[0398] Strain Cultivation

[0399] All the growth experiments were done using M9 medium that contains IX M9 salts, 0.1 mM CaCIz, 2 mM MgSO4, 0.2% glucose, IX trace minerals solution, IX Wolfe's vitamin solution, 25 pg / mL chloramphenicol. Supplementation is added as needed.

[0400] Creatine bioconversion medium (CBM)

[0401] M9 supplemented with 0.2% L-arginine, 0.2% L-glycine Guanidinoacetic acid bioconversion medium (GBM) M9 supplemented with 0.2% L-arginine, 0.2% L-glycine, 0.05% L-proline Auxotrophy selection medium (ASM) 83737PC01

[0402] 42

[0403] M9 supplemented with 0.05% L-arginine, 0.2% L-glycine

[0404] Ornithine growth selection medium (OSM)

[0405] M9 supplemented with 0.1% L-arginine, 0.2% L-glycine, 0.05% L-proline Wolfe's vitamin solution 1000X

[0406] 10 mg / L pyridoxine hydrochloride, 5 mg / L thiamine-HCI, 5 mg / L riboflavin, 5 mg / L nicotinic acid, 5 mg / L calcium D-(+)-pantothenate, 5 mg / L p-aminobenzoic acid, 5 mg / L thioctic acid, 2 mg / L biotin, 2 mg / L folic acid, 0.1 mg / L vitamin B12. Trace minerals solution 2000X

[0407] 33.4 g / L EDTA disodium salt dihydrate, 0.48 g / L copper(II) sulfate pentahydrate, 0.54 g / L cobalt chloride hexahydrate, 0.54 g / L zinc sulfate heptahydrate, 2 g / L calcium chloride dihydrate, 41.8 g / L iron trichloride hexahydrate, 0.3 g / L manganese(II) sulfate monohydrate, 0.5 g / L sodium molybdate dihydrate. Plasmids and strains

[0408] All plasmids are constructed using Golden Gate assembly or USER cloning. USER Enzyme and T4 DNA ligase were ordered from NEB. All synthetic gene blocks were codon optimized for E. coli and ordered from IDT.

[0409] PCR cleanup and PCR purification from agarose gel was done using the MACHEREY-NAGEL kit NucleoSpin Gel and PCR Clean-up. Plasmid miniprep was performed with the MACHEREY-NAGEL kit NucleoSpin Plasmid.

[0410] Error-prone PCR was performed according to the supplier instructions using the GeneMorph II random mutagenesis kit from Agilent Technologies. Colony PCR was performed using NEB OneTaq 2X Master Mix with Standard Buffer, or NEB Q5 High-Fidelity DNA Polymerase if the PCR product was subsequently sequenced. NEB BsaI-HF®v2 was used for Bsal restriction, including Golden Gate reactions. Constructed plasmids were cloned into One Shot™ TOP10 Chemically Competent E. coli cells according to supplier instructions. Mix2Seq kits for Sanger sequencing were purchased from Eurofins Genomics and used to verify open reading frames and Golden Gate junctions. Verified plasmids were transformed into relevant strains using either the chemical transformation or electroporation. Transformants were selected on LB agar supplemented with relevant antibiotics for selection. E. coli strains were grown at 37 °C on LB agar or in liquid LB broth at 225 rpm for overnight cultures. Lower temperatures were used for directed evolution experiments and growth profiler characterizations as described. 100 mg / L L- cysteine was added to all media used to grow strains derived from SDT651. 83737PC01

[0411] 43

[0412] Table 2: Plasmids 83737PC01

[0413] 44 83737PC01

[0414] 45 83737PC01

[0415] 46 83737PC01

[0416] 47 83737PC01

[0417] 48

[0418] Table 3. Strains 83737PC01

[0419] 49 83737PC01

[0420] 50 83737PC01

[0421] 51 83737PC01

[0422] 52 83737PC01

[0423] 53 83737PC01

[0424] 54 83737PC01

[0425] 55

[0426] Plasmid backbone construction pSD137 was constructed by amplifying an existing plasmid backbone with pl5A ori and AmpR and assembling it with a linker that contains two Bsal sites. pSD168 was constructed through Golden Gate assembly using pSD136, GAMT4, GAT8. pSD221 was constructed through ligation of PCR-amplified DNA fragments containing CmR and pSClOl ori, and annealed oligos for promoter J23100 and RBSU058.

[0427] Vector pSD249 was constructed by amplifying the backbone of pSD137 using SD_PR733 and SD_PR547. Promoter P2 and an RBS were incorporated into the tail of SD_PR733. The PCR product was purified from agarose gel and digested for 1 hour at 37 °C with Bsal. The digested PCR product was purified with a PCR cleanup column. A linker region for Golden Gate cloning was obtained by annealing oligos SD_PR495 and SD_PR496. The Bsal-digested PCR product and the compatible linker region were ligated using T4 DNA ligase.

[0428] GAT8 expression plasmids

[0429] PSD296-pSD302, used for varied expression level of wild-type GAT8 in E. coli, were constructed by Golden Gate assembly of the GAT8 gBIock and, pSD289- pSD295 respectively,. Plasmids used for expression of evolved GAT8 variants, pSD327-pSD329, were constructed by amplification of the evolved GAT8 variants from plasmids isolated from the selected population using primers SD_PR512 and SD_PR513, and Golden Gate assembly of this product to the vector pSD293. Plasmids used for co-expression of evolved GAT8 variants with GAMT4, pSD330- pSD332, were constructed in the same way, using primers SD_PR540 and SD_PR513 for insert amplification and pSD284 as the vector.

[0430] For co-expression of GAMT4 and the epPCR-generated variants of GAT8, all of pSD168 except the region encoding GAT8 was amplified using SD_PR692 and 83737PC01

[0431] 56

[0432] SD_PR531 to generate a Golden Gate compatible PCR product. A linker region was generated by annealing SD_PR716 and SD_PR496. GAMT4 co-expression vector pSD284 was constructed from PCR product and linker region using Golden Gate assembly.

[0433] Golden Gate compatible vectors pSD289-pSD295, used as vectors for GAT8, were constructed by amplification the backbone of pSD221 with SD_PR692 and SD_PR709. After Dpnl digestion and PCR cleanup, the backbone was digested with Bsal. For each promoter and RBS incorporated into a vector, two complementary single stranded oligos were ordered and annealed to generate dsDNA fragments consisting of: 1) the relevant promoter or RBS and 2) sticky ends compatible with either another annealed fragment or the Bsal-digested pSD221 backbone. A Golden Gate linker region was prepared by annealing SD_PR495 and SD_PR496. Finally, Thermo Scientific T4 DNA ligase was used to ligate the pSD221 backbone to the three fragments generated by oligo annealing. The features of the plasmids are shown in Table 2. Oligos used to generate fragments are listed in table 8.

[0434] OCD / OAT expression plasmids

[0435] Golden-Gate compatible gBIocks™ encoding codon-optimized OCD and OAT variants were ordered from Integrated DNA Technologies, Inc. (IDT). The gBIocks were cloned in front of the P2 promoter by Golden Gate assembly with pSD249 to construct OCD and OAT expression plasmids pSD260-pSD267. Plasmid libraries pSD276-pSD281 with randomly mutated variants of OCDv4, OCD_AGRT4, OCD_BRUSU, OCD_AGRFC, OAT_HUMAN and OAT_BACSU were generated by epPCR of the OCD / OAT gBIocks and Golden Gate assembly with pSD249. Refer to Table 8 for primers used in epPCR.

[0436] GAT8 expression plasmids

[0437] PSD296-pSD302, used for varied expression level of wild-type GAT8 in E. coli, were constructed by Golden Gate assembly of the GAT8 gBIock and, respectively, pSD289-pSD295. Plasmids used for expression of evolved GAT8 variants, pSD327-pSD329, were constructed by amplification of the evolved GAT8 variants from plasmids isolated from the selected population using primers SD_PR512 and SD_PR513, and Golden Gate assembly of this product to the vector pSD293. 83737PC01

[0438] 57

[0439] Plasmids used for co-expression of evolved GAT8 variants with GAMT4, pSD330- pSD332, were constructed in the same way, using primers SD_PR540 and SD_PR513 for insert amplification and pSD284 as the vector.

[0440] Plasmid library construction for GAT selection

[0441] Variants of the GAT8 gene were generated by error-prone PCR using either the wild-type GAT8 gBIock or lysate from previous selection rounds as template. Primers SD_PR512 and SD_PR513 were each used at a concentration of 160 nM, and template concentration was 20 pg / pL, which according to the GeneMorph II Random Mutagenesis Kit manual should result in a mutation frequency of 9-16 mutations / kb. The epPCR product was gel purified and cloned into pSD221, pSD292, pSD293 and pSD294 using Golden Gate assembly. To maximize the number of transformed GAT8 variants, the Golden Gate reactions were concentrated prior to transformation. This was done by removing salts and small DNA fragments using AMPure XP Bead purification (Beckman Coulter Life Sciences), followed by elution in Milli-Q water.

[0442] Transformation of GAT plasmid libraries

[0443] Populations used for selection for increased GAT activity were generated by electroporation of purified and concentrated Golden Gate reaction product into the AproB, AargA, AastABCDE strain SDT1066. To increase the number of distinct transformants, four transformations were performed to generate each population. For each population, after outgrowth in S.O.C. medium, recovered cells from four transformations were mixed in a 250 mL baffled Erlenmeyer flask. Serial dilutions of the mixture were plated on LB with ampicillin and chloramphenicol to estimate the number of transformed cells prior to amplification of the library. The mixture was then diluted 1:3 with 2xYT, chloramphenicol and ampicillin were added, and the population was incubated 30 °C and 225 rpm for 2 hours and 30 minutes for the first round of proline auxotrophy-based selection, or overnight for subsequent rounds. The amplified population was then stored at -70 °C as glycerol stocks until its revival for the subsequent selection experiment.

[0444] Auxotrophy-based growth-coupled selection

[0445] In the first round of proline auxotrophy-based selection, for each of the three populations used for selection, two vials of frozen glycerol stock with amplified 83737PC01

[0446] 58 population were thawed and washed twice in ASM. The washed populations were used in their entirety to inoculate 50 mL of ASM in 250 mL baffled Erlenmeyer flasks. The flasks were incubated at 30 °C and 225 rpm until ODeoo stopped increasing, at which point the culture was used to inoculate 50 mL fresh ASM medium. After the second flask stopped growing, the culture was passed once more, resulting in a total of three flasks per selected population. The outgrown, selected population from each flask was saved and used to screen for GAT variants by streaking for single colonies on LB with chloramphenicol and ampicillin.

[0447] In the second round of auxotrophy-based selection, lysate from flask 1 of the first round was used template for a second round of GAT epPCR. The PCR product was treated with Dpnl to remove template, purified, and cloned into pSD292 using Golden Gate assembly. Transformation of the resulting library, selection conditions, and number of passes to fresh media were all the same as in round 1. Ornithine biomass-based growth-coupled selection The population used for selection of improved GAT variants with biomass-based growth-coupled selection was created by transformation of SDT969 with epGAT8 cloned into pSD221. The template for epGAT8 was lysate from round 1 of auxotrophy-based selection. SD_PR512 and SD_PR513 were used as primers for epPCR. Inoculation of selection medium, passing to fresh medium and incubation conditions were the same as for auxotrophy-based selection. The selection medium used was OSM.

[0448] Identification of enriched GAT variants

[0449] Glycerol cryostocks with populations sampled from selection experiments were streaked out on LB with ampicillin and chloramphenicol. Colony PCR was performed on restreaked single colonies using NEB Q5® High-Fidelity 2X Master Mix with primers SD_PR710 and SD_PR543 to amplify variants of GAT. PCR products were treated with ThermoFisher ExoSAP-IT™ Express PCR Product Cleanup and Sanger sequenced using SD_PR710 and SD_PR543.

[0450] Characterization of GAT variants

[0451] Colonies with unique GAT variants identified by Sanger sequencing were used as template in colony PCR to amplify the GAT variants with overhangs compatible with Golden Gate cloning into pSD284, using primers SD_PR540 and SDPR513, and pSD293, using primers SD_PR512 and SD_PR513. Amplified GAT8 variants 83737PC01

[0452] 59 were cloned into the two plasmids, and the insert sequences and orientations in the resulting constructs were verified by Sanger sequencing.

[0453] Gene knockout

[0454] MAD7 gRNA plasmid pSD259 was constructed by replacing the gRNA targeting fhuA in pGE13 (in-house plasmid) with two gRNAs targeting argA and astD. The pGE13 backbone was amplified with Phusion U Hot Start DNA Polymerase using primers SD_PR728 and SD_PR729. These primers have tails encoding the two desired gRNAs as well as USER overhangs. The PCR product was digested with Dpnl and gel purified. To circularize the plasmid, 2 pL PCR product was mixed with 6 pL MilliQ water, 1 pL 10X T4 DNA ligase reaction buffer (NEB), and 1 pL NEB USER™ Enzyme, and incubated at 37 °C for one hour. To ligate the nicked, circular plasmid, 2.5 U of T4 DNA ligase (NEB) was added to the reaction and incubated at room temperature for another 15 minutes before transformation. To delete the argA gene and astABCDE operon, strain SDT640 was first transformed with the temperature-sensitive plasmid pGE3 carrying A- recombinases, MAD7 nuclease, and an ampicillin resistance gene. A preculture was grown overnight in LB with ampicillin at 30 °C and 225 rpm and used to inoculate 50 mL LB with ampicillin, which was incubated in a baffled Erlenmeyer flask until OD reached 0.5. The A Red system was induced by adding arabinose to a final concentration of 0.2%, and the culture was then incubated at 37 °C, 225 rpm for 45 minutes. The induced culture was then made electrocompetent by washing with glycerol. The repair templates SD_PR568 and SD_PR730, along with the MAD7 gRNA-expressing plasmid pSD259, were then simultaneously transformed by electroporation. After recovery, cells were transferred to 2 mL LB supplemented with ampicillin and chloramphenicol and incubated overnight at 30 °C, 225 rpm. Following outgrowth, culture was streaked onto LB with ampicillin and chloramphenicol to select for transformants. Deletion of argA and astABCDE was verified by colony PCR.

[0455] Plasmid curing

[0456] Plasmids from evolved CALE strains were cured using the pFREE plasmid curing system (Lauritsen et al.). CALE isolates transformed with pSD228, a modified version of the original pFREE plasmid, where kanR was replaced with ampR. Transformation was done using the TSS buffer method. After 2 hours of outgrowth 83737PC01

[0457] 60 in 1 mL LB with 100 mg / L cysteine at 30 °C, the pFREE system was induced by addition of 3 mL LB, 40 pL 20% rhamnose and 0.4 pL 2 mg / mL anhydrotetracycline. 4pL 100 mg / mL ampicillin was also added to select for transformation with pSD228. Additional cysteine was added to bring the concentration to 100 mg / L. The 4 mL induced cultures were transferred to 14 mL culture tubes and incubated for 24 hours at 30 °C. The cultures were then plated onto LB agar with 100 mg / L cysteine to obtain single colonies. To test for plasmid curing, individual colonies were checked for growth on LB agar plates containing the relevant antibiotic: ampicillin for pSD228, kanamycin for pHMll, and chloramphenicol for pSD158.

[0458] Chemical analysis

[0459] Creatine assay

[0460] Whole-cell bioconversions of arginine and glycine to GAA and creatine were used to assess improvements in enzyme activity from wild-type GAT8 to selected GAT variants.

[0461] Precultures were growth overnight in 2xYT at 30 °C and 225 rpm. Precultures were washed once with 0.85% NaCI and once with CBM. Before proceeding to bioconversion, a single sample of 0.9 UODeoo was taken from each washed preculture for intracellular GAA measurement.

[0462] Unless otherwise specified, triplicate bioconversions were performed for each preculture by resuspending washed precultures in three separate 2 mL volumes CBM to an ODeoo of 20. Bioconversion mixtures were incubated at 30 °C and 225 rpm for 18 hours.

[0463] After bioconversion, cells were separated from supernatant by centrifugation for 5 minutes at 4000xg. Creatine content in the culture supernatant was measured using "Creatine Assay Kit" from Sigma-Aldrich (MAK079) according to the manufacturer's protocol. The supernatant was diluted with water 100 times to fit in the linear range of the assay.

[0464] In the bioconversion assay performed on CALE strains retransformed with pHMll and pSD158, only a single bioconversion was performed per strain (n = l). The creatine assay was then performed on three samples from each bioconversion tube. 83737PC01

[0465] 61

[0466] GAA bioconversion

[0467] Precultures were growth overnight in 2xYT at 30 °C and 225 rpm. Precultures were washed once with 0.85% NaCI and once with GBM. Before proceeding to bioconversion, a single sample of 0.9 UODeoo was taken from each washed preculture for intracellular GAA measurement. Triplicate bioconversion was performed for each preculture by resuspending washed precultures in three separate 2 mL volumes GBM to an ODeoo of 20. Bioconversion mixtures were incubated at 30 °C and 225 rpm for 20 hours.

[0468] After bioconversion, cells were separated from supernatant by centrifugation for 5 minutes at 4000xg. Intracellular and extracellular GAA was measured by UHPLC- MS / MS.

[0469] Ultra-high performance liquid chromatography coupled to tandem mass spectrometry (UHPLC-MS / MS) for analysing guanidinoacetic acid (GAA)

[0470] A commercial standard for guanidinoacetic acid (GAA) was purchased from Sigma Aldrich (Merck Life Science A / S, Soborg, Denmark). In addition, LC-MS grade acetonitrile (ACN) and 99% pure formic acid (HCOOH) were purchased from VWR (Avantor, Radnor, PA, USA). A stock solution of the analyte was prepared in 20% HCOOH to a concentration of 5.4 mg mL . Working standard solutions of the stock solution were then prepared in H2O. A calibration curve in the concentrations of 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 25, 50, 100, 250, 500 and 1000 ng mL'1was prepared in H2O.

[0471] After cell harvest, the supernatant was transferred to a new 1.5 mL Eppendorf tube and centrifuged at 16000 x g for 15 minutes. The supernatant was transferred to a new 1.5 mL Eppendorf tube and centrifuged at 16000 x g for 15 minutes. The final supernatant was filtered into a new Eppendorf tube using a 0.2 pm syringe filter. Finally, the samples were diluted using a dilution factor of 1:500 before analysis. The dilution was done in two steps. First, a 1: 50 dilution was created by pipetting 980 pL of H2O and 20 pL of original sample into a glass vial. The solution was vortex mixed after which the final 1: 500 dilution was prepared by pipetting 900 pL of H2O into a glass vial and by adding 100 pL of the 1: 50 diluted sample to this vial. The solution was vortex mixed before analysis.

[0472] Before analyses, the intra-cellular samples were extracted according to the following procedure: The cell pellet was washed twice in ice-cold 0.85% NaCI. The 83737PC01

[0473] 62 volume of washed cell suspension, which would result in an OD-600-value of 1 when resuspended in 1 mL, was transferred to a new 1.5 mL Eppendorf tube and centrifuged at 5000 x g for 5 minutes. The supernatant was discarded, and 300 pL of 0.1% formic acid was added to the washed cell pellet. This results in an OD600 of 3. The samples were vortex mixed and sonicated in a Bioruptor® Plus sonication device for 20 cycles of 30 seconds ON and 30 seconds OFF at the HIGH setting, before they were centrifuged at 16000 x g (4 °C) for 5 minutes. The supernatant was diluted 1: 3 by adding 600 pL 1% formic acid. The diluted supernatant was then filtered through a 0.2 pm syringe filter into a 1.5 mL Eppendorf tube, and finally transferred to a glass vial with an insert before analysis.

[0474] The samples were randomized after sample preparation and analysed by ultra- high performance liquid chromatography (ExionLC, SCIEX, Framingham, MA, USA) coupled to tandem mass spectrometry (SCIEX Triple Quad 6500+, SCIEX) using electrospray ionization in positive ion mode (UHPLC-MS / MS). Selected reaction monitoring was used for quantifying GAA and declustering potential, entrance potential, collision energy and collision cell exit potential values were optimized for each ion transition m / z 118 72 and m / z 118 101). Nitrogen generated by an

[0475] Infinity 1031 nitrogen generator (Peak Scientific Instruments Ltd, Inchinnan, Scotland, UK) is used as the ion source gas, the curtain gas and as the collisionally activated dissociation gas. The analytes were separated chromatographically before they entered the mass spectrometer. This was done using a Luna® NH2 Column (2 mm x 100 mm, particle size 3 pm) from Phenomenex (Torrance, CA, USA) and H2O + 0.1% (v / v) HCOOH as eluent A and ACN + 0.1% (v / v) HCOOH as eluent B with a flow rate of 0.35 pL min-1. The elution gradient was as follows: 0-0.5 min 90% B, 0.5-1 min 90% to 5% B, 1-2.9 min 5% B, 2.9-3 min 5% to 90% B. After each run, the column was reequilibrated at 90% B for 3 min. The injection volume for each sample was 5 pL and the column oven temperature was maintained at 40 °C. All data was acquired and processed using the SCIEX OS software (version 3.0.0.3399).

[0476] Growth characterization

[0477] Growth curves were generated by cultivating strains in a Growth Profiler 960 (Enzyscreen) using 96-well CR1496dg half-deepwell microplates. G-values were converted to ODeoo using a calibration curve made with WT BW25113. Unless 83737PC01

[0478] 63 otherwise specified, strains were grown without replicates overnight in LB with appropriate antibiotics. Overnight cultures were washed in M9 medium and used to inoculate 250 pL medium in the 96-well plate to an initial OD of 0.05. The number of replicate wells was 2, 3 or 4 depending on the specific experiment. The growth profiler was set to 30 °C and 225 rpm, with pictures taken every 20 minutes.

[0479] Unless stated otherwise, all media used for growth profiler experiments were M9- based, containing the following: IX M9 salts, 0.1 mM CaC , 2 mM MgSO4, IX trace minerals solution, IX Wolfe's vitamin solution.

[0480] Adaptive laboratory evolution (ALE)

[0481] To obtain growth on proline or ornithine, we designed several ALE runs to achieve the target growth. First, we grew BW25113 AastA in M9 media supplemented with 4g / L proline and 0.6g / L glycerol. Glycerol was gradually reduced during evolution. The end point ALE population was transformed with either OCD epPCR library and selected by growing in 4g / L ornithine, resulting SDT913. Meanwhile, the end point ALE population was also transformed with plasmid pSD219 containing OCDwt.

[0482] One colony (SDT920) was selected by growthing in 4g / L ornithine.

[0483] In parallel, we performed another ALE run by growing SDT652 in M9 media with both Proline(4g / L) and Ornithine(4g / L), gradually removing glycerol (Proline_Ornithine ALE). No reproducible growth was observed in this ALE run. We plated the end point ALE culture on M9 agar plates containing 4g / L ornithine and identified one colony SDT844 that grows well on ornithine, resulted in OCDv3

[0484] To address toxicity from intracellular creatine, creatine tolerance ALE (CALE) was performed by culturing the strain SDT1004 in medium with increasing GAA concentration. Supplementing the medium with methionine enabled coupling of GAMT activity through, preventing deletion of the GAMT gene during evolution (Luo et al., 2019).

[0485] The GAA / creatine tolerance ALE experiment was conducted by cultivating SDT1004 (KanR, CmR) in M9 medium supplemented with 2 g / L glucose, 2.24 g / L L-methionine and required antibiotics. We fed the strain with an increased level of GAA concentrations, starting from 2.2475mM and reaching 29.653mM after over 83737PC01

[0486] 64

[0487] 100 generations. To further increase the stress level, we also added creatine in the media starting from 1.07g / L reaching 12.84g / L at the end.

[0488] One end isolate from five parallel ALE experiments were randomly picked and cured of pHMll and pSD158 (which may carry mutations) using the pFREE system, yielding SDT1238, SDT1239, SDT1241, SDT1242 and SDT1243. For creatine production tests compared to SDT1004, pHMll and pSD158 were reintroduced to the strains. For comparing creatine production with different GAT8 variants generated by selection, SDT1243 was transformed with each variant cloned into pSD293.

[0489] Example 2 - Creatine production in Escherichia coli

[0490] The biosynthesis of creatine from L-arginine and glycine requires the heterogenous enzyme glycine a midi notransferase (GAT, EC 2.1.4.1) to produce guanidinoacetate (GAA) and guanidinoacetate N-methyltransferase (GAMT, EC 2. 1.1.2) for the subsequent methylation to creatine (Figure 1A). We selected 8 GAT and 5 GAMT enzyme sequences from UniProt, covering a large diversity both in their sources and the sequences. GAT sequences were taken from prokaryotes and eukaryotes, including several classes of vertebrates, and GAMT from several vertebrate classes, since that enzyme is exclusively found in vertebrates. The enzyme sequences were subsequently codon-optimized for expression in E. coli and cloned into expression vectors.

[0491] The creatine biosynthetic pathway was establsihed in E. coli. To establish the creatine biosynthesis pathway in E. coli BW25113 after verifying that extracellular supplementation with GAA and creatine did not result in significant growth inhibition within solubility limits (Figure 2). We constructed co-expression plasmids, bearing both GAT and GAMT, for an initial assessment of the creatine production levels in E. coli. To this end, a single GAMT variant (GAMT4) was combined into co-expression cassettes with 6 different GAT variants. Introduction of the combined creatine production pathway into E. coli strain BW25113 resulted in accumulation of creatine in the supernatant (Figure lB)(creatine in cell lysate below the detection limit) . The different GAT variants were compared by measuring creatine levels after 25 hours of incubation. Of the six successfully constructed variants, CyrA from Cylindrospermopsis raciborskii AWT205 (UniProt B0LI36_9CYAN), displayed the highest creatine concentrations (Figure 1C). 83737PC01

[0492] 65

[0493] Therefore, we chose to use this enzyme as starting point for subsequent enzyme optimization efforts, from here on referred to as GAT8.

[0494] Example 3 - Design of a growth-coupled selection system for GAT

[0495] To circumvent the need for time-consuming GAA and creatine assays, we opted to couple the biosynthesis pathway to growth, which can be used as high-throughput selection and evolution experiments. Finding strategies that strictly couple a metabolic functionality to growth remains a challenging task due to the inherent complexity of metabolic networks and the high redundancy of their reactions. Constraint-based metabolic model approaches have successfully been used to deriving growth-coupling strain designs by harnessing the dense biochemical and phenotypical knowledge embedded in these models (Klamt and Mahadevan 2015; Alter and Ebert 2019). We utilized an automated workflow based on the gcOpt algorithm to find enzyme selection systems for GAT (Alter and Ebert, 2019). Here, we did not yet focus on the second enzyme in the pathway, a methyltransferase, but note that an existing coupling strategy could be used to improve this enzyme (Luo et al. 2019).

[0496] The computationally generated design for growth-coupling of amidinotransferase activity utilizeds arginine and glycine as sole carbon sources, which are funneled to ornithine via GAT. By heterologously expressing an ornithine cyclodeaminase (OCD) OCD and eliminating competing ornithine sinks, proline synthesis and therefore GAT activity become essential for growth. This L-ornithine is used as both carbon and energy source through a conversion to L-proline by a heterologously expressed ornithine cyclodeaminase (OCD) (Figure ID). The design is thus based on the model prediction that E. coli, possibly after some metabolic readjustments, could efficiently grow on L-proline as sole carbon and energy source. It is crucial that there exist no alternative L-arginine utilization pathways, that circumvented the amidinotransferase reaction, in this enzyme selection system. Two main pathways that could enable escape from the designed flux pattern were identified via modeling: AST and SPE (encoded by astABCDE and speAB respectively).

[0497] When this enzyme selection system is applied to GAT, an increase in its activity would lead to higher growth rates, within an operating window where GAT activity is a bottleneck for arginine utilization. An in silico analysis of the selection system 83737PC01

[0498] 66 design using flux balance analysis confirms that GAT activity is enforced for any given growth rate (Figure IE). The maximum growth rate is predicted at 0.32 Fr1for an L-arginine uptake rate of 8 mmol gDW-1h-1corresponding to a carbon uptake rate of 48 C-mmol gDW-1h’1. To partially relieve the inflicted metabolic burden on the mutant cell, glycine, the second GAT substrate besides L-arginine, could be co-fed. When employing the same total carbon uptake rate for L- arginine, cofeeding of glycine and L-arginine at an equal molar ratio (6 mmol gDW-1h-1) increases the maximum growth rate to 0.42 h-1while preserving growth-coupling of GAT, above a minimum growth rate. Altogether, these simulations confirm the feasibility of the GAT growth-coupling design and establish a method to alleviate flux requirements by co-feeding.

[0499] Example 4 - Establishment of growth-coupled design of GAT in E. coli The GAT selection system utilized in the strain design replied on the host organism to be able to grow on L-ornithine as sole carbon source and that an improvement in GAT activity is reflected in L-ornithine availability. In practice, this required the OCD enzyme and all downstream reactions to be able to carry sufficient metabolic flux to enable growth on ornithine while the selection pressure is left at the GAT reaction. We therefore selected a synthetic OCD variant (OCDK205G / M86K / TI62A from Pseudomonas putida, OCDvl), that was reported to have a 2.85-fold increased catalytic efficiency compared to the WT (OCDWT), to implement the selection system (Long et. Al. 2020). We subsequently performed growth test for E. coli BW25113 with and without OCDvl, a speA knockout and astA knockout, with all relevant carbon sources (Table 4). Some growth was observed when strains were grown on proline, which supports the possibility of using proline as intermediate for biomass, as found in the selection system. Nevertheless, when supplying ornithine to a strain expressing OCDvl, no growth was observed. This could be explained by the already inefficient growth on proline, insufficient OCD activity, or a combination thereof. Moreover, slight growth was observed for some strains in the presence of arginine and glycine, even when the SPE pathway was blocked, but not when the AST pathway was knocked out. We therefore concluded that the AST pathway has the potential to provide an escape for the selection system and accordingly used an AST deficient strain in the subsequent experiments. 83737PC01

[0500] 67

[0501] Table 4. Growth on different carbon / energy source (0.2%)*

[0502] We sought to verify whether the observed lack of growth on ornithine could be explained by low OCDvl activity. Therefore, a proline auxotroph (AproB) strain (SDT640) was employed, such that when ornithine is provided, an active OCD converts ornithine into proline to rescue cell growth. In most cases, however, OCDvl could not restore growth of the / a proline auxotroph when a combination of ornithine and glucose was supplemented (Figure 3A). However, with extended incubation time, one mutant arose at the late stage of growth, OCDK205G / M86T / TI62A (OCDv2), which showed better growth than OCDvl and used for later experiments. 83737PC01

[0503] 68

[0504] Improving growth on proline and ornithine

[0505] With more potent OCD candidates identified, the challenge of slow growth of E. coli on proline remained (Table 4). It was not immediately obvious which enzymatic step was the limiting factor for proline-based growth, we thus chose to use adaptive laboratory evolution (ALE) to overcome the proline-based growth challenge. During ALE, cultures of BW25113 and a astA strain were grown in the presence of proline and a low concentration of glycerol, with glycerol concentration being decreased each passage. After several passages, growth on proline was achieved for both BW25113 and the astA strain, with growth rates above 0.2 / h (Figure 3B).

[0506] .We subsequently used a directed evolution approach to identify improved OCD variants, that would increase the possible flux through the enzyme. Error-prone PCR (epPCR) was used on OCDWT to generate variants, which were cloned into an expression vector and transformed into a population of proline-growers obtained through ALE. The transformant library was plated onto a plate with ornithine as the sole carbon source. Colonies of various sizes showed up after 4-5 days. For the biggest colony, sequencing revealed four mutations (R65H, R184H, A185V, R244H; OCDv4) with respect to OCDWT.

[0507] In parallel, we also attempted to achieve growth directly on orthinine through ALE, using a astA strain carrying a plasmid expressing OCDv2 (SDT652). We performed ALE by feeding ornithine alone (Ornithine ALE) or co-feeding proline and ornithine (Proline-Ornithine ALE). Only Proline-Ornithine resulted one variant (OCDK205G / M86M / TI62A, OCDV3) that led to growth on ornithine.

[0508] To further improve the ornithine utilization, we transformed novel OCD variants with promoters and ribosome binding sites of different strengths into three genomic backgrounds from the proline and ornithine ALE experiments (Table 5). This resulted in different variants with a wide range of growth rates and lag times (Figure 3C). From the novel variant strains, we selected SDT969 to proceed with in selection experiments, because of its relatively high growth rate and low lag time.

[0509] Table 5. Example strains tested for growth on ornithine (*Proline_Ornithine ALE, **Proline ALE) 83737PC01

[0510] 69

[0511] GAT selection failure due to OCD inhibition by arginine 83737PC01

[0512] 70

[0513] With a suitable GAT selection system strain determined (SDT969), we set out to perform a selection run using an epPCR-generated GAT variant library. The library was constructed in on the pSD221 plasmid backbone with the pSClOl origin of replication, which was compatible with the pl5A origin in the OCDv4-bearing plasmid. The transformed library was first grown in non-selective media (2xYT) with chloramphenicol and ampicillin, and subsequently washed and transferred to both liquid media and agar plates, with arginine and glycine as the sole carbon sources. Finally, the transformed library was incubated, but even after approximately a week no growth could be observed.

[0514] The observed lack of growth could be explained by an exceptionally strict flux requirement for the GAT enzymes in the library and thus too much selection pressure, which could not successfully be alleviated by the supplemented glycine. Alternatively, there could be secondary effects that prevent successful selection using the current selection system implementation. Therefore, we first explored whether intracellular GAA could pose a problem. Although GAA was shown not to be toxic when added in the media, there is a risk that, if the cell has limited capacity to export the compound, it accumulates intracellularly. This could impose a bigger burden on the cell than when grown in the presence of GAA that has been supplemented extracellularly. This hypothesis was tested by growing SDT969 and SDT941 (a proline auxotroph, ! proB + OCDWT) strains transformed with GAT8 in permissive growth conditions (0.2% glucose for SDT969, 0.2% glucose and 0.1% proline for SDT941), with or without the supplementation of arginine and glycine. The strains are not dependent on arginine and glycine for carbon and energy under these conditions, which enables the detection of growth inhibition by GAA when arginine and glycine are supplemented to the GAT8 bearing strains. However, we did not observe any growth defect when arginine and glycine were supplemented (Figure 3A, 3B). This indicates that, even if GAA is accumulated in the cell, it is not disruptive to the selection system.

[0515] Another potential reason for the lack of growth observed for the full selection system could be inhibition by one of the metabolites on the heterologously expressed OCD enzyme. To investigate this, we devised growth experiments that study the effects of arginine and glycine supplementation on SDT969 growing on ornithine (Figure 4A) and SDT941 growing on glucose (Figure 4B). In both cases, conversion of ornithine to proline by OCD is a requirement for growth, which 83737PC01

[0516] 71 enables the detection of inhibitory effects of arginine or glycine on OCD. We observed that the presence of arginine severely inhibits growth when OCD is required, and higher arginine concentrations, from 0.001% to 0.2%, increase this effect. In contrast, the level of inhibition did not appear to depend on glycine supplementation, and no growth inhibition was seen when 0.2% glycine was supplemented without arginine. It can therefore be concluded that the OCD enzyme used in these experiments was inhibited by arginine. Similar phenomena were previously reported for some OCD homologs in an early enzymology study (Schindler, Sans, and Schroder 1989). We hypothesize that due to the highly similar molecular structures of ornithine and arginine, arginine can bind in the substrate binding site of OCD and block ornithine from being converted to proline. We attempted to find alternatives to OCDv4 that would not suffer from arginine inhibition by constructing a small library of OCD variants of different origins and two ornithine aminotransferases (OAT). OAT converts ornithine to pyrroline-5- carboxylate, a precursor of proline in E. coli, and could therefore be a substitute for OCD activity. Most of the OCD / OAT showed activity, some at levels comparable with OCDV4, based on their ability to rescue growth of ! proB in the absence of proline (Figure 4C). However, all of variants still suffered from apparent inhibition by arginine. We additionally created epPCR-based variant libraries for two OCD variants (OCD_AGRT4, OCD_AGRFC) and one OAT variant (OAT_HUMAN) that showed higher activity levels. The library was transformed into ! proB and grown in the presence of glucose and arginine (both at 0.2%), such that OCD / OAT variants with higher activity or more resistance to arginine would lead to better growth. Three enriched colonies were chosen for verification, but they did not display higher activity or resistance to arginine. The pervasive arginine inhibition on OCD and OAT may indicate that cross- recog nition of arginine is an inherent property of the catalytic mechanisms for similar enzymes.

[0517] In conclusion, after solving the bottlenecks of OCD activity and growth on proline, our attempt of building growth-couple with GAT selection still failed due to arginine inhibition on OCD. We then looked for a modified design and implementation with lower selection pressure. 83737PC01

[0518] 72

[0519] Example 5 - Evaluation of selection system based on proline auxotrophy

[0520] The inhibition of OCD by arginine posed a major challenge for the originally designed selection system. A high concentration of arginine was required to support the full biomass and energy needs of the cell. Yet, such a high level of arginine would also mean strong inhibition of OCD, a step in the pathway on which cell growth depends on. Since our efforts to reduce the inhibitory effect of arginine had been unsuccessful, we instead aimed to considerably reduce the concentration of arginine required for the selection system. This can be achieved by coupling GAT activity to the relatively low cellular proline requirement in a proline auxotroph, instead of full biomass and energy (Figure 5A). The auxotrophy-based selection system design follows the same principles as the initial design, where OCD was introduced to facilitate conversion of the ornithine generated by GAT to proline and the major alternative arginine utilization pathway was blocked (JXastABCDE). Additionally, the proline biosynthesis pathway is knocked out (AproB) to introduce the auxotrophy and the biosynthesis of ornithine from the main carbon source is prevented (AargA) to assure that GAT activity is the sole source for ornithine. Computational analysis of this design using flux balance analysis revealed that GAT activity was indeed growth-coupled, but with up to two orders of magnitude lower flux requirements at a growth rate of 0.2 / h (Figure 5B).

[0521] The auxotrophy-based selection system was implemented in the previously used proline auxotroph ! proB by the sequentially introduction of ! argA and / astABCDE, followed by transformation with a plasmid containing OCDv4 under control of P2-RBSU100 (Materials and Methods). The resulting strain, SDT1066 (J proB argA astABCDE + OCDv4), is an auxotroph for both proline and arginine, but can synthesize both arginine and proline when ornithine is available. To determine the operating window of the auxotrophy-based selection system, the strain was assessed under various conditions. We observed that the selection system strain SDT1066 reached a maximum OD when ornithine is supplemented at 0.02%, which is an indication of the cellular needs for both arginine and proline (Figure 6A). Similarly, we estimated the arginine requirement to be approximately 0.005%, based on the maximum OD observed when growing a ! argA ! astABCDE strain with a range of arginine concentrations (Figure 6B). In addition, we sought to profile the inhibitory effects of arginine on OCDv4 as a function of arginine concentration in the novel selection system. At a fixed ornithine concentration of 83737PC01

[0522] 73

[0523] 0.01%, arginine starts to severely inhibit growth of SDT1066 from 0.005%, with strong correlation between growth inhibition and arginine concentration (Figure 5C).

[0524] These concentration ranges provided an indication of a possible operating window for the auxotrophy-based selection system and elucidate around which concentrations certain effects start to become significant. To get a more integral view of how the different components interact, we assessed the growth of the selection system strain SDT1066 with and without the GAT8 enzyme for a range of arginine concentrations (Figure 5D). It was observed that in the presence of GAT8, at arginine concentrations below 0.025% lower arginine concentration led to lower or no growth, corresponding to insufficient supply to meet demands of arginine and thus proline. At concentrations above 0.05%, higher arginine concentrations lead to slower growth, which corresponds to a regime where inhibition of OCDv4 starts to significantly inhibit synthesis of proline from ornithine. In addition, since little to no growth was observed when GAT8 was absent, the experiment reveals that higher GAT activity could provide an additional advantage through rapid conversion of intracellular arginine, providing relief of OCD inhibition.

[0525] We simulated the effect of an improved GAT variant in the selection system to verify that cells with improved variants would indeed gain a selective advantage. Firstly, we studied the effect of an increased ornithine concentration, which simulates improved conversion of arginine and glycine by GAT. We grew the selection system strain with GAT8 supplemented with 0.025% or 0.05% of arginine, with or without 0.01% ornithine, and observed that additional ornithine indeed led to improved growth for both arginine concentrations (Figure 5E). This indicates that at the current GAT8 expression level (pSC101-J23100-RBSU058- GAT8, pSD242), higher GAT activity can provide growth advantage.

[0526] Secondly, we sought to simulate the effect of different GAT activity levels by different GAT8 expression levels. Six different promoters with a wide range of strengths and two RBSs were used to control the expression of GAT8, resulting in 8 expression levels of GAT8. Strains harboring these expression constructs were tested for growth, with 0.05% arginine as the sole source for proline (Figure 5F). Note that the growth result of higher GAT8 expression is compounded by both the benefit of having higher GAT activity, and the burden of higher GAT8 expression. The intermediate expression construct with J23110-RBSU58 led to the shortest lag 83737PC01

[0527] 74 phase and intermediate growth rate. Both stronger and weaker expression levels led to longer lag phase, and in some cases noticeably lower growth rate.

[0528] Example 6 - GAT selection using auxotrophy

[0529] Having established an operating window for the proline auxotrophy-based selection system, we set out to perform selection on a library of GAT variants and improve GAA production rates (Figure 7A). We chose three expression levels for the GAT library selection: 123110, 123109, and 123114, in order of decreasing predicted strength (Reis and Salis 2020; "Promoters / Catalog / Anderson - Parts. Igem.Org," n.d.), all with RBSU058 as the RBS. Error-prone PCR library of GAT8 under control of these promoters and RBS were constructed and transformed into the selection strain (SDT1066, AproB AargA AastABCDE + OCDv4). Based on colony counts of plates made immediately after recovery, there were 84 thousand, 11 million, and 303 thousand transformants in the 123109-, 123110-, and 123114-based populations, respectively. Selection flasks with 50 ml selection media (M9 with glucose and 0.05% arginine) were subsequently inoculated with the libraries. The initial OD was 0.01 for the 123109-based population, and 0.02 for the 123110 and 123114-based populations. A control flask with wild-type GAT8, instead of the library, was included for each expression level, with initial OD controlled to be the same for each library.

[0530] Growth was observed for the 123110-based library after four days, with the population flask reaching an ODeoo of 3 and the control flask an ODeoo of 1.3. The other two libraries, on the other hand, did not grow quicker than the corresponding control flasks. Two passages were performed with the 123110 group, inoculating population and control flask to the same ODeoo each time. Growth in the library flask was consistently higher than in the control flask: 24 hours after the first passage, ODeoo reached 1.85 in the population flask and 1.25 in the control flask. 24 hours after the second passage, ODeoo reached 2.2 and 1.3, respectively.

[0531] Colonies were streaked from the three flasks between passages for the 123110- controlled library and plasmids were purified from 10 colonies originating from the last flask with the 123110-based population. The region of the plasmid encoding GAT was Sanger sequenced, revealing that all 10 colonies expressed the same version of GAT8, with the amino acid substitutions M175T, Q211L and R232G (GAT8-V1). To investigate the dynamics of enrichment during selection progress, the GAT8 variants in 4 colonies from flask 1 and 12 colonies from flask 2 were 83737PC01

[0532] 75 also sequenced. In flask 1, 3 of the 4 colonies carried a version of GAT8 with amino acid substitutions G25E, S181N and R232G (GAT8-V2), while one of the colonies contained a different variant, with amino acid substitutions R232G and L242S. In flask 2, 9 of the 12 colonies carried GAT8-V1, while two carried GAT8- V2. One final colony carried a variant with amino acid substitutions M182L, and G281D (GAT8-V3).

[0533] The three selected GAT variants were recloned to appropriate plasmid backbones to test their effect on growth in the selection host (SDT1066) and on their GAA production rates. When replacing wild-type GAT8 with one of the three new variants in the strain used for proline auxotrophy-based selection, an improvement in growth rate and reduction in lag phase was observed when growing the strain in the same conditions as were used to select the GAT variants (Figure 7B). We subsequently conducted bioconversion assays, where arginine and glycine are converted into GAA by the novel GAT8 variants. This assay revealed improvements in GAA production of 90%, 94%, and 73% for GAT8-V1, V2 and V3, respectively, compared with GAT8 (Figure 7C). Importantly, these findings demonstrate that the growth rate increases observed in the selection experiment were directly linked to an improvement in catalytic capability of GAT.

[0534] Example 7 - Improvement of creatine tolerance in E. coli.

[0535] After obtaining improved GAT variants, we then continued to assemble the pathway for creatine production. Although no significant growth defect was seen when added externally (Figure 2), we noticed that creatine inhibits E. coli growth when produced intracellularly (Figure 8).

[0536] To solve this problem, we design another ALE to improve the tolerance of E. coli cells to creatine (CALE). Given that external creatine did not inhibit growth, likely due to transport limitation, we introduced the GAMT enzyme (SAM-dependent methyltransferase) and fed GAA so the cells can produce creatine intracellularly. One potential risk of such design is that the cells can lose the GAMT plasmid to avoid growth inhibition by creatine. We then utilized a previously reported growthcouple strain that requires SAM-dependent methyltransferase to be active in order to grow in a specific selection condition. This strain, derived from E. coli BW25113, has rewired SAM-cycle (encoded by pHMll) and is deficient of the natural cystine synthesis; instead, cystine is provided by SAM cycle and SAM- 83737PC01

[0537] 76 dependent methylation - in this case GAMT step (encoded in pSD158) (Luo et al., 2019). Without cysteine supplementation, the strain needs to maintain GAMT plasmid to survive.

[0538] We picked an efficient natural variant GAMT3 among the candidates (Figure 9) for the creatine tolerance ALE. The strain was evolved by feeding increasing concentrations of GAA (materials and methods). We compared creatine tolerance of selected end point ALE isolates by curing the 2 plasmids (materials and methods) and retransformed with pHMll and pSD158 containing GAMT3. All the strains showed improved growth and tolerance to creatine (Figure 10A). The strain that reached the highest OD (SDT1277) was selected, and its background strain SDT1243 was used as the "creatine tolerant" background for further experiments.

[0539] The mutations of selected ALE isolates were listed in Table 6. The isolates contain more than one mutation. One of the most common mutations is at MetC (cystathionine 0-lyase) which is involved in biosynthesis of methionine to cysteine. Mutations in this gene is most likely the consequence of the growth coupled- design of SAM-dependent methylation. Another common mutation is deactivation of ghxP, encoding guanine / hypoxanthine transporter GhxP. We suspect that GhxP is responsible for the reimport of creatine, thus knockout of GhxP might help reduce the intracellular concentration of creatine. We then took the advantage of Ghxp KO from KEIO collection strain and tested it tolerance to creatine. We cloned a plasmid containing GAT to BW25113 (wild type control) and GhxP KO. However, no significant difference was observed between these 2 strains (Figure 11), suggesting that the tolerance is the consequence of multiple mutations. Lastly, gene nusA, encoding transcription termination / antitermination protein NusA, also has mutation in multiple isolates. NusA is involved in the transcription termination of large number of genes, the impact of its mutation is more difficult to predict.

[0540] Table 6 83737PC01

[0541] 77

[0542] The new GAT variants lead to improved creatine production.

[0543] Although we did not identify the specific resistant mechanism, we successfully obtained creatine tolerant strains. The new improved GAT variants (GAT8_vl, v2, v3) were introduced to a creatine tolerant strain background SDT1243 together with GAMT4 resulting strain SDT1369, SDT1370 and SDT1371 respectively. Extracellular creatine concentrations were measured (materials and methods). As shown in Figure 10B, one of variants from GAT selection GAT8_vl led to over lOOmg / L creatine titre, which is 58% improvement compared to GAT8_wt. In summary, in this study we first designed and optimized GAT growth-coupled strategy to improve the efficiency of GAT. Next, we utilized a previously adaptive evolution to alleviate product inhibition of the cells. The resulting strain reached more than 50% titre increase compared to the starting strain. 83737PC01

[0544] 78

[0545] Example 8 - Sequence alignment of GAT variants.

[0546] To understand how beneficial mutations in GAT8 have consensus in other GAT sequences, we conducted an alignment of all GAT sequences that were successfully cloned in the study. The amino acid sequences of GAT3, GAT4, GAT5, GAT6, GAT8 and GAT9 were aligned using the 'Align' function of Uniprot (https : / / w )rot .org / align ) . The alignment was performed using 'Clustal

[0547] Omega' package (http: / / www.clustal.org / omega / ), and the results are shown in Table 7 and Figure 12.

[0548] Some mutations show consensus in other GAT sequences too. For example, G25 in GAT8 has correspondence in other GATs as G, N or K. E is similar to N in terms of amino acid chemistry. We thus predict that having E on this position is beneficial in other GAT sequences as well. Similarly, Q211L and G281D can be also extended to other GATs. R232G has no consensus in other GATs, but this mutation appeared 2 independent ALE runs, which is considered as a sign of parallel evolution and important for GAT8.

[0549] Table 7 GAT8 mutations and correspondence in other GATs

[0550] References

[0551] Alter, Tobias B., and Birgitta E. Ebert. 2019. "Determination of Growth-Coupling Strategies and Their Underlying Principles." BMC Bioinformatics 20 (1): 447. https: / / doi.org / 10.1186 / sl2859-019-2946-7.

[0552] Ebrahim, Ali, Joshua A. Lerman, Bernhard 0 Palsson, and Daniel R. Hyduke. 2013. "COBRApy: COnstraints-Based Reconstruction and Analysis for Python." BMC Systems Biology 7 (1): 74. https: / / doi.org / 10.1186 / 1752-0509-7-74. 83737PC01

[0553] 79

[0554] Klamt, Steffen, and Radhakrishnan Mahadevan. 2015. "On the Feasibility of Growth-Coupled Product Synthesis in Microbial Strains." Metabolic Engineering 30 (July) : 166-78. https: / / doi.Org / 10.1016 / j.ymben.2015.05.006.

[0555] Luo, Hao, Anne Sofie L. Hansen, Lei Yang, Konstantin Schneider, Mette Kristensen, Ulla Christensen, Hanne B. Christensen, et al. 2019. "Coupling S- Adenosylmethionine-Dependent Methylation to Growth : Design and Uses." PLOS Biology 17 (3) : e2007050. https: / / doi.org / 10.1371 / journal.pbio.2007050.

[0556] Lauritsen, I., Porse, A., Sommer, M. O. A., & Norholm, M. H. H. 2017. "A versatile one-step CRISPR-Cas9 based approach to plasmid-curing". Microbial cell factories, 16(1), 135. https: / / doi.org / 10.1186 / sl2934-017-0748-z

[0557] Long, M., Xu, M., Qiao, Z., Ma, Z., Osire, T., Yang, T., Zhang, X., Shao, M., & Rao, Z. 2020. "Directed Evolution of Ornithine Cyclodeaminase Using an EvoIvR-Based Growth-Coupling Strategy for Efficient Biosynthesis of l-Proline". ACS synthetic biology, 9(7), 1855-1863. https: / / doi.org / 10.1021 / acssynbio.0c00198

[0558] Phaneuf, P. V., Zielinski, D. C., Yurkovich, J. T., Johnsen, J., Szubin, R., Yang, L., Kim, S. H., Schulz, S., Wu, M., Dalldorf, C., Ozdemir, E., Lennen, R. M., Palsson, B. O., & Feist, A. M. 2021. "Escherichia coli Data-Driven Strain Design Using Aggregated Adaptive Laboratory Evolution Mutational Data". ACS synthetic biology, 10(12), 3379-3395. https: / / doi.org / 10.1021 / acssynbio. lc00337

[0559] Reis, Alexander C., and Howard M. Salis. 2020. "An Automated Model Test System for Systematic Development and Improvement of Gene Expression Models." ACS Synthetic Biology 9 (11) : 3145-56. https: / / doi.org / 10.1021 / acssynbio.0c00394. Schindler, U, N Sans, and J Schroder. 1989. "Ornithine Cyclodeaminase from Octopine Ti Plasmid Ach5: Identification, DNA Sequence, Enzyme Properties, and Comparison with Gene and Enzyme from Nopaline Ti Plasmid C58." Journal of Bacteriology 171 (2) : 847-54. https: / / doi.Org / 10.1128 / jb.171.2.847-854.1989.

[0560] Sequences

[0561] SEQ ID NO: 1 - GAT1 - DNA sequence ATGCAAGAATGCCCaGTtTGtTCCTACAATGAATGGGATCCACTTGAGGAGGTCATCGTCG GGCGTCCGGAGAACGCTAATGTTCCACCTTTTAGCGTTGAAGTCAAGGCGAACACATAC GAGAAGTACTGGCCGTTCTACCAAAAGCATGGAGGCCAAAGCTTCCCGGTCGATCACGT TAAAAAGGCCATTGAAGAAATCGAGGAGATGTGTAAAGTCCTGAAGCACGAGGGAGTAA TTGTGCAACGTCCTGAGGTGATTGATTGGTCTGTGAAGTATAAGACGCCCGATTTTGAGT 83737PC01

[0562] 80

[0563] CCACGGGGATGTACGCAGCCATGCCTCGCGACATTCTGTTGGTAGTGGGTAACGAAATT

[0564] ATTGAGGCCCCTATGGCTTGGCGTGCGCGTTTCTTCGAGTACCGTGCTTATCGTCCATTG

[0565] ATCAAGGATTACTTCCGTCGCGGCGCAAAATGGACTACCGCGCCAAAACCCACGATGGC

[0566] GGACGAACTGTACGACCAGGATTATCCAATTCGTACTGTCGAGGACCGCCACAAGTTGG

[0567] CAGCAATGGGTAAGTTCGTGACGACTGAATTCGAGCCCTGCTTCGATGCGGCAGACTTTA

[0568] TGCGCGCTGGACGTGATATCTTTGCACAGCGCTCACAGGTCACCAACTACCTGGGTATC

[0569] GAGTGGATGCGCCGTCATCTGGCCCCCGATTACATGAAAGTTCATATCATCTCGTTCAAA

[0570] GATCCCAATCCTATGCACATCGATGCGACCTTTAACATTATCGGTCCCGGCTTGGTATTAT

[0571] CGAACCCGGATCGTCCATGCCACCAAATCGAGTTGTTTAAGAAAGCTGGCTGGACGGTG

[0572] GTCACACCTCCCACGCCTTTAATTCCCGATAACCATCCTTTGTGGATGTCTTCAAAGTGGT

[0573] TAAGTATGAACGTTTTGATGTTAGATGAGAAGCGTGTGATGGTCGACGCTAACGAGACG

[0574] AGCATTCACAAAATGTTCGAAAAGCTTGGTATTAGTACCATTAAGGTGAACATCCGTCAC

[0575] GCCAACTCTTTAGGTGGGGGGTTCCACTGTTGGACGTGTGATATTCGCCGTCGCGGCAC

[0576] TTTACAGTCATATTTCCGCTAA

[0577] SEQ ID NO: 2 - GAT2 - DNA sequence

[0578] ATGAAGGAGTGTCCCGTCTGCAGCTACAACGAATGGGATCCCTTAGAAGAAGTGATTGT

[0579] CGGGCGTGCCGAAAATGCGTGCGTACCACCATTCAGCGTGGAGGTTAAGGCAAACACGT

[0580] ATGAAAAATATTGGGGCTTTTACCAGAAGTTCGGGGGAGAATCTTTTCCTAAAGATCATG

[0581] TCAAAAAAGCCATTGCCGAAATTGAGGAAATGTGTAATATCCTTAAAAAGGAAGGGGTGA

[0582] TTGTCAAGCGTCCCGACCCTATCGATTGGTCGGTGAAATATCGTACCCCGGACTTCGAAA

[0583] GCACCGGGATGTATGCGGCTATGCCACGCGACATCTTGCTTGTAGTAGGAAACGAGATC

[0584] ATTGAGGCACCGATGGCATGGCGTGCTCGTTTTTTCGAGTACCGCGCGTACCGCCGTAT

[0585] CATCAAAGATTATTTCAACAATGGTGCGAAGTGGACTACGGCGCCTAAACCAACAATGGC

[0586] AGACGAACTGTACGATCAAGATTATCCAATCCGCTCAGTAGAGGACCGTCACAAGTTAGC

[0587] TGCGCAGGGAAAATTTGTTACTACGGAATTCGAGCCCTGTTTTGATGCAGCCGATTTCAT

[0588] CCGCGCCGGACGTGACATTTTCGTTCAGCGTTCCCAGGTTACTAATTACATGGGCATTGA

[0589] GTGGATGCGTCGTCACTTAGCTCCAGACTATCGCGTGCACGTCATCAGTTTTAAAGATCC

[0590] GAATCCGATGCACATCGACACTACATTCAATATTATTGGGCCGGGTTTGGTCTTGAGTAA

[0591] CCCTGATCGTCCTTGCCACCAGATCGAACTTTTTAAAAAGGCCGGCTGGACTGTGATTCA

[0592] TCCCCCCGTGCCTCTTATTCCTGACGATCACCCCTTGTGGATGTCCTCAAAGTGGTTATC

[0593] GATGAATGTATTGATGCTTGATGAAAAGCGCGTGATGGTCGATGCGAACGAAACGTCAA

[0594] TTCAGAAGATGTTTGAGAACTTGGGCATCTCTACTATCAAGGTAAACATTCGCCATGCTAA

[0595] TTCTCTTGGAGGCGGATTTCATTGCTGGACTTGCGATATTCGCCGTCGCGGAACTTTGCA

[0596] ATCGTATTTCGATTAA

[0597] SEQ ID NO: 3 - GAT3 - DNA sequence 83737PC01

[0598] 81

[0599] ATGAAGGATTGTCCAGTCAGCTCCTTCAATGAGTGGGACCCGCTTGAAGAAGTCATTGTA

[0600] GGACGTGCAGAGAATGCGTGTGTGCCACCCTTTACGGTGGAGGTCAAAGCAAACACGTA

[0601] TGACAAGCATTGGCCATTCTACCAAAAATATGGTGGGTCATATTTTCCTAAAGACCATTTG

[0602] CAGAAGGCCGTAGCCGAAATCGAGGAAATGTGTAACATTCTTAAAATGGAAGGGGTGAC

[0603] AGTGCGTCGTCCTGACCCGATTGACTGGTCACTGAAGTATAAGACCCCCGATTTCGAAAG

[0604] CACCGGTTTATATGGAGCCATGCCACGCGATATTCTGATCGTTGTTGGCAACGAAATTAT

[0605] TGAAGCTCCGATGGCTTGGCGTGCACGTTTTTTTGAGTACCGTGCATACCGTACTATCAT

[0606] TAAGGACTACTTTCGCCGTGGTGCTAAATGGACAACCGCCCCGAAGCCTACCATGGCCG

[0607] ACGAACTGTACGATCAGGATTATCCAATCCACAGTGTTGAAGACCGTCACAAACTGGCAG

[0608] CTCAGGGCAAATTCGTTACAACGGAGTTTGAACCATGCTTTGACGCTGCAGACTTCATTC

[0609] GCGCAGGACGTGACATTTTTGTACAACGCTCCCAGGTGACAAATTATATGGGGATTGAGT

[0610] GGATGCGTAAACACCTTGCCCCTGATTACCGTGTGCATATTGTTTCGTTCAAAGATCCTAA

[0611] TCCGATGCATATTGACGCAACGTTTAACATCATCGGACCCGGACTTGTACTGTCGAATCC

[0612] CGATCGCCCATGTCACCAAATCGATTTATTCAAAAAAGCTGGGTGGACGATTGTGACACC

[0613] GCCGACACCGATTATTCCAGATGATCATCCGTTGTGGATGTCATCCAAGTGGTTGTCAAT

[0614] GAATGTGCTGATGTTAGATGAGAAACGTGTCATGGTAGATGCAAATGAAGTTCCGATTCA

[0615] GAAAATGTTCGAAAAATTAGGGATCTCCACTATTAAGGTATCTATTCGCAATGCGAACTCC

[0616] CTTGGCGGGGGGTTCCACTGCTGGACGTGTGATGTACGCCGCCGTGGCACCCTTCAATC

[0617] CTATTTTGATTAA

[0618] SEQ ID NO: 4 - GAT4 - DNA sequence

[0619] ATGTTGCAAGAGTGCCCTGTTTGCGCGTATAATGAATGGGACCCTCTGGAAGAGGTCATT

[0620] GTCGGACGCGCTGAGAACGCGCGTGTCCCGCCTTTCACGGTCGAGGTGAAGGCGAACA

[0621] CGTATGAAAAGCATTGGCCGTTCTACCAGAAGTACGGTGGTCAGTCCTTCCCTGAGGATC

[0622] ACCTGAAGAAGGCGGTGGCGGAGATCGAGGAAATGTGCAACATTCTTCGCATGGAGGGT

[0623] GTCACTGTCCAACGCCCTGAGCCAATGGACTGGAGTTTTGAGTATAACACTCCAGATTTT

[0624] ACTTCAACGGGGATGTATGCCGCTATGCCGCGCGATATTCTGATGGTGGTTGGAAACGA

[0625] GATTATTGAAGCTCCGATGGCGTGGCGCAGCCGCTTCTTCGAGTACCGCGCCTACCGTC

[0626] CCCTTATTAAGGAATATTTTCGCAAAGGTGCGAAGTGGACCACTCCCCCGAAACCTACTA

[0627] TGTCGGACGAGCTTTATGACCAAGAGTACCCTATCCGTACAGTCGAGGACCGTCATAAGC

[0628] TGGCAGCGCAGGGCAAATTTGTTACAACAGAGCATGAGCCGTGTTTCGATGCAGCAGAC

[0629] TTTATTCGCGCAGGGCGCGATCTGTTCGTACAACGCTCTCAAGTCACTAACTATATGGGA

[0630] ATTGAATGGATGCGCCGCCATTTGGCGCCCGACTATAAAGTTCATATCATTAGTTTTAAG

[0631] GACCCGAATCCGATGCATATTGACGCCACATTTAATATCATCGGGCCGGGACTTGTTTTG

[0632] TCGAATCCCGACCGTCCTTGCCGCCAAGTAGAGATGTTTGAGAAGGCGGGTTGGACTGT

[0633] AGTAAAACCGCCTACGCCGTTGATTCCAGACGATCATCCGCTTTGGATGAGCAGCAAATG 83737PC01

[0634] 82

[0635] GCTTAGCATGAATGTTCTTATGCTTGATCCCAAGCGCGTGATGTGTGATGCAAATGAACA

[0636] CACAATTCACAAGATGTTTGAGAACTTAGGAATCAAAACCATTAAAGTAAACATCCGTCAC

[0637] GCAAATTCCTTGGGCGGGGGTTTTCATTGCTGGACTACAGACGTCCGTCGCCGCGGATC

[0638] ACTGGAATCATACTTTCATTAA

[0639] SEQ ID NO: 5 - GAT5 - DNA sequence

[0640] ATGGCaGCttTGCCTGCtGTTCACGCTGATAATGAATGGTCGCCCCTGCGTGCGATTATCG

[0641] TGGGACGCGCGGGGAAATCTTGCTTTCCCTCTGCAAGTAAACGCATGGTTGAGGAaACC

[0642] ATGCCTAGTGCGCATGTACACCGCTTTCAAACTAAATCTCCATTCTCAGAGGGGCTGATT

[0643] GCGAAAGCTGAGGTTGAATTGGACCAATTTGCAGCTATCCTGGAGTCTGAGGGTATTAAA

[0644] GTCTACCGCCCTCCGAAAGACATTGATTGGCTTGCCGTCGACGGGTACACGGGAGCAAT

[0645] GCCTCGTGATGGGCTTATCTCCGTCGGAAACACGTTAATCGAAGCATGTTTCGCCTGGAA

[0646] ATGCCGTAGTCGTGAAATTGAACTTGGCTTTGCCCCTATCTTAGACGACTTAGCCCAGGA

[0647] CCCACAGGTTCGCATCGTCCGTCGTCCGGCCCACACATTCGCCGATACTTTAGACAACAA

[0648] TGAAGGGAAATGGGCAATTAACAACTCACGTCCGGCATTTGATACTGCCGATTTTATGCG

[0649] TTTTGGACACACACTGATCGGTCAGTACTCGCATGTAACAAATCAGGCAGGAGTAGACTA

[0650] TGTGCGCCAGCATTTACCTCCCGGATACCAAATTGAAATTTTAGATTCTAATGACCCTCAG

[0651] GCTATGCATATTGATGCGACAATTCTGCCCTTGCGCGAGGGACTTCTGATTTACCACCCC

[0652] CACAAGGTTACCGAGGATACTTTACGCCAGCACGCGGTCTTGGCAGACTGGGATCTTCG

[0653] CGCCTATCCTTTTATTCCTGAGGACCGTGATGAACCTCCGTTATATATGACTTCATCATGG

[0654] TTGTGCCTGAATGTGTTGGTCCTGGATGGGAAAAAAGTAGTTGTTGAAGCCGGTGACGA

[0655] ACGTACAGCAGGTTGGTTTGAAGAACTGGGAATGGAATGTATCCGTTGCCCTTTTCAACA

[0656] CGTAAACAGTATCGGAGGCTCATTCCACTGCGCAACGGTAGACCTGGTCCGTGAGGCGT AA

[0657] SEQ ID NO: 6 - GAT6 - DNA sequence

[0658] ATGTCCCTGGTTAGTGTCCATAACGAGTGGGATCCGTTGGAGGAGATCATTGTCGGGAC

[0659] TGCCGTGGGAGCACGCGTGCCACGTGCCGACCGCTCCGTATTCGCAGTTGAGTATGCGG

[0660] ATGAATATGATTCTCAGGACCAAGTACCTGCGGGACCTTACCCAGACCGTGTTTTGAAGG

[0661] AaACCGAAGAGGAACTGCATGTGTTGTCTGAAGAACTGACTAAGTTGGGGGTAACAGTT

[0662] CGCCGCCCCGGACAGCGTGATAATTCAGCGCTTGTGGCGACCCCAGATTGGCAAACAGA

[0663] CGGGTTCCACGACTACTGCCCACGTGATGGCCTTCTTGCCGTGGGACAGACCGTAATTG

[0664] AGAGCCCTATGGCACTTCGCGCGCGTTTCTTGGAGAGCCTGGCCTATAAGGATATCTTGC

[0665] TGGAATATTTCGCCAGCGGGGCACGTTGGTTATCCGCTCCTAAACCACGTCTTGCCGATG

[0666] AGATGTACGAGCCAACCGCTCCCGCGGGGCAGCGTCTGACCGATTTAGAGCCAGTCTTC

[0667] GACGCTGCGAACGTCCTTCGCTTCGGGACAGACCTTCTGTACCTGGTTAGTGATTCCGGT

[0668] AACGAATTAGGCGCAAAATGGTTGCAGTCCGCTCTTGGTAGTACCTACAAGGTGCATCCC 83737PC01

[0669] 83

[0670] TGTCGTGGGTTATATGCCTCCACGCATGTAGACTCAACTATTGTTCCCTTACGCCCCGGC CTGGTTTTGGTAAACCCAGCCCGCGTCAACGACGATAATATGCCTGATTTTCTGCGTTCG

[0671] TGGCAGACTGTTGTTTGCCCTGAGTTAGTTGACATTGGTTTTACAGGAGATAAACCTCATT

[0672] GTTCAGTCTGGATTGGTATGAATCTGCTTGTCGTGCGCCCAGACTTGGCTGTTGTCGACC

[0673] GCCGTCAGACGGGACTGATTAAGGTGTTGGAAAAACACGGTGTCGATGTATTGCCGCTG

[0674] CAATTGACACATTCTCGCACGCTTGGAGGCGGGTTCCATTGTGCGACCTTAGATGTTCGT

[0675] CGCACTGGTTCACTTGAaACCTACCGCTTCTAA

[0676] SEQ ID NO: 7 - GAT7 - DNA sequence

[0677] ATGAAGGAAATTCGCTCAATGCAGTTAAACGAGAAGGACAGCAACGTCATCCAGACGAG

[0678] CCCTGTCTGTTCATATACAGAATGGGATTTACTTGAGGAAATCATTGTTGGCGTGGTGGA

[0679] TGGTGCCTGTATCCCGCCCTGGCACGCGGCAATGGAACCTTGCCTTCCGACTCAACAGC

[0680] ACCAGTTTTTCCGTGATAACGCCGGGAAGCCCTTTCCTCAGGAACGTATTGACCTGGCCC

[0681] GCAAGGAATTGGACGAGTTCGCCCGCATCCTGGAGTGCGAAGGTGTAAAAGTCCGTCGC

[0682] CCCGAACCTAAAAACCAGTCATTGGTGTATGGGGCCCCAGGATGGTCTTCCACTGGGAT

[0683] GTACGCTGCCATGCCCCGCGATGTGCTTTTAGTCGTTGGGACTGACATCATTGAATGCCC

[0684] TTTAGCGTGGCGTTCCCGTTATTTCGAAACTGCGGCCTACAAGAAGCTGCTGAAAGAATA

[0685] TTTTCATGGAGGAGCTAAATGGTCTAGCGGCCCTAAGCCTGAACTGTCAGATGAGCAGTA

[0686] TGTAGATGGGTGGGTTGAAGACGAAGCTGCCACTTCAGCTAACCTGGTGATTACTGAGTT

[0687] CGAGCCTACATTTGATGCCGCTGATTTTACGCGTTTAGGAAAAGATATCATTGCCCAGAA

[0688] GAGTAATGTAACTAACGAGTTCGGCATTAACTGGTTGCAACGTCATTTGGGTGACGACTA

[0689] TAAAATCCACGTCTTGGAATTCAATGATATGCACCCGATGCATATTGATGCAACCTTGGTT

[0690] CCTCTTGCTCCTGGTAAATTGCTTATCAATCCTGAGCGCGTTCAGAAGATGCCTGAGATTT

[0691] TTCGCGGATGGGATGCTATTCATGCACCCAAGCCTATCATGCCAGATAGTCATCCGTTAT

[0692] ACATGACCTCTAAGTGGATCAACATGAACATCTTAATGTTGGATGAACGCCGCGTTGTCG

[0693] TAGAACGCCAAGACGAACCAATGATCAAGGCCATGAAAGGAGCCGGATTTGAGCCGATC

[0694] CTGTGCGATTTCCGCAATTTTAACTCTTTTGGTGGAAGCTTTCACTGCGCCACGGTGGAC

[0695] ATCCGTCGCCGCGGAAAACTGGAGAGCTATCTGGTGTAA

[0696] SEQ ID NO: 8 - CyrA from Cylindrospermopsis raciborskii AWT205 (UniProt

[0697] B0LI36_9CYAN) - GAT8 - DNA sequence

[0698] ATGCAGACACGTATCGTCAATAGTTGGAACGAGTGGGATGAACTGAAGGAGATGGTTGT

[0699] GGGGATTGCAGACGGAGCCTATTTTGAACCGACTGAGCCCGGCAACCGCCCTGCACTTC

[0700] GTGATAAAAATATTGCTAAAATGTTTAGCTTCCCGCGCGGACCCAAGAAGCAAGAAGTAA

[0701] CGGAAAAAGCCAATGAGGAATTAAACGGCCTGGTTGCGTTGTTGGAGTCTCAGGGAGTC

[0702] ACCGTACGTCGCCCGGAAAAGCATAACTTTGGTTTATCTGTCAAGACTCCGTTCTTTGAA GTCGAAAACCAGTACTGTGCTGTTTGTCCGCGTGATGTGATGATCACGTTTGGCAACGAG 83737PC01

[0703] 84

[0704] ATTTTAGAAGCGACTATGTCACGCCGTAGCCGTTTTTTTGAATACCTTCCATATCGCAAAC

[0705] TGGTGTATGAGTATTGGCATAAGGACCCGGATATGATTTGGAACGCCGCTCCGAAGCCT

[0706] ACCATGCAAAATGCGATGTATCGTGAGGACTTCTGGGAATGTCCTATGGAAGATCGTTTC

[0707] GAGAGTATGCACGACTTCGAGTTCTGTGTCACCCAGGACGAGGTTATCTTCGATGCAGCA

[0708] GACTGCTCCCGTTTTGGTCGTGATATTTTTGTGCAGGAGTCCATGACTACCAATCGTGCG

[0709] GGGATTCGCTGGCTTAAGCGCCATTTGGAGCCTCGTCGTTTTCGCGTTCACGATATTCAC

[0710] TTTCCTTTGGACATCTTCCCCTCCCATATCGACTGTACGTTTGTCCCGCTGGCACCGGGG

[0711] GTGGTCCTTGTTAATCCTGACCGCCCAATTAAAGAGGGTGAAGAGAAACTTTTTATGGAT

[0712] AATGGCTGGCAATTTATCGAAGCGCCACTTCCGACATCGACTGATGACGAGATGCCCATG

[0713] TTCTGTCAGAGTAGTAAGTGGCTGGCAATGAACGTACTTAGCATTAGCCCAAAAAAAGTT

[0714] ATCTGTGAGGAGCAGGAACATCCACTTCATGAGTTGTTAGACAAACATGGGTTTGAAGTG

[0715] TACCCGATCCCTTTCCGCAACGTGTTCGAATTTGGAGGGAGTTTACACTGTGCCACGTGG

[0716] GACATCCATCGCACGGGAACGTGCGAGGACTATTTTCCCAAATTAAATTATACCCCGGTG

[0717] ACCGCGTCCACTAATGGTGTCTCCCGCTTTATTATTTAA

[0718] SEQ ID NO: 9 - GAT9 - DNA sequence

[0719] ATGAAACTTAATTCTTTCGATGATTGGTCACCATTAAAGGAAATCATCGTGGGCACAGCA

[0720] CAAAACTATATCTCACACGAGCGCGACTTAAGCTTTGACATTTTCTTCCATGAGAATCTGT

[0721] TCAAGAGCGATTGGGCCTATCCACGCTTAAATAAAAGCTCGAAAGAGGTCAATGACTTAA

[0722] ATTGGCAAATCAAACAGAAGTACGTCGAAGAATTAAACGAGGACGTAGAGGATTTGGTG

[0723] AACCTGTTGAAAAAGTTCGATGTTACCGTACACCGCCCGATGATTTTACCCTTTAATCCGC

[0724] AAAATATCAAAGGCCTGGGCTGGGAGAGCGCACCTGTTCCTGCGCTTAATGTTCGTGATA

[0725] ATACGTTGATCTTAGGGGATGAGATCATTGAAACTCCACCTGCAATTCGCTCACGCTACT

[0726] TGGAGACACGCCTGTTGGCGCCAATTTTTATGAAGTATTTTGAAGCGGGCGCGACATGG

[0727] ACCACTATGCCTCGCCCCATCCTTACAGACATGTCGTTTGATTTGTCCTACGCACGTGAC

[0728] CTTGAAACCACCTTGGGAGGTCCGACGGAACCTATTGAGGATCCTGAAAGCAGTCCGTA

[0729] TGACGTTGGGTTCGAGATGATGCTGGACGGCGCACAATGTCTGCGCCTTGGTAAGGACA

[0730] TCATCGTCAACATTGCAAATCAAAACCACAAACTTGCGTGTGATTGGTTGGAGCGTCATG

[0731] TTGAAAACCGCTATCGTATCCATCGTGTATACCGTATGTCCGATAATCATATCGATAGCAT

[0732] GTTGTTAGCCTTACGCCCTGGCGTTTTCTTGGCCCGTCATAAAGATTTGAAAGATTTGCTT

[0733] CCGGAACCGTTTAAGAAATGGAAAATGATCGTCCCGCCAGAACCATCTGCTACCAATTTT

[0734] CCAACCTACGATGATGGGGATTTGCTGCTTACAAGTCCCTATATTGACCTGAACGTCCTTT

[0735] CGGTGAATCCGGAAACCGTTTTAGTCAATGAGGTATGCGTGGATCTTATTAAAACTTTGG

[0736] AAAAAGAGGGGTTCACCGTTGTTCCCGTGAAGCATCGTCATCGCCGCTTGTTCGGCGGT

[0737] GGATTCCATTGCTTTACCTTAGACACAGTTCGCGACGGAGGGTTGGAGGATTACACCATC

[0738] TAA 83737PC01

[0739] 85

[0740] SEQ ID NO: 10 - GAMT1 - DNA sequence

[0741] ATGTCAACTGCCCAACCCATCTTCTCGAAGGGCGAGGACTGCAAGGCGGGCTGGCACGA

[0742] TGCCTCTGCGGGTTACAATGAAACTGACACTCACCTTGAGATTTTCGGGAAACCTGTCAT

[0743] GGAACGCTGGGAaACCCCCTACATGCATTCACTGGCAACGATCGCTTCCAGCAAAGGTG

[0744] GCCGTGTGCTGGAGATCGGCTTCGGTATGGCTATTGCTGCGACGAAGGTGGAAAGTTTT

[0745] CCGATTGAAGAACATTGGATTATCGAATGTAATGACGGAGTGTTCGCTCGTCTGCAAGAG

[0746] TGGGCGAAAGCCCAGCCTCATCGTATTGTACCATTGAAGGGCCTGTGGGAACAAGTAGT

[0747] GGGGGGACTTCCCGACAACCATTTCGACGGTATTCTTTACGACACCTATCCGCTGTCCGA

[0748] GGATACTTGGCATACACATCAGTTCGACTTCATCAAAGGTCACGCTCATCGCCTGTTGAA

[0749] ACCGGGTGGGGTACTTACTTATTGTAACCTTACATCCTGGGGGGAGCTGCTGAAGTCCA

[0750] AGTATGACAATATTGATAATATGTTCCAAGAGACACAAATGCCCCACCTTTTGGAAGCGG

[0751] GTTTTAAGCGTGAAAAGATCTCAACGACAACAATGGACATTGTACCCCCTAGTGAGTGTA

[0752] AATACTATGCATTTCGTAAAATGATCACCCCCACGATCTTAAAAGAGTAA

[0753] SEQ ID NO: 11 - GAMT2 - DNA sequence

[0754] ATGTCCGCATTAAGTGATACAGCGCCCATTTTCACTGAGGGGGAAGATTGTAAGGCAGC

[0755] ATGGCAAGAGGCAACTGCTGCATATGACGCTCCGGATACCCACCTTGAGATCCTGGGAA

[0756] AACCCGTGATGGAGCGCTGGGAAACGCCATACATGCACTCATTGGCCACCGTTGCGGCT

[0757] TCTCGCGGGGGCCGCGTTCTGGAAGTAGGCTTTGGAATGGCGATTGCTGCAACTAAGGT

[0758] CCAGGAGTTTAACATTGAGGAACATTGGATTGTCGAATGCAACGACGGGGTATTCCAAC

[0759] GTCTGGAGGAATGGGCGCGTGTGCAACCGCACAAGGTCGTACCATTAAAAGGCTTATGG

[0760] GAAGATGTGGTACCTACGCTTCCTGACGGCCATTTCTCTGGAATTTTGTACGATACATATC

[0761] CATTGTCGGCTCAAACATGGCACACACACCAATTCGCCTTTATTAAAGACCACGCTTTTCG

[0762] CTTGTTGCGTCCAGGGGGGGTATTGACCTACTGTAACCTGACCTCTTGGGGCGAATTGTT

[0763] AAAAGGTAAGTATTCTGACATCGAAAAGATGTTCGAAGAGACACAGGTAGCACAGCTTCT

[0764] GGAAGCGGGGTTCCGCCGCGAAAACATCTCAACTGCTGTGATGGAACTGGTACCGCCCC

[0765] GCGAATGTCGCTATTATTCGTTCCCTCGTATGATTACACCGCGTGTTGTGAAACACTAA

[0766] SEQ ID NO: 12 - GAMT3 - DNA sequence

[0767] ATGAGTTCGGAAAAaATCTTcTTAGAAGGGGAGAGTTGTCAGTCTAGTTGGCACAACGCC

[0768] ACCGCTGGGTATGATGAGACAGATACTCACCTTGAAATCTTAGGGAAGCCTGTAATGGAA

[0769] CGTTGGGAGACGCCCTATATGCACTCCCTTGCTACAGTGGCCGCTTCCAAGGGGGGCCG

[0770] TGTCCTGGAGATTGGGTTCGGGATGGCAATCGCCGCCACCAAGTTGGAACAGTGCAACA

[0771] TCGAAGAACATTGGATTATTGAATGCAACGACGGCGTTTTCAAGCGTTTGCAAGAGTGGG

[0772] CCACTAAACAGCCGCACAAAATCGTTCCCTTGAAGGGCTTGTGGGAAGATGTGGTTCCG

[0773] ACATTGCCTAATGGACACTTCGACGGAATCCTGTACGACACATATCCGTTATCCGAGGAG

[0774] ACATGGCACACTCACCAGTTTAATTTCATTAAAGGTCATGCCTACCGTCTGCTGAAGCCT 83737PC01

[0775] 86

[0776] GGGGGGGTATTGACATATTGCAACTTGACCAGCTGGGGTGAATTATTAAAAACAAAATAT

[0777] AATGATATTGAGAAAATGTTTCAGGAGACGCAAACGCCTCAACTGGTAGACGCTGGGTTC

[0778] AAGTGTGAAAACATCTCCACCACCGTAATGAATTTAGTACCGCCAGAGGACTGCCGCTAT

[0779] TATAGCTTTAAgAAAATGATTACTCCTACGATCATTAAGGTTTAA

[0780] SEQ ID NO: 13 - GAMT4 - DNA sequence

[0781] ATGTCTGCTCCaGCTGCaACaCCCATTTTCGCTCCAGGCGAGAACTGTTCGCCAGCATGG

[0782] CGCGCAGCTCCTGCAGCATATGATGCTAGCGACACACATTTACAGATCTTAGGCAAACCT

[0783] GTCATGGAGCGTTGGGAGACGCCCTATATGCACGCCCTTGCTGCTGCTGCTGCATCACG

[0784] CGGGGGCCGTGTATTAGAGGTCGGATTCGGGATGGCGATTGCCGCTACTAAGGTTCAAG

[0785] AAGCGCCTATCGAGGAACATTGGATTATTGAATGTAATGAGGGCGTCTTTCAACGTCTGC

[0786] AAGATTGGGCGTTGCAACAACCCCACAAAGTTGTGCCACTTAAAGGTTTATGGGAAGAG

[0787] GTGGCTCCTACCTTGCCGGACTCACATTTCGACGGGATTCTGTATGATACCTATCCCCTG

[0788] AGCGAAGAAACATGGCACACCCACCAGTTCAACTTCATCCGTGATCACGCATTCCGTTTG

[0789] CTTAAACCAGGGGGTGTTCTGACTTACTGCAACCTGACTTCCTGGGGAGAGTTAATGAAG

[0790] ACCAAGTATAGTGACATCACGACCATGTTCGAAGAAACTCAGGTTCCAGCATTATTGGAG

[0791] GCTGGATTTCGTCGTGATAACATCCGCACTCAGGTGATGGAATTAGTACCTCCGGCCAAT

[0792] TGCCGTTACTACGCGTTCCCTCGCATGATTACACCGCTTGTTACTAAGCATTAA

[0793] SEQ ID NO: 14 - GAMT5 - DNA sequence

[0794] ATGTCtTCtTCAGAaGCaGCtTCACCTATCTTTCATCAAGGCGAAAACTGCAAGCCAAGTTG

[0795] GCGTGAGGCAGAAGCTGGATACAATCAAAAGGATACTCATCTGAAGATTTTGGGGAAAC

[0796] CGGTGATGGAACGTTGGGAAACTCCATACATGCATTCCTTGGCGACTGTAGCCGCGTCC

[0797] AAGGGAGGGCGTGTGCTTGAAGTCGGGTTCGGGATGGCGATTGCCGCGTCGAAGGTAG

[0798] AGGAGTTTAATATCGAAGAGCATTGGATTGTCGAGTGTAACGAAGGGGTTTTTAAGCGTC

[0799] TTGAAGAATGGGCGAAAAAACAGCCCCATAAGGTGGTTCCATTACGTGGATTGTGGGAG

[0800] GACGTAGTTCCTACCCTTCCCGATGGCCTGTTCGACGGAATCCTTTATGACACGTACCCA

[0801] CTTAGTGCCGAaACCTGGCATACGCATCAGTTTCACTTTATTAAAAGCCACGCTTTCCGCC

[0802] TGCTGAAGCCAGGTGGCGTCTTGACTTATTGCAACCTTACGTCTTGGGGTGAGTTACTTA

[0803] AAACTAAGTACACAGATATCGAGAAAATGTTTGAGGAGACTCAAGTCGGACACTTAGTGG

[0804] AGGCTGGATTTAAGCGCGAGAATATTTCAACTGCGGTCATGGACTTAACCCCGCCCCAG

[0805] GATTGCCGCTACTACGCTTTTCCGAAGATGATTACACCTACCATCGTTAAGCAATAA

[0806] SEQ ID NO: 15 - OCD_AGRT4 - DNA sequence

[0807] ATGCCGATTGACCCGAAACTGAACGTAGTGCCTTTCATCAGCGTCGATCACATGATGAAA

[0808] CTGGTTCTGAAAGTGGGCATTGATACTTTCCTGACCGAACTGGCCGCGGAGATTGAAAA

[0809] GGACTTCCGCCGTTGGCCGATTTTCGATAAAAAGCCGCGCGTTGGTAGCCACTCTCAGG

[0810] ATGGCGTAATCGAACTGATGCCTACTAGCGATGGTAGCCTGTACGGTTTCAAATATGTTA 83737PC01

[0811] 87

[0812] ACGGTCACCCGAAGAACACCCACCAGGGTCGTCAGACCGTGACTGCGTTCGGCGTTCTG

[0813] TCTGACGTAGGCAACGGCTACCCGCTGCTGCTGAGCGAAATGACCATCCTGACCGCCCT

[0814] GCGCACCGCAGCAACCTCCGCGCTGGCCGCGAAATATCTGGCCCGTCCGAACTCTAAAA

[0815] CCATGGCAATCATTGGCAACGGTGCGCAGAGCGAATTCCAGGCGCGTGCGTTCCGCGCT

[0816] ATTCTGGGTATTCAGAAACTGCGTCTGTTCGACATCGATACCAGCGCAACCCGCAAATGC

[0817] GCGCGTAACCTGACTGGTCCGGGTTTTGACATCGTCGAATGCGGTAGCGTTGCCGAAGC

[0818] TGTGGAAGGTGCTGACGTGATCACTACTGTAACCGCTGACAAACAGTTCGCAACTATCCT

[0819] GTCCGATAACCATGTGGGCCCGGGTGTTCATATTAATGCAGTTGGCGGCGACTGTCCTG

[0820] GTAAAACCGAAATCTCCATGGAGGTGCTGCTGCGTTCTGACATTTTTGTTGAATACCCGC

[0821] CGCAGACTTGGATCGAGGGTGACATTCAGCAGCTGCCGCGTACTCACCCAGTAACCGAG

[0822] CTGTGGCAGGTTATGACCGGTGAAAAAACCGGCCGTGTAGGTGACCGTCAGATCACCAT

[0823] GTTCGATTCCGTAGGTTTCGCTATTGAAGACTTCTCCGCACTGCGTTACGTGCGTGCTAA

[0824] AATCACTGATTTTGAAATGTTCACCGAACTGGACCTGCTGGCGGATCCGGATGAACCGCG

[0825] TGACCTGTATGGCATGCTGCTGCGCTGCGAAAAAAAACTGGAACCGACGGCGGTCGGTT AA

[0826] SEQ ID NO: 16 - OCD_BRUSU - DNA sequence

[0827] ATGCCAGCGCTGGCAAACCTGAATATCGTACCTTTCATCTCTGTAGAAAACATGATGGAT

[0828] CTGGCGGTGTCTACCGGTATCGAAAACTTCCTGGTCCAGCTGGCAGGCTACATCGAAGA

[0829] GGATTTCCGTCGCTGGGAATCTTTCGATAAAATTCCGCGTATCGCCTCCCATTCTCGTGA

[0830] TGGTGTGATCGAACTGATGCCGACCTCCGATGGTACCCTGTACGGCTTCAAATACGTAAA

[0831] CGGCCACCCGAAAAACACGAAATCTGGTCGTCAGACGGTCACTGCTTTCGGTGTTCTGA

[0832] GCGATGTTGACTCTGGTTATCCTCTGCTGCTGTCTGAAATGACCATCCTGACTGCGCTGC

[0833] GTACCGCAGCGACCTCTGCAATCGCTGCGAAATATCTGGCTCGTAAAGACTCTCGTACCA

[0834] TGGCACTGATTGGTAACGGTGCGCAGAGCGAATTTCAGGCACTGGCTTTTAAAGCTCTGA

[0835] TTGGCGTTGATCGCATTCGTCTGTATGACATCGATCCGGAAGCTACTGCGCGTTGCAGCC

[0836] GTAACCTGCAGCGTTTTGGCTTCCAGATCGAAGCGTGCACTTCTGCGGAACAGGCGGTG

[0837] GAAGGCGCGGATATCATCACCACTGCGACGGCGGACAAGCACAACGCTACTATTCTGTC

[0838] TGATAACATGATCGGTCCGGGTGTACATATCAACGGTGTGGGTGGCGATTGCCCGGGTA

[0839] AAACCGAAATGCACCGTGACATTCTGCTGCGTAGCGACATTTTTGTTGAATTCCCGCCGC

[0840] AGACCCGTATCGAAGGTGAGATCCAGCAGCTGGCTCCGGATCACCCGGTTACCGAACTG

[0841] TGGCGTGTTATGACCGGCCAGGACGTCGGTCGTAAAAGCGATAAACAGATTACCCTGTT

[0842] CGACAGCGTTGGTTTCGCAATTGAAGACTTCTCCGCACTGCGTTATGTGCGCGATCGTGT

[0843] TGAGGGTAGCTCTCACTCCTCTCCGCTGGATCTGCTGGCTGATCCGGACGAACCGCGTG

[0844] ACCTGTTCGGCATGCTGCTGCGTCGCCAGGCTTTTCGCCGTCTGGGTGGTTAA 83737PC01

[0845] 88

[0846] SEQ ID NO: 17 - OCD_AGRFC - DNA sequence

[0847] ACTGGTCTCCTCTAATGCCAGCGCTGGCAAACCTGAATATCGTACCTTTCATCTCTGTAGA

[0848] AAACATGATGGATCTGGCGGTGTCTACCGGTATCGAAAACTTCCTGGTCCAGCTGGCAG

[0849] GCTACATCGAAGAGGATTTCCGTCGCTGGGAATCTTTCGATAAAATTCCGCGTATCGCCT

[0850] CCCATTCTCGTGATGGTGTGATCGAACTGATGCCGACCTCCGATGGTACCCTGTACGGCT

[0851] TCAAATACGTAAACGGCCACCCGAAAAACACGAAATCTGGTCGTCAGACGGTCACTGCTT

[0852] TCGGTGTTCTGAGCGATGTTGACTCTGGTTATCCTCTGCTGCTGTCTGAAATGACCATCC

[0853] TGACTGCGCTGCGTACCGCAGCGACCTCTGCAATCGCTGCGAAATATCTGGCTCGTAAA

[0854] GACTCTCGTACCATGGCACTGATTGGTAACGGTGCGCAGAGCGAATTTCAGGCACTGGC

[0855] TTTTAAAGCTCTGATTGGCGTTGATCGCATTCGTCTGTATGACATCGATCCGGAAGCTACT

[0856] GCGCGTTGCAGCCGTAACCTGCAGCGTTTTGGCTTCCAGATCGAAGCGTGCACTTCTGC

[0857] GGAACAGGCGGTGGAAGGCGCGGATATCATCACCACTGCGACGGCGGACAAGCACAAC

[0858] GCTACTATTCTGTCTGATAACATGATCGGTCCGGGTGTACATATCAACGGTGTGGGTGGC

[0859] GATTGCCCGGGTAAAACCGAAATGCACCGTGACATTCTGCTGCGTAGCGACATTTTTGTT

[0860] GAATTCCCGCCGCAGACCCGTATCGAAGGTGAGATCCAGCAGCTGGCTCCGGATCACCC

[0861] GGTTACCGAACTGTGGCGTGTTATGACCGGCCAGGACGTCGGTCGTAAAAGCGATAAAC

[0862] AGATTACCCTGTTCGACAGCGTTGGTTTCGCAATTGAAGACTTCTCCGCACTGCGTTATG

[0863] TGCGCGATCGTGTTGAGGGTAGCTCTCACTCCTCTCCGCTGGATCTGCTGGCTGATCCG

[0864] GACGAACCGCGTGACCTGTTCGGCATGCTGCTGCGTCGCCAGGCTTTTCGCCGTCTGGG

[0865] TGGTTAAGACTAGAGACCTTC

[0866] SEQ ID NO: 18 - RBSU004 - DNA sequence

[0867] GGGCCCAAGTTCACTTAAAAAGGAGATCAACAATGAAAGCAATTTTCGTACTGAAACATC

[0868] TTAATCATGCGATGGACGGTT

[0869] SEQ ID NO: 19 - RBSU058 - DNA sequence

[0870] GGGCCCAAGTTCACTTAAAAAGGAGATCAACAATGAAAGCAATTTTCGTACTGAAACATC

[0871] TTAATCATGCTGCGGAGGGTT

[0872] SEQ ID NO: 20 - RBSU100 - DNA sequence

[0873] GGGCCCAAGTTCACTTAAAAAGGAGATCAACAATGAAAGCAATTTTCGTACTGAAACATC

[0874] TTAATCATGCTAAGGAGGTTT

[0875] SEQ ID NO: 21 - OAT_HUMAN - DNA sequence

[0876] ATGTTCAGCAAACTGGCGCATCTGCAGCGTTTTGCTGTTCTGAGCCGTGGTGTACACTCC

[0877] TCTGTGGCGTCTGCCACGAGCGTTGCGACCAAAAAAACCGTTCAGGGTCCGCCGACCAG

[0878] CGATGATATCTTTGAACGTGAATACAAATATGGCGCTCACAACTATCATCCGCTGCCTGTT

[0879] GCGCTGGAACGTGGTAAAGGCATTTATCTGTGGGACGTGGAAGGTCGTAAATATTTTGA

[0880] CTTCCTGTCCAGCTATTCTGCTGTTAACCAGGGCCACTGTCACCCGAAAATCGTTAACGC 83737PC01

[0881] 89

[0882] CCTGAAATCTCAAGTGGACAAACTGACCCTGACCTCCCGTGCTTTCTATAACAACGTGCT

[0883] GGGCGAATACGAGGAATACATTACCAAACTGTTTAACTATCACAAAGTTCTGCCGATGAA

[0884] CACTGGTGTGGAAGCGGGCGAAACTGCCTGTAAACTGGCCCGCAAATGGGGTTATACCG

[0885] TTAAGGGCATCCAGAAGTACAAAGCAAAAATCGTTTTTGCCGCTGGTAACTTTTGGGGCC

[0886] GTACTCTGTCCGCAATCTCTTCTAGCACTGACCCGACCAGCTATGACGGCTTCGGTCCGT

[0887] TCATGCCGGGTTTCGACATCATCCCGTATAACGATCTGCCGGCGCTGGAACGCGCACTG

[0888] CAGGACCCGAACGTTGCAGCATTCATGGTAGAACCGATTCAGGGTGAAGCTGGCGTAGT

[0889] CGTCCCGGACCCGGGTTACCTGATGGGTGTGCGTGAGCTGTGCACCCGTCATCAGGTAC

[0890] TGTTCATTGCCGACGAAATCCAGACTGGTCTGGCACGTACCGGTCGTTGGCTGGCCGTG

[0891] GATTATGAAAACGTTCGCCCGGATATCGTTCTGCTGGGTAAAGCGCTGTCCGGCGGTCT

[0892] GTACCCGGTATCTGCGGTACTGTGCGACGATGACATCATGCTGACTATCAAACCGGGCG

[0893] AACATGGTTCTACTTACGGCGGTAACCCGCTGGGCTGTCGTGTGGCAATTGCGGCACTG

[0894] GAAGTACTGGAAGAGGAAAACCTGGCGGAGAACGCTGACAAACTGGGTATCATCCTGCG

[0895] CAACGAACTGATGAAACTGCCGTCCGATGTTGTGACCGCGGTCCGTGGCAAAGGTCTGC

[0896] TGAACGCTATCGTAATCAAGGAAACCAAAGACTGGGATGCGTGGAAAGTTTGCCTGCGC

[0897] CTGCGTGATAACGGTCTGCTGGCAAAACCGACTCACGGCGATATTATCCGCTTTGCACCG

[0898] CCGCTGGTGATTAAAGAAGATGAACTGCGCGAATCCATTGAAATTATTAACAAAACCATT

[0899] CTGAGCTTTTAA

[0900] SEQ ID NO: 22 - OAT_BACSU - DNA sequence

[0901] ATGACCGCACTGTCTAAGTCTAAAGAAATCATTGACCAGACGAGCCACTATGGCGCAAAC

[0902] AACTACCACCCTCTGCCGATTGTTATCTCTGAAGCCCTGGGCGCTTGGGTGAAAGATCCG

[0903] GAAGGTAACGAATACATGGATATGCTGAGCGCGTATAGCGCAGTTAACCAGGGTCATCG

[0904] TCACCCGAAAATCATTCAGGCCCTGAAAGACCAAGCGGATAAAATCACCCTGACGTCCCG

[0905] TGCTTTCCACAACGATCAGCTGGGTCCGTTCTATGAAAAGACCGCTAAACTGACTGGCAA

[0906] AGAAATGATCCTGCCGATGAATACCGGCGCGGAAGCGGTTGAAAGCGCGGTTAAAGCTG

[0907] CGCGCCGTTGGGCTTATGAGGTAAAAGGTGTGGCAGACAACCAGGCTGAAATTATTGCT

[0908] TGCGTAGGCAACTTCCACGGTCGTACGATGCTGGCAGTTTCCCTGTCTTCCGAAGAAGAA

[0909] TACAAACGTGGCTTTGGTCCAATGCTGCCGGGTATCAAACTGATCCCGTACGGTGACGTT

[0910] GAAGCTCTGCGCCAGGCGATCACTCCGAACACTGCTGCATTCCTGTTCGAACCGATCCA

[0911] GGGCGAAGCGGGCATCGTTATTCCGCCGGAAGGTTTCCTGCAAGAAGCGGCGGCAATCT

[0912] GTAAAGAAGAGAACGTACTGTTTATTGCAGACGAGATCCAAACCGGCCTGGGTCGTACT

[0913] GGTAAGACCTTCGCGTGCGATTGGGATGGTATTGTTCCGGACATGTATATTCTGGGTAAA

[0914] GCGCTGGGTGGTGGTGTTTTCCCGATCAGCTGTATCGCTGCAGATCGCGAAATCCTGGG

[0915] CGTTTTCAATCCGGGTTCCCACGGTAGCACCTTCGGTGGTAACCCACTGGCCTGCGCAGT

[0916] ATCTATCGCGTCCCTGGAAGTCCTGGAAGATGAAAAGCTGGCAGACCGTTCTCTGGAGC 83737PC01

[0917] 90

[0918] TGGGTGAATACTTCAAATCCGAACTGGAAAGCATCGACAGCCCGGTGATCAAAGAAGTT

[0919] CGCGGCCGCGGTCTGTTCATCGGTGTTGAACTGACTGAAGCAGCACGTCCGTATTGCGA

[0920] ACGTCTGAAAGAAGAAGGTCTGCTGTGTAAAGAAACCCATGATACTGTTATCCGTTTCGC

[0921] GCCGCCGCTGATCATCAGCAAAGAGGATCTGGACTGGGCGATTGAGAAAATTAAACACG

[0922] TGCTGCGTAATGCTTAA

[0923] SEQ ID NO: 23 - GAT8-varl - DNA sequence

[0924] ATGCAGACACGTATCGTCAATAGTTGGAACGAGTGGGATGAACTGAAGGAGATGGTTGT

[0925] GGGGATTGCAGACGGAGCCTATTTTGAACCGACTGAGCCCGGCAACCGCCCTGCACTTC

[0926] GTGATAAAAATATTGCTAAAATGTTTAGCTTCCCGCGCGGACCCAAGAAGCAAGAAGTAA

[0927] CGGAAAAAGCCAATGAGGAATTAAACGGCCTGGTTGCGTTGTTGGAGTCTCAGGGAGTC

[0928] ACCGTACGTCGCCCGGAAAAGCATAACTTTGGTTTGTCTGTCAAGACTCCGTTCTTTGAA

[0929] GTCGAAAACCAGTACTGTGCTGTTTGTCCGCGTGATGTGATGATCACGTTTGGCAACGAG

[0930] ATTTTAGAAGCGACTATGTCACGCCGTAGCCGTTTTTTTGAATACCTTCCATATCGCAAAC

[0931] TGGTGTATGAGTATTGGCATAAGGACCCTGATATGATTTGGAACGCCGCTCCGAAGCCTA

[0932] CCATGCAAAATGCGATGTATCGTGAGGACTTTTGGGAATGTCCTACGGAAGATCGTTTCG

[0933] AGAGTATGCACGACTTCGAGTTCTGTGTCACCCAGGACGAGGTTATCTTCGATGCAGCAG

[0934] ACTGCTCCCGTTTTGGTCGTGATATTTTTGTGCTGGAGTCCATGACTACCAATCGTGCGG

[0935] GGATTCGCTGGCTTAAGCGCCATTTGGAGCCTCGTGGTTTTCGCGTTCACGATATTCACT

[0936] TTCCTTTGGACATCTTCCCCTCCCATATCGACTGTACGTTTGTCCCGCTGGCACCGGGGG

[0937] TGGTCCTTGTTAATCCTGACCGCCCAATTAAAGAGGGAGAAGAGAAACTTTTTATGGATA

[0938] ATGGCTGGCAATTTATCGAAGCGCCACTTCCGACATCGACTGATGACGAGATGCCCATGT

[0939] TCTGTCAGAGTAGTAAGTGGCTGGCAATGAACGTACTTAGCATTAGCCCAAAAAAAGTTA

[0940] TCTGTGAGGAGCAGGAACATCCACTTCATGAGTTGTTAGACAAACATGGGTTTGAAGTGT

[0941] ACCCGATCCCTTTCCGCAACGTGTTCGAATTTGGAGGGAGTTTACACTGTGCCACGTGGG

[0942] ACATCCATCGCACGGGAACGTGCGAGGACTATTTTCCCAAATTAAATTACACCCCGGTGA

[0943] CCGCGTCCACTAATGGTGTCTCCCGCTTTATTATTTAA

[0944] SEQ ID NO: 24 - GAT8-var2 - DNA sequence

[0945] ATGCAGACACGTATCGTCAATAGTTGGAACGAGTGGGATGAACTGAAGGAGATGGTTGT

[0946] GGGGATTGCAGACGAAGCCTATTTTGAACCGACTGAGCCCGGCAACCGCCCTGCACTTC

[0947] GTGATAAAAATATTGCTAAAATGTTTAGCTTCCCGCGCGGACCCAAGAAGCAAGAAGTAA

[0948] CGGAAAAAGCCAATGAGGAATTAAACGGCCTGGTTGCGTTGTTGGAGTCTCAGGGAGTC

[0949] ACCGTACGTCGCCCGGAAAAGCATAACTTTGGTTTATCTGTCAAGACACCGTTCTTTGAA

[0950] GTCGAAAACCAGTACTGTGCTGTTTGTCCGCGTGATGTGATGATCACGTTTGGCAACGAG

[0951] ATTTTAGAAGCGACTATGTCACGCCGTAGCCGTTTTTTTGAATACCTTCCATATCGCAAAC

[0952] TGGTGTATGAGTATTGGCATAAGGACCCGGATATGATTTGGAACGCCGCTCCGAAGCCT 83737PC01

[0953] 91

[0954] ACCATGCAAAATGCGATGTATCGTGAGGACTTCTGGGAATGTCCTATGGAAGATCGTTTC

[0955] GAGAATATGCACGACTTCGAGTTCTGTGTCACCCAGGACGAGGTTATCTTCGATGCAGCA

[0956] GACTGCTCCCGTTTTGGTCGTGATATTTTTGTGCAGGAGTCCATGACTACCAATCGTGCG

[0957] GGGATTCGCTGGCTTAAGCGCCATTTGGAGCCTCGTGGTTTTCGCGTTCACGATATTCAC

[0958] TTTCCTTTGGACATCTTCCCCTCCCATATCGACTGTACGTTTGTCCCGCTGGCACCGGGG

[0959] GTGGTCCTTGTTAATCCTGACCGCCCAATTAAAGAGGGTGAAGAGAAACTTTTTATGGAT

[0960] AATGGCTGGCAATTTATCGAAGCGCCACTTCCGACATCGACTGATGACGAGATGCCCATG

[0961] TTCTGTCAGAGTAGTAAGTGGCTGGCAATGAACGTACTTAGCATTAGCCCAAAAAAAGTT

[0962] ATCTGTGAGGAGCAGGAACATCCACTTCATGAGTTGTTAGACAAACATGGGTTTGAAGTG

[0963] TACCCGATCCCTTTCCGCAACGTGTTCGAATTTGGAGGGAGTTTACACTGTGCCACGTGG

[0964] GACATCCATCGCACGGGAACGTGCGAGGACTATTTTCCCAAATTAAATTATACCCCGGTG

[0965] ACCGCGTCCACTAATGGTGTCTCCCGCTTTATTATTTAA

[0966] SEQ ID NO: 25 - GAT8-var3 - DNA sequence

[0967] ATGCAGACACGTATCGTCAATAGTTGGAACGAGTGGGATGAACTGAAGGAGATGGTTGT

[0968] GGGGATTGCAGACGGAGCCTATTTTGAACCGACGGAGCCCGGCAACCGCCCTGCACTTC

[0969] GTGATAAAAATATTGCTAAAATGTTTAGCTTCCCGCGCGGACCCAAGAAGCAAGAAGTAA

[0970] CGGAAAAAGCCAATGAGGAATTAAACGGCCTGGTTGCGTTGTTGGAGTCTCAGGGAGTC

[0971] ACCGTACGTCGCCCGGAAAAGCATAACTTTGGTTTATCTGTCAAGACTCCGTTCTTTGAA

[0972] GTCGAAAACCAGTACTGTGCTGTTTGTCCGCGTGATGTGATGATCACGTTTGGCAACGAG

[0973] ATTTTAGAAGCGACTATGTCACGCCGTAGCCGTTTTTTTGAATACCTTCCATATCGCAAAC

[0974] TGGTGTATGAGTATTGGCATAAGGACCCGGATATGATTTGGAACGCCGCTCCGAAGCCT

[0975] ACCATGCAAAATGCGATGTATCGTGAGGACTTCTGGGAATGTCCTATGGAAGATCGTTTC

[0976] GAGAGTTTGCACGACTTCGAGTTCTGTGTCACCCAGGACGAGGTTATCTTCGATGCAGCA

[0977] GACTGCTCCCGTTTTGGTCGTGATATTTTTGTGCAGGAGTCCATGACTACCAATCGTGCG

[0978] GGGATTCGCTGGCTTAAGCGCCATTTGGAGCCTCGTCGTTTTCGCGTTCACGATATTCAC

[0979] TTTCCTTTGGACATCTTCCCCTCCCATATCGACTGTACGTTTGTCCCGCTGGCACCGGGG

[0980] GTGGTCCTTGTTAATCCTGACCGCCCAATTAAAGAGGGAGAAGAGAAACTTTTTATGGAT

[0981] AATGACTGGCAATTTATCGAAGCGCCACTTCCGACATCGACTGATGACGAGATGCCCATG

[0982] TTCTGTCAGAGTAGTAAGTGGCTGGCAATGAACGTACTTAGCATTAGCCCAAAAAAAGTT

[0983] ATCTGTGAGGAGCAGGAACATCCACTTCATGAGTTGTTAGACAAACATGGGTTTGAAGTG

[0984] TACCCGATCCCTTTCCGCAACGTGTTCGAATTTGGAGGGAGTTTACACTGTGCCACGTGG

[0985] GACATCCATCGCACGGGAACGTGCGAGGACTATTTTCCCAAATTAAATTATACCCCGGTG

[0986] ACCGCGTCCACTAATGGTGTCTCCCGCTTTATTATTTAA

[0987] SEQ ID NO: 26 - J23100 - DNA sequence

[0988] TTGACGGCTAGCTCAGTCCTAGGTACAGTGCTAGC 83737PC01

[0989] 92

[0990] SEQ ID NO: 27 - P2 - DNA sequence

[0991] AAAAAGAGTATTGACTTCGCATCTTTTTGTACCTATAATGTGTGGA

[0992] SEQ ID NO: 28 - Pl - DNA sequence

[0993] AAAAAGAGTATTGACTTAAAGTCTAACCTATAGGTATAATGTG

[0994] SEQ ID NO: 29 - P5 - DNA sequence

[0995] TTGACAATTAATCATCCGGCTCGTAATTTATG

[0996] SEQ ID NO: 30 - Pll - DNA sequence

[0997] TTGACATCAGGAAAATTTTTCTGTATAATGTGTGGA

[0998] SEQ ID NO: 31 - J23109 - DNA sequence

[0999] TTTACAGCTAGCTCAGTCCTAGGGACTGTGCTAGC

[1000] SEQ ID NO: 32 - J23114 - DNA sequence

[1001] TTTATGGCTAGCTCAGTCCTAGGTACAATGCTAGC

[1002] SEQ ID NO: 33 - P13 - DNA sequence

[1003] TTCCCTATTAATCATCCGGCTCGTATAATGTGAATC

[1004] SEQ ID NO: 34 - J23110 - DNA sequence

[1005] TTTACGGCTAGCTCAGTCCTAGGTACAATGCTAGC

[1006] SEQ ID NO: 35 - RBSU100 - DNA sequence

[1007] GGGCCCAAGTTCACTTAAAAAGGAGATCAACAATGAAAGCAATTTTCGTACTGAAACATC

[1008] TTAATCATGCTAAGGAGGTTT

[1009] SEQ ID NO: 36 - BCD1 / RBS - DNA sequence

[1010] GGGCCCAAGTTCACTTAAAAAGGAGATCAACAATGAAAGCAATTTTCGTACTGAAACATC

[1011] TTAATCATGCTAAGGAGGTTT

[1012] SEQ ID NO: 37 - Linker Region - DNA sequence

[1013] TCTAGGAGACCGGATCCGTCGACCTGCAGGCATGCAAGCTTACTAGTAGATCTGGTCTC

[1014] GGACT

[1015] Table 8: Primers used in this study

[1016] SEQ ID NO 141-142 are primers used for USER cloning that contain a deoxyuridine residue instead of a normal deoxynucleotide. 83737PC01

[1017] 93 83737PC01

[1018] 94 83737PC01

[1019] 95 83737PC01

[1020] 96 83737PC01

[1021] 97 83737PC01

[1022] 98 83737PC01

[1023] 99 83737PC01

[1024] 100 83737PC01

[1025] 101 83737PC01

[1026] 102

[1027] SEQ ID NO: 176 - OCDwt - DNA sequence 83737PC01

[1028] 103 atgacgtatttcattgatgttccaaccatgtcggacttggtgcatgacattggcgtagcgccgtttattggcgagcttg ccgctgccctgcgggacgatttcaaacgctggcaggcgtttgacaagtccgcgcgcgttgccagccactcggaagt gggcgtgatcgaactgatgccggtcgccgacaaaagccgctatgccttcaaatacgtcaacggccacccggccaat actgcacgcaacctgcatacggtgatggccttcggtgtactggccgatgtcgactccggttacccggtgctgctgtcc g a a ctg a cca teg cca ccg ccctg eg ca ccg ca g eg a etteg etg a tg g ca g ceca g g ca etg g cccg cccg a a c gcgcgcaagatggcgctgatcggcaacggcgcgcaaagcgaattccaggccctggccttccacaagcacctgggc atcgaagaaatcgtcgcctacgacaccgacccgctggccaccgccaagcttatcgccaacctcaaggaatacagcg gcttgaccatccgccgggccagctcggtagccgaggcagtgaaaggcgtggacatcatcaccacggtaacggccg acaaggcctacgccaccatcattacccccgacatgctggagcccggtatgcacctgaacgccgtgggtggtgactg ccctggcaaaaccgagctgcacgccgatgtgctgcgcaacgcgcgggtgttcgtcgagtacgagccacaaacccg tatcgaaggtgagatccagcagctgccggcggactttccggtggtcgacctgtggcgcgtgctgcgcggagaaacc gagggccgccaaagcgacagccaggtcactgtattcgattcggtgggcttcgccctcgaagattacaccgtactac ggtacgtactacagcaagccgaaaagcgcgggatgggcaccaagatcgacctggtgccgtgggtggaggacga cccgaaagacctgttcagccacacccgaggccgcgctggcaaaaggcgtatccgacgggttgcctga SEQ ID NO: 177 - GAT1 - Protein sequence MQECPVCSYNEWDPLEEVIVGRPENANVPPFSVEVKANTYEKYWPFYQKHGGQSFPVDHVK KAIEEIEEMCKVLKHEGVIVQRPEVIDWSVKYKTPDFESTGMYAAMPRDILLVVGNEIIEAPM AWRARFFEYRAYRPLIKDYFRRGAKWTTAPKPTMADELYDQDYPIRTVEDRHKI-AAMGKFVT TEFEPCFDAADFMRAGRDIFAQRSQVTNYLGIEWMRRHLAPDYMKVHIISFKDPNPMHIDAT FNIIGPGLVLSNPDRPCHQIELFKKAGWTVVTPPTPLIPDNHPLWMSSKWLSMNVLMLDEKR VMVDANETSIHKMFEKLGISTIKVNIRHANSLGGGFHCWTCDIRRRGTLQSYFR SEQ ID NO: 178 - GAT2 - Protein sequence MKECPVCSYNEWDPLEEVIVGRAENACVPPFSVEVKANTYEKYWGFYQKFGGESFPKDHVK KAIAEIEEMCNILKKEGVIVKRPDPIDWSVKYRTPDFESTGMYAAMPRDILLVVGNEIIEAPM AWRARFFEYRAYRRIIKDYFNNGAKWTTAPKPTMADELYDQDYPIRSVEDRHKLAAQGKFV TTEFEPCFDAADFIRAGRDIFVQRSQVTNYMGIEWMRRHI-APDYRVHVISFKDPNPMHIDTT FNIIGPGLVLSNPDRPCHQIELFKKAGWTVIHPPVPLIPDDHPLWMSSKWLSMNVLMLDEKR VMVDANETSIQKMFENLGISTIKVNIRHANSLGGGFHCWTCDIRRRGTLQSYFD SEQ ID NO: 179 - GAT3 - Protein sequence MKDCPVSSFNEWDPLEEVIVGRAENACVPPFTVEVKANTYDKHWPFYQKYGGSYFPKDHLQ KAVAEIEEMCNILKMEGVTVRRPDPIDWSLKYKTPDFESTGLYGAMPRDILIVVGNEIIEAPM AWRARFFEYRAYRTIIKDYFRRGAKWTTAPKPTMADELYDQDYPIHSVEDRHKI-AAQGKFVT TEFEPCFDAADFIRAGRDIFVQRSQVTNYMGIEWMRKHLAPDYRVHIVSFKDPNPMHIDATF NIIGPGLVLSNPDRPCHQIDLFKKAGWTIVTPPTPIIPDDHPLWMSSKWLSMNVLMLDEKRV MVDANEVPIQKMFEKLGISTIKVSIRNANSLGGGFHCWTCDVRRRGTLQSYFD 83737PC01

[1029] 104

[1030] SEQ ID NO: 180 - GAT4 - Protein sequence

[1031] MLQECPVCAYNEWDPLEEVIVGRAENARVPPFTVEVKANTYEKHWPFYQKYGGQSFPEDHL KKAVAEIEEMCNILRMEGVTVQRPEPMDWSFEYNTPDFTSTGMYAAMPRDILMVVGNEIIEA PMAWRSRFFEYRAYRPLIKEYFRKGAKWTTPPKPTMSDELYDQEYPIRTVEDRHKLAAQGKF VTTEHEPCFDAADFIRAGRDLFVQRSQVTNYMGIEWMRRHI-APDYKVHIISFKDPNPMHIDA TFNIIGPGLVLSNPDRPCRQVEMFEKAGWTVVKPPTPLIPDDHPLWMSSKWLSMNVLMLDP

[1032] KRVMCDANEHTIHKMFENLGIKTIKVNIRHANSLGGGFHCWTTDVRRRGSLESYFH

[1033] SEQ ID NO: 181 - GAT5 - Protein sequence

[1034] MAALPAVHADNEWSPLRAIIVGRAGKSCFPSASKRMVEETMPSAHVHRFQTKSPFSEGLIAK

[1035] AEVELDQFAAILESEGIKVYRPPKDIDWI-AVDGYTGAMPRDGLISVGNTLIEACFAWKCRSR EIELGFAPILDDLAQDPQVRIVRRPAHTFADTLDNNEGKWAINNSRPAFDTADFMRFGHTLI GQYSHVTNQAGVDYVRQHLPPGYQIEILDSNDPQAMHIDATILPLREGLLIYHPHKVTEDTLR

[1036] QHAVLADWDLRAYPFIPEDRDEPPLYMTSSWLCLNVLVLDGKKVVVEAGDERTAGWFEELG MECIRCPFQHVNSIGGSFHCATVDLVREA

[1037] SEQ ID NO: 182 - GAT6 - Protein sequence

[1038] MSLVSVHNEWDPLEEIIVGTAVGARVPRADRSVFAVEYADEYDSQDQVPAGPYPDRVLKET

[1039] EEELHVLSEELTKLGVTVRRPGQRDNSALVATPDWQTDGFHDYCPRDGLI-AVGQTVIESPM ALRARFLESI-AYKDILLEYFASGARWLSAPKPRLADEMYEPTAPAGQRLTDLEPVFDAANVLR

[1040] FGTDLLYLVSDSGNELGAKWLQSALGSTYKVHPCRGLYASTHVDSTIVPLRPGLVLVNPARV NDDNMPDFLRSWQTVVCPELVDIGFTGDKPHCSVWIGMNLLVVRPDLAVVDRRQTGLIKVL EKHGVDVLPLQLTHSRTLGGGFHCATLDVRRTGSLETYRF

[1041] SEQ ID NO: 183 - GAT7 - Protein sequence

[1042] MKEIRSMQLNEKDSNVIQTSPVCSYTEWDLLEEIIVGVVDGACIPPWHAAMEPCLPTQQHQ

[1043] FFRDNAGKPFPQERIDI-ARKELDEFARILECEGVKVRRPEPKNQSLVYGAPGWSSTGMYAA MPRDVLLVVGTDIIECPI-AWRSRYFETAAYKKLLKEYFHGGAKWSSGPKPELSDEQYVDGW

[1044] VEDEAATSANLVITEFEPTFDAADFTRLGKDIIAQKSNVTNEFGINWLQRHLGDDYKIHVLEF NDMHPMHIDATLVPI-APGKLLINPERVQKMPEIFRGWDAIHAPKPIMPDSHPLYMTSKWINM NILMLDERRVVVERQDEPMIKAMKGAGFEPILCDFRNFNSFGGSFHCATVDIRRRGKLESYL

[1045] V

[1046] SEQ ID NO: 184 - GAT8 - Protein sequence

[1047] MQTRIVNSWNEWDELKEMVVGIADGAYFEPTEPGNRPALRDKNIAKMFSFPRGPKKQEVTE KANEELNGLVALLESQGVTVRRPEKHNFGLSVKTPFFEVENQYCAVCPRDVMITFGNEILEAT MSRRSRFFEYLPYRKLVYEYWHKDPDMIWNAAPKPTMQNAMYREDFWECPMEDRFESMHD FEFCVTQDEVIFDAADCSRFGRDIFVQESMTTNRAGIRWLKRHLEPRRFRVHDIHFPLDIFPS HIDCTFVPI_APGVVLVNPDRPIKEGEEKLFMDNGWQFIEAPLPTSTDDEMPMFCQSSKWI_A 83737PC01

[1048] 105

[1049] MNVLSISPKKVICEEQEHPLHELLDKHGFEVYPIPFRNVFEFGGSLHCATWDIHRTGTCEDYF

[1050] PKLNYTPVTASTNGVSRFII

[1051] SEQ ID NO: 185 - GAT9 - Protein sequence

[1052] MKLNSFDDWSPLKEIIVGTAQNYISHERDLSFDIFFHENLFKSDWAYPRLNKSSKEVNDLN

[1053] WQIKQKYVEELNEDVEDLVNLLKKFDVTVHRPMILPFNPQNIKGLGWESAPVPALNVRDNTL

[1054] ILGDEIIETPPAIRSRYLETRLI-APIFMKYFEAGATWTTMPRPILTDMSFDLSYARDLETTLGGP

[1055] TEPIEDPESSPYDVGFEMMLDGAQCLRLGKDIIVNIANQNHKI-ACDWLERHVENRYRIHRVY

[1056] RMSDNHIDSMLLALRPGVFI-ARHKDLKDLLPEPFKKWKMIVPPEPSATNFPTYDDGDLLLTS

[1057] PYIDLNVLSVNPETVLVNEVCVDLIKTLEKEGFTVVPVKHRHRRLFGGGFHCFTLDTVRDGGL

[1058] EDYTI

[1059] SEQ ID NO: 186 - GAMT1 - Protein sequence

[1060] MSTAQPIFSKGEDCKAGWHDASAGYNETDTHLEIFGKPVMERWETPYMHSLATIASSKGG

[1061] RVLEIGFGMAIAATKVESFPIEEHWIIECNDGVFARLQEWAKAQPHRIVPLKGLWEQVVGGL

[1062] PDNHFDGILYDTYPLSEDTWHTHQFDFIKGHAHRLLKPGGVLTYCNLTSWGELLKSKYDNID

[1063] NMFQETQMPHLLEAGFKREKISTTTMDIVPPSECKYYAFRKMITPTILKE

[1064] SEQ ID NO: 187 - GAMT2 - Protein sequence

[1065] MSALSDTAPIFTEGEDCKAAWQEATAAYDAPDTHLEILGKPVMERWETPYMHSI-ATVAASR

[1066] GGRVLEVGFGMAIAATKVQEFNIEEHWIVECNDGVFQRLEEWARVQPHKVVPLKGLWEDV

[1067] VPTLPDGHFSGILYDTYPLSAQTWHTHQFAFIKDHAFRLLRPGGVLTYCNLTSWGELLKGKY

[1068] SDIEKMFEETQVAQLLEAGFRRENISTAVMELVPPRECRYYSFPRMITPRVVKH

[1069] SEQ ID NO: 188 - GAMT3 - Protein sequence

[1070] MSSEKIFLEGESCQSSWHNATAGYDETDTHLEILGKPVMERWETPYMHSI-ATVAASKGGRV

[1071] LEIGFGMAIAATKLEQCNIEEHWIIECNDGVFKRLQEWATKQPHKIVPLKGLWEDVVPTLPN

[1072] GHFDGILYDTYPLSEETWHTHQFNFIKGHAYRLLKPGGVLTYCNLTSWGELLKTKYNDIEKM

[1073] FQETQTPQLVDAGFKCENISTTVMNLVPPEDCRYYSFKKMITPTIIKV

[1074] SEQ ID NO: 189 - GAMT4 - Protein sequence

[1075] MSAPAATPIFAPGENCSPAWRAAPAAYDASDTHLQILGKPVMERWETPYMHALAAAAASRG

[1076] GRVLEVGFGMAIAATKVQEAPIEEHWIIECNEGVFQRLQDWALQQPHKVVPLKGLWEEVAP

[1077] TLPDSHFDGILYDTYPLSEETWHTHQFNFIRDHAFRLLKPGGVLTYCNLTSWGELMKTKYSD

[1078] ITTMFEETQVPALLEAGFRRDNIRTQVMELVPPANCRYYAFPRMITPLVTKH

[1079] SEQ ID NO: 190 - GAMT5 - Protein sequence

[1080] MSSSEAASPIFHQGENCKPSWREAEAGYNQKDTHLKILGKPVMERWETPYMHSI-ATVAAS

[1081] KGGRVLEVGFGMAIAASKVEEFNIEEHWIVECNEGVFKRLEEWAKKQPHKVVPLRGLWEDV

[1082] VPTLPDGLFDGILYDTYPLSAETWHTHQFHFIKSHAFRLLKPGGVLTYCNLTSWGELLKTKYT

[1083] DIEKMFEETQVGHLVEAGFKRENISTAVMDLTPPQDCRYYAFPKMITPTIVKQ 83737PC01

[1084] 106

[1085] SEQ ID NO: 191 - OCD_AGRT4 - Protein sequence

[1086] MPIDPKLNVVPFISVDHMMKLVLKVGIDTFLTELAAEIEKDFRRWPIFDKKPRVGSHSQDGVI ELMPTSDGSLYGFKYVNGHPKNTHQGRQTVTAFGVLSDVGNGYPLLLSEMTILTALRTAATS ALAAKYLARPNSKTMAIIGNGAQSEFQARAFRAILGIQKLRLFDIDTSATRKCARNLTGPGFD IVECGSVAEAVEGADVITTVTADKQFATILSDNHVGPGVHINAVGGDCPGKTEISMEVLLRS DIFVEYPPQTWIEGDIQQLPRTHPVTELWQVMTGEKTGRVGDRQITMFDSVGFAIEDFSALR YVRAKITDFEMFTELDLLADPDEPRDLYGMLLRCEKKLEPTAVG

[1087] SEQ ID NO: 192 - OCD_BRUSU - Protein sequence

[1088] MTQPNLNIVPFVSVDHMMKLVLRVGVETFLKELAGYVEEDFRRWQNFDKTPRVASHSKEGV I ELM PTS DGTLYG FKYVNG H PKNTRDG LQTVTAFG VLAN VGSGYPM LLTEMTI LTALRTAAT SAVAAKHLAPKNARTMAIIGNGAQSEFQALAFKAILGVDKLRLYDLDPQATAKCIRNLQGAG FDIVACKSVEEAVEGADIITTVTADKANATILTDNMVGAGVHINAVGGDCPGKTELHGDILR RSDIFVEYPPQTRIEGEIQQLPEDYPVNELWEVITGRIAGRKDARQITLFDSVGFATEDFSAL RYVRDKLKDTGLYEQLDLLADPDEPRDLYGMLLRHEKLLQSESTKPAA

[1089] SEQ ID NO: 193 - OCD_AGRFC - Protein sequence

[1090] MPALANLNIVPFISVENMMDLAVSTGIENFLVQLAGYIEEDFRRWESFDKIPRIASHSRDGVI ELMPTSDGTLYGFKYVNGHPKNTKSGRQTVTAFGVLSDVDSGYPLLLSEMTILTALRTAATS AIAAKYLARKDSRTMALIGNGAQSEFQALAFKALIGVDRIRLYDIDPEATARCSRNLQRFGFQ IEACTSAEQAVEGADIITTATADKHNATILSDNMIGPGVHINGVGGDCPGKTEMHRDILLRS DIFVEFPPQTRIEGEIQQLAPDHPVTELWRVMTGQDVGRKSDKQITLFDSVGFAIEDFSALR YVRDRVEGSSHSSPLDLLADPDEPRDLFGMLLRRQAFRRLGG

[1091] SEQ ID NO: 194 - OAT_HUMAN - Protein sequence

[1092] MFSKLAHLQRFAVLSRGVHSSVASATSVATKKTVQGPPTSDDIFEREYKYGAHNYHPLPVAL ERGKGIYLWDVEGRKYFDFLSSYSAVNQGHCHPKIVNALKSQVDKLTLTSRAFYNNVLGEYE EYITKLFNYHKVLPMNTGVEAGETACKLARKWGYTVKGIQKYKAKIVFAAGNFWGRTLSAIS SSTDPTSYDGFGPFMPGFDIIPYNDLPALERALQDPNVAAFMVEPIQGEAGVVVPDPGYLMG VRELCTRHQVLFIADEIQTGLARTGRWLAVDYENVRPDIVLLGKALSGGLYPVSAVLCDDDI MLTIKPGEHGSTYGGNPLGCRVAIAALEVLEEENLAENADKLGIILRNELMKLPSDVVTAVRG

[1093] KGLLNAIVIKETKDWDAWKVCLRLRDNGLLAKPTHGDIIRFAPPLVIKEDELRESIEIINKTIL SF

[1094] SEQ ID NO: 195 - OAT_BACSU - Protein sequence

[1095] MTALSKSKEIIDQTSHYGANNYHPLPIVISEALGAWVKDPEGNEYMDMLSAYSAVNQGHRH PKIIQALKDQADKITLTSRAFHNDQLGPFYEKTAKLTGKEMILPMNTGAEAVESAVKAARRW AYEVKGVADNQAEIIACVGNFHGRTMLAVSLSSEEEYKRGFGPMLPGIKLIPYGDVEALRQAI TPNTAAFLFEPIQGEAGIVIPPEGFLQEAAAICKEENVLFIADEIQTGLGRTGKTFACDWDGIV 83737PC01

[1096] 107

[1097] PDMYILGKALGGGVFPISCIAADREILGVFNPGSHGSTFGGNPLACAVSIASLEVLEDEKLAD RSLELGEYFKSELESIDSPVIKEVRGRGLFIGVELTEAARPYCERLKEEGLLCKETHDTVIRFAP PLIISKEDLDWAIEKIKHVLRNA

[1098] SEQ ID NO: 196 - GAT8-varl - Protein sequence

[1099] MQTRIVNSWNEWDELKEMVVGIADGAYFEPTEPGNRPALRDKNIAKMFSFPRGPKKQEVTE KANEELNGLVALLESQGVTVRRPEKHNFGLSVKTPFFEVENQYCAVCPRDVMITFGNEILEAT MSRRSRFFEYLPYRKLVYEYWHKDPDMIWNAAPKPTMQNAMYREDFWECPTEDRFESMHD FEFCVTQDEVIFDAADCSRFGRDIFVLESMTTNRAGIRWLKRHLEPRGFRVHDIHFPLDIFPS HIDCTFVPLAPGVVLVNPDRPIKEGEEKLFMDNGWQFIEAPLPTSTDDEMPMFCQSSKWLA MNVLSISPKKVICEEQEHPLHELLDKHGFEVYPIPFRNVFEFGGSLHCATWDIHRTGTCEDYF PKLNYTPVTASTNGVSRFII

[1100] SEQ ID NO: 197 - GAT8-var2 - Protein sequence

[1101] MQTRIVNSWNEWDELKEMVVGIADEAYFEPTEPGNRPALRDKNIAKMFSFPRGPKKQEVTE KANEELNGLVALLESQGVTVRRPEKHNFGLSVKTPFFEVENQYCAVCPRDVMITFGNEILEAT MSRRSRFFEYLPYRKLVYEYWHKDPDMIWNAAPKPTMQNAMYREDFWECPMEDRFENMH DFEFCVTQDEVIFDAADCSRFGRDIFVQESMTTNRAGIRWLKRHLEPRGFRVHDIHFPLDIF PSHIDCTFVPLAPGVVLVNPDRPIKEGEEKLFMDNGWQFIEAPLPTSTDDEMPMFCQSSKWL AMNVLSISPKKVICEEQEHPLHELLDKHGFEVYPIPFRNVFEFGGSLHCATWDIHRTGTCEDY FPKLNYTPVTASTNGVSRFII

[1102] SEQ ID NO: 198 - GAT8-var3 - Protein sequence

[1103] MQTRIVNSWNEWDELKEMVVGIADGAYFEPTEPGNRPALRDKNIAKMFSFPRGPKKQEVTE KANEELNGLVALLESQGVTVRRPEKHNFGLSVKTPFFEVENQYCAVCPRDVMITFGNEILEAT MSRRSRFFEYLPYRKLVYEYWHKDPDMIWNAAPKPTMQNAMYREDFWECPMEDRFESLHD FEFCVTQDEVIFDAADCSRFGRDIFVQESMTTNRAGIRWLKRHLEPRRFRVHDIHFPLDIFPS HIDCTFVPLAPGVVLVNPDRPIKEGEEKLFMDNDWQFIEAPLPTSTDDEMPMFCQSSKWLA MNVLSISPKKVICEEQEHPLHELLDKHGFEVYPIPFRNVFEFGGSLHCATWDIHRTGTCEDYF PKLNYTPVTASTNGVSRFII

[1104] SEQ ID NO: 199 - OCDwt - Protein sequence

[1105] MTYFIDVPTMSDLVHDIGVAPFIGELAAALRDDFKRWQAFDKSARVASHSEVGVIELMPVAD KSRYAFKYVNGHPANTARNLHTVMAFGVLADVDSGYPVLLSELTIATALRTAATSLMAAQAL ARPNARKMALIGNGAQSEFQALAFHKHLGIEEIVAYDTDPLATAKLIANLKEYSGLTIRRASS VAEAVKGVDIITTVTADKAYATIITPDMLEPGMHLNAVGGDCPGKTELHADVLRNARVFVEY EPQTRIEGEIQQLPADFPVVDLWRVLRGETEGRQSDSQVTVFDSVGFALEDYTVLRYVLQQA EKRGMGTKIDLVPWVEDDPKDLFSHTRGRAGKRRIRRVA 83737PC01

[1106] 108

[1107] Items

[1108] 1. A bacterial cell comprising :

[1109] • a nucleic acid sequence encoding an enzyme, such as Glycine a midi notransferase (GAT) or a functional variant thereof, capable of catalysing the reaction arginine < = > ornithine, wherein the enzyme is an enzyme selected from any of the EC numbers EC 2.1.4.1, EC 2.1.4.3, or EC 3.5.3.1; and

[1110] • a nucleic acid encoding o an enzyme, such as Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of catalysing the reaction ornithine < = > proline, wherein the enzyme is an enzyme according to EC 4.3.1.12, or o an enzyme, such as Ornithine aminotransferase (OAT) or a functional variant thereof, capable of catalysing the reaction ornithine < = > glutamate+ semialdehyde, wherein the enzyme is an enzyme according to EC 2.6.1.13, and wherein pathways for generating proline and ornithine are downregulated or deleted.

[1111] 2. The bacterial cell of item 1, wherein the pathways for generating proline are selected from the group consisting of:

[1112] • proB (N-acetylglutamate synthase);

[1113] • proA (glutamate-5-semialdehyde dehydrogenase); and

[1114] • proC (pyrroline-5-carboxylate reductase), preferably proB is the pathway for generating proline.

[1115] 3. The bacterial cell of any of the preceding items, wherein the pathways for generating ornithine are selected from the group consisting of:

[1116] • argA (N-acetylglutamate synthase);

[1117] • argB (acetylglutamate kinase);

[1118] • argC (N-acetylglutamylphosphate reductase);

[1119] • argD, astC, and gabT (N-acetylornithine aminotransferase); and

[1120] • argE (acetylornithine deacetylase) 83737PC01

[1121] 109 preferably, argA is the pathway for generating ornithine.

[1122] 4. The bacterial cell of any of the preceding items, wherein at least one enzyme encoded by the pathway encoded by the operon astABCDE is downregulated or deleted.

[1123] 5. The bacterial cell of item 1, wherein the downregulated or deleted pathways for generating proline or ornithine are astABCDE, proB, and argA.

[1124] 6. The bacterial cell of any of the preceding items, wherein the nucleic acid sequence encoding an enzyme capable of catalysing the reaction arginine < = > ornithine is a nucleic acid sequence encoding glycine amidinotransferase (GAT) having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 8, or a degenerate sequence thereof, such as at least 85

[1125] % sequence identity, such as at least 90 % sequence identity, such as at least 95

[1126] % sequence identity, such as at least 96 % sequence identity, such as at least 97

[1127] % sequence identity, such as at least 98 % sequence identity, such as at least 99

[1128] % sequence identity.

[1129] 7. The bacterial cell of any of the preceding items, wherein the encoded GAT has between 1 and 4 mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 8, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 8.

[1130] 8. The bacterial cell of any of the preceding items, wherein GAT is selected from the group consisting of GAT3 as encoded by the sequence according to SEQ ID NO: 3, GAT8 as encoded by the sequence according to SEQ ID NO: 8, GAT9 as encoded by the sequence according to SEQ ID NO: 9, preferably GAT8 as encoded by the sequence according to SEQ ID NO: 8.

[1131] 9. The bacterial cell of any of the preceding items, comprising mutations in GAT, such as one or more of the mutations M175T, Q211L, R232G, G25E, S181N, M182L and G281D, preferably at least R232G. 83737PC01

[1132] 110

[1133] 10. The bacterial cell of any of the preceding items, wherein the GAT comprises a combination of mutations selected from the group of combinations consisting of:

[1134] • M175T, Q211L and R232G (GAT8-V1);

[1135] • G25E, S181N and R232G (GAT8-V2); and

[1136] • M182L, and G281D (GAT8-V3).

[1137] 11. The bacterial cell of any of the preceding items, wherein GAT is selected from the group consisting of GAT8-V1 as encoded by the sequence according to SEQ ID NO: 23, GAT8-V2 as encoded by the sequence according to SEQ ID NO: 24, GAT8-V3 as encoded by the sequence according to SEQ ID NO: 25 and wherein the encoded GAT has between 1 and 4 additional mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 23, 24, or 25, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 23, 24, or 25.

[1138] 12. The bacterial cell of any of the preceding items, wherein GAT is selected from the group consisting of GAT8-V1 as encoded by the sequence according to SEQ ID NO: 23, GAT8-V2 as encoded by the sequence according to SEQ ID NO: 24, GAT8-V3 as encoded by the sequence according to SEQ ID NO: 25.

[1139] 13. The bacterial cell of any of the preceding items, wherein the expression of GAT is regulated by a promoter selected from J23110, J23109, and J23114, preferably the J23110 promoter.

[1140] 14. The bacterial cell of any of the preceding items, wherein the expression of GAT is regulated by RBSU058 as the RBS.

[1141] 15. The bacterial cell of any of the preceding items, wherein the expression of GAT is regulated by a promoter selected from J23110, J23109, and J23114, preferably the J23110 promoter and RBSU058 as the RBS.

[1142] 16. The bacterial cell of any of the preceding items, wherein the nucleic acid sequence encoding the enzyme capable of catalysing the reaction ornithine < = > proline is a nucleic acid sequence Ornithine cyclodeaminase (OCD) having at least 83737PC01

[1143] 111

[1144] 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 176, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[1145] 17. The bacterial cell of any of the preceding items, wherein OCD is selected from the group consisting of OCDvl, OCDv2, OCDv3, OCDv4, OCDWT, OCD_AGRT4, and OCD_AGRFC, wherein OCDv4 is preferred.

[1146] 18. The bacterial cell of any of the preceding items, wherein the encoded OCD has between 1 and 4 mutations with respect to the protein encoded by the sequence according to SEQ ID NO: 176, such as between 1 and 3 mutations, such as between 1 and 2 mutations, such as 1 mutation with respect to the protein encoded by the sequence according to SEQ ID NO: 176.

[1147] 19. The bacterial cell of any of the preceding items, wherein the OCD comprises a methionine at position 86, with respect to the amino acid sequence encoded by OCDWT.

[1148] 20. The bacterial cell of any of the preceding items, comprising mutations in OCD, such as one or more of the mutations K205G, T162A, R65H, R184H, A185V, or R244H.

[1149] 21. The bacterial cell of any of the preceding items, comprising a combination of mutations in OCD, wherein a combination is selected from the group of combinations consisting of:

[1150] • K205G, M86K, T162A;

[1151] • K205G, M86T, T162A;

[1152] • K205G, M86M, T162A; and

[1153] • R65H, R184H, A185V, R244H.

[1154] 22. The bacterial cell of any of the preceding items, wherein OAT is selected from OAT_HUMAN, or OAT_BACSU. 83737PC01

[1155] 112

[1156] 23. The bacterial cell of any of the preceding items, wherein the nucleic acid sequence encoding the enzyme capable of catalysing the reaction ornithine< = > glutamate+ semialdehyde is a nucleic acid sequence encoding Ornithine aminotransferase (OAT) having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 21, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[1157] 24. The bacterial cell of any of the preceding items, wherein the nucleic acid sequence encoding the enzyme capable of catalysing the reaction ornithine< = > glutamate+ semialdehyde is a nucleic acid sequence encoding Ornithine aminotransferase (OAT) having at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 21, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[1158] 25. The bacterial cell of any of the preceding items, wherein the bacterial cell is selected from the group consisting of:

[1159] • Escherichia coli;

[1160] • Geobacter metallireducens;

[1161] • Synechococcus elongatus;

[1162] • Synechocystis sp;

[1163] • Saccharomyces cerevisiae;

[1164] • Shigella dysenteriae; and

[1165] • Phaeodactylum tricornutum : preferably Escherichia coli.

[1166] 26. The bacterial cell of any of the preceding items, wherein the bacterial cell is Escherichia coli and the strain is selected from the group consisting of BW25113 and Escherichia coli str. K-12 substr. MG1655. 83737PC01

[1167] 113

[1168] 27. An e. coli, wherein the e. coli comprises a nucleic acid sequence encoding guanidinoacetate methyltransferase (GAMT) or a functional variant thereof such as an enzyme capable of catalysing the reaction guanidinoacetate + S-adenosyl-L- methionine < = > creatine + H(+) + S-adenosyl-L-homocysteine.

[1169] 28. The e. coli of item 27, wherein GAMT has at least 80 % sequence identity to the protein, GAMT3, encoded by the sequence according to SEQ ID NO: 12, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

[1170] 29. The e. coli of items 27-28, wherein GAMT is selected from the group consisting of

[1171] • GAMT1 as encoded by the sequence according to SEQ ID NO: 10;

[1172] • GAMT2 as encoded by the sequence according to SEQ ID NO: 11;

[1173] • GAMT3 as encoded by the sequence according to SEQ ID NO: 12;

[1174] • GAMT4 as encoded by the sequence according to SEQ ID NO: 13; or

[1175] • GAMT5 as encoded by the sequence according to SEQ ID NO: 14.

[1176] 30. The e. coli of items 27-29, wherein the bacterial cell has rewired SAM-cycle and is deficient of the natural cysteine synthesis.

[1177] 31. The e. coli of items 27-30, wherein cysteine is provided by the SAM cycle and SAM-dependent methylation.

[1178] 32. The e. coli of items 27-31, wherein the rewired SAM-cycle is rewired by the introduction of a nucleic acid encoding sacB, CYS3, CYS4.

[1179] 33. The e. coli of item 27-32, wherein the bacterial strain comprises one or more of the following features; i. A rpoA D305X mutation, wherein X may be any amino acid, preferably H or Y; ii. A nusA mutation, such as A234E R104L, and / or R228P; 83737PC01

[1180] 114 iii. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; iv. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or v. a combination of the above.

[1181] 34. The e. coli of any preceding items, wherein the E. coli is the strain BW25113, preferably comprising one or more of the features according to item 33.

[1182] 35. A bacterial cell, such as e. coli, comprising one or more of the following features; i. a nucleic acid encoding a GAT as identified in any of items 6-15; ii. a nucleic acid encoding a GAMT as identified in any of items 28-29; iii. A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; iv. A nusA mutation, such as A234E R104L, and / or R228P; v. Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; vi. One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or vii. a combination of the above.

[1183] 36. The bacterial cell according to item 35, wherein the bacterial cell is selected from the group consisting of

[1184] • E. coli;

[1185] • Pseudomonas putida;

[1186] • Vibrio natriegens;

[1187] • Bacillus subtilis;

[1188] • Corynebacterium glutamicum;

[1189] • Cyanobacteria, such as Synechocystis sp. PCC 6803 and) Synechococcus elongatus PCC 7942;

[1190] • Streptomyces spp;

[1191] • Lactobacillus spp; and 83737PC01

[1192] 115

[1193] • Saccharomyces cerevisiae;

[1194] • Clostridium butyclicum;

[1195] • Aspergillus niger; and

[1196] • Bacillus licheniformis.

[1197] 37. Use of the bacterial cell according to any of the items 35-36 in a fermentation process, such as a fermentation process for producing creatine.

[1198] 38. A process for selecting mutants of GAT and / or OCD / OAT, the process comprising: i. Providing a bacterial host cell comprising

[1199] ■ a nucleic acid sequence encoding an enzyme, such as Glycine a midi notransferase (GAT) or a functional variant thereof, capable of catalysing the reaction arginine < = > ornithine, wherein the enzyme is an enzyme selected from any of the EC numbers EC 2.1.4.1, EC 2.1.4.3, or EC 3.5.3.1; and

[1200] ■ a nucleic acid encoding

[1201] • an enzyme, such as Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of catalysing the reaction ornithine < = > proline, wherein the enzyme is an enzyme according to EC 4.3.1.12, or

[1202] • an enzyme, such as Ornithine aminotransferase (OAT) or a functional variant thereof, capable of catalysing the reaction ornithine < = > glutamate+ semialdehyde, wherein the enzyme is an enzyme according to EC 2.6.1.13, ii. optionally, the GAT as identified in any of items 6-15 and 1) a nucleic acid encoding an OCD as identified in any of items 16-21, or 2) a nucleic acid encoding an OAT as identified in any of items 22- 24; iii. Downregulating or deleting pathways for generating proline or ornithine, wherein the pathways for generating proline or ornithine according to any of items 2-5, or optionally if only mutants of OCD are selected the pathways may be encoded by astABCDE and speAB; 83737PC01

[1203] 116 iv. Growing the cells for a time sufficient to obtain improved variants of GAT and / or OCD / OAT; v. Identifying said improved variants.

[1204] 39. The process according to 33, wherein the cells are grown in a medium comprising arginine between 0.002 % and 0.020 %, such as between 0.002 % and 0.010 %, such as between 0.003 % and 0.010 %, preferably at about 0.005 %.

[1205] 40. A process for selecting bacterial strains with increased tolerance to creatine, the process comprising: i. providing an e. coli with a rewired SAM-cycle and being deficient of the natural cysteine synthesis; ii. introducing into the bacterial cell: o a nucleic acid sequence encoding GAT according to any of items 6-15; and o a nucleic acid sequence encoding GAMT according to any of items 28-29; iii. optionally, the e. coli may comprise mutations selected from the group consisting of; o A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y; o A nusA mutation, such as A234E R104L, and / or R228P; o Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308; o One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or o a combination of the above. iv. Growing the bacterial cell in conditions with substantially no cysteine being present, for a time sufficient to obtain strains with increased tolerance to creatine; and v. optionally, isolating and / or identifying said strains with increased tolerance to creatine. 83737PC01

[1206] 117

[1207] 41. Glycine amidinotransferase (GAT) according to any of items 6-15.

[1208] 42. A nucleic acid sequence encoding GAT according to item 41.

[1209] 43. A vector, such as a plasmid, comprising the nucleic acid sequence according to item 42.

[1210] 44. A host cell, such as a bacterial host cell, comprising the nucleic acid or vector according to any of items 42-43, optionally wherein the host cell is a host cell for reproduction of the vector.

[1211] 45. Ornithine cyclodeaminase (OCD) according to any of items 16-21.

[1212] 46. A nucleic acid sequence encoding OCD according to item 45.

[1213] 47. A vector, such as a plasmid, comprising the nucleic acid sequence according to item 47.

[1214] 48. A host cell, such as a bacterial host cell, comprising the nucleic acid or vector according to any of items 44-47, optionally wherein the host cell is a host cell for reproduction of the vector.

[1215] 49. Ornithine aminotransferase (OAT) according to any of items 22-24.

[1216] 50. A nucleic acid sequence encoding OCD according to item 49.

[1217] 51. A vector, such as a plasmid, comprising the nucleic acid sequence according to item 50.

[1218] 52. A host cell, such as a bacterial host cell, comprising the nucleic acid or vector according to any of items 50-51, optionally wherein the host cell is a host cell for reproduction of the vector.

[1219] 53. Use of GAT, the nucleic acid, the vector, or the cell according to any of items 41-44, in a process for producing creatine. 83737PC01

[1220] 118

[1221] 54. Use of OCD or OAT, the nucleic acid, the vector, or the cell according to any of items 45-52 in a process for selecting improved variants of GAT.

[1222] 55. Use of an e.coli, such as the strain BW25113, comprising mutations selected from the group consisting of;

[1223] • A rpoA D305X mutation, wherein X may be any amino acid, such as H or Y;

[1224] • A nusA mutation, such as A234E R104L, and / or R228P;

[1225] • Knockout of ghxP, such as a 11 bp deletion at 449-459 or a 5 bp deletion at 304-308;

[1226] • One or more of the mutations metC G341D, metC Albp (1022 / 1188 nt), murA A119V, trpA Q243P, metC G241C, folM A6 bp (52-57 / 723 nt), metC S345*, metC A207T; or

[1227] • a combination of the above; in a process for producing creatine.

[1228] 56. The use according to item 55, wherein the e. coli comprises a combination of mutations selected from the group of combinations consisting of:

[1229] • metC G341D, and rpoA D305Y;

[1230] • metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt);

[1231] • trpA Q243P, metC G241C, and nusA A234E;

[1232] • folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and

[1233] • metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt).

[1234] 57. The use according to any of the items 56-57, wherein the e. coli comprises a gene encoding GAMT, such as GAMT according to any of items 28-29.

[1235] 58. A method of producing creatine, the method comprising: i. providing a cell, such as a bacterial cell, such as e. coli; ii. introducing into the cell: o a nucleic acid sequence encoding GAT according to any of items 6- 15; and o a nucleic acid sequence encoding GAMT according to any of items 28-29; 83737PC01

[1236] 119 iii. optionally, wherein the cell is e. coli and the e. coli comprises genomic mutations as defined in any of items 55-56; iv. growing the cell under conditions suitable for the production of creatine; and v. optionally, purifying creatine from the supernatant.

[1237] 59. The method according to item 58, wherein the cell is grown in a fermentation broth comprising L-arginine and optionally glycine.

Claims

83737PC01120Claims1. A genetically modified Escherichia coli comprising: a) a nucleic acid sequence encoding Glycine a midi notransferase (GAT) or a functional variant thereof, capable of increasing the reaction rate of conversion between arginine and ornithine; and b) a nucleic acid sequence encoding■ Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of increasing the reaction rate of conversion between ornithine and proline, or■ Ornithine aminotransferase (OAT) or a functional variant thereof, capable of increasing the reaction rate of conversion between ornithine and glutamate 5-semialdehyde, c) one or more genetic modifications for deleting or downregulating the pathways for generating ornithine, or proline and ornithine, wherein the growth of the cell is dependent on the production of ornithine from the enzyme in a), wherein the one or more genetic modifications for deleting or downregulating the pathway for generating proline are introduced into a gene selected from the group consisting of:• proB (N-acetylglutamate synthase);• proA (glutamate-5-semialdehyde dehydrogenase); and• proC (pyrroline-5-carboxylate reductase), and wherein the one or more genetic modifications for deleting or downregulating the pathway for generating ornithine are introduced into a gene selected from the group consisting of:• argA (N-acetylglutamate synthase);• argB (acetylglutamate kinase);• argC (N-acetylglutamylphosphate reductase);• argD, astC, and gabT (N-acetylornithine aminotransferase); and• argE (acetylornithine deacetylase), wherein the one or more genetic modifications for deleting or downregulating the pathways for generating ornithine, or proline and ornithine reduces the transcription or translation of the gene so that the levels of functional protein, encoded by the gene are significantly reduced in the host cell by at least 95%, as compared to a control.83737PC011212. The genetically modified E. coli of claim 1, wherein at least one enzyme encoded by the operon astABCDE is downregulated or deleted.

3. The genetically modified E. coli of any of the preceding claims, wherein the nucleic acid sequence encoding glycine amidinotransferase (GAT) has at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 8, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

4. The genetically modified E. coli of claim 3, wherein GAT is selected from the group consisting of GAT8-V1 as encoded by the sequence according to SEQ ID NO: 23, GAT8-V2 as encoded by the sequence according to SEQ ID NO: 24, GAT8-V3 as encoded by the sequence according to SEQ ID NO: 25.

5. The genetically modified E. coli of any of the preceding claims, wherein the nucleic acid sequence encoding Ornithine cyclodeaminase (OCD) has at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 176, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity.

6. The genetically modified E. coli of any of the preceding claims, wherein the nucleic acid sequence encoding Ornithine cyclodeaminase (OCD) has the sequence according to SEQ ID NO: 176, and comprising a combination of mutations in OCD, wherein a combination is selected from the group of combinations consisting of:• R65H, R184H, A185V, R244H• K205G, M86K, T162A;• K205G, M86T, T162A; and83737PC01122K205G, M86M, T162A.

7. The genetically modified E. coli of any of the claim 1-4, wherein the nucleic acid sequence encoding OAT has at least 80 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 21 (OAT_HUMAN) or SEQ ID NO: 22 (OAT_BACSU) or any degenerate sequences thereof, such as at least 85% sequence identity, such as at least 90 % sequence identity, such as at least 95% sequence identity, such as at least 96 % sequence identity, such as at least 97% sequence identity, such as at least 98 % sequence identity, such as at least 99% sequence identity.

8. A bacterial cell, wherein the bacterial cell is E. coli, comprising; i. a nucleic acid encoding a GAT as identified in any of claims 3-4; ii. a nucleic acid encoding a protein having at least 80 % sequence identity to the protein, GAMT3, encoded by the sequence according to SEQ ID NO: 12, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity, or a nucleic acid encoding the protein encoded by the sequence according to SEQ ID NO: 12; and a combination of mutations selected from the group: i. metC G341D, and rpoA D305Y; ii. metC Albp (1022 / 1188 nt), murA A119V, rpoA D305H, and ghxP All bp (449 459 / 1350 nt); iii. trpA Q243P, metC G241C, and nusA A234E; iv. folM A6 bp (52-57 / 723 nt), metC S345*, and nusA R104L; and v. metC A207T, nusA R228P, and ghxP A5 bp (304-308 / 1350 nt).

9. Use of the bacterial cell according to claim 8 in a fermentation process, such as a fermentation process for producing creatine.83737PC0112310. A process for selecting mutants of GAT and / or OCD / OAT, the process comprising: i. Providing an E. coli cell comprising■ a nucleic acid sequence encoding Glycine amidinotransferase (GAT) or a functional variant thereof, capable of increasing the reaction rate of conversion between arginine and ornithine; and■ a nucleic acid encoding• Ornithine cyclodeaminase (OCD) or a functional variant thereof, capable of increasing the reaction rate of conversion between ornithine and proline, or• Ornithine aminotransferase (OAT) or a functional variant thereof, capable of increasing the reaction rate of conversion between ornithine and glutamate 5- semialdehyde, ii. optionally, the GAT as identified in any of claims 3-4 and a. a nucleic acid encoding an OCD as identified in any of claims 5-6, or b. a nucleic acid encoding an OAT as identified in claim 7; iii. Downregulating or deleting pathways for generating proline or ornithine, or if only mutants of OCD are selected downregulating or deleting the pathways astABCDE and speAB; iv. Growing the cells for a time sufficient to obtain improved variants of GAT and / or OCD / OAT; v. Identifying said improved variants, wherein the cells are grown in a medium comprising arginine between 0.002 % and 0.020 %, such as between 0.002 % and 0.010 %, such as between 0.003 % and 0.010 %, preferably at about 0.005 %, wherein the pathways for generating proline are downregulated or deleted by one or more genetic modifications introduced into a gene selected from the group consisting of:• proB (N-acetylglutamate synthase);• proA (glutamate-5-semialdehyde dehydrogenase); and• proC (pyrroline-5-carboxylate reductase),83737PC01124 and wherein the pathways for generating ornithine are downregulated or deleted by one or more genetic modifications introduced into a gene selected from the group consisting of:• argA (N-acetylglutamate synthase);• argB (acetylglutamate kinase);• argC (N-acetylglutamylphosphate reductase);• argD, astC, and gabT (N-acetylornithine aminotransferase); and• argE (acetylornithine deacetylase), wherein the deleting or downregulating of pathways reduces the transcription or translation of the gene so that the levels of functional protein, encoded by the gene are significantly reduced in the host cell by at least 95%, as compared to a control.

11. A method of producing creatine, the method comprising: i. providing an E. coli; ii. introducing into the cell: o a nucleic acid sequence encoding a GAT selected from the group consisting of GAT8-V1 as encoded by the sequence according to SEQ ID NO: 23, GAT8-V2 as encoded by the sequence according to SEQ ID NO: 24, GAT8-V3 as encoded by the sequence according to SEQ ID NO: 25; and o a nucleic acid encoding a protein having at least 80 % sequence identity to the protein, GAMT3, encoded by the sequence according to SEQ ID NO: 12, or a degenerate sequence thereof, such as at least 85 % sequence identity, such as at least 90 % sequence identity, such as at least 95 % sequence identity, such as at least 96 % sequence identity, such as at least 97 % sequence identity, such as at least 98 % sequence identity, such as at least 99 % sequence identity, or a nucleic acid encoding the protein encoded by the sequence according to SEQ ID NO: 12; iii. optionally, wherein the E. coli comprises a combination of mutations as defined in claim 8; iv. growing the cell under conditions suitable for the production of creatine; and v. optionally, purifying creatine from the supernatant.83737PC0112512. The method according to claim 11, wherein the cell is grown in a fermentation broth comprising L-arginine and optionally glycine.

13. A Glycine a midi notransferase (GAT) having at least 99 % sequence identity to the protein encoded by the sequence according to SEQ ID NO: 8, or a degenerate sequence thereof, wherein the GAT comprises one or more of the mutations M175T, Q211L, R232G, G25E, S181N, M182L and / or G281D, preferably at least R232G.

14. The GAT according to claim 13, wherein the GAT comprises a combination of mutations selected from the group of combinations consisting of:• M175T, Q211L and R232G;• G25E, S181N and R232G; and • M182L, and G281D.

15. A nucleic acid sequence encoding the GAT according to any of claims 13-14, such as a nucleic acid sequence according to any of SEQ ID NO: NO: 23, 24, or 25, or any degenerate sequence thereof.