Glucagon-like peptide 1 agonists
Patent Information
- Application Number
- EP2024802179
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-11-04
- Publication Date
- 2026-09-09
AI Technical Summary
Current GLP-1 agonists have a short half-life, requiring frequent administration and being challenging to manufacture at high production levels.
Development of alternatively modified GLP-1 [7-37] analogues with non-canonical amino acids (ncAAs) inserted at position 7 and/or position 8, using orthogonal aaRS/tRNA pairs for production, allowing for longer half-life and easier manufacturing.
The modified GLP-1 analogues demonstrate enhanced binding affinity to the GLP-1 receptor and comparable or extended half-life to semaglutide, while being easier to produce in large quantities without chemical processing.
Smart Images

Figure EP2024081074_08052025_PF_FP_ABST
Abstract
Description
[0001] Glucagon-Like peptide 1 agonists
[0002] Field of the Invention
[0003] The present application relates to compounds that are agonists of the glucagon-like peptide-1 (GLP-1 ) receptor which have a higher system half-life than GLP-1 or other GLP-1 agonists.
[0004] Glucagon-like peptide 1 (GLP-1 ) is a gut enteroendocrine cell-derived hormone and one of two prominent endogenous physiological incretins. The effects of incretins were first discovered in the 1930s in studies comparing the oral and intravenous administration of glucose. The action of GLP-1 was elucidated later, for its ability to stimulate insulin secretion from pancreatic beta cells (Drucker et al., Proc Natl Acad Sci USA. 1987;
[0005] 84(10):3434-3438). GLP-1 improves glycemic control by stimulating glucose-dependent insulin secretion in response to nutrients (glucose), inhibits glucagon secretion from the pancreatic alpha-cells, slows gastric emptying, and induces body weight loss primary by decreasing food consumption.
[0006] GLP-1 agonists, also known as incretin mimetics, are a class of drugs used primarily in the treatment of type 2 diabetes mellitus (T2DM). They mimic the effects of the native GLP-1 hormone, which stimulates insulin secretion and inhibits glucagon release in a glucose-dependent manner. More recently, GLP-1 agonists such as semaglutide have been prescribed for appetite suppression and weight loss.
[0007] GLP-1 agonists have a notoriously short half-life, requiring frequent administration, often by injection.
[0008] Exenatide, the first GLP-1 receptor agonist, was approved in the U.S. in 2005. It is a synthetic form of a peptide derived from the saliva of the Gila monster, a lizard native to the southwestern U.S. Due to its short half-life, exenatide requires twice-daily subcutaneous injections (DeFronzo et al., Diabetes Care. 2005;28(5):1092-1100).
[0009] Liraglutide, a once-daily GLP-1 agonist, was introduced in 2009. Derived from human GLP-1 , liraglutide has a longer half-life due to a fatty acid side chain that delays its absorption and degradation (Garber et al., Lancet. 2009;373(9662):473-481 ).
[0010] More recently, longer-acting GLP-1 agonists have been developed. Dulaglutide (weekly injection) and semaglutide (weekly injection or oral form) are examples of such longer- acting GLP-1 agonists. These agents offer the advantage of less frequent dosing (Nauck et al., Diabetes Care. 2014;37(8):2149-2158; Marso et al., N Engl J Med.
[0011] 2016;375(19):1834-1844; Marso et al., N Engl J Med. 2016;375(4):311-322). Semaglutide is a 94% structural analogue of native human GLP-1 , consisting of residues 7-37 of human GLP-1 with the sequence H-His-Aib-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val- Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys(C18diacid-y-Glu-OEG-OEG)]-Glu-Phe-lle-Ala- Trp-Leu-Val-Arg-Gly-Arg-Gly-OH. It contains several modifications:
[0012] 1. An alanine at the 8th position (position 2 in semaglutide itself is position 8 in GLP-1 ) is substituted with an aminoisobutyric acid (Aib) to enhance resistance against dipeptidyl peptidase-4 (DPP-4), an enzyme responsible for the rapid degradation of native GLP-1 .
[0013] 2. A spacer and C18 fatty di-acid chain are attached to the lysine at position 26 via a glutamic acid linker. This modification promotes albumin binding, which in turn reduces renal clearance and extends the circulation time of the drug.
[0014] Semaglutide, as a GLP-1 agonist, has the following effects:
[0015] 1. Glucose-Dependent Insulin Secretion: Increases insulin secretion from pancreatic beta cells when glucose levels are elevated.
[0016] 2. Decrease in Glucagon Secretion: Reduces glucagon release from pancreatic alpha cells, which decreases hepatic glucose production.
[0017] 3. Delay in Gastric Emptying: Slows gastric emptying, which moderates postprandial glucose excursions.
[0018] 4. Appetite Suppression: Acts on the brain to induce satiety, leading to reduced food intake and aiding weight loss.
[0019] See Lau J, et al., J Med Chem. 2015;58(18):7370-7380; Wilding et al., N Engl J Med. 2021 ;384(11 ):989-1002.
[0020] Semaglutide is currently marketed under different brand names (e.g., Ozempic for the injectable version and Rybelsus for the oral version) and is available in different formulations, including a once-weekly injectable form and a daily oral tablet. However, manufacturing semaglutide is a bottleneck because introduction of 2-AIB within the peptide is a chemical process which is difficult to scale up to high levels of production.
[0021] There remains a need for alternative GLP-1 agonist formulations which have a half-life as long or longer than semaglutide, which can be administered with comparable frequency to semaglutide, and which are easier to manufacture.
[0022] Using orthogonal aaRS / tRNA pairs, it is possible to produce GLP-1 analogues containing a single ncAA substituent using stop codon suppression. Multiple substitutions, however, will generally require recoded cells. Genome-wide replacement of a target codon with synonymous codons (synonymous codon compression) may provide a foundation for reassigning sense codons to non- canonical amino acids (or other monomers) to facilitate the in vivo biosynthesis of genetically encoded non-canonical biopolymers (Chin, J.W., 2017. Nature, 550(7674), 53-60).
[0023] Site-directed mutagenesis approaches have been used to replace up to 321 amber stop codons in the E. co / / genome (Mukai, T., et al., 2015. Scientific reports, 5, p.9699). However, sense codons are commonly orders of magnitude more abundant than stop codons, and genome synthesis, rather than mutagenesis, may be the preferred route to tackling sense codon removal in many cases.
[0024] Genome synthesis has enabled the creation of Mycoplasma with synthetic genomes (Gibson, D.G., et al., 2010. Science, 329(5987), 52-56) and the creation of nine strains of S. cerevisiae in which the DNAfor one or two of the sixteen chromosomes is replaced by synthetic DNA (Zhang, W., et al., 2017. Science, 355(6329), eaaf3981 ; and Richardson, S.M., et al., 2017. Science, 355(6329), 1040-1044). These experiments have replaced up to 1 Mb of DNA (0.99 Mb, yeast; 1.08 Mb, Mycoplasma) in individual strains. Replicon excision for enhanced genome engineering through programmed recombination (REXER) has been reported for replacing > 100 kb of the E. co / / genome with synthetic DNA in a single step. Moreover, it has been shown that REXER can be iterated via genome stepwise interchange synthesis (GENESIS) to replace 220 kb of the E. coli genome with 230 kb of synthetic DNA (Wang, K., et al., 2016. Nature, 539(7627), 59-64; WO 2018 / 020248).
[0025] Synonymous codons have been altered in individual genes (Napolitano, M.G., et al., 2016. PNAS, 113(38), E5588-E5597), genomic regions and essential operons (Wang, K., et al., 2016. Nature, 539(7627), 59-64; and Lau, Y.H., et al. 2017. Nucleic acids research, 45(11 ), 6971-6980). For instance, Wang et al. used defined 'recoding schemes' to replace a 20 kb region of the E. coli genome rich in both essential genes and target codons.
[0026] In WO2020 / 229592 and WO2022 / 248061 , we describe the development of synthetic organisms with multiple codon replacements for the introduction of ncAAs into expressed polypeptides.
[0027] Summary of the Invention
[0028] The present invention provides an alternatively modified GLP-1 [7-37] analogue, in which an ncAAis inserted at position 7 and / or position 8, wherein if the amino acid at position 7 is Histidine, the substituent at position 8 is selected from a-methyl-L-valine or Isovaline. Preferably, the ncAA at position 7 has an aromatic sidechain.
[0029] Preferably, the ncAA substituent at position 7 is selected from napA, 2-F-Phe and 2-CI- Phe.
[0030] In a preferred embodiment, a substitution is made at position 8, adjacent to an ncAA substituent at position 7. The substitution at position 8 may be a canonical or non- canonical amino acid.
[0031] Canonical amino acid substituents at position 8 may be selected from Glu, Asp or Gin, or from Asn, Gin or Asp.
[0032] In one embodiment, it is a canonical amino acid, and is preferably selected from Glu and Asn. Accordingly, the invention provides a GLP-1 agonist comprising amino acids 7-37 of human GLP-1 , wherein position 7 is substituted with a ncAA, which may be selected from napA, 2-F-Phe and 2-CI-Phe and position 8 is substituted with Glu or Asn.
[0033] Preferably, the ncAA at position 7is napA or 2-F-Phe.
[0034] In one embodiment, the GLP-1 analogue of the invention is [napA]-Glu-Glu-Gly-Thr-Phe- Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe-lle-Ala-Trp-Leu-Val- Arg-Gly-Arg-Gly.
[0035] In another embodiment, the GLP-1 analogue is [2F-Phe]-Asn-Glu-Gly-Thr-Phe-Thr-Ser- Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe-lle-Ala-Trp-Leu-Val-Arg-Gly- Arg-Gly.
[0036] In a further embodiment, the GLP-1 analogue comprises a substitution at position 8 selected from a-methyl-L-valine and Isovaline. Preferably, position 7 is Histidine.
[0037] In one embodiment, the GLP-1 analogue of the invention is His-[a-Me-Val-OH]-Glu-Gly- Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe-lle-Ala-Trp- Leu-Val-Arg-Gly-Arg-Gly.
[0038] In another embodiment, the GLP-1 analogue of the invention is His-[lva]-Glu-Gly-Thr- Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe-lle-Ala-Trp-Leu- Val-Arg-Gly-Arg-Gly.
[0039] The GLP-1 agonists of the invention bind to the GLP-1 receptor with greater affinity than semaglutide, and have a comparable half-life. Moreover, because they can be produced by expression, for example using stop-codon suppression in non-recoded cells, or codon replacement in recoded cells comprising orthogonal aminoacyl-tRNA synthetases (aaRSs) and orthogonal tRNAs which encode the ncAA, the GLP-1 agonists of the invention can be produced in large quantities without chemical processing.
[0040] In a second aspect, therefore, the invention provides a method for preparing a GLP-1 analogue comprising a ncAA, wherein the GLP-1 analogue is expressed in a host cell which comprises an orthogonal amino acyl-tRNA synthetase / tRNA pair which incorporates the ncAA into a nascent polypeptide, wherein
[0041] (1 ) The GLP-1 analogue is expressed in a host cell together with a stop codon suppressor tRNA, and the ncAA is encoded by the cognate stop codon; or
[0042] (2) The GLP-1 analogue is expressed in a recoded cell, and the ncAA is encoded by a recoded codon.
[0043] In one embodiment, a suppressor tRNA suppresses the TAG stop codon, and the ncAA is incorporated in resepone to the TAG stop codon. In another embodiment, the stop codon can be recoded; for example, the recoded codon can be TAG. Sense codons may also be recoded; multiple codons may be recoded, permitting the incorporation of multiple different ncAAs.
[0044] The invention accordingly provides a method for preparing a GLP-1 analogue comprising a ncAA, said method comprising: providing a host cell; providing a nucleic acid encoding the GLP-1 agonist; providing a tRNA-tRNA synthetase pair orthogonal to said host cell, wherein the orthogonal tRNA-tRNA synthetase pair is a cognate pair capable of interacting with the ncAA; adding the ncAA, wherein said ncAA is a substrate for said orthogonal tRNA synthetase; and incubating to allow incorporation of said ncAA into the GLP-1 analogue via the orthogonal tRNA-tRNA synthetase pair.
[0045] In embodiments, the recoded organism comprises a mutant eRF1 , said mutant eRF1 having amino acid sequence having at least 60% sequence identity to the human wild type eRF1 sequence.
[0046] Mutant amino-acyl tRNA synthetases are known in the art, and can comprise synthetases derived such as from archaea or bacteria.
[0047] The aminoacyl-tRNA synthetase may be genetically engineered.
[0048] Brief Description of the Figures Figure 1 illustrates the three ncAA used in the examples of the present invention, namely (2-naphthyl)alanine (napAor 2Nal), 2-fluoro-L-phenylalanine (2-F-Phe), 2-chloro-L- phenylalanine (2-CI-Phe), a-methyl-L-valine (a-Me-Val-OH) and Isovaline (Iva).
[0049] Figure 2 sets forth BLI binding affinity traces for: A: semaglutide (2-AiB), and the peptides of the invention, B: napA-GLP-1 , C: 2-F-Phe-GLP-1 and D: 2-CI-Phe-GLP-1 . The respective Kd are 55.5nM for semaglutide, 37.4nM for napA-GLP-1 , 43.5nM for 2-F- Phe-GLP-1 , and 41 ,5nM for 2-CI-Phe-GLP-1 .
[0050] Figure 3 illustrates the half-life of semaglutide, wt GLP-1 and the analogues of the invention in the presence of DPP-4 protease. 2-F-Phe-GLP-1 had a half-life of 13 hours in this experiment, which is identical to semaglutide. napA-GLP-1 was stable in this experiment (labelled [2Nal]-GLP-1 in Figure 3).
[0051] Figure 4 is a chart of half-lives of the peptides of the invention, in addition to wt GLP-1 and semaglutide, as determined in Figure 3.
[0052] Figure 5 is a concentration-response curve for agonist-induced cAMP accumulation in CHO-K1 / Ga15 cells expressing human GLP-1 R, comparing the effects of semaglutide, CBP0015 and CBP0019.
[0053] Detailed Description of the Invention
[0054] The terms “comprising", “comprises" and “comprised of as used herein are synonymous with “including” or “includes"; or “containing” or “contains", and are inclusive or open- ended and do not exclude additional, non-recited members, elements or steps. The terms “comprising”, “comprises" and “comprised of also include the term “consisting of.
[0055] GLP-1 , as referred to herein, is a peptide hormone derived from proglucagon.
[0056] The tissue-specific cleavage of proglucagon by selective expression of the prohormone convertase (PC1 ) enzymes results in the liberation of GLP-1 , GLP-2, glicentin, oxyntomodulin, and IP2.
[0057] In contrast, PC2 expression results in cleavage of Gcg into “pancreatic type" glucagon, GRPP, MPGF, and a small intervening peptide (IP1 ).
[0058] Several forms of GLP-1 are processed from proglucagon and vary in their ability to enhance glucose-induced insulin secretion. The different forms include GLP-1 (1-37) (or 1-36amide) and two “truncated” forms, GLP-1 (7-36amide) (“amidated GLP-1 ”) and GLP- 1 (7-37) (“glycine-extended GLP-1”). In humans, nearly all circulating GLP-1 is one of the truncated forms, with about 80% of GLP-1 immunoreactivity corresponding to GLP-1 (7- 36amide) and about 20% to the glycine-extended GLP-1 (7-37). As used herein, “GLP-1” is human GLP-1 (7-37), and has the sequence His-Ala-Glu-Gly- Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe-lle-Ala-Trp- Leu-Val-Arg-Gly-Arg-Gly.
[0059] Native GLP-1 has a very short half-life, of about 1-2 min, and results from two causes: (a) the action of the enzyme dipeptidylpeptidase-4 (DPP-4) and (b) renal elimination. DPP-4 cleaves GLP-1 (7-37) at the N-terminal dipeptide to generate GLP-1 (9-37), which has a low affinity for the GLP-1 receptor. This peptide and other GLP-1 metabolites are also rapidly cleared from the circulation via the kidneys.
[0060] In contrast to semaglutide, in which N-terminal dipeptide is modified by the insertion of the ncAA Aib at position 8, we have shown that a N-terminal ncAA at position 7 and replacement of Ala 8 with Glu or Asn, or substitution at position 8 with Iva or a-Me-Val-OH provides superior stability and resistance to digestion by DPP-4.
[0061] The term "analogue" generally refers to a peptide, the sequence of which has one or more amino acid changes when compared to a reference amino acid sequence. A GLP-1 analogue, therefore, is a peptide the sequence of which comprises one or more amino acid changes respective to human GLP1 (7-37).
[0062] As used herein, “non-canonical amino acids” or ncAAs are amino acids that are not naturally encoded or found in the genetic code. Despite the use of only 22 amino acids by the translational machinery to assemble proteins (the proteinogenic amino acids - 20 in the standard genetic code and an additional 2 that can be incorporated by special translation mechanisms), over 140 amino acids are known to occur naturally in proteins and thousands more may occur in nature or be synthesized in the laboratory. Thus, non- canonical amino acids may comprise any amino acid excluding L-alanine, L-cysteine, L- aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L- threonine, L-valine, L-tryptophan and L-tyrosine. Unnatural amino acids additionally exclude L-pyrrolysine and L-selenocysteine.
[0063] In some embodiments, the non-canonical amino acids are unnatural amino acids (UAAs).
[0064] Suitable non-canonical amino acid and UAAs will be well known to those of skill in the art, for example those disclosed in Neumann, H., 2012. FEBS letters, 586(15), pp.2057-2064; and Liu, C.C. and Schultz, P.G., 2010. Annual review of biochemistry, 79, pp.413-444. In some embodiments the non-proteinogenic amino acid and / or UAAs for use in the position 7 substitution are selected from one or more of: p-Acetylphenylalanine, m- Acetylphenylalanine, O-allyltyrosine, Phenylselenocysteine, p-Propargyloxyphenylalanine, p-Azidophenylalanine, p-Boronophenylalanine, O-methyltyrosine, p-Aminophenylalanine, p-Cyanophenylalanine, m-Cyanophenylalanine, p-Fluorophenylalanine, p- lodophenylalanine, p-Bromophenylalanine, p-Nitrophenylalanine, L-DOPA, 3- Aminotyrosine, 3-lodotyrosine, p-lsopropylphenylalanine, 3-(2-Naphthyl)alanine, Biphenylalanine, Homoglutamine, D-tyrosine, p-Hydroxyphenyllactic acid, 2-Aminocaprylic acid, Bipyridylalanine, HQ-alanine, p-Benzoylphenylalanine, o-Nitrobenzylcysteine, o- Nitrobenzylserine, 4,5-Dimethoxy-2-nitrobenzylserine, o-Nitrobenzyllysine, o- Nitrobenzyltyrosine, 2-Nitrophenylalanine, Dansylalanine, p-Carboxymethylphenylalanine, 3-Nitrotyrosine, Sulfotyrosine, Acetyllysine, Methylhistidine, 2-Aminononanoic acid, 2- Aminodecanoic acid, Pyrrolysine, Cbz-lysine, Boc-lysine and Allyloxycarbonyllysine.
[0065] Preferred ncAAs are ncAAs with aromatic sidechains. Examples include (2- naphthyl)alanine (napA), 2-Fluorophenylaalanine (2-F-Phe) and 2-Chlorophenylalanine (2-CI-Phe). NapA and 2-F-Phe are preferred.
[0066] “Binding affinity" is defined as the strength of the interaction between the GLP-1 analogue and its binding partner, the GLP-1 receptor, as measured by the equilibrium dissociation constant (Kd).
[0067] The equilibrium dissociation constant (Kd) is the concentration of ligand at which half of the available binding sites are occupied in a system at equilibrium. Preferably, the Kd for the analogue of the invention is 50nM or below. Preferably, the Kd is below 45nM. Preferably it is below 38nM.
[0068] The sequence of the analogue of the invention may be substituted at one or more positions, in addition to the incorporation of the ncAA at position 7. Any position may be substituted. In a preferred embodiment, the length of the analogue is not changed, that is no amino acids are added or removed without being substituted.
[0069] A preferred position for a substitution is position 8. The substituent selected for position 8 will depend on the ncAA incorporated at position 7; if the ncAAis napA, the substituent at position 8 is preferably selected from Glu, Asp or Gin. Glu is preferred.
[0070] If position 7 is unsubstituted, that is His is retained, then the position 8 substituent is selected from Iva and a-Me-Val-OH.
[0071] If the position 7 ncAA is 2-F-Phe, the substituent at position 8 is preferably selected from Asn.GIn or Asp. Asn is preferred.
[0072] Further substitutions may be made to the analogue, for example as described in WO2022 / 018815.
[0073] Analogues according to the invention may synthetically produced. Advantageously, they are expressed in recoded cells which can incorporate ncAAs into polypeptides. A recoded host cell, as referred to herein, is a cell in which the endogenous genome has bene recoded to make a codon available for incorporation of a ncAA. Preferably, an endogenous tRNA has been deleted, and / or an endogenous elongation factor has been deleted, to prevent misincorporation of the ncAA into endogenous essential genes.
[0074] The recoded host cell may be a eukaryotic cell, or a prokaryotic cell. Preferably, a prokaryotic cell is used. The cells may be non-recoed, or recoded. In one example, a non-recoded prokaryotic cell is used. In another example a recoed prokaryotic cell is used.
[0075] Typically a prokaryotic cell will be produced by genetic modification of a pre-existing (i.e. “parental” or “parent") cell. Thus, a prokaryotic cell may be derived from a parent cell, i.e. be identical to a parent cell, except for comprising one or more genetic modifications. The skilled person will be able to readily identify the parent cell on which a prokaryotic cell is based and the genetic modifications carried out. As used herein, a “parent cell" may be any naturally-occurring, commercially-available, deposited, catalogued or otherwise well- known cell, or derivative thereof.
[0076] A prokaryote is a unicellular organism that lacks a membrane-bound nucleus, mitochondria, or any other membrane-bound organelle. Prokaryotes are divided into two domains, Archaea and Bacteria. The genome of prokaryotic organisms generally is a circular, double-stranded piece of DNA, multiple copies of which may exist at any time.
[0077] Preferably, the prokaryotic cell of the present invention is a bacterial cell. Preferably the prokaryotic cell is suitable for heterologous protein production, in particular the production of polypeptides comprising one or more canonical or non-canonical amino acids (for instance those described by Ferrer-Miralles, N. and Villaverde, A., 2013. Microbial Cell Factories, 12:113). Suitable bacterial cells include: Escherichia (e.g. Escherichia coll), caulobacteria (e.g. Caulobacter crescentus), phototrophic bacteria (e.g. Rodhobacter sphaeroides), cold adapted bacteria (e.g. Pseudoalteromonas haloplanktis, Shewanella sp. strain Ac10), pseudomonads (e.g. Pseudomonas fluorescens, Pseudomonas putida, Pseudomonas aeruginosa), halophilic bacteria (e.g. Halomonas elongate, Chromohalobacter salexigens), streptomycetes (e.g. Streptomyces lividans, Streptomyces griseus), nocardia (e.g. Nocardia lactamdurans), mycobacteria (e.g. Mycobacterium smegmatis), coryneform bacteria (e.g. Corynebacterium glutamicum, Corynebacterium ammoniagenes, Brevibacterium lactofermentum), bacilli (e.g. Bacillus subtilis, Bacillus brevis, Bacillus megaterium, Bacillus licheniformis, Bacillus amyloliquefaciens), and lactic acid bacteria (e.g. Lactococcus lactis, Lactobacillus plantarum, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri) cells. In some embodiments the prokaryotic cell is a gram-negative bacterial cell.
[0078] Preferably, the prokaryotic cell of the present invention is an Escherichia coli, Salmonella enterica, or Shigella dysenteriae cell. These are phylogenetically related species as disclosed by Lukjancenko, O., et al., 2010. Microbial ecology, 60(4), pp.708-720; and Karberg, K.A., et al., 2011. PNAS, 108(50), pp.20154-20159.
[0079] More preferably, the prokaryotic cell of the present invention is an E. coli cell. The parent cell may be any suitable E. coli, including K-12, MG1655, BL21 , BL21(DE3), AD494, Origami, HMS174, BLR(DE3), HMS174(DE3), Tuner(DE3), Origami2(DE3), Rosetta2(DE3), Lemo21 (DE3), NiCo21(DE3), T7 Express, SHuffle Express, C41(DE3), C43(DE3), and m15 pREP4 or derivatives thereof (Rosano, G.L. and Ceccarelli, E.A., 2014. Frontiers in microbiology, 5, p.172). Most preferably, the parent cell is MDS42, MG1655, or BL21 or a derivative thereof. MG1655 is considered as the wild type strain of E coli. The GenBank ID of genomic sequence of this strain is U00096. BL21 is widely available commercially. For example, it can be purchased from New England BioLabs with catalog number C2530H (https: / / www.neb.com / products / c2530-bl21-competent-e-coli).
[0080] Preferably one or more tRNA or release factors may be deleted from the prokaryotic cell and the cell may remain viable. For example, a tRNA which decodes only the one or more sense codons that have been replaced (or deleted) may be dispensable. Similarly, a tRNA which decodes the one or more sense codons that have been replaced (or deleted) may be dispensable if the remaining sense codons that it decodes may also be decoded by an alternative tRNA. For example, serT, encoding tRNASerUGA, is the only tRNA that decodes TCA codons in E. coli, and is therefore normally essential. However, if the genome of the prokaryotic cell does not contain TCA codons then serT may be dispensable.
[0081] Methods for modifying bacterial cells for the production of polymers comprising non- canonical amino acids are set forth in WO2020229592. Such methods are useful for producing prokaryotic cells useful for the production of GLP-1 analogues comprising ncAA.
[0082] When the genome of a cell has been modified, the prokaryotic cell preferably does not display a substantially decreased growth rate. Thus, preferably the prokaryotic cell does not have a substantially decreased growth rate relative to the host cell comprising the parent genome. In some embodiments the prokaryotic cell has a doubling time less than 4 times, 3 times, 2 times, or about 1 .6 times, slower than the parent cell. The doubling time can be determined by any method known to those of skill in the art. In some embodiments the doubling time is determined at 37°C, 25°C or 42°C, in LB media. When the genome of a cell has been modified, the cell advantageously does not have any substantial phenotypical changes. Thus, preferably the prokaryotic cell does not have any substantial phenotypical changes relative to the host cell. In some embodiments the host cell comprising the synthetic prokaryotic genome has a mean cell length less than 100%, 50%, or about 20% greater than the host cell comprising the parent genome. For example, the cell length may be about 1 .5 to 3 microns. The cell length can be determined by any method known to those of skill in the art.
[0083] Codon replacement
[0084] As used herein, a “sense codon” is a nucleotide triplet that codes for an amino acid. Thus, sense codons may be identified in the genome of a cell by gene prediction, i.e. by identifying regions of the genome that code for proteins (i.e. genes) and the corresponding open reading frames (ORFs). Typically, genomes naturally comprise 61 sense codons: GCT, GCC, GCA, GCG, CGT, CGC, CGA, CGG, AGA, AGG, AAT, AAC, GAT, GAC, TGT, TGC, CAA, CAG, GAA, GAG, GGT, GGC, GGA, GGG, CAT, CAC, ATT, ATC, ATA, TTA, TTG, CTT, CTC, CTA, CTG, AAA, AAG, ATG, TTT, TTC, CCT, CCC, CCA, CCG, TCT, TCC, TCA, TCG, AGT, AGC, ACT, ACC, ACA, ACG, TGG, TAT, TAC, GTT, GTC, GTA, and GTG (read from 5’ to 3’ on the coding strand of DNA). The standard genetic code encodes the 20 canonical amino acids using the 61 triplet codons. 18 of the 20 amino acids are encoded by more than one synonymous codon (see Figure 17). The one or more sense codons may be one or more native sense codons, i.e. sense codons which are present in the parent genome.
[0085] The 61 sense codons in DNA are transcribed into corresponding mRNA and subsequently decoded by one or more tRNAs. tRNAs carry an amino acid to a ribosome as directed by the sense codons in the mRNA. The tRNAs can recognise one or more sense codons via a complementary anticodon. A sequence of sense codons is subsequently translated into a polypeptide (i.e. a sequence of amino acids).
[0086] Preferably, the genome-wide removal of the one or more sense codons, but not other sense codons, enables all the cognate tRNA corresponding to said one or more sense codons to be deleted without removing the ability to decode the one or more sense codons remaining in the genome. Thus, the one or more sense codons may be selected from: TCG, TCA, AGT, AGC, GCG, GCA, GTG, GTA, CTG, CTA, TTG, TTA, ACG, ACA, CCG, CCA, CGG, CGA, CGT, CGC, AGG, AGA, GGG, GGA, GGT, GGC, ATT, and ATC.
[0087] Aminoacyl-tRNA synthetases for serine, leucine and alanine do not recognize the anticodons of their cognate tRNAs. This may facilitate the assignment of codons within these boxes to new amino acids through the introduction of tRNAs bearing cognate anticodons that do not direct mis-aminoacylation by endogenous synthetases. Thus, the one or more sense codons may be selected from: TCG, TCA, TOT, TOO, AGT, AGO, GCG, GCA, GCT, GCC, CTG, CTA, CTT, CTC, TTG, and TTA.
[0088] Preferably, the one or more sense codons fulfill both these criteria, thus the one or more sense codons may be selected from: TCG, TCA, AGT, AGC, GCG, GCA, CTG, CTA, TTG, and TTA. More preferably, the one of more sense codons are selected from TCG, TCA, AGT, AGC, TTG, TTA, GCG and GCA. Most preferably, the one of more sense codons are TCG and / or TCA.
[0089] Preferably, one or more sense codons are removed such that the genome of the prokaryotic cell is compatible with codon reassignment to non-canonical amino acids. Thus, the one or more sense codons may comprise one or more of TCA, CTA, or TTA. Alternatively, two or more sense codons are removed, wherein the two or more sense codons comprise one or more of the sense codon pairs, selected from the group consisting of: GCG and GCA; GCT and GCC; TCG and TCA; AGT and AGC; TCT and TCC; CTG and CTA; TTG and TTA; and CTT and CTC. Preferably, two or more sense codons are removed, wherein the two or more sense codons comprise one or more of the sense codon pairs, selected from the group consisting of: GCG and GCA; TCG and TCA; AGT and AGC; CTG and CTA; and TTG and TTA. More preferably, the two or more sense codons comprise TCG and TCA.
[0090] To achieve removal of sense codons they may be replaced with synonymous sense codons. This is preferable to ensure that the encoded protein sequence of the cell genome is not changed. The person skilled in the art is able to deduce suitable synonymous sense codon replacements. For example, in E. coli, typically TCG, TCA, TCT, TCC, AGT and AGC all encode serine; typically GCG, GCA, GCT and GCC all encode alanine; typically CTG, CTA, CTT, CTC, TTG and TTA all encode leucine.
[0091] In some embodiments, the replacement is a defined replacement, i.e. one sense codon is replaced with a single synonymous sense codon.
[0092] For example, the defined replacement may be: GCG replaced with either GCT or GCC; GCA replaced with either GCT or GCC; TCG replaced with any one of TCT, TCC, AGT, or AGC; TCA replaced with any one of TCT, TCC, AGT, or AGC; AGT replaced with any one of TCG, TCA, TCT, or TCC; AGC replaced with any one of TCG, TCA, TCT, or TCC; CTG replaced with any one of CTT, CTC, TTG or TTA; CTA replaced with any one of CTT, CTC, TTG or TTA; TTG replaced with any one of CTG, CTA, CTT or CTC; or TTA replaced with any one of CTG, CTA, CTT or CTC. Preferably the one or more defined sense codon replacements are selected from one or more of: GCG to either GCT or GCC; GCA to either GCT or GCC; TOG to either AGT or AGO; TCAto either AGT or AGO; AGT to either TCA or TOT; AGO to either TOG or TGC or TCA; TTG to CTT; and TTAto CTC. More preferably, TCG and / or TCA are replaced with AGC and / or AGT. Most preferably, TCG is replaced with AGC and / or TCA is replaced with AGT. Preferably, the defined replacement is such that the genome is compatible with codon reassignment to ncAAs. For example: (i) GCG may be replaced with either GCT or GCC, and GCA may be replaced with either GCT or GCC; (ii) TCG may be replaced with any of TCT, TCC, AGT, or AGC, and TCA may be replaced with any of TCT, TCC, AGT, or AGC; (iii) AGT may be replaced with any of TCG, TCA, TCT, or TCC, and AGC may be replaced with any of TCG, TCA, TCT, or TCC; (iv) CTG may be replaced with any of CTT, CTC, TTG or TTA, and CTA may be replaced with any of CTT, CTC, TTG or TTA; or (v) TTG may be replaced with any of CTG, CTA, CTT or CTC, and TTA may be replaced with any of CTG, CTA, CTT or CTC.
[0093] Preferably, the defined replacement scheme is one or more of those listed in the table below:
[0094] Preferably, none of these codon replacements affect ribosomal binding sites (AGGAGG), which are highly conserved regulatory sequences in E. coli. The selected codon replacements may be tested on a small test region (e.g. a 20 kb region of the genome rich in both essential target genes and target codons) to assess viability. If the codon replacements are not viable on the small test region they may be disregarded.
[0095] When replacement of one or more sense codons in the parent cell genome with defined replacement synonymous sense codons does not result in a viable genome, alternative replacement synonymous sense codons may be used. For instance, 99.9% of the occurrences of one or more sense codons in the parent cell genome may be replaced with a defined (i.e. single) synonymous sense codon, and the remaining 0.1% with alternative synonymous sense codons. For example, 99.9% of the occurrences of TCG may be replaced with AGO and 0.1% replaced with TOT, TOO, AGT or AGO; and / or 99.9% of the occurrences of TCA may be replaced with AGT and 0.1% replaced with TOT, TOO, AGT or AGO.
[0096] Preferably 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of the one or more sense codons in the unmodified prokaryotic cell are replaced with synonymous sense codons. In some embodiments 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCG and / or TCA in the unmodified prokaryotic cell are replaced with AGC and / or AGT, most preferably 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCG in the unmodified parent prokaryotic cell are replaced with AGC and / or 90%, 95%, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, 99.9% or more, or 100% of the occurrences of TCA in the unmodified parent prokaryotic cell are replaced with AGT.
[0097] As used herein, a “stop codon” or “nonsense codon" is a nucleotide triplet that codes for termination of translation into proteins. Typically, genomes naturally comprise 3 stop codons: TAA (“ochre”), TGA(“opal” or “umber”) and TAG (“amber”).
[0098] In some embodiments the prokaryotic cell genome further comprises 10 or fewer, 5 or fewer, or no occurrences of one or two stop codons, preferably 10 or fewer, 5 or fewer, or no occurrences of the amber stop codon (TAG). Preferably wherein 90% or more, 95% or more, 98% or more, 99% or more, or all of the occurrences of TAG in the parent prokaryotic genome are replaced with TAA (the ochre stop codon). In preferred embodiments the synthetic prokaryotic genome comprises no occurrences of the amber stop codon (TAG), optionally wherein all of the occurrences of TAG in the parent prokaryotic genome are replaced with TAA (the ochre stop codon).
[0099] Accordingly, in preferred embodiments the prokaryotic cell genome comprises no occurrences of one or more, or two or more sense codons and no occurrences of one stop codon, preferably the amber stop codon (TAG). In more preferred embodiments the prokaryotic cell genome of the present invention comprises no occurrences of two sense codons, preferably TOG and TCA, and no occurrences of the amber stop codon (TAG), optionally wherein TOG, TCA and TAG in the parent prokaryotic genome are replaced with synonymous codons, for example 99.9% or more of the occurrences of TCG in the parent prokaryotic cell genome are replaced with AGC, 99.9% or more of the occurrences of TCA in the parent prokaryotic cell genome are replaced with AGT and all of the occurrences of TAG in the parent prokaryotic cell genome are replaced with TAA.
[0100] Codon reassignment and orthogonal tRNA synthetases
[0101] In some embodiments the one or more sense codons (i.e. those removed from the parent genome) are reassigned to encode alternative canonical amino acids. For example, if TCG and TCA have been removed, one or both may be reassigned to encode a monomer other than serine.
[0102] For instance, the genome of the prokaryotic cell of the present invention substantially or completely lacks one or more sense codons. Therefore, one or more tRNA or release factors may be deleted from the synthetic genome. For instance, a tRNA which decodes the one or more sense codons that have been replaced (or deleted) may be deleted from the synthetic prokaryotic genome. A tRNA which decodes one or more sense codons that have been replaced (or deleted) may be deleted and the prokaryotic cell will remain viable if the tRNA decodes only the one or more sense codons that have been replaced (or deleted); or alternatively if the tRNA decodes one or more sense codons that have been replaced (or deleted) and one or more sense codons that have not been replaced (or deleted), if the tRNA is dispensable for the one or more sense codons that have not been replaced (or deleted) (i.e. the one or remaining sense codons which the tRNA decodes are decoded by one or more alternative tRNAs). For example, if the prokaryotic cell genome lacks TCA sense codons, serT, encoding tRNASeruGA, may be deleted and / or if the prokaryotic cell genome lacks TCG sense codons, serU, encoding tRNASerCGA, may be deleted. The deletion of one or more tRNAs may be used, for instance, in combination with an orthogonal aminoacyl-tRNA synthetase / tRNA pair to reassign the one or more sense codons to a ncAAor alternative amino acid. For example, if TCG and TCA have been removed from the synthetic prokaryotic genome, serT, encoding tRNASerUGA, and serU, encoding tRNASerCGA, may be deleted from the synthetic prokaryotic genome, and either the IRNACGA can be reassigned (e.g. to tRNAAlacGA) an orthogonal aminoacyl-tRNA synthetase / tRNAcGA pair may be introduced to the host cell (e.g. by a heterologous nucleic acid or by incorporation into the synthetic prokaryotic genome) to reassign TCG to a ncAA. Thus, the prokaryotic cell of the present invention further comprises one or more reassigned tRNAs and / or one or more heterologous nucleotides (e.g. plasmids) encoding one orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. In some embodiments the host cell of the present invention further comprises a plasmid encoding an orthogonal aminoacyl-tRNA synthetase (aaRS)- tRNA pair. Alternatively, the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair may be introduced into the host cell by incorporation into the genome of the prokaryotic cell. Thus, in some embodiments the prokaryotic cell genome encodes an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair, preferably wherein the gene encoding the native tRNA has been deleted from the prokaryotic cell genome. In preferred embodiments the prokaryotic cell of the present invention further comprises one or more reassigned tRNAs. Methods for reassigning tRNAs will be well known to those of skill in the art.
[0103] Thus, the present invention provides for use of a recoded cell for producing GLP-1 analogues one or more ncAAs, two or more ncAAs, ory three or more ncAAs.
[0104] Genetic code expansion uses an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair to direct the incorporation of non-proteinogenic amino acids into proteins, in response to an unassigned codon (e.g. the amber stop codon, UAG) introduced at the desired site in a gene of interest. The orthogonal synthetase does not recognize endogenous tRNAs, and specifically aminoacylates an orthogonal cognate tRNA (which is not an efficient substrate for endogenous synthetases) with the monomer provided to (or synthesized by) the cell (Chin, J.W., 2017. Nature, 550(7674), 53-60). The person skilled in the art would be able to identify and / or generate suitable orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pairs (e.g. Elliott, T. S. et al., 2014. Nat Biotechnol 32, 465-472; Elliott, T. S., et al., 2016. Cell Chem Biol 23, 805-815; and Krogager, T. R et al., 2018. Nat Biotechnol 36, 156-159). Thus, in some embodiments, the prokaryotic cell of the present invention further comprises one or more heterologous nucleotides (e.g. plasmids) encoding one orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. In preferred embodiments the prokaryotic cell of the present invention further comprises a plasmid encoding an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. Alternatively, the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair may be introduced into the prokaryotic cell by incorporation into the prokaryotic cell genome. Thus, in some embodiments the prokaryotic genome encodes an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair, preferably wherein the gene encoding the native tRNA has been deleted from the parent prokaryotic genome.
[0105] In some embodiments the prokaryotic cell genome lacks genes encoding release factors (e.g. RF1 ) and / or the host cell lacks release factors (e.g. RF1 ) to increase the efficiency of incorporation of non-proteinogenic amino acids.
[0106] In order to incorporate two ncAA, it is necessary to use two different tRNA synthetase / tRNA pairs, which must be mutually orthogonal so as to avoid cross-incorporation of different monomers. Thus, a first orthogonal tRNA synthetase - tRNA pair and a second orthogonal tRNA synthetase - tRNA pair may be introduced into a prokaryotic cell of the invention. The first tRNA may decode one of the sense codons removed from the genome, and the second tRNA may decode another sense codon removed from the genome, and / or a nonsense codon removed from the genome. The orthogonal tRNA synthetases specifically charge their orthogonal cognate tRNA with a ncAA. Hence, the prokaryotic cell may contain a system wherein two or more sense codons have been repurposed to code for ncAAs.
[0107] Orthogonal tRNA synthetase enzymes and paired tRNAs may be obtained from any suitable source. Organisms from which aaRS may be derived include Methanosarcina mazei, Archeoglobus fulgidus, Methanomethylophilus sp, Methanocaldococcus jannaschii and Methanosarcina barkeri. For example, orthogonal aaRS / tRNA pairs include Methanosarcina mazei (Mm)PylRS / MmtRNAPylcGA, Archeoglobus fulgidus (Af)TyrRS(plF) / AftRNATyr(A01)cuA, Methanomethylophilus sp. 1 R26
[0108] (1 R26)PylRS(CbzK) / AlvtRNAANPyl(8)cuA and Methanocaldococcus jannaschii ( / W / )TyrRS(Nap) / / WjtRNATyrcuA. Methanosarcina barkeri PylRS has also been used to incorporate ncAAs.
[0109] Preferably, a suitable aaRS is selected for its ability to interact with the intended ncAA. The aaRS may then be mutagenized to improve its specificity for a desired ncAA.
[0110] Advantageously, the aaRS is mutagenized in order to select variants which are able to incorporate different monomers. For example, a library of aaRS can be constructed by randomising positions M300, L301 , A302, M344 and N346 in MmPyIRS with degenerate codons. Further mutagenesis can be used to select mutants better able to incorporate ncAAs having aromatic side chains; preferably, mutations are selected for at positions C348, V401 and W417, which delimit the pocket to the enzyme which binds to the substrate sidechain. Other mutations can be introduced, based on the principle of variation of the enzyme in the region which interacts with the sidechain of the monomer. Reassignment schemes
[0111] Once a codon has been selected for reassignment, the aaRS and cognate tRNA can be selected in order to determine which monomer is incorporated at the defined codon position. Thus, although the nucleic acid encoding the polymer will determine the sequence of the monomers in the polymer, the aaRS and cognate tRNA which incorporate the monomer will determine the identity of the incorporated monomer.
[0112] It has been shown that certain aaRS are more selective for certain monomers, and in addition certain aaRS / tRNA pairs are mutually orthogonal in their acylation activity. Thus reassignment schemes can be devised which promote incorporation of different monomers at desired locations in the polymer without cross-reacting to incorporate undesired monomers.
[0113] Advantageously, therefore, the first and second reassigned codons are recognised by orthogonal tRNAs, which in turn are acylated by aaRS enzymes which are mutually orthogonal in their monomer specificity.
[0114] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made to the Examples, which are not intended to limit the invention in any way.
[0115] Examples
[0116] Analogue design
[0117] We performed computational thermodynamic simulations on GLP-l-like peptides containing non- canonical amino acids at the N-terminus. We provided a database of >150 non-canonical amino acids. The initial search returned at least 3 potential hits of peptides containing ncAAs at position 7, expected to bind GLP-1 more tightly than semaglutide. The non-canonical amino acids are napA, 2-F-Phe and 2-CI-Phe. See Figure 1. Subsequently, we identified two possible ncAA substituents at position 8. These are Iva and a-Met-Val-OH.
[0118] Binding to GLP-1
[0119] The GLP-1 analogues were synthesised and binding to GLP-1 R was assessed by BLI. All tested peptides returned higher affinity for GLP-1 R than semaglutide. See Figure 2.
[0120] Analogue stability
[0121] The stability of the GLP-1 analogues was tested by exposure to DPP-4 in solution, and examination of HPLC digestion profiles (see Figure 3). Both [2-F-Phe] GLP-1 and [napA]-GLP-1 have superior stability to semaglutide, and vastly superior stability to wt GLP-1 . napA-GLP1 is stable in the experiment shown in Figure 3. 2-CI-Phe-GLP-1 was poorly soluble in the protease solution and could not produce reliable data. However, some the available data from all of the peptides was used to chart half-life as shown in Figure 4. The half-life values are as follows:
[0122] NapA-GLP-1 and 2-F-Phe-GLP-1 have a half-life extending beyond that semaglutide. The half-life of napA-GLP-1 is significantly longer, which would suggest that this peptide can be administered more infrequently than semaglutide. lva-GLP-1 and a-Met-Val-OH-GLP1 also have half lives comparable to semaglutide. cAMP accumulation assay
[0123] The analogues of the invention were tested in a cAMP accumulation GLP-1 assay. This assay measures the activation of the glucagon-like peptide-1 (GLP-1 ) receptor by measuring the accumulation of cyclic adenosine monophosphate (cAMP): when GLP-1 binds to its receptor, it activates the G protein complex, which stimulates adenylate cyclase activity and increases cAMP levels.
[0124] GLP-1 R agonists mimic the action of GLP-1 by binding to the receptor and triggering receptor activation. This results in increase cAMP accumulation, which can be used to monitor the GLP-1 R activation process, and thus the biological activity of GLP-1 analogues. The accumulation of intracellular cAMP can be measured through a competitive immunoassay based on HTRF technology (cAMP Gs dynamic kit, Revvity, cat. no 62AM4PE) where native cAMP produced by cells competes with d2-labelled cAMP for binding to a monoclonal Europium Cryptate-labelled antibody, and the detected signal drops as a function of increasing intracellular cAMP concentrations. The accumulation of intracellular cAMP can then be plotted as function of the concentration of agonist present.
[0125] CHO-K1 cells are incubated with the peptides of the invention, then an agonist challenge is added. After a period of incubation, cellular cAMP levels are measured. The data from the experiment is tabulated below; CBP0015 comprises a-Me-Val-OH at position 8, whilst
[0126] CBP0019 comprises Iva at position 8.
[0127] Figure 5 shows a concentration-response curve for agonist-induced cAMP accumulation in CHO-K1 / Ga15 cells expressing human GLP-1 R. Each data point represents a mean of replicates from three independent experiments with errors (mean ± SEM), except for GLP1(7-37) which was tested in a single assay run only (mean of technical duplicated ± SEM).
Claims
Claims:
1. A GLP-1 [7-37] analogue comprising an ncAA at position 7 and / or position 8, wherein if the amino acid at position 7 is Histidine, the substituent at position 8 is selected from a-methyl-L-valine or Isovaline.
2. The GLP-1 analogue of claim 1 , wherein the ncAA at position 7 has an aromatic sidechain.
3. The GLP-1 analogue of claim 1 or claim 2, wherein the ncAA at position 7 is selected from napA, 2-F-Phe and 2-CI-Phe.
4. The GLP-1 analogue according to any preceding claim, wherein a substitution is made at position 8, adjacent to an ncAA residue at position 7.
5. The GLP-1 analogue according to claim 4, wherein the substitution at position 8 is a canonical amino acid.
6. The GLP-1 analogue according to claim 5, wherein the amino acid substituents at position 8 may be selected from Glu, Asp or Gin, or from Asn, Gin or Asp.
7. The GLP-1 analogue according to claim 6, wherein the ncAA is napA and the substituent at position 8 is Glu.
8. The GLP-1 analogue according to claim 6, wherein the ncAA is 2-F-Phe, and the substituent at position 8 is Asn.
9. The GLP-1 analogue according to claim 7, having the sequence [napA]-Glu-Glu- Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe- lle-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly.
10. The GLP-1 analogue according to claim 8, having the sequence [2F-Phe]-Asn- Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu- Phe-lle-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly.
11. The GLP-1 analogue according to claim 1, having the sequence His-[a-Me-Val- OH]-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys- Glu-Phe-lle-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly.
12. The GLP-1 analogue according to claim 1, having the sequence His-[lva]-Glu-Gly- Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-GIn-Ala-Ala-Lys-Glu-Phe-lle- Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly.
13. The GLP-1 analogue according to any preceding claim, which has a higher affinity for the GLP-1 receptor than semaglutide.
14. The GLP-1 analogue according to any preceding claim, which has higher stability in the presence of DPP-4 protease than semaglutide.
15. A method for preparing a GLP-1 analogue comprising a ncAA, wherein the GLP-1 analogue is expressed in a host cell which comprises an orthogonal amino acyl- tRNA synthetase / tRNA pair which incorporates the ncAA into a nascent polypeptide, wherein a. The GLP-1 analogue is expressed in a host cell together with a stop codon suppressor tRNA, and the ncAA is encoded by the cognate stop codon; or b. The GLP-1 analogue is expressed in a recoded cell, and the ncAA is encoded by a recoded codon.
16. A method accrdong to claim 15, wherein a suppressor tRNA suppresses the TAG stop codon, and the ncAA is incorporated in response to the TAG stop codon.
17. A method according to claim 15, wherein the stop codon can be recoded and / or one or more sense codons is recoded.
18. A method according to claim 15, wherein the GLP-1 analogue is an analogue according to any one of claims 1 to 12.
19. A method according to claim 15 or claim 16, comprising: a. providing a host cell; b. providing a nucleic acid encoding the GLP-1 analogue; c. providing a tRNA-tRNA synthetase pair orthogonal to said host cell, wherein the orthogonal tRNA-tRNA synthetase pair is a cognate pair capable of interacting with the ncAA; d. adding the ncAA, wherein said ncAA is a substrate for said orthogonal tRNA synthetase; and e. incubating to allow incorporation of said ncAA into the GLP-1 analogue via the orthogonal tRNA-tRNA synthetase pair.