Somatostatin derivatives, processes for their preparation and use

By constructing semaglutide precursor fusion protein using the Fmoc orthogonal protection method and green fluorescent protein folding units, the problems of difficult-to-obtain raw materials and high costs in semaglutide synthesis have been solved, achieving efficient preparation of semaglutide, improving yield and purity, and reducing production costs.

CN115667318BActive Publication Date: 2025-10-21NINGBO KUNPENG BIOTECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202180041125.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-24
Filing Date
2021-06-11
Publication Date
2025-10-21
Estimated Expiration
2041-06-11

AI Technical Summary

Technical Problem

Existing technologies for synthesizing semaglutide suffer from problems such as difficulty in obtaining raw materials, high costs, low coupling efficiency, and the generation of impurities, making it difficult to achieve industrial-scale production.

Method used

Semaglutide was prepared using the Fmoc orthogonal protection method. A semaglutide precursor fusion protein was constructed using green fluorescent protein folding units. The purification and synthesis conditions were optimized. The semaglutide precursor fusion protein was prepared by fermentation with recombinant bacteria. Enzymatic digestion and side chain reactions were performed to obtain the Fmoc and Boc modified semaglutide backbone.

Benefits of technology

It improved the yield of semaglutide, reduced production costs, simplified the production process, reduced the generation of racemic impurities, and improved product purity and yield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

Provided are a somatostatin derivative and a method for preparing the same. Specifically, provided is a fusion protein comprising a green fluorescent protein folding unit and a somatostatin precursor or an active fragment thereof, and the expression amount of the fusion protein is significantly improved. Also, the green fluorescent protein folding unit in the fusion protein can be digested into small fragments by a protease, and the molecular weight difference is large compared to the target protein, and it is easy to separate. Also provided is a method for preparing somatostatin and an intermediate using the fusion protein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedicine, and more particularly to a semaglutide derivative and application thereof. Background Art

[0002] Diabetes is a major global health threat. In China, with changing lifestyles and an aging population, the prevalence of diabetes is rapidly increasing. Acute and chronic complications of diabetes, especially chronic complications that affect multiple organs, can lead to high rates of disability and mortality, severely impacting patients' physical and mental health and placing a heavy burden on individuals, families, and society.

[0003] Semaglutide is a hypoglycemic drug developed by Novo Nordisk. It can significantly reduce glycated hemoglobin (HbA1c) levels and reduce weight in patients with type 2 diabetes, while also significantly reducing the risk of hypoglycemia. Semeglutide is obtained by modifying and remodeling GLP-1 (7-37). Compared with liraglutide, semeglutide has a longer fatty chain and increased hydrophobicity. However, semeglutide has been modified with short-chain PEG, which greatly enhances its hydrophilicity. PEG modification not only allows it to bind tightly to albumin and mask the DPP-4 enzymatic hydrolysis site, but also reduces renal excretion, prolonging its biological half-life and achieving a long-circulation effect.

[0004] The CAS number of semaglutide is 910463-68-2, and its English name is Semaglutide. Its sequence is as follows: H-His1-Aib2-Glu3-Gly4-Thr5-Phe6-Thr7-Ser8-Asp9-Val10-Ser11-Ser12-Tyr13-Leu14-Glu15-Gly16-Gln17-Ala18-Ala19-Lys20(PEG2-PEG2-γ-Glu-Octadecanedioic acid)-Glu21-Phe22-Ile23-Ala24-Trp25-Leu26-Val27-Arg28-Gly29-Arg30-Gly31-OH.

[0005] Patent application number CN201611095162 uses a fragment condensation method to synthesize fully protected semaglutide, and then obtains crude semaglutide peptide after cleavage. Because this method uses fragments for condensation, its raw materials are difficult to obtain and are expensive. In addition, the main chain is first condensed to Thr at position 5, and then the side chain protecting group Alloc of Lys at position 20 is removed to condense the side chain. This method is prone to polycondensation of the fragment 2 resin during the synthesis process, greatly reducing the coupling efficiency of the amino acid after Lys at position 20 and fragment 1, and is prone to produce racemic impurities, which is not conducive to industrial production.

[0006] Patent application number CN201511027176 describes a solid-phase synthesis method for obtaining fully protected semaglutide resin. Crude semaglutide is then cleaved and purified to obtain refined semaglutide. This method involves first condensing the backbone, then removing the Alloc group protecting the Lys side chain, and then condensing the side chains. This method is prone to polycondensation of the resin during the synthesis process, significantly reducing coupling efficiency and generating racemic impurities, particularly the racemization of the final amino acid, His. This significantly reduces product yield and increases production costs.

[0007] Therefore, those skilled in the art have devoted themselves to new methods for producing semaglutide. Summary of the Invention

[0008] The object of the present invention is to provide a semaglutide derivative and application thereof.

[0009] In the first aspect of the present invention, a semaglutide precursor fusion protein is provided. The semaglutide precursor fusion protein has a structure shown in Formula I from N-terminus to C-terminus:

[0010] A-FP-TEV-EK-G(I)

[0011] Where,

[0012] “-” represents a peptide bond;

[0013] A is none or a leader peptide sequence,

[0014] FP is the green fluorescent protein folding unit;

[0015] TEV is the first restriction enzyme cleavage site, preferably the TEV enzyme cleavage site (as shown in the sequence ENLYFQG, SEQ ID NO: 8);

[0016] EK is the second restriction enzyme cleavage site, preferably an enterokinase cleavage site (as shown in the sequence DDDDK, SEQ ID NO: 9);

[0017] G is a semaglutide precursor or a fragment thereof;

[0018] Wherein, the green fluorescent protein folding unit comprises 2-6 β-folding units selected from the following group:

[0019] β-pleated sheet unit Amino acid sequence u1 VPILVELDGDVNG (SEQ ID NO: 11) u2 HKFSVRGEGEGDAT (SEQ ID NO: 12) u3 KLTLKFICTT (SEQ ID NO: 13) u4 YVQERTISFKD (SEQ ID NO: 14) u5 TYKTRAEVKFEGD (SEQ ID NO: 15) u6 TLVNRIELKGIDF (SEQ ID NO: 16) u7 HNVYITADKQ (SEQ ID NO: 17) u8 GIKANFKIRHNVED (SEQ ID NO: 18) u9 VQLADHYQQNTPIG (SEQ ID NO: 19) u10 HYLSTQSVLSKD (SEQ ID NO: 20) u11 HMVLLEFVTAAGI (SEQ ID NO:21).

[0020] In another preferred embodiment, the green fluorescent protein folding unit is u2-u3, u4-u5, u1-u2-u3, u3-u4-u5 or u4-u5-u6.

[0021] In another preferred embodiment, the G is a Boc-modified semaglutide precursor, the semaglutide precursor lacks 2-7 amino acids at the N-terminus of the semaglutide main chain, and the lysine contained in the semaglutide precursor is modified by Boc.

[0022] In another preferred embodiment, the epsilon amino group of the Boc-modified lysine is modified with a tert-butyloxycarbonyl group.

[0023] In another preferred embodiment, the amino acid sequence of the semaglutide backbone is shown in SEQ ID NO: 3.

[0024] In another preferred embodiment, the semaglutide precursor comprises:

[0025] Position 18 is a Boc-modified first precursor of semaglutide, the amino acid sequence of which is shown in SEQ ID NO: 1;

[0026] Alternatively, the second precursor of semaglutide with a Boc modification at position 17, the amino acid sequence of the second precursor being as shown in SEQ ID NO: 2;

[0027] Alternatively, a third precursor of semaglutide modified with Boc at position 16, wherein the amino acid sequence of the third precursor is shown in SEQ ID NO: 23;

[0028] Alternatively, a fourth precursor of semaglutide modified with Boc at position 15, wherein the amino acid sequence of the fourth precursor is shown in SEQ ID NO: 24;

[0029] Alternatively, the fifth precursor of semaglutide modified with Boc at position 14, the amino acid sequence of the fifth precursor is shown in SEQ ID NO: 25.

[0030] SEQ ID NO:1:EGTFTSDVSSYLEGQAA K EFIAWLVRGRG

[0031] SEQ ID NO:2:GTFTSDVSSYLEGQAA K EFIAWLVRGRG

[0032] SEQ ID NO:23:TFTSDVSSYLEGQAA K EFIAWLVRGRG

[0033] SEQ ID NO:24: FTSDVSSYLEGQAA K EFIAWLVRGRG

[0034] SEQ ID NO:25:TSDVSSYLEGQAA K EFIAWLVRGRG

[0035] (underlined K is Boc-modified lysine)

[0036] In another preferred embodiment, the fourth amino acid at the C-terminus of the semaglutide precursor is arginine or lysine.

[0037] In another preferred embodiment, the 4th arginine at the C-terminus of the semaglutide precursor can be replaced by lysine.

[0038] In another preferred embodiment, the fourth amino acid at the C-terminus of the fusion protein is arginine or lysine.

[0039] In another preferred embodiment, the 4th arginine at the C-terminus of the fusion protein can be replaced by lysine.

[0040] In this application, the complete semaglutide sequence (H(Aib)EGTFTSDVSSYLEGQAAKEFIAWLVRGRG, SEQ ID NO: 3) is defined as the semaglutide backbone, and the semaglutide with the N-terminal amino acid deleted is defined as the semaglutide precursor. For the Fmoc-modified semaglutide backbone, the H at its N-terminus is modified with Fmoc; for the Boc-modified semaglutide backbone, the 20th lysine is Nε-(tert-butyloxycarbonyl)-lysine.

[0041] In another preferred embodiment, the green fluorescent protein folding unit is u3-u4-u5.

[0042] In another preferred embodiment, the amino acid sequence of the leader peptide is shown in SEQ ID NO: 7.

[0043] In another preferred embodiment, the 14th, 15th, 16th, 17th or 18th position of the semaglutide precursor is Nε-(tert-butyloxycarbonyl)-lysine (i.e., the amino acid at position 20 of the semaglutide backbone of each semaglutide precursor is Nε-(tert-butyloxycarbonyl)-lysine).

[0044] In the second aspect of the present invention, a semaglutide backbone modified with Fmoc and Boc is provided, wherein position 20 of the semaglutide backbone is a protected lysine, the protected lysine is Nε-(tert-butyloxycarbonyl)-lysine, and the N-terminus of the semaglutide backbone is an Fmoc-modified histidine.

[0045] In another preferred embodiment, the Fmoc is fluorenylmethoxycarbonyl.

[0046] In another preferred embodiment, the amino acid sequence of the semaglutide backbone is shown in SEQ ID NO: 3.

[0047] In the third aspect of the present invention, a Boc-modified semaglutide precursor is provided, wherein the semaglutide precursor comprises:

[0048] Position 18 is a Boc-modified first precursor of semaglutide, the amino acid sequence of which is shown in SEQ ID NO: 1;

[0049] Alternatively, the second precursor of semaglutide with a Boc modification at position 17, the amino acid sequence of the second precursor being as shown in SEQ ID NO: 2;

[0050] Alternatively, a third precursor of semaglutide modified with Boc at position 16, wherein the amino acid sequence of the third precursor is shown in SEQ ID NO: 23;

[0051] Alternatively, a fourth precursor of semaglutide modified with Boc at position 15, wherein the amino acid sequence of the fourth precursor is shown in SEQ ID NO: 24;

[0052] Alternatively, the fifth precursor of semaglutide modified with Boc at position 14, the amino acid sequence of the fifth precursor is shown in SEQ ID NO: 25.

[0053] In a fourth aspect of the present invention, an Fmoc-modified semaglutide backbone is provided, wherein the N-terminus of the semaglutide backbone is an Fmoc-modified histidine, and the amino acid sequence of the semaglutide backbone is shown in SEQ ID NO: 3.

[0054] In a fifth aspect of the present invention, a method for preparing semaglutide is provided, comprising the steps of:

[0055] (A) using recombinant bacteria for fermentation to prepare semaglutide precursor fusion protein,

[0056] (B) using the semaglutide precursor fusion protein to prepare semaglutide,

[0057] Wherein, the semaglutide precursor fusion protein is as described in the first aspect of the present invention.

[0058] In another preferred embodiment, the step (B) further comprises the steps of:

[0059] (i) performing enzymatic digestion on the semaglutide precursor fusion protein to obtain a Boc-modified semaglutide precursor, wherein the Boc-modified semaglutide precursor lacks X amino acids at the N-terminus of the semaglutide backbone, wherein X is an integer of 2-7;

[0060] (ii) connecting an Fmoc complex to the N-terminus of the Boc-modified semaglutide precursor to prepare an Fmoc- and Boc-modified semaglutide backbone,

[0061] Wherein, the Fmoc complex comprises X amino acids at the N-terminus of the semaglutide backbone, and the N-terminal amino acid of the Fmoc complex is modified by Fmoc;

[0062] (iii) performing a Boc removal treatment on the Fmoc- and Boc-modified semaglutide backbone, and reacting the backbone with the semaglutide side chain to obtain Fmoc-modified semaglutide; and

[0063] (iv) performing a de-Fmoc treatment on the Fmoc-modified semaglutide to obtain de-Fmoc semaglutide;

[0064] (v) performing side chain OtBu removal treatment on the de-Fmoc semaglutide to obtain semaglutide.

[0065] In another preferred embodiment, in step (i), the enzyme cleavage is performed using enterokinase.

[0066] In another preferred embodiment, the Boc-modified semaglutide precursor comprises:

[0067] A first precursor of semaglutide modified with Boc at position 18, wherein the amino acid sequence of the first precursor is shown in SEQ ID NO: 1;

[0068] Alternatively, a second precursor of semaglutide modified with Boc at position 17, wherein the amino acid sequence of the second precursor is shown in SEQ ID NO: 2;

[0069] Alternatively, a third precursor of semaglutide modified with Boc at position 16, wherein the amino acid sequence of the third precursor is shown in SEQ ID NO: 23;

[0070] Alternatively, a fourth precursor of semaglutide modified with Boc at position 15, wherein the amino acid sequence of the fourth precursor is shown in SEQ ID NO: 24;

[0071] Alternatively, the fifth precursor of semaglutide modified with Boc at position 14, the amino acid sequence of the fifth precursor is shown in SEQ ID NO: 25.

[0072] In another preferred embodiment, the Fmoc complex is Fmoc-H-Aib, Fmoc-H-Aib-E, Fmoc-H-Aib-EGTF, Fmoc-H-Aib-EGT or Fmoc-H-Aib-EG.

[0073] In another preferred embodiment, in step (i) and step (ii), the value of X is the same.

[0074] In another preferred embodiment, the Fmoc and Boc modified semaglutide backbone is as described in the second aspect of the present invention.

[0075] In another preferred embodiment, the reaction of step (ii) is as follows:

[0076]

[0077] In another preferred embodiment, the side chain of semaglutide is as follows:

[0078]

[0079] In another preferred embodiment, in step (ii), Fmoc complex (activated ester), DIPEA (N,N-diisopropylethylamine) and DMF (N,N-dimethylformamide) are added to connect the Fmoc complex to the N-terminus of the Boc-modified semaglutide precursor.

[0080] In another preferred embodiment, the Fmoc complex is an Fmoc complex in the form of an activated ester formed by activation with HOSu / DCC, HoBt / DIC, or TBTU / DIPEA.

[0081] In another preferred embodiment, the molar ratio of the added Fmoc complex (activated ester), DIPEA and Boc-modified semaglutide precursor is (1.0-3.0):(10-14):(0.8-1.2), preferably (2-2.8):(11-13):(0.8-1.2).

[0082] In another preferred embodiment, between step (ii) and step (iii), the method further comprises the step of purifying the prepared Fmoc- and Boc-modified semaglutide backbones.

[0083] In another preferred embodiment, the purification treatment is to add an organic solvent to the reaction solution to obtain a solid product, and more preferably, the organic solvent is a mixture of tertiary methyl ether and petroleum ether.

[0084] In another preferred embodiment, in step (iii), the method further comprises the following steps:

[0085] (a) Compound 2 (Fmoc- and Boc-modified semaglutide backbone) was added to a pre-cooled TFA solution at 0±5°C, stirred, and subjected to Boc removal treatment to obtain a Boc-free product;

[0086] (b) adding an organic solvent to the reaction solution of step (a) to obtain a solid Boc-free product, preferably the organic solvent is a mixture of tertiary methyl ether and petroleum ether;

[0087] (c) The Boc-free product is mixed with the side chain of semaglutide to prepare Fmoc-modified semaglutide.

[0088] In another preferred embodiment, in step (c), the solid Boc-free product is mixed with the semaglutide side chain in DMF and reacted at room temperature.

[0089] In another preferred embodiment, in step (c), the reaction system further comprises DIPEA.

[0090] In another preferred embodiment, in step (iv), a DMF solution containing piperidine is added to perform a de-Fmoc treatment to obtain de-Fmoc semaglutide.

[0091] In another preferred embodiment, in step (v), a mixed solution of TFA, TIS and DCM is added to remove the side chain OtBu protecting group to obtain semaglutide.

[0092] In another preferred embodiment, step (v) includes a step of purifying the prepared semaglutide.

[0093] In another preferred embodiment, the Boc-modified semaglutide precursor is prepared using genetic recombination technology.

[0094] In another preferred embodiment, in step (A), semaglutide precursor fusion protein inclusion bodies are isolated from the fermentation broth of the recombinant bacteria, and the inclusion bodies are renatured and enzymatically digested to obtain semaglutide precursor fusion protein.

[0095] In another preferred embodiment, before and after step (i), a purification step is further included, preferably reverse phase chromatography.

[0096] In another preferred embodiment, the recombinant bacteria contains or integrates an expression cassette for expressing the semaglutide precursor fusion protein.

[0097] In another preferred embodiment, the method is as follows:

[0098]

[0099] In another preferred embodiment, the method comprises the steps of:

[0100] (i) providing the semaglutide precursor fusion protein described in the first aspect of the present invention, and enzymatically cleaving the semaglutide precursor fusion protein to obtain compound 1,

[0101] (ii) linking compound 1 to an Fmoc-H-Aib complex (preferably in the form of an activated ester) to produce compound 2 (Fmoc and Boc-modified semaglutide backbone),

[0102] (iii) performing a Boc removal treatment on the compound 2 and reacting it with the side chain of semaglutide to obtain compound 4; and

[0103] (iv) performing Fmoc removal treatment on compound 4 to obtain compound 5;

[0104] (v) Compound 5 is subjected to side chain OtBu removal treatment to obtain semaglutide represented by compound 6.

[0105] In another preferred embodiment, the method comprises the steps of:

[0106] (i) providing the semaglutide precursor fusion protein described in the first aspect of the present invention, and enzymatically digesting the semaglutide precursor fusion protein to obtain compound 7, compound 8, compound 9, or compound 10;

[0107] (ii) linking compound 7 to the Fmoc-H-Aib-E complex (preferably in the activated ester form),

[0108] Alternatively, compound 8 is linked to a Fmoc-H-Aib-EG complex (preferably in the activated ester form),

[0109] Alternatively, compound 9 is linked to a Fmoc-H-Aib-EGT complex (preferably in the activated ester form),

[0110] Alternatively, compound 10 is linked to the Fmoc-H-Aib-EG-TF complex (preferably in the activated ester form);

[0111] Thus, compound 2 was obtained.

[0112] (iii) performing a Boc removal treatment on the compound 2 and reacting it with the side chain of semaglutide to obtain compound 4; and

[0113] (iv) performing Fmoc removal treatment on compound 4 to obtain compound 5;

[0114] (v) Compound 5 is subjected to side chain OtBu removal treatment to obtain semaglutide represented by compound 6.

[0115] In the sixth aspect of the present invention, an isolated polynucleotide is provided, which encodes the semaglutide precursor fusion protein described in the first aspect of the present invention, the Fmoc and Boc modified semaglutide backbone described in the second aspect of the present invention, the Boc modified semaglutide precursor described in the third aspect of the present invention, or the Fmoc modified semaglutide backbone described in the fourth aspect of the present invention.

[0116] The seventh aspect of the present invention provides a vector comprising the polynucleotide described in the sixth aspect of the present invention.

[0117] In another preferred embodiment, the vector is selected from the group consisting of DNA, RNA, plasmid, lentiviral vector, adenoviral vector, retroviral vector, transposon, or a combination thereof.

[0118] The eighth aspect of the present invention provides a host cell, wherein the host cell contains the vector described in the seventh aspect of the present invention, or the exogenous polynucleotide described in the sixth aspect of the present invention integrated into the chromosome.

[0119] In another preferred embodiment, the host cell is Escherichia coli, Bacillus subtilis, yeast cells, insect cells, mammalian cells or a combination thereof.

[0120] The ninth aspect of the present invention provides a preparation comprising the semaglutide precursor fusion protein described in the first aspect of the present invention, the Fmoc and Boc modified semaglutide backbone described in the second aspect of the present invention, the Boc modified semaglutide precursor described in the third aspect of the present invention, or the Fmoc modified semaglutide backbone described in the fourth aspect of the present invention.

[0121] The tenth aspect of the present invention provides a semaglutide preparation, which is prepared using the method described in the fifth aspect of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0122] Figure 1 A map of plasmid pBAD-FP-TEV-EK-GLP-1 (18) is shown.

[0123] Figure 2 A map of plasmid pBAD-FP-TEV-EK-GLP-1 (17) is shown.

[0124] Figure 3 A map of plasmid pEvol-pylRs-pylT is shown.

[0125] Figure 4 The SDS-PAGE electrophoresis of Boc-semaglutide precursor fusion protein inclusion bodies is shown.

[0126] Figure 5 The HPLC detection spectrum of Boc-semaglutide precursor 1 is shown.

[0127] Figure 6 The HPLC detection spectrum of Boc-semaglutide precursor 3 is shown.

[0128] Figure 7 A map of plasmid pBAD-FP-TEV-EK-GLP-1 (16) is shown. DETAILED DESCRIPTION

[0129] After extensive and in-depth research, the inventors have discovered a new method and process for preparing semaglutide products. Specifically, the method utilizes the Fmoc orthogonal protection method to perform the side chain addition step in the preparation process of semaglutide, and optimizes the purification and synthesis conditions in the preparation process. The method of the present invention does not require expensive solid-phase synthesis equipment, shortens the production cycle, has a simple production process, and improves product purity and yield. In addition, the present invention also provides a novel precursor fusion protein (Formula I) and corresponding intermediates (i.e., Fmoc and Boc modified semaglutide main chain and Fmoc modified semaglutide main chain). The precursor fusion protein of the present invention can be used to efficiently prepare the intermediate. Using the intermediate of the present invention, on the one hand, the condensation reaction between Lys at position 20 and the side chain is efficient and the conditions are mild; on the other hand, the protection effect of the N-terminal Fmoc is good, and the removal conditions are mild, and it will not cause the racemization of the N-terminal His, and it is not easy to produce racemic impurities. Studies have shown that the use of the optimized precursor fusion protein and optimized intermediate of the present invention can greatly improve the yield of semaglutide and reduce production costs. The present invention was completed on this basis.

[0130] Semaglutide

[0131] Semaglutide, developed by Novo Nordisk and marketed as Semaglutide (CAS number: 204656-20-2), is a human glucagon-like peptide-1 (GLP-1) analog with the following sequence: H-His1-Aib2-Glu3-Gly4-Thr5-Phe6-Thr7-Ser8-Asp9-Val10-Ser11-Ser12-Tyr13-Leu14-Glu15-Gly16-Gln17-Ala18-Ala19-Lys20(PEG2-PEG2-γ-Glu-Octadecanedioic acid)-Glu21-Phe22-Ile23-Ala24-Trp25-Leu26-Val27-Arg28-Gly29-Arg30-Gly31-OH. It shares 97% sequence homology with native human GLP-1.

[0132] Semaglutide is a hypoglycemic drug developed by Novo Nordisk. This product can significantly reduce the level of glycated hemoglobin (HbA1c) and reduce weight in patients with type 2 diabetes, while greatly reducing the risk of hypoglycemia. Semeglutide is obtained by modifying and transforming GLP-1 (7-37). Compared with Liraglutide, Semeglutide has a longer fatty chain and increased hydrophobicity, but Semeglutide has been modified with a short chain PEG, and its hydrophilicity has been greatly enhanced. After PEG modification, it can not only bind tightly to albumin and mask the DPP-4 enzyme hydrolysis site, but also reduce renal excretion, prolong the biological half-life, and achieve a long circulation effect. It can significantly reduce fasting or postprandial blood sugar in patients with type 2 diabetes to regulate blood sugar levels in the body, while also reducing the patient's weight and reducing the risk of death in patients with cardiovascular disease.

[0133] Protein of the present invention

[0134] As used herein, the term "protein of the present invention" includes the precursor fusion protein of the present invention and the corresponding intermediates, specifically including the semaglutide precursor fusion protein described in the first aspect of the present invention, the Fmoc and Boc modified semaglutide backbone described in the second aspect of the present invention, the Boc-modified semaglutide precursor described in the third aspect of the present invention, or the Fmoc-modified semaglutide backbone described in the fourth aspect of the present invention.

[0135] As used herein, “intermediates of the present invention” include the Fmoc- and Boc-modified semaglutide backbones described in the second aspect of the present invention, the Boc-modified semaglutide precursors described in the third aspect of the present invention, or the Fmoc-modified semaglutide backbones described in the fourth aspect of the present invention.

[0136] Fusion protein

[0137] As used herein, the terms "fusion protein of the present invention", "precursor fusion protein of the present invention", and "semaglutide precursor fusion protein of the present invention" are used interchangeably to refer to the semaglutide precursor fusion protein having the structure of Formula I described in the first aspect of the present invention.

[0138] In the present invention, the inventors constructed a semaglutide precursor fusion protein using the green fluorescent protein folding unit, as described in the first aspect of the present invention.

[0139] The green fluorescent protein folding unit FP contained in the fusion protein of the present invention comprises 2-6, preferably 2-3 β-folding units selected from the following group:

[0140] Amino acid sequence u1 VPILVELDGDVNG (SEQ ID NO: 11) u2 HKFSVRGEGEGDAT (SEQ ID NO: 12) u3 KLTLKFICTT (SEQ ID NO: 13) u4 YVQERTISFKD (SEQ ID NO: 14) u5 TYKTRAEVKFEGD (SEQ ID NO: 15) u6 TLVNRIELKGIDF (SEQ ID NO: 16) u7 HNVYITADKQ (SEQ ID NO: 17) u8 GIKANFKIRHNVED (SEQ ID NO: 18) u9 VQLADHYQQNTPIG (SEQ ID NO: 19) u10 HYLSTQSVLSKD (SEQ ID NO: 20) u11 HMVLLEFVTAAGI (SEQ ID NO:21).

[0141] In another preferred embodiment, the green fluorescent protein folding unit FP can be selected from: u8, u9, u2-u3, u4-u5, u8-u9, u1-u2-u3, u2-u3-u4, u3-u4-u5, u5-u6-u7, u8-u9-u10, u9-u10-u11, u3-u5-u7, u3-u4-u6, u4-u7-u10, u6-u8-u10, u1-u2-u3-u4, u2-u3-u4-u5, u3-u4-u3-u4, u3-u5-u7-u9, u5-u6-u7-u8, u1-u3-u7-u9, u2-u2-u7-u8, u7-u2-u5 -u11, u3-u4-u7-u10, u1-I-u2, u1-I-u5, u2-I-u4, u3-I-u8, u5-I-u6, or u10-I-u11.

[0142] In another preferred embodiment, the green fluorescent protein folding unit is u3-u4-u5 or u4-u5-u6.

[0143] As used herein, the term "fusion protein" also includes variant forms having the above-mentioned activities. These variant forms include (but are not limited to): deletion, insertion and / or substitution of 1-3 (usually 1-2, more preferably 1) amino acids, and addition or deletion of one or several (usually within 3, preferably within 2, more preferably within 1) amino acids at the C-terminus and / or N-terminus. For example, in the art, when amino acids with similar or similar properties are substituted, the function of the protein is generally not changed. For another example, adding or deleting one or several amino acids at the C-terminus and / or N-terminus generally does not change the structure and function of the protein. In addition, the term also includes monomeric and multimeric forms of the polypeptides of the present invention. The term also includes linear and non-linear polypeptides (such as cyclic peptides).

[0144] The present invention also includes active fragments, derivatives and analogs of the above-mentioned fusion proteins. As used herein, the terms "fragment", "derivative" and "analog" refer to polypeptides that substantially retain the function or activity of the fusion protein of the present invention. The polypeptide fragments, derivatives or analogs of the present invention can be (i) polypeptides in which one or more conservative or non-conservative amino acid residues (preferably conservative amino acid residues) are substituted, or (ii) polypeptides having a substituent group in one or more amino acid residues, or (iii) polypeptides formed by fusion with another compound (such as a compound that extends the half-life of the polypeptide, such as polyethylene glycol), or (iv) polypeptides formed by fusion of an additional amino acid sequence to this polypeptide sequence (fusion proteins formed by fusion with tag sequences such as a leader sequence, a secretory sequence or 6His). According to the teachings herein, these fragments, derivatives and analogs are within the scope known to those skilled in the art.

[0145] A preferred class of active derivatives refers to polypeptides in which no more than three, preferably no more than two, and more preferably no more than one amino acid is replaced by an amino acid with similar or similar properties compared to the amino acid sequence of the present invention. These conservative variant polypeptides are preferably generated by making amino acid substitutions according to Table A.

[0146] Table A

[0147] Initial residue Representative replacement Preferred substitutions Ala(A) Val; Leu; Ile Val Arg(R) Lys; Gln; Asn Lys Asn(N) Gln; His; Lys; Arg Gln Asp(D) Glu Glu

[0148] Cys(C) Ser Ser Gln(Q) Asn Asn Glu(E) Asp Asp Gly(G) Pro; Ala Ala His(H) Asn; Gln; Lys; Arg Arg Ile(I) Leu; Val; Met; Ala; Phe Leu Leu(L) Ile; Val; Met; Ala; Phe Ile Lys(K) Arg; Gln; Asn Arg Met(M) Leu; Phe; Ile Leu Phe(F) Leu; Val; Ile; Ala; Tyr Leu Pro(P) Ala Ala Ser(S) Thr Thr Thr(T) Ser Ser Trp(W) Tyr; Phe Tyr Tyr(Y) Trp; Phe; Thr; Ser Phe Val(V) Ile;Leu;Met;Phe;Ala Leu

[0149] The present invention also provides analogs of the fusion proteins of the present invention. These analogs may differ from the polypeptides of the present invention in terms of amino acid sequence, in terms of modifications that do not affect the sequence, or in terms of both. Analogs also include analogs having residues other than naturally occurring L-amino acids (e.g., D-amino acids), as well as analogs having non-naturally occurring or synthetic amino acids (e.g., β- and γ-amino acids). It should be understood that the polypeptides of the present invention are not limited to the representative polypeptides exemplified above.

[0150] In addition, the fusion proteins of the present invention may also be modified. Modifications (generally without altering the primary structure) include: chemical derivatization of the polypeptide in vivo or in vitro, such as acetylation or carboxylation. Modifications also include glycosylation, such as those produced by glycosylation during polypeptide synthesis and processing or in further processing steps. Such modifications can be accomplished by exposing the polypeptide to a glycosylation enzyme (such as a mammalian glycosylase or deglycosylase). Modifications also include sequences containing phosphorylated amino acid residues (such as phosphotyrosine, phosphoserine, and phosphothreonine). Also included are polypeptides that have been modified to increase their resistance to proteolysis or optimize their solubility.

[0151] The term "polynucleotide encoding the fusion protein of the present invention" may include a polynucleotide encoding the fusion protein of the present invention, or may also include additional coding and / or non-coding sequences.

[0152] The present invention also relates to variants of the aforementioned polynucleotides, including fragments, analogs, and derivatives encoding polypeptides or fusion proteins having the same amino acid sequence as the present invention. These nucleotide variants include substitution variants, deletion variants, and insertion variants. As is known in the art, an allelic variant is an alternative form of a polynucleotide, which may contain one or more nucleotide substitutions, deletions, or insertions that do not substantially alter the function of the encoded fusion protein.

[0153] The present invention also relates to polynucleotides that hybridize to the above-mentioned sequences and have at least 50%, preferably at least 70%, and more preferably at least 80% identity between the two sequences. The present invention particularly relates to polynucleotides that hybridize to the polynucleotides of the present invention under stringent conditions (or stringent conditions). In the present invention, "stringent conditions" refer to: (1) hybridization and elution at relatively low ionic strength and relatively high temperature, such as 0.2×SSC, 0.1% SDS, 60°C; or (2) the addition of a denaturing agent during hybridization, such as 50% (v / v) formamide, 0.1% calf serum / 0.1% Ficoll, 42°C; or (3) hybridization occurs only when the identity between the two sequences is at least 90%, more preferably at least 95%.

[0154] The fusion proteins and polynucleotides of the present invention are preferably provided in an isolated form, and more preferably, purified to homogeneity.

[0155] The full-length sequences of the polynucleotides of the present invention can generally be obtained by PCR amplification, recombinant methods, or synthetic methods. For PCR amplification, primers can be designed based on the nucleotide sequences disclosed herein, particularly the open reading frame sequences, and commercially available cDNA libraries or cDNA libraries prepared by conventional methods known to those skilled in the art can be used as templates to amplify the relevant sequences. For long sequences, two or more PCR amplifications are often required, followed by splicing the fragments amplified in the correct order.

[0156] Once the relevant sequence is obtained, it can be obtained in large quantities by recombinant methods. This is usually done by cloning it into a vector, then transferring it into cells, and then isolating the relevant sequence from the propagated host cells by conventional methods.

[0157] In addition, the sequences can also be synthesized by artificial synthesis, especially when the fragment length is shorter. Usually, a long fragment can be obtained by synthesizing multiple small fragments and then connecting them.

[0158] Currently, DNA sequences encoding proteins of the present invention (or fragments thereof, or derivatives thereof) can be obtained entirely by chemical synthesis. This DNA sequence can then be introduced into various existing DNA molecules (or vectors) and cells known in the art.

[0159] Methods using PCR techniques to amplify DNA / RNA are preferably used to obtain the polynucleotides of the present invention. In particular, when full-length cDNA is difficult to obtain from a library, the RACE method (RACE - rapid amplification of cDNA ends) is preferably used. Primers used for PCR can be appropriately selected based on the sequence information of the present invention disclosed herein and can be synthesized using conventional methods. The amplified DNA / RNA fragments can be separated and purified using conventional methods, such as gel electrophoresis.

[0160] expression vector

[0161] The present invention also relates to a vector comprising the polynucleotide of the present invention, a host cell produced by genetic engineering using the vector of the present invention or the coding sequence of the fusion protein of the present invention, and a method for producing the polypeptide of the present invention by recombinant technology.

[0162] The polynucleotide sequences of the present invention can be used to express or produce recombinant fusion proteins using conventional recombinant DNA techniques. Generally, the following steps are involved:

[0163] (1) Transforming or transducing a suitable host cell with a polynucleotide (or variant) encoding the fusion protein of the present invention, or a recombinant expression vector containing the polynucleotide;

[0164] (2) Host cells cultured in a suitable culture medium;

[0165] (3) Isolate and purify proteins from culture medium or cells.

[0166] In the present invention, the polynucleotide sequence encoding the fusion protein can be inserted into a recombinant expression vector. The term "recombinant expression vector" refers to bacterial plasmids, bacteriophages, yeast plasmids, plant cell viruses, mammalian cell viruses such as adenoviruses, retroviruses, or other vectors well known in the art. Any plasmid or vector can be used as long as it can replicate and be stable in the host. An important feature of an expression vector is that it generally contains an origin of replication, a promoter, a marker gene, and translation control elements.

[0167] Methods well known to those skilled in the art can be used to construct expression vectors containing the DNA sequence encoding the fusion protein of the present invention and appropriate transcriptional / translational control signals. These methods include in vitro recombinant DNA techniques, DNA synthesis techniques, in vivo recombination techniques, and the like. The DNA sequence can be operatively linked to an appropriate promoter within the expression vector to direct mRNA synthesis. Representative examples of such promoters include the lac or trp promoters of Escherichia coli; the lambda phage PL promoter; eukaryotic promoters including the CMV immediate early promoter, the HSV thymidine kinase promoter, the early and late SV40 promoter, retroviral LTRs, and other known promoters that control gene expression in prokaryotic or eukaryotic cells or their viruses. The expression vector also includes a ribosome binding site for translation initiation and a transcription terminator.

[0168] In addition, the expression vector preferably contains one or more selectable marker genes to provide a phenotypic trait for selection of transformed host cells, such as dihydrofolate reductase, neomycin resistance, and green fluorescent protein (GFP) for eukaryotic cell culture, or tetracycline or ampicillin resistance for Escherichia coli.

[0169] A vector containing the above-mentioned appropriate DNA sequence and an appropriate promoter or control sequence can be used to transform an appropriate host cell to enable it to express the protein.

[0170] Host cells can be prokaryotic cells, such as bacterial cells; lower eukaryotic cells, such as yeast cells; or higher eukaryotic cells, such as mammalian cells. Representative examples include: Escherichia coli, Streptomyces; bacterial cells of Salmonella typhimurium; fungal cells such as yeast, and plant cells (such as ginseng cells).

[0171] When the polynucleotides of the present invention are expressed in higher eukaryotic cells, transcription will be enhanced if an enhancer sequence is inserted into the vector. Enhancers are cis-acting DNA factors, typically about 10 to 300 base pairs in length, that act on promoters to increase gene transcription. Examples include the SV40 enhancer (100 to 270 base pairs on the late replication origin side), the polyoma enhancer on the late replication origin side, and adenovirus enhancers.

[0172] Those skilled in the art will appreciate how to select appropriate vectors, promoters, enhancers and host cells.

[0173] Transformation of host cells with recombinant DNA can be performed using conventional techniques well known to those skilled in the art. When the host is a prokaryotic organism such as Escherichia coli, competent cells capable of absorbing DNA can be harvested after the exponential growth phase and treated with CaCl2, using procedures well known in the art. Another method is to use MgCl2. If desired, transformation can also be performed using electroporation. When the host is a eukaryotic organism, the following DNA transfection methods can be used: calcium phosphate coprecipitation, conventional mechanical methods such as microinjection, electroporation, liposome packaging, etc.

[0174] The obtained transformants can be cultured using conventional methods to express the polypeptide encoded by the gene of the present invention. Depending on the host cell used, the culture medium used can be selected from various conventional culture media. Culture is carried out under conditions suitable for the growth of the host cells. After the host cells grow to an appropriate cell density, the selected promoter is induced using a suitable method (such as temperature conversion or chemical induction), and the cells are cultured for a period of time.

[0175] The recombinant polypeptide in the above method can be expressed intracellularly, on the cell membrane, or secreted extracellularly. If necessary, the recombinant protein can be isolated and purified by various separation methods utilizing its physical, chemical, and other properties. These methods are well known to those skilled in the art. Examples of these methods include, but are not limited to, conventional renaturation treatment, treatment with a protein precipitant (salting out method), centrifugation, osmotic sterilization, ultrafiltration, ultracentrifugation, molecular sieve chromatography (gel filtration), adsorption chromatography, ion exchange chromatography, high performance liquid chromatography (HPLC), and various other liquid chromatography techniques and combinations of these methods.

[0176] Construction of semaglutide expression vector

[0177] A FP-TEV-EK-GLP1 fragment (Boc modification at position 18, 17, 16, 15, or 14) containing the target gene was synthesized. The fragment had recognition sites for restriction endonucleases NcoⅠ and XhoⅠ at both ends. This sequence was codon-optimized to achieve high-level expression of the functional protein in E. coli. After expression, the expression vector "pBAD / His A (Kana R )" and a plasmid containing the target gene "FP-TEV-EK-GLP1 (18, 17, 16, 15 or 14)", the enzyme digestion products were separated by agarose electrophoresis, then extracted using an agarose gel DNA recovery kit, and finally the two DNA fragments were ligated using T4 DNA ligase. The ligation product was chemically transformed into Escherichia coli Top10 cells, and the transformed cells were cultured overnight on LB agar medium (10 g / L yeast peptone, 5 g / L yeast extract, 10 g / L NaCl, 1.5% agar) containing 50 μg / mL kanamycin. Three viable colonies were picked and plated in 5 mL of liquid LB medium (10 g / L yeast peptone, 5 g / L yeast extract, 10 g / L The cells were cultured overnight in 4% NaCl (1% NaCl) and plasmids were extracted using a plasmid miniprep kit. The extracted plasmids were then sequenced using the sequencing oligonucleotide primer 5'-ATGCCATAGCATTTTTATCC-3' (SEQ ID NO: 15) to confirm correct insertion. The resulting plasmid was named "pBAD-FP-TEV-EK-GLP1 (18, 17, 16, 15, or 14)."

[0178] Fmoc modification

[0179] In the field of biomedicine, the use of peptides is becoming more and more extensive. Amino acids are the basic raw materials for peptide synthesis technology. Amino acids contain α-amino and carboxyl groups, and some also contain active side chain groups, such as hydroxyl, amino, guanidine and heterocycle. Therefore, the amino group and the active side chain groups need to be protected in the peptide grafting reaction, and the protecting groups should be removed after the peptide is synthesized, otherwise the mis-grafting of amino acids and many side reactions will occur.

[0180] Fluorenylmethyloxycarbonyl (Fmoc) is a base-sensitive protecting group that can be removed in concentrated ammonia or dioxane-methanol-4N NaOH (30:9:1) as well as 50% dichloromethane solution of ammonia such as piperidine, ethanolamine, cyclohexylamine, 1,4-dioxane, and pyrrolidone.

[0181] Under weak alkaline conditions such as sodium carbonate or sodium bicarbonate, the Fmoc protecting group is generally introduced using Fmoc-Cl or Fmoc-OSu. Compared to Fmoc-Cl, Fmoc-OSu is easier to control reaction conditions and has fewer side reactions.

[0182] Fmoc has strong UV absorption, with maximum absorption wavelengths at 267nm (ε18950), 290nm (ε5280), and 301nm (ε6200). Therefore, it can be detected by UV absorption, bringing significant convenience to automated peptide synthesis. Furthermore, it is compatible with a wide range of solvents and reagents, has high mechanical stability, and can be used with a variety of carriers and activation methods. Therefore, the Fmoc protecting group is the most commonly used in peptide synthesis today.

[0183] Fmoc-OSu (9-fluorenylmethyl succinimidyl carbonate)

[0184]

[0185] Semaglutide side chain

[0186] tBuO-Ste-Glu(AEEA-AEEA-OSu)-OtBu is the side chain of semaglutide.

[0187]

[0188] Semaglutide is prepared by first using genetic recombination technology to obtain a semaglutide precursor with Boc-protected lysine at positions 14, 15, 16, 17 or 18, and then connecting the semaglutide side chain tBuO-Ste-Glu(AEEA-AEEA-OSu)-OtBu to obtain semaglutide.

[0189] Preparation of Semaglutide

[0190] The present invention provides five synthetic routes for semaglutide, as shown in the following formulas A, B, C, D, and E, respectively. Compound 2 modified with an Fmoc complex is prepared from a Boc-semaglutide precursor (compounds 1, 7, 8, 9, and 10). Compound 2 is deprotected with Boc to obtain compound 3. Compound 3 reacts with the activated semaglutide side chain tBuO-Ste-Glu(AEEA-AEEA-OSu)-OtBu to obtain compound 4, which is then subjected to an Fmoc removal reaction to obtain compound 5. The OtBu protecting group of the side chain is removed to finally obtain semaglutide compound 6.

[0191]

[0192] Formula A

[0193]

[0194] Formula B

[0195]

[0196] Formula C

[0197]

[0198] Formula D

[0199]

[0200] Formula E.

[0201] Specifically, the present invention provides a method for preparing semaglutide, comprising the steps of:

[0202] (i) providing a Boc-modified semaglutide precursor;

[0203] (ii) modifying the Boc-modified semaglutide precursor with an activated Fmoc complex to obtain an Fmoc- and Boc-modified semaglutide backbone;

[0204] (iii) performing a Boc removal treatment on the Fmoc- and Boc-modified semaglutide backbone, and reacting the backbone with the semaglutide side chain to obtain Fmoc-modified semaglutide; and

[0205] (iv) performing Fmoc removal and side chain OtBu removal treatment on the Fmoc-modified semaglutide to obtain semaglutide.

[0206] The main advantages of the present invention include:

[0207] (1) The present invention produces a Boc-modified semaglutide precursor without the need for dilution, ultrafiltration, or other methods to remove excess inorganic salts from the fermentation broth supernatant. In the method of the present invention, a chromatography column is used to separate the Boc-semaglutide precursor, and the one-step yield is over 70%, which is three times higher than that of the conventional method. The yield of the Boc-semaglutide precursor is approximately 800-1000 mg / L. In addition, the method of the present invention can remove most pigments, shorten the original multi-step process, and reduce process time and equipment investment costs.

[0208] (2) Due to the protection of Boc-lysine at position 20, the present invention can directly utilize the orthogonal reaction with Fmoc protection to synthesize semaglutide.

[0209] (3) The semaglutide synthesized by the method of the present invention does not contain impurities acylated at the N-terminal fatty acid, which is beneficial for downstream purification and reduces costs.

[0210] (4) Compared with solid-phase synthesis, the method of the present invention does not produce racemic impurity polypeptides, does not require the use of large amounts of modified amino acids, does not use large amounts of organic reagents, has less environmental pollution, and is lower in cost;

[0211] (5) The fusion protein of the present invention contains a high proportion of the semaglutide main chain (increased fusion ratio), and the FP or A-FP in the fusion protein contains arginine and lysine, which can be digested by proteases into small fragments. Compared with the target protein, the molecular weight difference is large and the separation is easy.

[0212] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the invention. The experimental methods in the following examples, for which specific conditions are not specified, are generally based on conventional conditions or the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are calculated by weight.

[0213] Example 1 Construction of Semaglutide Expression Strain

[0214] The construction of the semaglutide expression plasmid was described in the examples of Chinese patent application No. 201910210102.9. The DNA fragments of the fusion protein FP1-TEV-EK-GLP-1 (18, 17, 16, 15, 14) were cloned into the NcoI-XhoI site downstream of the araBAD promoter of the expression vector plasmid pBAD / His A (purchased from NTCC, kanamycin resistance) to obtain the plasmid pBAD-FP1-TEV-EK-GLP-1 (18) or pBAD-FP2-TEV-EK-GLP-1 (17), pBAD-FP2-TEV-EK-GLP-1 (16), pBAD-FP2-TEV-EK-GLP-1 (15), and pBAD-FP2-TEV-EK-GLP-1 (14). Among them, the plasmid map of pBAD-FP1-TEV-EK-GLP-1 (18) or pBAD-FP2-TEV-EK-GLP-1 (17) is as follows Figure 1 , as shown in 2.

[0215] Based on the semaglutide precursors with 2-7 amino acids deleted at the N-terminus as shown in SEQ ID NOs: 1, 2, 23, 24, and 25, respectively, fusion proteins 1, 2, 3, 4, and 5 were constructed.

[0216] The amino acid sequence of fusion protein 1 is shown in SEQ ID NO: 4:

[0217]

[0218] The amino acid sequence of fusion protein 2 is shown in SEQ ID NO: 5:

[0219]

[0220] The amino acid sequence of fusion protein 3 is shown in SEQ ID NO: 26:

[0221]

[0222] The amino acid sequence of fusion protein 4 is shown in SEQ ID NO: 27:

[0223]

[0224] The amino acid sequence of fusion protein 5 is shown in SEQ ID NO: 28:

[0225]

[0226] Among them, the leader peptide sequence is MVSKGEELFTGV (SEQ ID NO: 7)

[0227] The sequence of the green fluorescent protein (FP) folding unit is

[0228] FP1:KLTLKFICTTYVQERTISFKDTYKTRAEVKFEGD(SEQ ID NO:6,U3-U4-U5)

[0229] FP2:YVQERTISFKDTYKTRAEVKFEGDTLVNRIELKGIDF(SEQ ID NO:10,U4-U5-U6)

[0230] The TEV enzyme cleavage site is ENLYFQG (SEQ ID NO: 8);

[0231] Enterokinase cleavage site is DDDDK (SEQ ID NO: 9)

[0232] The amino acid sequences of the semaglutide precursors with 2-7 amino acids deleted at the N-terminus are shown in SEQ ID NOs: 1, 2, 23, 24, and 25, respectively.

[0233] SEQ ID NO:1:EGTFTSDVSSYLEGQAA K EFIAWLVRGRG

[0234] SEQ ID NO:2:GTFTSDVSSYLEGQAA K EFIAWLVRGRG

[0235] SEQ ID NO:23:TFTSDVSSYLEGQAA K EFIAWLVRGRG

[0236] SEQ ID NO:24: FTSDVSSYLEGQAA K EFIAWLVRGRG

[0237] SEQ ID NO:25:TSDVSSYLEGQAA K EFIAWLVRGRG

[0238] ( K is Boc-modified lysine)

[0239] The pylRs DNA sequence was then cloned into the expression vector plasmid pEvol-pBpF (purchased from NTCC, chloramphenicol-resistant) at the SpeI-SalI site downstream of the araBAD promoter. Simultaneously, the DNA sequence for lysyl-tRNA synthetase tRNA (pylTcua) was inserted downstream of the proK promoter using PCR. This plasmid was named pEvol-pylRs-pylT. The plasmid map is shown in Figure 2. Figure 3 shown.

[0240] The constructed plasmids pBAD-FP1-TEV-EK-GLP-1(18) and pEvol-pylRs-pylT were co-transformed into Escherichia coli TOP10 strain, and the recombinant strain expressing the semaglutide precursor fusion protein FP-TEV-EK-GLP-1(18) was screened.

[0241] The constructed plasmids pBAD-FP2-TEV-EK-GLP-1(17) and pEvol-pylRs-pylT were co-transformed into Escherichia coli TOP10 strain, and the recombinant strain expressing the semaglutide precursor fusion protein FP2-TEV-EK-GLP-1(17) was screened.

[0242] The constructed plasmids pBAD-FP1-TEV-EK-GLP-1(16) and pEvol-pylRs-pylT were co-transformed into Escherichia coli TOP10 strain, and the recombinant strain expressing the semaglutide precursor fusion protein FP-TEV-EK-GLP-1(16) was screened.

[0243] Example 2 Expression of Boc-semaglutide precursor

[0244] The three recombinant Escherichia coli seed solutions were inoculated into fermentation medium (yeast peptone, yeast extract powder, glycerol, Boc-L-lysine, buffer and trace elements) at an inoculum size of 5% (V / V), and cultured in batches until the pH rose to 7.05. The carbon and nitrogen sources were fed separately, and the carbon and nitrogen sources were added according to the constant pH method. After feeding, 7.5M ammonia water was automatically added, and the pH was controlled at 7.0-7.2. After 4-6 hours of cultivation, 2.5g / L of L-arabinose was added for induction, and the induction was continued for 14±2h. Three fermentation broths containing semaglutide precursor fusion protein were obtained.

[0245] Example 3 Preparation of Boc-semaglutide precursor inclusion bodies

[0246] After centrifugation of the three fermentation broths obtained in Example 2, the wet cells were mixed with a lysis buffer (0.5-1.5% (ml / ml) Tween 80, 1 mmol / L EDTA-2Na, and 100 mmol / L NaCl) at a volume ratio of 1:1, suspended for 3 hours, and then lysed using a high-pressure homogenizer (800±50 bar, 6-20°C). After lysis, the inclusion bodies were collected by centrifugation, washed with a buffer, and weighed after washing. The yields of inclusion bodies for fusion proteins 1, 2, and 3 were 39-43 g / L, 41-45 g / L, and 40-43 g / L, respectively. The SDS-PAGE electrophoresis results of fusion protein 1 are shown in FIG. Figure 4 shown.

[0247] Example 4 Renaturation and enzymatic cleavage of Boc-semaglutide precursor inclusion bodies

[0248] 8 mol / L urea dissolution buffer was added to the inclusion bodies obtained in Example 3 at a weight-to-volume ratio of 1:15, and the solution was dissolved by stirring at room temperature. The protein concentration was determined by Bradford method to control the total protein concentration of the inclusion body solution to be around 20 mg / ml. The pH was adjusted to 9.0±1.0 with NaOH. The inclusion body solution was added dropwise to a refolding buffer containing 5-20 mmol / L sodium carbonate, 5-20 mmol / L glycine, and 0.3-0.5 mmol / L EDTA-2Na. The inclusion body solution was diluted 5-10 times for refolding. The pH value of the fusion protein refolding solution was maintained at 9.0-10.0, the temperature was controlled at 4-8°C, and the refolding time was 10-20 h.

[0249] The results showed that after dissolution, the proportions of fusion protein 1 and fusion protein 2 were approximately 30% and 33%, and the proportion of fusion protein 3 was approximately 31%.

[0250] Example 5 Preliminary Purification of Boc-Semaglutide Precursor Fusion Protein

[0251] The fusion protein refolding solution obtained in Example 4 was taken and filtered through a 0.45 μm filter membrane to remove undissolved substances; based on the difference in isoelectric points of proteins, the fusion protein was preliminarily purified using a Q anion exchange column.

[0252] The experimental results showed that after anion exchange chromatography, the purity of Boc-semaglutide precursor fusion proteins 1, 2 and 3 reached more than 65%, the loading capacity was about 18 mg / mL, and the yield was greater than 80%.

[0253] Example 6 Enzymatic cleavage of Boc-semaglutide precursor fusion protein

[0254] The Boc-semaglutide precursor fusion protein initially purified in Example 5 was desalted, the pH was adjusted to 7.5-8.5, the temperature was controlled at 18-25° C., and 0.3-0.5 U / mg enterokinase was added for enzymatic digestion for 8-24 h to obtain Boc-semaglutide precursors. The concentrations of Boc-semaglutide precursor 1, precursor 2, and precursor 3 were approximately 0.9 g / L, 1.2 g / L, and 1.0 g / L, respectively, with an enzymatic digestion efficiency of ≥95%.

[0255] Example 7 Reverse Phase Chromatography of Boc-Semaglutide Precursor

[0256] Based on the difference in hydrophobicity between peptides and proteins, C8 reverse phase chromatography was used to purify the Boc-semaglutide precursor to remove most of the impurities.

[0257] 3 M hydrochloric acid was added to the enzymatic cleavage solution of Boc-semaglutide precursor 1, precursor 2, and precursor 3 obtained in Example 6 to adjust the pH of the sample to 2.0-3.0. Acetonitrile was added to the sample to make the acetonitrile concentration in the sample 10% (v / v), and the sample was filtered through a 0.45 μm filter membrane for later use, followed by reverse phase chromatography separation and purification.

[0258] An aqueous solution containing trifluoroacetic acid was used as mobile phase A; an acetonitrile solution containing trifluoroacetic acid was used as mobile phase B. The Boc-semaglutide precursor was combined with the filler, and the loading amount of the Boc-semaglutide precursor was controlled to be no more than 10 mg / mL. The Boc-semaglutide precursor was collected by gradient elution. The experimental results showed that the purity of Boc-semaglutide precursors 1, 2, and 3 collected by reverse phase chromatography was ≥90%, and the yield was greater than 80%. The HPLC detection spectrum of the purified Boc-semaglutide precursor 1 is shown in FIG. Figure 5 The HPLC detection spectrum of precursor 3 is shown in Figure 6 The molecular weights of Boc-semaglutide precursor 1, precursor 2, and precursor 3 measured by mass spectrometry were consistent with the theoretical values.

[0259] Example 8 Preparation of Semaglutide (Fmoc-H-Aib, Route 1) Using Boc-Semaglutide Precursor 1

[0260] Take the Boc-semaglutide precursor 1 (Compound 1) obtained in Example 7 (the molar ratio of the feed is 30 mg as an example), add activated Fmoc-H-Aib, DIPEA and DMF according to the molar ratio in Table 1, and react for 8-12 hours to obtain the Fmoc and Boc protected semaglutide backbone. Among them, Fmoc-H-Aib is an activated ester form of Fmoc-H-Aib formed by HOSu / DCC activation, and an OSu group is connected to its Aib amino acid. Subsequently, a mixed solution of tert-methyl ether / petroleum ether (3:1) at 0±5°C was added to the reaction solution, the precipitate was centrifuged, and the precipitate was washed 2-3 times with tert-methyl ether to obtain Fmoc protected compound 2: Fmoc-GLP-1 (Lys 20 Boc).

[0261] Table 1 Molar ratio of feed

[0262] Boc-semaglutide Fmoc-H-Aib DIPEA DMF Equivalent or volume 1.0eq 2.5eq 12eq 1V

[0263] Compound 2 was added to the pre-cooled 0±5°C TFA solution and stirred for 0.5-2.0h. A 15-20-fold volume of a 0±5°C methyl tert-butyl ether and petroleum ether (3:1) mixture was added to the reaction solution. The precipitate was centrifuged and washed 2-3 times with the mixture to obtain the solid compound 3 after Boc removal: Fmoc-GLP-1 (Lys 20 NH2).

[0264] Take compound 3 after Boc removal, add DMF and 12eq of DIPEA, and stir gently at room temperature for 5 minutes. Dissolve 2.5eq of tBuO-Ste-Glu(AEEA-AEEA-OSu)-OtBu in DMF solution and add it to the resulting mixture. Gently shake the reaction mixture at room temperature for 2-3 hours. Add 15-20 times the volume of the reaction system of a mixed solution of 0±5℃ methyl tert-butyl ether and petroleum ether (3:1) to the reaction system, precipitate and centrifuge, wash the solid with the mixed solution 2-3 times, and vacuum dry to obtain compound 4: Fmoc-GLP-1-(tBuO-Ste-Glu(AEEA-AEEA)-OtBu)(20).

[0265] Compound 4 was added to a DMF solution containing 20% ​​piperidine and reacted at room temperature for 0.5-2.0 hours. A 3:1 mixed solvent of methyl tert-ether and petroleum ether at 0±5°C was then added to the reaction system. The precipitate was centrifuged and the solid was washed 3-5 times with a 3:1 mixed solvent of methyl tert-ether and petroleum ether to obtain compound 5 after removal of Fmoc: GLP-1-(tBuO-Ste-Glu(AEEA-AEEA)-OtBu)(20).

[0266] Compound 5 was added to a mixture of TFA (trifluoroacetic acid), TIS (triisopropylsilane), and DCM (dichloromethane) (90% TFA:10% TIS:DCM = 1:2). The reaction was shaken at room temperature for 2-4 hours to remove the side chain OtBu protecting group. A 10-20-fold volume of a 0±5°C mixed solvent of tert-butyl methyl ether and petroleum ether (3:1) was added to the reaction system. The precipitate was centrifuged, and the solid was washed three times with a mixed solvent of tert-butyl methyl ether and petroleum ether (3:1) to obtain the final product. After HPLC purification, semaglutide with a purity greater than 98% was obtained.

[0267] Example 9 Preparation of Semaglutide (Fmoc-H-Aib-E, Route 2) Using Boc-Semaglutide Precursor 2

[0268] Take the Boc-semaglutide precursor 2 (Compound 7) obtained in Example 7 (the molar ratio of the feed is 30 mg as an example), add activated Fmoc-H-Aib-E, DIPEA and DMF according to the molar ratio in Table 2, and react for 8-12 hours to obtain the Fmoc and Boc protected semaglutide backbone. Among them, Fmoc-H-Aib-E is an activated ester form formed by HOSu / DCC activation. A mixed solution of tertiary methyl ether and petroleum ether (3:1) at 0±5°C was added to the reaction solution, the precipitate was centrifuged, and the precipitate was washed 2-3 times to obtain Fmoc protected compound 2: Fmoc-GLP-1 (Lys 20 Boc).

[0269] Table 2 Molar ratio of feed

[0270] Boc-semaglutide Fmoc-H-Aib-E DIPEA DMF Equivalent or volume 1.0eq 2.5eq 12eq 1V

[0271] Compound 2 was added to the pre-cooled 0±5°C TFA solution and stirred for 0.5-2.0h. A mixture of methyl tert-butyl ether and petroleum ether (3:1) at 0±5°C in an amount 15-20 times the volume of the reaction system was added to the reaction solution. The precipitate was centrifuged and washed 2-3 times with the mixture to obtain the solid compound 3 after Boc removal: Fmoc-GLP-1 (Lys 20 NH2).

[0272] Take compound 3 after Boc removal, add DMF and 12eq of DIPEA, and stir gently at room temperature for 5 minutes. Dissolve 2.5eq of tBuO-Ste-Glu(AEEA-AEEA-OSu)-OtBu in DMF solution and add it to the resulting mixture. Gently shake the reaction mixture at room temperature for 2-3 hours. Add 15-20 times the volume of the reaction system of a mixed solution of 0±5℃ methyl tert-butyl ether and petroleum ether (3:1) to the reaction system, precipitate and centrifuge, wash the solid with the mixed solution 2-3 times, and vacuum dry to obtain compound 4: Fmoc-GLP-1-(tBuO-Ste-Glu(AEEA-AEEA)-OtBu)(20).

[0273] Compound 4 was added to a DMF solution containing 20% ​​piperidine and reacted at room temperature for 0.5-2.0 hours. A 0±5°C mixed solvent of tert-methyl ether and petroleum ether (3:1) was then added to the reaction system. The precipitate was centrifuged and the solid was washed 3-5 times with a mixed solvent of tert-methyl ether and petroleum ether (3:1) to obtain compound 5 after removal of Fmoc: GLP-1-(tBuO-Ste-Glu(AEEA-AEEA)-OtBu)(20).

[0274] Compound 5 was added to a mixture of TFA (trifluoroacetic acid), TIS (triisopropylsilane), and DCM (dichloromethane) (90% TFA:10% TIS:DCM = 1:2). The reaction was shaken at room temperature for 2-4 hours to remove the side chain OtBu protecting group. A 10-20-fold volume of a 0±5°C mixed solvent of tert-butyl methyl ether and petroleum ether (3:1) was added to the reaction system. The precipitate was centrifuged, and the solid was washed three times with a mixed solvent of tert-butyl methyl ether and petroleum ether (3:1) to obtain the final product. After HPLC purification, semaglutide with a purity greater than 98% was obtained.

[0275] Example 10 Preparation of Semaglutide (Fmoc-H-Aib-EG, Route 3) Using Boc-Semaglutide Precursor 2

[0276] Take the Boc-semaglutide precursor 3 (compound 8, 30 mg of the feed is used as an example) obtained in Example 7, add activated Fmoc-H-Aib-EG, DIPEA and DMF according to the molar ratio in Table 3, and react for 8-12 hours to obtain the Fmoc and Boc protected semaglutide backbone. Among them, Fmoc-H-Aib-EG is in the form of an activated ester formed by HOSu / DCC activation. Subsequently, a mixed solution of tert-methyl ether / petroleum ether (3:1) at 0±5°C was added to the reaction solution, the precipitate was centrifuged, and the precipitate was washed 2-3 times with tert-methyl ether for crude purification to obtain Fmoc and Boc protected compound 2: Fmoc-GLP-1 (Lys 20 Boc).

[0277] Table 3 Molar ratio of feed

[0278] Boc-semaglutide precursor Fmoc-H-Aib-EG DIPEA DMF Equivalent or volume 1.0eq 2.5eq 12eq 1V

[0279] Compound 2 was added to the pre-cooled 0±5°C TFA solution and stirred for 0.5-2.0h. A mixture of methyl tert-butyl ether and petroleum ether (3:1) at 0±5°C in an amount 15-20 times the volume of the reaction system was added to the reaction solution. The precipitate was centrifuged and washed 2-3 times with the mixture to obtain the solid compound 3 after Boc removal: Fmoc-GLP-1 (Lys 20 NH2).

[0280] Take compound 3 after Boc removal, add DMF and 12eq of DIPEA, and stir gently at room temperature for 5 minutes. Dissolve 2.5eq of tBuO-Ste-Glu(AEEA-AEEA-OSu)-OtBu in DMF solution and add it to the resulting mixture. Gently shake the reaction mixture at room temperature for 2-3 hours. Add 15-20 times the volume of the reaction system of a mixed solution of 0±5℃ methyl tert-butyl ether and petroleum ether (3:1) to the reaction system, precipitate and centrifuge, wash the solid with the mixed solution 2-3 times, and vacuum dry to obtain compound 4: Fmoc-GLP-1-(tBuO-Ste-Glu(AEEA-AEEA)-OtBu)(20).

[0281] Compound 4 was added to a DMF solution containing 20% ​​piperidine and reacted at room temperature for 0.5-2.0 hours. A 0±5°C mixed solvent of tert-methyl ether and petroleum ether (3:1) was then added to the reaction system. The precipitate was centrifuged and the solid was washed 3-5 times with a mixed solvent of tert-methyl ether and petroleum ether (3:1) to obtain compound 5 after removal of Fmoc: GLP-1-(tBuO-Ste-Glu(AEEA-AEEA)-OtBu)(20).

[0282] Compound 5 was added to a mixture of TFA (trifluoroacetic acid), TIS (triisopropylsilane), and DCM (dichloromethane) (90% TFA:10% TIS:DCM = 1:2). The reaction was shaken at room temperature for 2-4 hours to remove the side chain OtBu protecting group. A 10-20-fold volume of a 0±5°C mixed solvent of tert-butyl methyl ether and petroleum ether (3:1) was added to the reaction system. The precipitate was centrifuged, and the solid was washed three times with a mixed solvent of tert-butyl methyl ether and petroleum ether (3:1) to obtain the final product. After HPLC purification, semaglutide with a purity greater than 98% was obtained.

[0283] Comparative Example

[0284] A similar method to that of Examples 1-3 was used to construct and express the fusion protein expression strain, the only difference being that the amino acid sequence of the fusion protein used for expression is shown in SEQ ID NO: 22.

[0285]

[0286] The above fusion protein contains a gIII signal peptide. The results show that the yield of inclusion bodies is 30g of inclusion bodies by wet weight. The above results show that the expression level of the fusion protein of the present invention is significantly increased compared with the expression of conventional structure fusion proteins.

[0287] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the claims appended hereto. Sequence Listing <110> Ningbo Kunpeng Biotechnology Co., Ltd. <120> A semaglutide derivative and its preparation method and application <130> P2022-2986 <150> CN202010531568.1 <151> 2020-06-11 <150> CN202010724452.X <151> 2020-07-24 <160> 28 <170> PatentIn version 3.5 <210> 1 <211> 29 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide prodrug <400> 1 Glu Gly Thr Phe Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala 1 5 10 15 Ala Lys Glu Phe Ile Ala Trp Leu Val Arg Gly Arg Gly 20 25 <210> 2 <211> 28 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide prodrug <400> 2 Gly Thr Phe Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala 1 5 10 15 Lys Glu Phe Ile Ala Trp Leu Val Arg Gly Arg Gly 20 25 <210> 3 <211> 31 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide backbone <400> 3 His Xaa Glu Gly Thr Phe Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly 1 5 10 15 Gln Ala Ala Lys Glu Phe Ile Ala Trp Leu Val Arg Gly Arg Gly 20 25 30 <210> 4 <211> 87 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide precursor fusion protein <400> 4 Met Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Lys Leu Thr Leu 1 5 10 15 Lys Phe Ile Cys Thr Thr Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys 20 25 30 Asp Thr Tyr Lys Thr Arg Ala Glu Val Lys Phe Glu Gly Asp Glu Asn 35 40 45 Leu Tyr Phe Gln Gly Asp Asp Asp Asp Lys Glu Gly Thr Phe Thr Ser 50 55 60 Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe Ile Ala 65 70 75 80 Trp Leu Val Arg Gly Arg Gly 85 <210> 5 <211> 89 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide precursor fusion protein <400> 5 Met Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Tyr Val Gln Glu 1 5 10 15 Arg Thr Ile Ser Phe Lys Asp Thr Tyr Lys Thr Arg Ala Glu Val Lys 20 25 30 Phe Glu Gly Asp Thr Leu Val Asn Arg Ile Glu Leu Lys Gly Ile Asp 35 40 45 Phe Glu Asn Leu Tyr Phe Gln Gly Asp Asp Asp Asp Lys Gly Thr Phe 50 55 60 Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe 65 70 75 80 Ile Ala Trp Leu Val Arg Gly Arg Gly 85 <210> 6 <211> 34 <212> PRT <213> Artificial Sequence <220> <223> Green fluorescent protein folding unit <400> 6 Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr Tyr Val Gln Glu Arg Thr 1 5 10 15 Ile Ser Phe Lys Asp Thr Tyr Lys Thr Arg Ala Glu Val Lys Phe Glu 20 25 30 Gly Asp <210> 7 <211> 12 <212> PRT <213> Artificial Sequence <220> <223> leader peptide <400> 7 Met Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val 1 5 10 <210> 8 <211> 7 <212> PRT <213> Artificial Sequence <220> <223> TEV enzyme cleavage site <400> 8 Glu Asn Leu Tyr Phe Gln Gly 1 5 <210> 9 <211> 5 <212> PRT <213> Artificial Sequence <220> <223> Enterokinase cleavage site <400> 9 Asp Asp Asp Asp Lys 1 5 <210> 10 <211> 37 <212> PRT <213> Artificial Sequence <220> <223> Green fluorescent protein folding unit <400> 10 Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys Asp Thr Tyr Lys Thr Arg 1 5 10 15 Ala Glu Val Lys Phe Glu Gly Asp Thr Leu Val Asn Arg Ile Glu Leu 20 25 30 Lys Gly Ile Asp Phe 35 <210> 11 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 11 Val Pro Ile Leu Val Glu Leu Asp Gly Asp Val Asn Gly 1 5 10 <210> 12 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 12 His Lys Phe Ser Val Arg Gly Glu Gly Glu Gly Asp Ala Thr 1 5 10 <210> 13 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 13 Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr 1 5 10 <210> 14 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 14 Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys Asp 1 5 10 <210> 15 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 15 Thr Tyr Lys Thr Arg Ala Glu Val Lys Phe Glu Gly Asp 1 5 10 <210> 16 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 16 Thr Leu Val Asn Arg Ile Glu Leu Lys Gly Ile Asp Phe 1 5 10 <210> 17 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 17 His Asn Val Tyr Ile Thr Ala Asp Lys Gln 1 5 10 <210> 18 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 18 Gly Ile Lys Ala Asn Phe Lys Ile Arg His Asn Val Glu Asp 1 5 10 <210> 19 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 19 Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr Pro Ile Gly 1 5 10 <210> 20 <211> 12 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> 20 His Tyr Leu Ser Thr Gln Ser Val Leu Ser Lys Asp 1 5 10 <210> twenty one <211> 13 <212> PRT <213> Artificial Sequence <220> <223> β-pleated sheet unit <400> twenty one His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly Ile 1 5 10 <210> twenty two <211> 85 <212> PRT <213> Artificial Sequence <220> <223> Fusion protein <400> twenty two Met Lys Lys Leu Leu Phe Ala Ile Pro Leu Val Val Pro Phe Tyr Ser 1 5 10 15 His Ser Thr Met Glu Leu Glu Ile Cys Ser Trp Tyr His Met Gly Ile 20 25 30 Arg Ser Phe Leu Glu Gln Lys Leu Ile Ser Glu Glu Asp Leu Asn Ser 35 40 45 Ala Val Asp Asp Asp Asp Asp Lys Glu Gly Thr Phe Thr Ser Asp Val 50 55 60 Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe Ile Ala Trp Leu 65 70 75 80 Val Arg Gly Arg Gly 85 <210> twenty three <211> 27 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide prodrug <400> twenty three Thr Phe Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys 1 5 10 15 Glu Phe Ile Ala Trp Leu Val Arg Gly Arg Gly 20 25 <210> twenty four <211> 26 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide prodrug <400> twenty four Phe Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu 1 5 10 15 Phe Ile Ala Trp Leu Val Arg Gly Arg Gly 20 25 <210> 25 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide prodrug <400> 25 Thr Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe 1 5 10 15 Ile Ala Trp Leu Val Arg Gly Arg Gly 20 25 <210> 26 <211> 88 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide precursor fusion protein <400> 26 Met Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Tyr Val Gln Glu 1 5 10 15 Arg Thr Ile Ser Phe Lys Asp Thr Tyr Lys Thr Arg Ala Glu Val Lys 20 25 30 Phe Glu Gly Asp Thr Leu Val Asn Arg Ile Glu Leu Lys Gly Ile Asp 35 40 45 Phe Glu Asn Leu Tyr Phe Gln Gly Asp Asp Asp Asp Lys Thr Phe Thr 50 55 60 Ser Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe Ile 65 70 75 80 Ala Trp Leu Val Arg Gly Arg Gly 85 <210> 27 <211> 87 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide precursor fusion protein <400> 27 Met Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Tyr Val Gln Glu 1 5 10 15 Arg Thr Ile Ser Phe Lys Asp Thr Tyr Lys Thr Arg Ala Glu Val Lys 20 25 30 Phe Glu Gly Asp Thr Leu Val Asn Arg Ile Glu Leu Lys Gly Ile Asp 35 40 45 Phe Glu Asn Leu Tyr Phe Gln Gly Asp Asp Asp Asp Lys Phe Thr Ser 50 55 60 Asp Val Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe Ile Ala 65 70 75 80 Trp Leu Val Arg Gly Arg Gly 85 <210> 28 <211> 83 <212> PRT <213> Artificial Sequence <220> <223> Semaglutide precursor fusion protein <400> 28 Met Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Lys Leu Thr Leu 1 5 10 15 Lys Phe Ile Cys Thr Thr Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys 20 25 30 Asp Thr Tyr Lys Thr Arg Ala Glu Val Lys Phe Glu Gly Asp Glu Asn 35 40 45 Leu Tyr Phe Gln Gly Asp Asp Asp Asp Lys Thr Ser Asp Val Ser Ser 50 55 60 Tyr Leu Glu Gly Gln Ala Ala Lys Glu Phe Ile Ala Trp Leu Val Arg 65 70 75 80 Gly Arg Gly

Claims

1. A semaglutide precursor fusion protein, characterized in that: The structure of the semaglutide precursor fusion protein from N-terminus to C-terminus is shown in Formula I: A-FP-TEV-EK-G(I) Where, "-" represents a peptide bond; A is none or a leader peptide sequence, FP is the green fluorescent protein folding unit; TEV is the first restriction site; EK is the second restriction enzyme cutting site; G is a semaglutide precursor, and the semaglutide precursor is: Position 18 is the first precursor of Boc-modified semaglutide, the amino acid sequence of which is shown in SEQ ID NO:

1. Alternatively, the second precursor of semaglutide with Boc modification at position 17, the amino acid sequence of the second precursor is shown in SEQ ID NO: 2, Alternatively, the third precursor of semaglutide modified by Boc at position 16, the amino acid sequence of the third precursor is shown in SEQ ID NO: 23, Alternatively, the fourth precursor of semaglutide modified with Boc at position 15, the amino acid sequence of the fourth precursor is shown in SEQ ID NO: 24, Alternatively, a fifth precursor of semaglutide modified with Boc at position 14, wherein the amino acid sequence of the fifth precursor is shown in SEQ ID NO: 25; Wherein, the green fluorescent protein folding unit is 2-3 β-folding units selected from the following group: u2-u3, u4-u5, u8-u9, u1-u2-u3, u3-u4-u5, u4-u5-u6, u8-u9-u10, u9-u10-u11 or u3-u5-u7, 2. The fusion protein according to claim 1, wherein The green fluorescent protein folding unit is u2-u3, u4-u5, u1-u2-u3, u3-u4-u5 or u4-u5-u6.

3. The fusion protein according to claim 1, wherein The TEV is a TEV enzyme cleavage site, and its sequence is shown as ENLYFQG (SEQ ID NO: 8).

4. The fusion protein according to claim 1, wherein The EK is an enterokinase cleavage site, and its sequence is shown as DDDDK (SEQ ID NO: 9).

5. The fusion protein according to claim 1, wherein The amino acid sequences of the fusion proteins are shown in SEQ ID NOs: 4, 5, 26, 27, and 28.

6. A method for preparing Fmoc and Boc modified semaglutide backbone, characterized in that: Including steps: (i) using recombinant bacteria for fermentation to prepare the semaglutide precursor fusion protein according to claim 1, (ii) performing enzymatic digestion on the semaglutide precursor fusion protein to obtain a Boc-modified semaglutide precursor, wherein the Boc-modified semaglutide precursor lacks X amino acids at the N-terminus of the semaglutide backbone; (iii) connecting an Fmoc complex to the N-terminus of the Boc-modified semaglutide precursor to prepare an Fmoc- and Boc-modified semaglutide backbone; wherein the Fmoc complex comprises X amino acids at the N-terminus of the semaglutide backbone, and the N-terminal amino acid of the Fmoc complex is modified by Fmoc; The 20th position of the semaglutide main chain is a protected lysine, and the protected lysine is Nε-(tert-butyloxycarbonyl)-lysine. In addition, the N-terminus of the semaglutide main chain is an Fmoc-modified histidine. The structural formula of the semaglutide main chain is shown in Formula 2 below:

7. A method for preparing a Boc-modified semaglutide precursor, characterized in that: Including steps: (i) using recombinant bacteria for fermentation to prepare the semaglutide precursor fusion protein according to claim 1, (ii) performing enzymatic digestion on the semaglutide precursor fusion protein to obtain a Boc-modified semaglutide precursor; The semaglutide precursor is: Position 18 is a Boc-modified first precursor of semaglutide, the amino acid sequence of which is shown in SEQ ID NO: 1; Alternatively, the second precursor of semaglutide with a Boc modification at position 17, the amino acid sequence of the second precursor being as shown in SEQ ID NO: 2; Alternatively, a third precursor of semaglutide modified with Boc at position 16, wherein the amino acid sequence of the third precursor is shown in SEQ ID NO: 23; Alternatively, a fourth precursor of semaglutide modified with Boc at position 15, wherein the amino acid sequence of the fourth precursor is shown in SEQ ID NO: 24; Alternatively, the fifth precursor of semaglutide modified with Boc at position 14, the amino acid sequence of the fifth precursor is shown in SEQ ID NO:

25.

8. A method for preparing an Fmoc-modified semaglutide backbone, characterized in that: Including steps: (i) using recombinant bacteria for fermentation to prepare the semaglutide precursor fusion protein according to claim 1, (ii) performing enzymatic digestion on the semaglutide precursor fusion protein to obtain a Boc-modified semaglutide precursor, wherein the Boc-modified semaglutide precursor lacks X amino acids at the N-terminus of the semaglutide backbone; (iii) connecting an Fmoc complex to the N-terminus of the Boc-modified semaglutide precursor to prepare an Fmoc- and Boc-modified semaglutide backbone; wherein the Fmoc complex comprises X amino acids at the N-terminus of the semaglutide backbone, and the N-terminal amino acid of the Fmoc complex is modified by Fmoc; (iv) performing a Boc removal treatment on the Fmoc- and Boc-modified semaglutide backbone to obtain an Fmoc-modified semaglutide backbone; The N-terminus of the semaglutide main chain is an Fmoc-modified histidine, and the amino acid sequence of the semaglutide main chain is shown in SEQ ID NO: 3, and its structural formula is shown in Formula 3 below:

9. An isolated polynucleotide, characterized in that The polynucleotide encodes the semaglutide precursor fusion protein according to claim 1.

10. A carrier, characterized in that The vector comprises the polynucleotide of claim 9.

11. A host cell, characterized in that The host cell contains the vector according to claim 10, or has the exogenous polynucleotide according to claim 9 integrated into the chromosome, or expresses the semaglutide precursor fusion protein according to claim 1.

12. A method for preparing semaglutide, characterized in that: The method comprises the steps of: (A) using recombinant bacteria to ferment and prepare the semaglutide precursor fusion protein according to claim 1, (B) using the semaglutide precursor fusion protein to prepare semaglutide, Wherein, said step (B) further comprises the steps of: (i) performing enzymatic digestion on the semaglutide precursor fusion protein to obtain a Boc-modified semaglutide precursor, wherein the Boc-modified semaglutide precursor lacks X amino acids at the N-terminus of the semaglutide backbone, wherein X is an integer of 2-7; (ii) connecting an Fmoc complex to the N-terminus of the Boc-modified semaglutide precursor to prepare an Fmoc- and Boc-modified semaglutide backbone, Wherein, the Fmoc complex comprises X amino acids at the N-terminus of the semaglutide backbone, and the N-terminal amino acid of the Fmoc complex is modified by Fmoc, and the Fmoc complex is Fmoc-H-Aib, Fmoc-H-Aib-E, Fmoc-H-Aib-EGTF, Fmoc-H-Aib-EGT or Fmoc-H-Aib-EG; (iii) performing a Boc removal treatment on the Fmoc- and Boc-modified semaglutide backbone, and reacting the backbone with the semaglutide side chain to obtain Fmoc-modified semaglutide; and (iv) performing a de-Fmoc treatment on the Fmoc-modified semaglutide to obtain de-Fmoc semaglutide; (v) performing side chain OtBu removal treatment on the de-Fmoc semaglutide to obtain semaglutide.

13. The method according to claim 12, wherein: In step (i), the enzyme digestion is performed using enterokinase.

14. The method according to claim 12, wherein: In step (ii), Fmoc complex, DIPEA (N,N-diisopropylethylamine) and DMF (N,N-dimethylformamide) are added to link the Fmoc complex to the N-terminus of the Boc-modified semaglutide precursor.

15. The method according to claim 14, wherein The Fmoc complex is an Fmoc complex in the form of an activated ester formed by activation with HOSu / DCC, HoBt / DIC, or TBTU / DIPEA.

16. The method according to claim 12, wherein In step (iii), the method further comprises the steps of: (a) adding Fmoc- and Boc-modified semaglutide backbones to a pre-cooled TFA solution, stirring, and performing Boc removal treatment to obtain a Boc-free product; (b) adding an organic solvent to the reaction solution of step (a) to obtain a solid Boc-free product; (c) The Boc-free product is mixed with the side chain of semaglutide to prepare Fmoc-modified semaglutide.

17. The method according to claim 16, wherein The organic solvent is a mixture of tertiary methyl ether and petroleum ether.

Citation Information

Patent Citations

  • A method for synthesizing semaglutide

    CN106749613B

  • Method for preparing semaglutide

    CN106928343A

  • Fusion protein containing fluorescent protein fragment and application of fusion protein

    CN111718417A

  • Fusion proteins of superfolder green fluorescent protein and use thereof

    CN104619726A

  • Preparation method of sermaglutide

    CN110294800A