GLP1R activating polypeptide or pharmaceutically acceptable salt thereof and application thereof

The short-sequence GLP1R activating peptide, optimized using a deep learning model, addresses the short half-life issue of existing GLP-1R agonists, achieving high-affinity activation of the GLP1R receptor and improving the efficacy and stability of treatments for metabolic disorders.

CN121494929APending Publication Date: 2026-02-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411096807.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing GLP-1R agonists suffer from short half-lives and are easily degraded, which limits their efficacy and ease of use in treating diseases such as diabetes and obesity.

Method used

A short-sequence GLP1R-activating peptide optimized by a deep learning model has been designed. It has a unique amino acid arrangement that can activate the GLP1R receptor with high affinity for the treatment of metabolic disorder-related diseases.

Benefits of technology

This peptide can effectively activate the GLP1R receptor, improving its stability and bioavailability in vivo, and can be used to treat metabolic disorders such as diabetes, obesity, dyslipidemia, and fatty liver disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004989591780000071
    Figure BDA0004989591780000071
  • Figure BDA0004989591780000112
    Figure BDA0004989591780000112
  • Figure BDA0004989591780000113
    Figure BDA0004989591780000113
Patent Text Reader

Abstract

The invention discloses a polypeptide or a pharmaceutically acceptable salt thereof. The polypeptide has an amino acid sequence as shown in SEQ ID NO: 1 or an amino acid sequence in a conservative modification form thereof. The polypeptide or the pharmaceutically acceptable salt thereof disclosed by the invention can be combined with GLP1R and is used for effectively activating the GLP1R. The GLP1R activating polypeptide is obtained through precise calculation and deep learning model optimization, has a short sequence, has unique amino acid arrangement compared with a traditional GLP-1 analogue, can effectively activate a GLP1R receptor, can also be used for treating or preventing metabolic disorder related diseases, and has a good application prospect. For example, obesity, diabetes, dyslipidemia related diseases, fatty liver diseases, metabolic syndromes, non-alcoholic fatty liver diseases and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biopharmaceutical technology, specifically relating to a GLP1R activating peptide or a pharmaceutically acceptable salt thereof and its uses, and more specifically to a GLP1R activating peptide designed based on a deep generative model. Background Technology

[0002] Glucagon-like peptide-1 receptor (GLP-1R) is a hormone primarily produced by intestinal L cells and belongs to the incretin class. Incretins induce approximately 50-70% of total insulin secretion, and their stimulation of insulin secretion is glucose-dependent. Currently identified incretins in the human body are mainly glucose-dependent insulinotropic peptide (GIP) and glucagon-like peptide-1 (GLP-1). As of 2021, all clinically used incretin drugs are based on GLP-1. For example, semaglutide can significantly reduce glycated hemoglobin (HbA1c); liraglutide injection can be used in adult patients with poorly controlled type 2 diabetes, in combination with other oral hypoglycemic agents on top of diet and exercise to improve glycemic control.

[0003] However, the aforementioned GLP-1R agonists that mimic endogenous GLP-1 may also suffer from problems such as short half-life and easy degradation. Therefore, developing a new GLP-1R agonist is of great importance in the treatment of diabetes, obesity, and other related diseases. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems existing in the prior art. To this end, this application provides a GLP1R activating peptide or a pharmaceutically acceptable salt thereof.

[0005] In a first aspect, this application provides a polypeptide or a pharmaceutically acceptable salt thereof. According to embodiments of this application, the polypeptide has an amino acid sequence as shown in SEQ ID NO:1 or an amino acid sequence of a conserved modified form thereof. The polypeptide or a pharmaceutically acceptable salt thereof can bind to GLP1R to effectively activate GLP1R and can be used to treat or prevent metabolic disorder-related diseases.

[0006] In a second aspect of this application, a polypeptide derivative or a pharmaceutically acceptable salt thereof is provided. According to embodiments of this application, the polypeptide derivative or a pharmaceutically acceptable salt thereof comprises: the polypeptide described in the first aspect, and a modifying group, wherein the polypeptide and the modifying group are linked. Using a polypeptide derivative containing the aforementioned polypeptide or a pharmaceutically acceptable salt thereof can bind to GLP1R for effective activation of GLP1R, and can be used to treat or prevent metabolic disorder-related diseases.

[0007] In a third aspect, this application proposes a fusion protein. According to embodiments of this application, the fusion protein comprises the polypeptide described in the first aspect. The polypeptide in the fusion protein of this application can bind to GLP1R for detecting or activating GLP1R, and can also be used to treat or prevent metabolic disorder-related diseases.

[0008] In a fourth aspect, this application provides a reagent or kit. According to embodiments of this application, the reagent or kit comprises the polypeptide described in the first aspect or the fusion protein described in the second aspect. The polypeptide in the reagent or kit of this application can bind to GLP1R for the detection of GLP1R.

[0009] In a fifth aspect, this application provides a pharmaceutical composition. According to embodiments of this application, the pharmaceutical composition comprises the polypeptide described in the first aspect or a pharmaceutically acceptable salt thereof, the polypeptide derivative described in the second aspect or a pharmaceutically acceptable salt thereof, or the fusion protein described in the third aspect. The polypeptide in the pharmaceutical composition of this application can bind to GLP1R and can be used to effectively activate GLP1R, and can be used to treat or prevent metabolic disorder-related diseases.

[0010] In a sixth aspect of this application, the use of the polypeptide described in the first aspect or a pharmaceutically acceptable salt thereof, the polypeptide derivative described in the second aspect or a pharmaceutically acceptable salt thereof, the fusion protein described in the third aspect, or the pharmaceutical composition described in the fifth aspect in the preparation of a medicament for the treatment or prevention of metabolic disorder-related diseases is provided.

[0011] In a seventh aspect of this application, a method for detecting GLP-1R is provided. According to embodiments of this application, the method includes: contacting a sample to be tested with the polypeptide described in the first aspect, the fusion protein described in the third aspect, and the reagent described in the fourth aspect to form an immune complex; and determining whether the sample to be tested contains GLP-1R based on the signal from the immune complex. The method of this application detects GLP-1R by binding the polypeptide to it.

[0012] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0013] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0014] Figure 1 This is a schematic diagram of the peptide design and screening process in this application;

[0015] Figure 2 The purity results of peptide GLP1R-TXP were determined by high performance liquid chromatography (HPLC) in Example 2 of this application.

[0016] Figure 3 This is the mass spectrometry report result of detecting peptide GLP1R-TXP using mass spectrometry analysis (MS) in Example 2 of this application;

[0017] Figure 4 This is the detection result of cAMP release from the GLP1R target activated by peptide GLP1R-TXP in Example 2 of this application;

[0018] Figure 5 This is a three-dimensional structural diagram of the peptide GLP1R-TXP binding to the GLP1R target in Example 2 of this application. Detailed Implementation

[0019] The embodiments of this application are described in detail below. The embodiments described below are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0020] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more.

[0021] In this document, the terms “comprising” or “including” are open-ended expressions, meaning they include the contents specified in this invention but do not exclude other aspects.

[0022] In this document, the terms “optionally,” “optionally,” or “optionally” generally refer to an event or condition that may, but may not, occur, and the description includes both cases in which the event or condition occurs and cases in which the event or condition does not occur.

[0023] In this article, the term "target protein" refers to a protein that plays a crucial role in an organism and is often the target of drug development. By binding to a target protein, drugs can regulate its biological activity, thereby achieving the goal of treating diseases.

[0024] In this article, the term "peptide" refers to a biological macromolecule composed of many amino acids, typically 50 or fewer amino acids linked together by peptide bonds. Peptides perform a variety of functions in living organisms, including participating in various physiological processes as enzymes, hormones, and antibodies. Furthermore, due to their excellent biocompatibility, selectivity, and high biological activity, peptides are widely studied for drug discovery and the treatment of various diseases, such as anti-tumor, antiviral, antibacterial, and anti-inflammatory drugs.

[0025] In this article, the term "glucagon-like peptide-1 receptor (GLP1R)" refers to a G protein-coupled receptor belonging to the class B G protein-coupled receptor family. It plays a crucial physiological role in the human body, particularly in regulating blood glucose levels. Activation of GLP1R promotes insulin secretion and inhibits glucagon release, thereby lowering blood glucose. Furthermore, GLP1R is involved in regulating various physiological processes, including gastrointestinal motility, appetite control, and energy balance. Activation of GLP1R leads to the release of the α subunit of the G protein and activation of adenylate cyclase (AC), a membrane-bound enzyme that catalyzes the conversion of ATP (adenosine triphosphate) to cAMP.

[0026] In this article, the term "cyclic adenosine monophosphate (cAMP)" refers to a small molecule ubiquitous in living organisms, derived from ATP (adenosine triphosphate) through the catalytic action of adenylate cyclase. As a key mediator of intracellular signal transduction, cAMP plays a crucial role in regulating various cellular functions and physiological processes.

[0027] In this document, amino acids are referred to by their conventional single-letter and three-letter codes for natural amino acids, as well as the generally accepted three-letter codes for other α-amino acids. Unless otherwise specified, in this application, uppercase letters represent amino acids with the L-configuration and lowercase letters represent amino acids with the D-configuration.

[0028] In this paper, the structural formula for the term "αMeK" is:

[0029] In this article, the structural formula for the term "HoK" is:

[0030] In this article, the structural formula for the term "N-Me-K" is:

[0031] In this article, the structural formula for the term "Orn" is:

[0032] In this article, the structural formula for the term "Dab" is:

[0033] In this article, the structural formula for the term "Dap" is:

[0034] In this document, "conservatively modified amino acid sequences" refers to amino acid modifications that do not significantly affect or alter the binding properties of antibodies containing that amino acid sequence. These modifications include amino acid substitutions, additions, and deletions. Conservative amino acid substitutions involve replacing an amino acid residue with an amino acid residue having a similar side chain. Families of amino acid residues with similar side chains have been identified in the art.

[0035] In this document, the term "modifying group" should be interpreted broadly, and can refer to chemical groups or amino acid fragments. The specific type is not limited, and all are within the scope of protection of this application.

[0036] In this document, the term "pharmaceutical acceptable" means that a substance or composition must be chemically and / or toxicologically compatible with other components comprising a polypeptide or its derivatives and / or with the mammals to which it is treated. Preferably, "pharmaceutical acceptable" as used herein means approved by a federal regulatory agency or national government, or listed in the United States Pharmacopeia or other generally recognized pharmacopoeia for use in animals, particularly in humans.

[0037] In this document, the term "pharmaceutically acceptable salt" refers to the organic and inorganic salts of the polypeptides or their derivatives of the present invention. Pharmaceutically acceptable salts are well-known in the field, as described in the literature: SMBerge et al., describe pharmaceutically acceptable salts in detail in J. Pharmaceutical Sciences, 1977, 66: 1-19. Pharmaceutically acceptable salts formed from non-toxic acids include, but are not limited to, inorganic acid salts (such as hydrochlorides, hydrobromic acids, phosphates, sulfates, and perchlorates) formed by reaction with amino groups, and organic acid salts (such as acetates, oxalates, maleates, tartrates, citrates, succinates, and malonates), or salts obtained by other methods described in the literature, such as ion exchange.

[0038] In this document, "pharmaceutical composition" can refer to a drug for the treatment of a disease or for use in in vitro cell culture experiments. When used for the treatment of a disease, the term "pharmaceutical composition" generally refers to a unit dose form and can be prepared by any method well known in the pharmaceutical industry. All methods involve the step of combining the active ingredient with excipients that constitute one or more adjunct components. Typically, compositions are prepared by uniformly and adequately combining an active polypeptide or its derivative or revitalizer with a liquid excipient, a finely pulverized solid excipient, or both.

[0039] In this document, the term "pharmaceuticalally acceptable excipient" may include any solvent, solid excipient, diluent, or other liquid excipient, etc., suitable for a particular target dosage form. The use of any conventional excipients is also within the scope of this invention, except for any range of incompatibilities with the polypeptides or derivatives thereof, pharmaceutical compositions, or drugs containing them, such as any adverse biological effects or harmful interactions with any other component of a pharmaceutically acceptable composition.

[0040] In addition to any conventional excipients, the use of polypeptides or their derivatives, pharmaceutical compositions or pharmaceuticals containing them that are incompatible with the present invention, such as any adverse biological effects or harmful interactions with any other component of a pharmaceutically acceptable composition, is also within the scope of this invention.

[0041] The pharmaceutical compositions disclosed herein include formulations suitable for parenteral administration. The formulations can be conveniently available in unit dosage forms and can be prepared by any method known in the pharmaceutical field. The amount of active ingredient in a single-dose form, which can be prepared in combination with excipients, is generally the amount of the polypeptide or a derivative thereof that produces the therapeutic effect.

[0042] In this paper, the term "agonist" refers to a substance (ligand) that activates the type of receptor.

[0043] In this document, the term "treatment" means used to refer to achieving a desired pharmacological and / or physiological effect. This effect may be preventative in terms of complete or partial prevention of disease or its symptoms, and / or therapeutic in terms of partial or complete cure of disease and / or adverse effects caused by disease. As used herein, "treatment" covers diseases in mammals, particularly humans, including: (a) prevention of disease or the onset of disease in individuals susceptible to disease but not yet diagnosed with the disease; (b) inhibition of disease, such as blocking disease progression; or (c) relief of disease, such as reducing disease-related symptoms. As used herein, "treatment" encompasses any administration of a drug or compound to an individual to treat, cure, relieve, improve, reduce, or inhibit the individual's disease, including but not limited to administration of a drug containing a compound described herein to an individual in need.

[0044] In this article, the term “non-alcoholic fatty liver disease (NAFLD)” generally refers to a clinicopathological syndrome characterized by excessive fat deposition in hepatocytes excluding alcohol and other clearly defined liver-damaging factors. It is an acquired metabolic stress-induced liver injury closely associated with insulin resistance and genetic susceptibility, including but not limited to simple fatty liver (SFL), non-alcoholic steatohepatitis (NASH), and its associated cirrhosis.

[0045] This application discloses a GLP1R activating peptide or a pharmaceutically acceptable salt thereof and its uses, which will be described in detail below.

[0046] polypeptides or their pharmaceutically acceptable salts

[0047] In one aspect of this application, a polypeptide or a pharmaceutically acceptable salt thereof is provided. According to embodiments of this application, the polypeptide has an amino acid sequence as shown in SEQ ID NO:1 or an amino acid sequence of a conserved modified form thereof.

[0048] EGRLTFDSVMDMDWLAGG (SEQ ID NO: 1).

[0049] Currently, most existing GLP1R agonists are GLP-1 analog peptides. Although they have some therapeutic effect in regulating blood glucose levels, they still have some significant limitations: 1. Sequence length: Existing GLP-1 analog peptide sequences are relatively long, which may lead to complex synthesis processes, higher costs, and potentially limited stability and bioavailability in vivo. 2. Stability issues: Longer peptide sequences may have shorter half-lives in vivo and are easily degraded by enzymes, limiting the duration of their therapeutic effect and ease of application. 3. The fact that most existing agonists are GLP-1 analogs limits the exploration of peptides with stronger binding affinity and greater stability.

[0050] However, the GLP1R activating peptide of this application was obtained through precise computation and deep learning model optimization. It can bind to GLP1R and has a unique amino acid sequence compared with traditional GLP-1 analogs, which can effectively activate the GLP1R receptor. Furthermore, the peptide is a short sequence, which was obtained through precise computation and deep learning model optimization. It has high affinity and biological activity for the GLP1R receptor and can effectively activate the GLP1R receptor for the treatment of metabolic disorder-related diseases (such as diabetes).

[0051] According to embodiments of this application, the above-mentioned polypeptide or its pharmaceutically acceptable salt may further include at least one of the following technical features:

[0052] According to embodiments of this application, the conservatively modified amino acid is selected from amino acid X whose side chain contains -NH2.

[0053] According to embodiments of this application, the amino acid X is selected from K, k, αMeK, HoK, Dap, Dab, Orn, or N-Me-K.

[0054] In an optional embodiment of this application, the conservative modification includes substituting or adding amino acid sequences as shown in SEQ ID NO:1, such that amino acid X is present in the amino acid sequence shown in SEQ ID NO:1. In an optional embodiment of this application, amino acid X may be coupled to subsequent modifying groups via the -NH2 group on its side chain.

[0055] According to embodiments of this application, the number of amino acids in the conservative modification is one or more.

[0056] In this paper, the term "number of amino acids in a conservative modification" refers to the number of amino acids that are substituted or added.

[0057] polypeptide derivatives or their pharmaceutically acceptable salts

[0058] In a second aspect of this application, a polypeptide derivative or a pharmaceutically acceptable salt thereof is provided. According to embodiments of this application, the polypeptide derivative or a pharmaceutically acceptable salt thereof comprises: the polypeptide described in the first aspect, and a modifying group, wherein the polypeptide and the modifying group are linked. Using a polypeptide derivative containing the aforementioned polypeptide or a pharmaceutically acceptable salt thereof can effectively activate GLP1R for the treatment or prevention of metabolic disorder-related diseases, such as obesity, diabetes, dyslipidemia-related diseases, fatty liver disease, metabolic syndrome, and non-alcoholic fatty liver disease.

[0059] According to embodiments of this application, the above-mentioned polypeptide derivatives or pharmaceutically acceptable salts thereof may further include at least one of the following technical features:

[0060] According to embodiments of this application, the polypeptide is an amino acid sequence in a conserved modified form as shown in SEQ ID NO:1.

[0061] According to an embodiment of this application, the modifying group is linked to the amino acid X-NH2.

[0062] According to embodiments of this application, the modifying group has at least one of the following structures:

[0063]

[0064] In this article, the "" in the description of chemical groups "" is used to describe the position of a group substitution. That is, the above chemical group is substituted by... It is linked to the amino acid X-NH-.

[0065] Fusion protein

[0066] In a third aspect, this application proposes a fusion protein. According to embodiments of this application, the fusion protein comprises the polypeptide described in the first aspect. The polypeptide in the fusion protein containing the aforementioned polypeptide can bind to GLP1R, effectively activating GLP1R, and can also be used to treat or prevent metabolic disorder-related diseases, such as obesity, diabetes, dyslipidemia-related diseases, fatty liver disease, metabolic syndrome, and non-alcoholic fatty liver disease.

[0067] In an optional embodiment of this application, the fusion protein further includes a functional fragment. The functionally active fragment in this application can be used to exert its effects in vivo or in vitro. Exemplarily, when the functionally active fragment is used to exert its effects in vivo, it can be used to prevent and / or treat diseases; when the functionally active fragment is used to exert its effects in vitro, it can be used to specifically bind to a substance, detect that substance, or diagnose diseases in vitro.

[0068] In an optional embodiment of this application, the functional fragment refers to an amino acid fragment, which may be a functionally active fragment or a protein tag. The specific type is not limited and is within the protection scope of this application.

[0069] In an optional embodiment of this application, the protein tag refers to a short peptide co-expressed with the target protein, which facilitates the expression, detection, tracing, or purification of the polypeptide of this application. Exemplarily, the protein tag includes at least one of the following: His tag, Flag tag, GST tag, MBP tag, SUMO tag, and C-Myc tag.

[0070] reagents or kits

[0071] In a fourth aspect, this application provides a reagent or kit. According to embodiments of this application, the reagent or kit comprises the polypeptide described in the first aspect or the fusion protein described in the second aspect. The polypeptide in the reagent or kit of this application can bind to GLP1R and can be used to detect GLP1R.

[0072] Pharmaceutical Composition

[0073] In a fifth aspect of this application, a pharmaceutical composition is provided. According to embodiments of this application, the pharmaceutical composition comprises the polypeptide described in the first aspect or a pharmaceutically acceptable salt thereof, the polypeptide derivative described in the second aspect or a pharmaceutically acceptable salt thereof, or the fusion protein described in the third aspect. The polypeptide in the pharmaceutical composition can bind to GLP1R and can be used to effectively activate GLP1R, thereby treating or preventing metabolic disorder-related diseases such as obesity, diabetes, dyslipidemia-related diseases, fatty liver disease, metabolic syndrome, and non-alcoholic fatty liver disease.

[0074] According to embodiments of this application, the pharmaceutical composition further comprises pharmaceutically acceptable excipients.

[0075] In one alternative embodiment of this application, pharmaceutically acceptable excipients refer to pharmaceutical excipients that are conventional in the pharmaceutical field, such as diluents, buffer solutions, osmotic pressure regulators, pH regulators, etc.

[0076] In one alternative embodiment of this application, a pharmaceutically acceptable carrier refers to a drug carrier conventional in the pharmaceutical field, such as a protectant.

[0077] In one alternative embodiment of this application, pharmaceutically acceptable mediators refer to pharmaceutical mediators conventional in the pharmaceutical field, such as solutions (e.g., water) and liposomes.

[0078] In one alternative embodiment of this application, examples of suitable pharmaceutically acceptable carriers, excipients, and mediators are well known in the art. Pharmaceutical compositions comprising such carriers, excipients, and mediators can be formulated using known conventional methods.

[0079] In some alternative embodiments of this application, the pharmaceutical composition of this application may also contain other active ingredients for treatment.

[0080] The pharmaceutical composition of this application can be administered via various routes, such as enterally, orally (e.g., liquid solution), or by injection (e.g., intravenously, subcutaneously, intramuscularly, intraperitoneally, intradermally). Preferably, the pharmaceutical composition of this application is in the form of a lyophilized preparation or an aqueous solution. The clinical dosing regimen will be determined by the attending physician and clinical factors. As is known in the medical field, the dosage for any given patient depends on many factors, including the patient's physique, body surface area, age, the drug to be administered, sex, time and route of administration, general health, and other concurrently administered drugs. The pharmaceutical composition of this application can be administered topically or systemically. Preferably, it can be administered intravenously or subcutaneously.

[0081] use

[0082] In a sixth aspect of this application, the use of the polypeptide described in the first aspect or a pharmaceutically acceptable salt thereof, the polypeptide derivative described in the second aspect or a pharmaceutically acceptable salt thereof, the fusion protein described in the third aspect, or the pharmaceutical composition described in the fifth aspect in the preparation of a medicament for the treatment or prevention of metabolic disorder-related diseases is provided.

[0083] According to embodiments of this application, the metabolic disorder-related diseases include at least one of obesity, diabetes, dyslipidemia-related diseases, fatty liver disease, metabolic syndrome, and non-alcoholic fatty liver disease.

[0084] method

[0085] In a seventh aspect of this application, a method for detecting GLP-1R is provided. According to embodiments of this application, the method includes: contacting a test sample with the polypeptide described in the first aspect, the fusion protein described in the third aspect, and the reagent described in the fourth aspect to form an immune complex; and determining whether the test sample contains GLP-1R based on the signal from the immune complex. The polypeptide used in the above method can bind to GLP-1R for detection.

[0086] In an eighth aspect of this application, a method for treating or preventing metabolic disorder-related diseases is provided. According to embodiments of this application, the method includes administering to a subject a pharmaceutically acceptable dose of the polypeptide described in the first aspect or a pharmaceutically acceptable salt thereof, the polypeptide derivative described in the second aspect or a pharmaceutically acceptable salt thereof, the fusion protein described in the third aspect, or the pharmaceutical composition described in the fifth aspect.

[0087] In one alternative embodiment of this application, the pharmaceutically acceptable dose may be selected from the effective dose (or effective amount).

[0088] The effective amount of the compound described in this application may vary depending on the administration method and the severity of the disease to be treated. A preferred effective amount can be determined by those skilled in the art based on various factors (e.g., through clinical trials). These factors include, but are not limited to: pharmacokinetic parameters of the active ingredient, such as bioavailability, metabolism, and half-life; the severity of the disease to be treated, the patient's weight, the patient's immune status, and the route of administration. For example, due to the urgency of the treatment condition, several separate doses may be administered daily, or the dose may be reduced proportionally.

[0089] The polypeptides, pharmaceutically acceptable salts thereof, polypeptide derivatives thereof, or pharmaceutical compositions of this application may be incorporated into suitable pharmaceuticals, which may be prepared in various forms, such as liquid, semi-solid, and solid dosage forms, including but not limited to solid dosage forms, semi-solid dosage forms, liquid dosage forms, and gaseous dosage forms. Various routes of administration of the polypeptides, pharmaceutically acceptable salts thereof, polypeptide derivatives thereof, pharmaceutical compositions, or pharmaceutical compositions of this application are contemplated, including peritoneal, intravenous, intramuscular, subcutaneous, dermal, oral, topical, nasal, pulmonary, rectal, and topical administration; however, this application is not limited to these exemplified routes of administration.

[0090] According to embodiments of this application, the metabolic disorder-related diseases include at least one of obesity, diabetes, dyslipidemia-related diseases, fatty liver disease, metabolic syndrome, and non-alcoholic fatty liver disease.

[0091] The following will explain the solution of this application with reference to embodiments. Those skilled in the art will understand that the following embodiments are for illustrative purposes only and should not be considered as limiting the scope of this application. Where specific techniques or conditions are not specified in the embodiments, they are performed according to the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be obtained commercially.

[0092] Example 1:

[0093] The schematic diagram of the peptide design and screening process in this application is shown below. Figure 1 It covers a series of steps from GLP1R sequence generation to screening and optimization.

[0094] I. The production and screening of GLP1R sequences in this embodiment refer to the patent application number 202410192333.2, application date 2024-02-20, entitled "Training of Ligand Information Generation Model, Method and Apparatus for Generating Ligand Information". This patent discloses an innovative deep learning method, TPDiffusion, which can generate peptide sequences that can bind to the amino acid sequence of a target protein. This method of converting the target protein sequence into a peptide sequence can be seen as a process of providing a specific answer to a specific problem. By training the conditional diffusion model TPDiffusion, the relationship rules between the target protein and the peptide sequence are learned, realizing the generation of specific peptides (i.e., GLP1R agonist peptides) for a specific target protein sequence (i.e., GLP1R). The training of the peptide sequence generation model TPDiffusion mainly includes forward diffusion and backward diffusion processes.

[0095] 1. The forward diffusion process includes the following steps:

[0096] 1-1 Encoding the Amino Acid Sequence: To model the joint feature space of the target protein-peptide pair, the target protein and peptide sequences are concatenated into a single unit. Simultaneously, an embedding transformation function (EMB(w)) is introduced to map each discrete amino acid word or character w to a continuous vector encoding space. Specifically, for any target protein-peptide pair, the pair includes target information x and peptide sequence y. Assuming the target information x includes m amino acid texts and the peptide sequence y includes n amino acid texts, after mapping the target protein pair using the embedding transformation function EMB(w), the sample receptor information (i.e., the target protein information) EMB(x1), ..., EMB(x2) can be obtained. m ) and sample ligand information (i.e., peptide sequence information) EMB(y1),...,EMB(y n This mapping process can be characterized as follows:

[0097] 1-2 Stepwise noise addition to the peptide sequence: Gaussian noise is progressively added to the peptide portion of the vector encoding based on a Markov chain (i.e., multiple noise additions are performed) until the peptide sequence is completely destroyed. This includes the following steps:

[0098] 1) For the first noise addition, the sample ligand information is noise-added for the first time to obtain the noise addition result of the first noise addition.

[0099] In this embodiment, noise data with initial noise addition can be randomly generated, and the noise data with initial noise addition can be fused with the sample ligand information to achieve the addition of noise data with initial noise addition to the sample ligand information, thereby obtaining the noise addition result with initial noise addition.

[0100] 2) For non-first noise addition, the noise addition result of the previous noise addition is subjected to non-first noise addition to obtain the noise addition result of non-first noise addition. The noise addition result of the last noise addition is used as the reference noise data.

[0101] In this embodiment, a second round of noise data can be randomly generated. The noise data from the second round is then fused with the noise result from the first round, adding the second round of noise data to the first round to obtain a second noise-added result. Next, a third round of noise data can be randomly generated. This third round is then fused with the noise result from the second round, adding the third round of noise data to the second round to obtain a third noise-added result. Following this noise-adding principle, the sample ligand information is subjected to multiple rounds of noise addition sequentially to obtain a final noise-added result, which serves as the reference noise data. Optionally, this noise-adding process can be expressed as the following formula:

[0102]

[0103] Among them, l t The data obtained after adding noise, i.e., the noise result of the first noise addition or the noise result of subsequent noise additions mentioned above, can be described as the noise result of the t-th noise addition. t-1 The data before noise addition, i.e., the sample ligand information mentioned above or the noise addition result of the previous noise addition (not the first noise addition), can be described as the sample ligand information or the noise addition result of the (t-1)th noise addition. q(l t |l t-1 Characterizing the l t-1 The l obtained by adding noise t The distribution that is satisfied. Symbols representing a Gaussian distribution. β represents the mean of a Gaussian distribution. t I represents the variance of the Gaussian distribution. Where β t I and I are two hyperparameters.

[0104] 2. After the forward diffusion process is completed, a clean set of target protein and noise polypeptide sequences is obtained. Then, the polypeptide sequence is restored through the reverse diffusion process, which mainly includes the following steps:

[0105] 2-1 Stepwise Denoising of the Peptide Sequence: A denoising network was constructed to estimate the distribution of noise data added during the forward diffusion process, and the peptide portion was progressively denoised until the peptide sequence was completely recovered. The denoising network used the BERT (Devlin et al. 2018) model. The BERT model architecture mainly consists of a multi-layer Transformer (Vaswanie et al. 2017) encoder, with each Transformer block containing two parts: self-attention and a feed-forward neural network. Through self-attention processing, target sequence information is introduced during the reverse diffusion process to recover the peptide sequence, thus implicitly modeling the relationship between the target protein and the peptide sequence, achieving a mapping from the target protein to the peptide sequence. During the generation process, given an arbitrary target protein sequence, the model first randomly samples from Gaussian noise and then performs a reverse diffusion process. Guided by the target protein sequence, the noise is progressively eliminated over fixed time steps, ultimately generating a binding peptide sequence for the given target. Specifically, the following steps are included:

[0106] 2-1-1 For the first denoising network, the reference noise data is denoised based on the sample receptor information to obtain the denoising result of the first denoising network. The first denoising network includes an attention network and a feedforward network.

[0107] 1) Attention processing is performed on the sample receptor information and reference noise data using an attention network to obtain the attention processing result. Details are as follows:

[0108] For any first sub-data point, an attention network performs a linear or non-linear mapping on it to obtain the corresponding query (Queue, Q) vector. For any content information, the attention network performs two different linear or non-linear mappings on it to obtain the corresponding key (Key, K) vector and value (Value, V) vector. Multiple content information points include multiple receptor composition information points and multiple first sub-data points. That is, one receptor composition information point is one content information point, and one first sub-data point is also one content information point.

[0109] Next, based on the query vector corresponding to the first sub-data and the key vector corresponding to the content information, the attention weight is determined. The attention weight is then multiplied by the value vector corresponding to the content information to obtain the processing result for that content information. Optionally, the processing result for the content information can be determined according to the formula shown below.

[0110]

[0111] Here, A represents the processing result of the content information. Attention represents the attention function used for attention processing, which includes three parameters: Q, K, and V. Q represents the query vector corresponding to the first sub-data, K represents the key vector corresponding to the content information, and V represents the value vector corresponding to the content information. Softmax is a normalized exponential function, and T represents the transpose matrix. K The dimensions representing the query vector and key vector. Among them, Characterize attention weights.

[0112] It is understandable that there are multiple pieces of content information. Any first sub-data can be processed as described above with each piece of content information to obtain the processing results of each piece of content information. Then, the processing results of each piece of content information are summed, averaged, weighted, and calculated to obtain the second sub-data corresponding to the first sub-data.

[0113] The aforementioned reference noise data includes first sub-data X to first sub-data M, and the sample receptor information includes receptor composition information P to receptor composition information F. The first sub-data X is mapped using an attention network to obtain the query vector Q of the first sub-data X. xFor each piece of content information i (i takes any value from X to M and from P to F) in the first sub-data X to the first sub-data M and the receptor composition information P to the receptor composition information F, the content information i is mapped through an attention network to obtain the key vector K of the content information i. i Sum vector V i The query vector Q based on the first sub-data X x The key vector K of content information i i Determine the attention weight a xi And the attention weight a xi and the value vector V of content information i i Multiply the results to obtain the processed content information i. Average the processed results of each content information to obtain the second sub-data K corresponding to the first sub-data X.

[0114] 2) The attention processing results are mapped through a feedforward network to obtain the mapping results; the mapping results or set features are mapped through a feedforward network to obtain the denoising results of the first denoising network.

[0115] The feedforward network consists of a first mapping layer, a selection layer, and a second mapping layer. The first mapping layer can be either a linear or non-linear layer. It performs a linear or non-linear mapping on the attention processing result to obtain the mapping result. The selection layer selects the maximum or minimum value from the mapping result and a set feature. This maximum or minimum value is also either a mapping result or a set feature. The second mapping layer can also be either a linear or non-linear layer. It performs a linear or non-linear mapping on the maximum or minimum value to obtain the processing result of the feedforward network. Finally, the denoising result of the first denoising network is determined based on the processing result of the feedforward network.

[0116] The set feature is a predefined feature, which can be a number or a matrix, etc. For example, the processing result of the feedforward network is determined according to the formula shown below, where the set feature is the number 0.

[0117] FFN(A) = max(0, AW1 + b1)W2 + b2

[0118] FFN(A) represents the processing result of the feedforward network. `max` represents the maximum value function used by the selection layer. `AW1+b1` represents the mapping result obtained by the first mapping layer, where A represents the attention processing result, W1 represents the weight parameters in the mapping function used by the first mapping layer, and b1 represents the bias parameters in the mapping function used by the first mapping layer. `W2` represents the weight parameters in the mapping function used by the second mapping layer, and b2 represents the bias parameters in the mapping function used by the second mapping layer.

[0119] 2-1-2 For non-first denoising networks, the denoising result of the previous denoising network is denoised based on the sample receptor information, and the denoising result of the non-first denoising network is obtained.

[0120] The denoising result of the first denoising network can be input into the second denoising network. Following the denoising principle of the first network, the second network denoises the result based on the sample receptor information, resulting in the denoising result of the second network. The denoising process is not detailed here. Next, the denoising result of the second network can be input into the third denoising network. Following the same denoising principle, the third network denoises the result based on the sample receptor information, resulting in the denoising result of the third network. This process is also not detailed here. This process is repeated to obtain the denoising result of the final denoising network.

[0121] Combining steps 2-1-1 to 2-1-2, it can be seen that by constructing multiple cascaded denoising networks, the reference noise data is progressively denoised until the denoising result of the last denoising network is obtained. This denoising can be expressed by the following formula:

[0122]

[0123] For a specific denoising network, l t The input to this denoising network is represented by the reference noise data mentioned above or the denoising result of the previous denoising network, and t represents the number of denoising iterations for the denoising network. t-1 This represents the output of the denoising network, i.e., the denoising result of the denoising network mentioned above. θ (l t-1 |l t The characterization is performed on the input l through a denoising network. t The output l obtained after denoising t-1 The distribution satisfied by θ, where θ represents the network parameters of the denoising network. The symbol μ represents a Gaussian distribution. θ (l t ,t) represents the mean of the Gaussian distribution, ∑ θ (l t ,t) represents the variance of the Gaussian distribution.

[0124] 2-1-3 The predicted ligand information is determined based on the denoising result of the last denoising network. Specifically:

[0125] For any given predictive sub-feature, based on the feature similarity between any given predictive sub-feature and multiple candidate ligand composition information, multiple target ligand composition information are selected from the multiple candidate ligand composition information; sampling is performed from the multiple target ligand composition information to obtain the ligand composition information corresponding to any given predictive sub-feature; based on the ligand composition information corresponding to each predictive sub-feature, the predicted ligand information is determined.

[0126] 2-2 Calculate the loss function: The probability distribution of the model output predicted ligand information is used. The mean squared error loss function is used to calculate the difference between the model prediction result (i.e., predicted ligand information) and the true result (i.e. sample ligand information). Backpropagation is then performed to update the model parameters to obtain the trained TPDiffusion.

[0127] 3. Using the trained TPDiffusion, the sequence of the GLP1R target was used as input to generate a batch of candidate peptide sequences that may have high affinity for the GLP1R target. These peptide sequences were further screened by dry experiments to identify peptides with high affinity for the GLP1R target.

[0128] II. Screening the candidate peptide sequences obtained above. Refer to patent application number 202410193088.7, application date 2024-02-20, entitled "Training Method, Device, Equipment and Storage Medium for Molecular Binding Prediction Model". This patent discloses a high-precision deep learning model, TPBinder, for predicting the binding probability of target protein-peptide sequences. It can perform high-throughput screening of candidate peptide sequences, retaining peptide sequences with a binding probability greater than 0.5, significantly improving screening efficiency and accuracy. Specifically, it includes the following steps:

[0129] 1. The pre-trained protein characterization model ESM-2 and peptide characterization model ESM-pep were used to extract sequence characterizations of the target sequence and peptide sequence, respectively.

[0130] 2. The interaction between target protein and peptide sequence features is extracted through cross-attention, where one sequence feature serves as the query input and the other as both the key and value input. Specifically, if the features of the target protein and peptide are represented as x and y, respectively, the formula for the attention weight matrix between them is as follows:

[0131]

[0132] Among them, Q x The query matrix (i.e., Q) represents the feature-corresponding feature of the target protein. x =W Q1 X), K y The bond matrix (i.e., K) represents the characteristic of the polypeptide.y =W K1 Y), V y The value matrix (i.e., V) representing the characteristics of a polypeptide. y =W V1 Y), W Q1 W K1 W V1 is an adjustable parameter matrix; softmax is the normalization function; and d is the dimension of the bond matrix.

[0133] 3. After calculating the interaction features between the target protein and peptide, these features are concatenated and fused with the initial features extracted from the pre-trained model. The fused features are then input into a multi-layer dense connective block for final prediction. Each dense connective block contains a linear connective layer, an inactivation layer, and a non-linear ReLU activation layer. The final model can predict the binding probability of the target protein and peptide with high throughput, retaining peptide sequences with a binding probability greater than 0.5 for subsequent evaluation.

[0134] The screened peptide sequences were further analyzed using AlphaFold for structural prediction to assess their three-dimensional structural patterns and stability when bound to GLP1R. Furthermore, molecular docking techniques were employed to analyze the binding energy between the peptides and the target protein, thereby evaluating their potential affinity.

[0135] Based on a comprehensive evaluation using structural prediction and molecular docking, the peptide sequence EGRLTFDSVMDMDWLAGG with the best performance was ultimately selected.

[0136] (Glu-Gly-Arg-Leu-Thr-Phe-Asp-Ser-Val-Met-Asp-Met-Asp-Trp-Leu-Ala-Gly-Gly). This peptide sequence not only exhibits good binding probability and energy score, but also shows high complementarity with the GLP1R target in the simulated binding structure.

[0137] Example 2:

[0138] This embodiment aims to evaluate the cAMP activation capacity of the candidate peptide GLP1R-TXP (amino acid sequence EGRLTDSVMDMDWLAGG) on huGLP1R-FL-CRE-HEK293-A5 cells. By quantitatively measuring cAMP production levels, this experiment can accurately assess the activation effect of the candidate sample on the GLP1R receptor, providing important bioactivity data for drug development for metabolic diseases such as diabetes. The experimental procedure is as follows:

[0139] Peptide Synthesis and Purification: Based on a designed specific amino acid sequence, the peptide GLP1R-TXP was synthesized using solid-phase peptide synthesis (Fmoc-SPPS) technology. Purification was performed using high-performance liquid chromatography (HPLC), and the purity results are as follows: Figure 2 As shown, the synthesized peptide has high purity and no obvious impurities. Further verification of the peptide's precise molecular weight was achieved by mass spectrometry (MS), as reported in the mass spectrometry report. Figure 3 As shown, this ensures the accuracy of the polypeptide sequence.

[0140] Peptide dissolution conditions: The purified peptide GLP1R-TXP was dissolved in DMSO (dimethyl sulfoxide) solvent to prepare a 1 mg / mL stock solution for subsequent experiments.

[0141] Cell starvation treatment: The huGLP1R-FL-CRE-HEK293-A5 cell line was constructed, and highly viable huGLP1R-FL-CRE-HEK293-A5 cells were selected and starved in DMEM medium containing 0% fetal bovine serum (FBS) to standardize experimental conditions and eliminate interference from other factors in serum.

[0142] Cell digestion, washing, and plating: Adherent cells were gently digested with Accutase enzyme and washed with PBS buffer to remove residual culture medium and enzymes. The cell suspension was precisely plated into HTRF 96-well low-volume assay plates to ensure a consistent cell count per well, providing uniform starting conditions for subsequent cAMP activity assays.

[0143] cAMP activity assay: In 5 μL of cells (density 3 × 10⁶ cells / mL)... 3 After seeding (cells / well), precisely diluted candidate peptide GLP1R-TXP (see the x-axis of Table 4 for specific concentrations) was added to each well, and the plate was incubated at 37°C for 20 minutes to simulate the in vivo environment and activate the GLP1R receptor. Then, 5 μL of cAMP d2 reagent and anti-cAMP Eu Cryptate antibody (PerkinElmer, catalog number 62AM4PEB) working solution were added to each well, and the plate was incubated at room temperature for 1 hour. This step allows cAMP to bind to the fluorescently labeled antibody, forming an energy transfer complex.

[0144] Signal detection: Using TR-FRET technology, with an excitation wavelength of 340 nm, the emitted signals were detected at 665 nm and 620 nm, respectively. This technology features high sensitivity and anti-interference capability. The FRET signal value was calculated as the ratio of the light intensity at 665 nm to that at 620 nm, reflecting the change in cAMP concentration; the change in the ratio is inversely proportional to the cAMP concentration. Detection results are as follows: Figure 4 As shown in Table 1.

[0145] Table 1

[0146] sample <![CDATA[EC 50 (μM)]]> peptide GLP1R-TXP 2.119

[0147] Experimental results: The synthesized peptide GLP1R-TXP exhibits good cAMP activation activity, and its EC50... 50 The value was 6.873 μM, indicating that it has high biological activity.

[0148] Furthermore, the crystal structure or molecular dynamics of the target protein-peptide complex was analyzed using AlphaFold and HPEPDOCK methods. The simulation results of the crystal structure or molecular dynamics of the target protein-peptide complex can be found in [link to simulation results]. Figure 5 As shown in the figure. The results demonstrate that the peptide can precisely bind to the active site of the GLP1R receptor and exert its biological effects.

[0149] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0150] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A polypeptide or a pharmaceutically acceptable salt thereof, characterized in that, The polypeptide has an amino acid sequence as shown in SEQ ID NO:1 or an amino acid sequence of its conserved modified form.

2. The polypeptide or a pharmaceutically acceptable salt thereof according to claim 1, characterized in that, The conservatively modified amino acid is selected from amino acid X containing -NH2 in its side chain; Optionally, the amino acid X is selected from K, k, αMeK, HoK, Dap, Dab, Orn, or N-Me-K; Optionally, the number of amino acids in the conservative modification is one or more.

3. A polypeptide derivative or a pharmaceutically acceptable salt thereof, characterized in that, include: The polypeptide according to any one of claims 1 to 2, and Modifying groups, wherein the polypeptide is linked to the modifying groups.

4. The polypeptide derivative or a pharmaceutically acceptable salt thereof according to claim 3, characterized in that, The polypeptide has a conserved modified amino acid sequence as shown in SEQ ID NO:1; Optionally, the modifying group is attached to the amino acid X-NH2; Optionally, the modifying group has at least one of the following structures:

5. A fusion protein, characterized in that, Includes the polypeptide described in any one of claims 1 to 2.

6. A reagent or kit, characterized in that, It includes the polypeptide described in any one of claims 1 to 2, or the fusion protein described in claim 5.

7. A pharmaceutical composition, characterized in that, Includes the polypeptide of any one of claims 1-2 or a pharmaceutically acceptable salt thereof, the polypeptide derivative of any one of claims 3-4 or a pharmaceutically acceptable salt thereof, or the fusion protein of claim 5, and Optional pharmaceutically acceptable excipients.

8. Use of the polypeptide of any one of claims 1 to 2 or a pharmaceutically acceptable salt thereof, the polypeptide derivative of any one of claims 3 to 4 or a pharmaceutically acceptable salt thereof, the fusion protein of claim 5, or the pharmaceutical composition of claim 7 in the preparation of a medicament for the treatment or prevention of metabolic disorder-related diseases.

9. The use according to claim 8, characterized in that, The metabolic disorder-related diseases include at least one of obesity, diabetes, dyslipidemia-related diseases, fatty liver disease, metabolic syndrome, and non-alcoholic fatty liver disease.

10. A method for detecting GLP-1R, characterized in that, include: The sample to be tested is contacted with the polypeptide of any one of claims 1 to 2, the fusion protein of claim 5, and the reagent of claim 6 to form an immune complex; Based on the signal from the immune complex, it is determined whether the sample to be tested contains GLP-1R.

Citation Information

Patent Citations

  • Training method and device of molecule binding prediction model, equipment and storage medium

    CN120526840A

  • Ligand information generation model training method and device and ligand information generation method and device

    CN120526881A